Inference is part of the job cost

13 July 2026 · Updated 12 August 2026 · 4 min read · Moji Team


A production AI task is rarely one model call. It may retrieve evidence, classify it, draft an output, check the result and recover a failed step. The useful cost unit is the complete job because that is the unit the customer needs delivered.

Provider dashboards record calls and tokens. Those records are necessary, but they do not show whether the job finished, whether the output passed, or how much recovery it needed. A production team needs those facts attached to the same run.

Set the budget around the work

A repeatable line of work can be expressed as a multi-step structure. Each step has a model and effort level, a predicted price, an acceptance check and a recovery rule. The step prices roll up to a budget for one item moving through that line.

That budget can also roll up across the week. The team can see how many jobs fit inside the committed spend, which lines are using the capacity and which jobs are likely to breach a threshold before the week closes.

Measure delivery with cost

A cheaper call has little value when the job fails later. The operating record therefore keeps cost beside first-pass yield, exceptions, cycle time and budget reliability. A failed step is recovered in place, with the completed state of the run preserved, and its extra cost remains visible.

This is the basis for a useful production decision: how many jobs can the workflow complete this week while holding the budget, acceptance criteria, reliability and latency thresholds the team committed?

Start from observed work

Moji connects to one service or repository and observes repeated runs. It groups those runs into lines of work, proposes the structure and prices it from the connected workflow. A person commits the budget before live enforcement begins.

The result is a cost model tied to delivered work. See how Moji controls a production workflow or bring us one workflow to measure.