How a JevAgent run is metered
JevAgent Usage and Pricing
This page offers a clear overview of JevAgent Usage and Pricing. It helps you explore this topic and find a useful next step.
1. System 1 and Jev are priced separately
A JevAgent run can call two kinds of model. System 1 is the generative provider and model you set on `JevAgentSettings`; it writes replies and drives the tool loop. System 2 is Jev, TypeSafe's decision model, which you configure on `JevRuntimeSettings.decision`; it answers the structured questions behind preflight, tool selection, specialist routing, and done checks.
The SDK prices both with one pricebook, `ModelPricingRegistry`. A generative call is priced from its provider's usage fields at that model's input, output, and cached-input rates. A Jev call reports only `input_tokens` and `output_tokens`: input is billed at TypeSafe's rate and output is free. TypeSafe publishes no cached-token rate, so every Jev input token bills at the same rate.
| Model | Input | Output | Cached input |
|---|---|---|---|
| Jev (`jev-latest`, `jev-preview`, `jev-1.13`) | $0.042 per million tokens | Free | No cache tier |
| System 1 generative model | That model's pricebook rate | That model's pricebook rate | Provider-specific, where the provider reports it |
Rates come from the SDK's built-in pricebook (`vidbyte/lib/registries/pricing.py`). A model that has no pricebook entry is recorded with `cost_usd=None`, and the run's `cost_complete` becomes False.
2. `agent.get_usage()` covers only the main agent loop
`agent.get_usage()` returns a `UsageRollup` for the main agent's own generative calls, including any extra turns a continuation sends it back for. Each call is a priced `UsageRecord`, and the rollup carries token totals, `cost_usd`, `cost_complete`, and `recording_integrity`. The same rollup is attached to the reply as `metadata["usage_rollup"]`.
Jev decision calls and JevAgent's helper agents are not added to that rollup. They are reported where the feature that made the call records its result, mostly on `agent.response`. To know what a whole run cost, read each source in the table below.
| Run phase | Model | Where its usage is recorded | In `get_usage()`? |
|---|---|---|---|
| Main agent loop and continuations | System 1 | `agent.get_usage()` and the reply's `metadata["usage_rollup"]` | Yes, priced |
| Preflight gate: presets and the specialist question, sent as one Jev request | Jev | `agent.response.usage`, as token counts (`JevUsage`) | No |
| Tool selector | Jev | The reply's `metadata["jev_tool_selector"]["usage"]`, as input and output tokens | No |
| Clarification agent | System 1 | `agent.response.clarification.usage`, a separate `UsageRollup` | No |
| Run state, reviewer, and handoff helpers | System 1 | `agent.response.run_state.usage`, `.review.usage`, and `.handoff.usage` | No |
| Done checks | Jev | `agent.response.done[check].usage`, as token counts | No |
| A chosen specialist | The specialist's model | That specialist agent's own `get_usage()` | No |
Every enabled done check is asked in one Jev request per finish attempt, so each check's result carries that same usage. Count it once, not once per check. `agent.response` keeps the latest finish attempt's done, review, and handoff records, so usage from earlier finish attempts is not retained there.
3. Add up a run's cost yourself
Until the SDK merges these sources into one rollup, sum them after the run. Price Jev token counts with the same pricebook the SDK uses, and read the helper agents' rollups from `agent.response`.
Total one JevAgent run
from vidbyte import ModelPricingRegistry
from vidbyte.lib.enums import ModelProvider
reply = agent.run("Summarize the renewal options for this account.")
response = agent.response
# System 1: the main loop is already priced.
main = agent.get_usage()
# Jev: usage arrives as token counts, so price it with the SDK pricebook.
jev_rate = ModelPricingRegistry.default().resolve(ModelProvider.TYPESAFE, "jev-latest")
preflight_jev = response.usage.cost_usd(jev_rate) if response.usage else 0.0
# All done checks share one Jev request per finish attempt: count it once.
done_usage = next((result.usage for result in response.done.values() if result.usage), None)
done_jev = done_usage.cost_usd(jev_rate) if done_usage else 0.0
# Helper agents keep their own generative rollups.
helpers = [part.usage for part in (response.clarification, response.run_state, response.review, response.handoff) if part and part.usage]
helper_cost = sum(rollup.cost_usd or 0.0 for rollup in helpers)
print(main.cost_usd, main.cost_complete, preflight_jev + done_jev, helper_cost)4. Who pays TypeSafe for Jev calls
Bring your own key (BYOK): set `TYPESAFE_API_KEY`, or pass `JevRuntimeSettings(decision=DecisionModelConfig(api_key=...))`. The SDK calls `api.typesafe.ai` directly, TypeSafe bills your TypeSafe account, and the key and request never reach Vidbyte. Nothing is charged to a Vidbyte wallet. Generative providers work the same way: the key you configure for System 1 is billed by that provider.
Managed Jev gateway: Vidbyte's backend exposes `POST /api/v1/models/typesafe/systemone`, which calls TypeSafe with Vidbyte's own key. It needs a Vidbyte API key (`vb_live_...`) with the `models:invoke` scope, and a wallet balance of at least one cent before any call is forwarded. Vidbyte bills only the usage TypeSafe reports in its own response, priced with the same SDK pricebook. A failed call, or an answer with no usage, bills nothing.
Gateway charges are summed per run in fractions of a cent, and a 10% markup is applied to the run total, not to each call. An open run is billed in whole cents as they accrue; closing the run bills the final part-cent once. Send an `X-Vidbyte-Run-Id` header to group calls into one run. Calls without it share one run per API key per UTC day. A run with no calls for one hour is closed automatically, and `POST /api/v1/models/runs/{run_id}/close` closes it explicitly and returns its calls, tokens, provider cost, fee, and billed cents.
| Aspect | BYOK | Managed gateway |
|---|---|---|
| Key | Your `TYPESAFE_API_KEY` | Your Vidbyte API key with `models:invoke` |
| Endpoint | `https://api.typesafe.ai/v1` | `/api/v1/models/typesafe` on the Vidbyte API |
| Who bills you | TypeSafe, directly | Your Vidbyte wallet |
| Price | TypeSafe's rate | TypeSafe's rate plus a 10% markup on the run total |
| Vidbyte sees the request | No | Yes, to forward and meter it |
The SDK has no managed-mode switch yet. You reach the gateway by setting `DecisionModelConfig(endpoint=..., api_key=...)` yourself, and the SDK does not send `X-Vidbyte-Run-Id`, so those calls fall into the per-day run.
5. What usage tracking does not do yet
`agent.get_usage()` does not include Jev calls, helper agents, or a chosen specialist.
Jev usage is reported as token counts on `agent.response`; its dollar cost is not computed for you.
Usage from earlier finish attempts' done checks, review, and handoff is not retained on `agent.response`.
The SDK cannot yet route Jev calls through Vidbyte's managed gateway on its own or tag them with a run id.
Models without a pricebook entry are recorded with no cost, so a run that uses one reports `cost_complete=False`.