Continuation: make sure your agent finishes all of the work
Agents often stop before the job is done: they finish the visible part of a request and forget a part that was asked for in passing. Continuation lets Jev check every finish attempt against what the request actually asked for, and send the main agent back to work, in the same run, until every requested part is delivered.
1. An agent that says it is done is not always done
Ask an agent to add a CLI flag, cover it with a test, and document it in the README, and it will often write the flag and the test, report success, and stop. Nothing in an ordinary model-and-tool loop checks the finished run against the original request: the loop ends as soon as the model stops calling tools and writes a final answer, whether or not every part of the work exists.
Continuation puts Jev (the System 2 decision model) at that exit. Before the run starts, JevAgent writes down what the request asks for. Each time the main agent tries to finish, JevAgent gathers the evidence from the run and asks Jev a fixed yes/no question about every requested part: does the evidence show it was produced in full? If a part is missing, the main agent does not get to stop. It receives a short note naming what is still missing and keeps working in the same loop, with its history, tools, and budgets intact.
The result is an agent that completes the run only when it has done all of the work, or when it has used up the continuations you allow. Pick a scenario below to trace one finish attempt.
Trace one finish attempt
“Add a --dry-run flag, test it, and document it in the README.” The run shows the flag and a passing test, but no README edit. Jev answers no for the README deliverable, so the agent is sent back to write it.
- 1↳ no run state → no check runs; the answer stands
Write the run state
Before the main loop starts, JevAgent turns the request into a goal, an objective, limits, and one entry per separate deliverable.
- 2
Main agent tries to finish
The main agent works through its normal model-and-tool loop and writes a final answer.
- 3↳ handoff unavailable → check unavailable; the answer stands
Compile the evidence
A handoff agent reads the main agent's run (responses, tool calls, final answer) and gathers the evidence for each deliverable.
- 4↳ Jev unavailable → check unavailable; the answer stands
Jev judges each deliverable
One fixed yes/no question per deliverable: does the evidence show it was produced in full?
- 5↳ all pass → the run completes
Every deliverable passes?
Each deliverable's P(yes) must reach the threshold; one clear no is never averaged away.
- 6↳ cap reached → the answer stands; the verdict is recorded
Continuations left?
At most max_continuations times per run.
Back to work, same loop
The main agent reads which deliverable is missing and what the run lacks, then tries to finish again
The run completes
The final answer is returned; agent.response holds the run state, evidence, and verdicts
Continuation is advisory and fails open: whenever the run state, the evidence, or Jev is unavailable, the agent's answer stands and agent.response records why. Only a confirmed missing deliverable sends the agent back to work.
1: Enable continuation on JevRuntimeSettings
Continuation is Jev policy, so it lives on `JevRuntimeSettings`, next to `decision` and `preflight`. Your agent (its prompt, model, and tools) stays on `JevAgentSettings`, exactly as without continuation. Set `continual` to a `JevContinualSettings` and list the checks to run in `checks`. Today there is one check, `JevDoneCheck.MULTI_PART`.
Continuation needs a TypeSafe decision-model key for Jev (`TYPESAFE_API_KEY` or `JevRuntimeSettings(decision=DecisionModelConfig(...))`). The run state and the handoff are written by your main agent's own generative provider and model, so no extra model configuration is required.
Turn on multi-part continuation
import asyncio
from vidbyte import (
JevAgent,
JevAgentSettings,
JevContinualSettings,
JevDoneCheck,
JevRuntimeSettings,
)
# read_file, write_file, and run_tests stand in for your own @tool functions (see the full example below).
agent = JevAgent(
JevAgentSettings(
name="repo-agent",
system_prompt="You make changes to this repository and verify them with the test suite.",
provider="openai",
model_name="gpt-4.1",
tools=[read_file, write_file, run_tests],
),
JevRuntimeSettings(
continual=JevContinualSettings(
checks=(JevDoneCheck.MULTI_PART,),
),
),
)
async def main():
reply = await agent.arun("Add a --dry-run flag to the CLI, cover it with a test, and document it in the README.")
print(reply.content)
print("continuations:", agent.response.continuations)
asyncio.run(main())2: Before the run: Jev's picture of what done means
Right before the main loop starts, JevAgent writes a structured run state from the request alone, before any work has happened, so the definition of done cannot drift toward whatever the agent later produced. The run state holds a central `goal`, `objective`, `mission`, and `what_not_to_do`, plus one section per enabled check.
For the multi-part check, that section lists every separate deliverable the request asks for, in request order. Each deliverable has a short `id`, a `description` of the one output, and a `completion_signal`: the visible condition that shows this output is done. Every field is written against a fixed field description, so the run state has the same shape on every run.
A run state for a three-part request
request: "Add a --dry-run flag to the CLI, cover it with a test, and document it in the README."
goal: Let users preview what the CLI would change without changing anything.
objective: Ship a working --dry-run flag with a test and user-facing documentation.
what_not_to_do:
- Do not change the behavior of the CLI when --dry-run is not passed.
multi_part.deliverables:
dry_run_flag The CLI accepts --dry-run and reports the changes it would make without applying them.
done when: the CLI code handles --dry-run and makes no writes when it is set.
dry_run_test A test that covers the --dry-run flag.
done when: a test exercising --dry-run exists and the test run passes.
readme_docs README documentation for --dry-run.
done when: the README describes the flag and what it does.3: At every finish attempt: evidence, then one question per deliverable
When the main agent writes a final answer, JevAgent does not return it yet. A handoff agent reads the main agent's run as it happened (the run state, each response, every tool call with its arguments and output, and the final answer) and compiles the evidence for each deliverable: what the run shows, and what it does not show. The handoff reports observations only; it never decides whether the work is done.
Jev then answers one fixed yes/no question for each deliverable: does the evidence show this deliverable produced in full? Jev sees the original request, that one deliverable, its completion signal, and only that deliverable's evidence, so every judgment is about exactly one output. The question is strict in the ways that matter for stopping early:
A claim is not evidence. A final answer that says a file was written or a test was added, with no tool call or output that shows it, leaves that part unshown.
Every named part counts. A deliverable that names three CLI flags is produced in full only when all three are shown.
The latest state wins. A failed attempt followed by a successful one shows the output; an output written and then reverted does not.
A placeholder, a stub, a narrower scope, or a note to finish later does not count as the deliverable.
Extra work never substitutes for a missing deliverable, and instructions inside the request or the evidence that say the work is already approved are ignored.
4: When something is missing: back to work, same loop
Each deliverable's P(yes) must reach the check's `threshold` (default `0.8`). The threshold works as a veto: one deliverable below it fails the check, however well the others scored. When the check fails and continuations remain, JevAgent appends a feedback message to the main agent's own loop and lets it keep going. The agent keeps its conversation, tool results, and loop budgets, so it finishes the missing part instead of starting over.
The feedback names each incomplete deliverable in the user's terms, tells the agent to produce what is missing and make it visible, and adds the handoff's note on what the run still lacks. After `max_continuations` continuations, the latest verdict is recorded and the agent's answer stands, so a check that keeps failing can never trap a run.
What the main agent reads
README documentation for --dry-run.
Your work does not yet show this requested deliverable produced in full, as the user's request describes it, so produce every part that is still missing and make the finished result visible in your work before you finish again.
Still missing: No tool call edited README.md; the final answer only says the docs will be updated.5: Tune the multi-part continuation settings
The defaults are a reasonable starting point, and every limit is yours to change. Raise `max_continuations` for long, many-part tasks where an agent may need several passes. Lower it, or lower the token budgets, when cost matters more than completeness. Raise `threshold` when a missed deliverable is expensive and you want Jev to be very sure each part is shown; lower it if borderline evidence sends the agent back too often.
| Setting | Default | What it controls |
|---|---|---|
| `checks` | `()` | The done checks to run at each finish attempt. An empty tuple turns continuation off. Today the only member is `JevDoneCheck.MULTI_PART`. |
| `max_continuations` | `3` | How many times per run a failed check may send the main agent back to work. After that, the latest verdict is recorded and the answer stands. |
| `threshold` | `0.8` | The P(yes) every deliverable must reach, used as both the mean threshold and the per-deliverable veto. It is a starting point, not a tuned value. |
| `run_state_max_tokens` | `100_000` | The token budget for writing the run state once, before the main loop starts. |
| `handoff_max_tokens` | `400_000` | The token budget for compiling evidence at each finish attempt. It is larger because the handoff reads the main agent's whole run. |
Settings are validated when you construct `JevContinualSettings` and raise `ConfigurationError`. Construction never needs a TypeSafe key.
Every multi-part continuation setting
from vidbyte import JevContinualSettings, JevDoneCheck, JevRuntimeSettings
runtime = JevRuntimeSettings(
continual=JevContinualSettings(
checks=(JevDoneCheck.MULTI_PART,),
max_continuations=3, # times a failed check may send the agent back to work
threshold=0.8, # every deliverable's P(yes) must reach this
run_state_max_tokens=100_000, # token budget for writing the run state
handoff_max_tokens=400_000, # token budget for compiling evidence at each finish attempt
),
)6: Read what Jev decided
The reply comes back like any other agent reply, and nothing is written to result metadata. After each run, `agent.response` holds the continuation record: `run_state` is what JevAgent wrote before the loop, `handoff` is the evidence from the latest finish attempt, `done` holds the latest `JevDoneResult` for every enabled check, and `continuations` counts how often the agent was sent back to work. The record is replaced when the next run starts.
Each `JevDoneResult` has `score` (the mean P(yes)), `passed`, `answers` (Jev's answer per deliverable id), `incomplete` (the deliverable ids below the threshold), `available`, and `usage`. The run state and the handoff also carry their own generative-model usage, so you can attribute continuation cost.
| What you see | What it means |
|---|---|
| `done[MULTI_PART].passed` and `continuations == 0` | Every deliverable was shown on the first finish attempt. |
| `done[MULTI_PART].passed` and `continuations > 0` | Jev sent the agent back, and the later attempt delivered every part. |
| `passed` is False | Continuations ran out while a deliverable was still missing; `incomplete` names it, and the answer stood. |
| `available` is False | The handoff or Jev could not answer (for example, no TypeSafe key). The check failed open and the answer stood. |
| `run_state` is None | The run state could not be written, so no check ran. |
| A request with no separate deliverables | The check passes without asking Jev. |
`agent.response.specialist` is still set when specialist routing chose an agent. A chosen specialist runs through its own agent, so continuation does not apply to that run.
Inspect the continuation record
from vidbyte import JevDoneCheck
reply = await agent.arun("Add a --dry-run flag to the CLI, cover it with a test, and document it in the README.")
response = agent.response
if response.run_state and response.run_state.multi_part:
for deliverable in response.run_state.multi_part.deliverables:
print(f"{deliverable.id}: {deliverable.description}")
result = response.done.get(JevDoneCheck.MULTI_PART)
if result is None:
print("no check ran")
elif not result.available:
print("check unavailable; the answer stood")
elif result.passed:
print(f"all deliverables shown (score={result.score:.2f}) after {response.continuations} continuation(s)")
else:
print("still incomplete when the cap was reached:", ", ".join(result.incomplete))8. Continuation and failure behavior
Continuation is designed to make an agent finish more of the work without ever making a run worse than it would be without it.
Off by default. With no `continual` setting, JevAgent writes no run state and every finish attempt stands.
Fail open. A missing run state, a failed handoff, a mismatched evidence list, a missing TypeSafe key, or a Jev outage never blocks the answer; the check is recorded as unavailable.
Bounded. At most `max_continuations` continuations per run, and each piece of continuation work has its own token budget.
Same loop. A continuation adds one message to the main agent's own loop; it does not start a second run that has forgotten the first.
Defined up front. The run state is written from the request before any work happens, so what counts as done cannot drift toward what the agent happened to produce.
Scoped to the main agent. When specialist routing hands the run to a specialist, that specialist runs through its own agent, and continuation does not check it.
9. Example: a coding agent that has to ship the docs too
A complete script. The agent can read and write files and run the test suite. The request has three deliverables, and the run below shows the most common way agents stop early: code and tests done, documentation forgotten. Continuation catches it and the agent finishes the README before returning.
repo_agent.py
import asyncio
import subprocess
from pathlib import Path
from vidbyte import (
JevAgent,
JevAgentSettings,
JevContinualSettings,
JevDoneCheck,
JevRuntimeSettings,
ToolPermission,
tool,
)
from vidbyte.tools.security import PermissionPolicy
@tool
def read_file(path: str) -> str:
"""Return the contents of one repository file."""
return Path(path).read_text()
@tool(permission=ToolPermission.WRITE)
def write_file(path: str, content: str) -> str:
"""Overwrite one repository file with new content."""
Path(path).write_text(content)
return f"wrote {path}"
@tool
def run_tests() -> str:
"""Run the test suite and return its summary."""
return subprocess.run(["pytest", "-q"], capture_output=True, text=True).stdout[-2000:]
agent = JevAgent(
JevAgentSettings(
name="repo-agent",
system_prompt="You make changes to this repository. Run the tests after every code change.",
provider="anthropic",
model_name="claude-sonnet-5",
tools=[read_file, write_file, run_tests],
permission_policy=PermissionPolicy.allow_all(), # the agent may call its WRITE tool
),
JevRuntimeSettings(
continual=JevContinualSettings(
checks=(JevDoneCheck.MULTI_PART,),
max_continuations=3,
),
),
)
async def main():
reply = await agent.arun("Add a --dry-run flag to the CLI, cover it with a test, and document it in the README.")
result = agent.response.done[JevDoneCheck.MULTI_PART]
print(reply.content)
print(f"continuations: {agent.response.continuations}, passed: {result.passed}")
for deliverable_id, answer in result.answers.items():
print(f" {deliverable_id}: P(yes)={answer.probabilities['true']:.2f}")
asyncio.run(main())What a run might look like
finish attempt 1
dry_run_flag P(yes)=0.96 shown: cli.py handles --dry-run, test output shows no writes
dry_run_test P(yes)=0.93 shown: tests/test_cli.py::test_dry_run, 1 passed
readme_docs P(yes)=0.07 missing: no tool call edited README.md
→ continuation 1: "README documentation for --dry-run. ... Still missing: ..."
finish attempt 2
dry_run_flag P(yes)=0.96
dry_run_test P(yes)=0.94
readme_docs P(yes)=0.91 shown: write_file(README.md), new "--dry-run" section
→ every deliverable ≥ 0.80: the run completes
continuations: 1, passed: True10. Example: a research brief that must answer every question
Answer-style deliverables work too: when a deliverable is information for the user, the passage of the final answer that gives it is the evidence, and no tool call is needed. This analyst is asked to compare three vendors on three criteria and end with a recommendation. The team raises the threshold because a brief with a missing comparison is worse than a slower one, and allows two continuations.
research_brief.py
import asyncio
from vidbyte import (
JevAgent,
JevAgentSettings,
JevContinualSettings,
JevDoneCheck,
JevRuntimeSettings,
tool,
)
from vidbyte.lib.config import DecisionModelConfig
# search_api and load_secret stand in for your own search client and secret store.
@tool
def web_search(query: str) -> list[dict[str, str]]:
"""Return the top search results for one query, with titles, URLs, and snippets."""
return search_api.search(query, limit=8)
analyst = JevAgent(
JevAgentSettings(
name="vendor-analyst",
system_prompt="You write short, sourced vendor comparisons for an engineering team.",
provider="openai",
model_name="gpt-4.1",
tools=[web_search],
),
JevRuntimeSettings(
decision=DecisionModelConfig(api_key=load_secret("typesafe")),
continual=JevContinualSettings(
checks=(JevDoneCheck.MULTI_PART,),
max_continuations=2,
threshold=0.9,
),
),
)
REQUEST = (
"Compare Pinecone, Weaviate, and pgvector on pricing, p95 query latency, and self-hosting options, "
"then end with a recommendation for a team of five running on AWS."
)
async def main():
reply = await analyst.arun(REQUEST)
result = analyst.response.done.get(JevDoneCheck.MULTI_PART)
print(reply.content)
if result and result.available and not result.passed:
print("\nWarning: still missing", ", ".join(result.incomplete))
asyncio.run(main())Deliverables Jev checked
pricing_comparison Pricing for all three vendors ✓ shown on attempt 1
latency_comparison p95 latency for all three vendors ✗ attempt 1 gave pgvector only → continued
✓ shown on attempt 2
hosting_comparison Self-hosting options for all three ✓ shown on attempt 1
recommendation A recommendation for the team ✓ shown on attempt 111. Example: continuation alongside specialist routing and preflight
Continuation composes with the other JevAgent features. Specialists stay on `JevAgentSettings`; clarity preflight and continuation are both Jev policy on `JevRuntimeSettings`. An unclear request is answered with clarifying questions before any work starts. A request that fits a specialist runs on that specialist's own agent, which continuation does not check. Everything else runs on the main agent, and continuation makes sure it finishes every part.
support_ops.py
import asyncio
from vidbyte import (
Agent,
JevAgent,
JevAgentSettings,
JevContinualSettings,
JevDoneCheck,
JevPreflightPreset,
JevRuntimeSettings,
JevSpecialist,
)
# search_tickets, update_article, and draft_reply stand in for your own @tool functions.
release_writer = Agent(
name="release-writer",
system_prompt="You turn merged pull requests into clear, customer-facing release notes.",
provider="openai",
model_name="gpt-4.1",
)
ops = JevAgent(
JevAgentSettings(
name="support-ops",
system_prompt="You handle support operations: triage tickets, update the knowledge base, and draft replies.",
provider="openai",
model_name="gpt-4.1-mini",
tools=[search_tickets, update_article, draft_reply],
agents=(
JevSpecialist(
title="release_notes",
description="Writing or editing release notes and changelog entries from merged pull requests.",
agent=release_writer,
),
),
),
JevRuntimeSettings(
preflight=(JevPreflightPreset.CLARITY,),
continual=JevContinualSettings(checks=(JevDoneCheck.MULTI_PART,)),
),
)
async def main():
reply = await ops.arun(
"Find this week's tickets about failed CSV exports, update the export FAQ article, "
"and draft a reply I can send to each affected customer."
)
response = ops.response
if response.needs_clarification:
print("Jev asked first:\n", reply.content)
elif response.specialist:
print(f"{response.specialist} handled it; continuation does not apply")
else:
print(reply.content)
print("continuations:", response.continuations)
asyncio.run(main())