Trajectories
trajectories reads agent behavior straight off Inspect eval samples; the .eval is the single source of truth.
A trajectory is the recorded behavior of one sample:
- environment/main-task/side-task context (
SampleMetadataon the sample) - action/tool-call history
- timing/cost/monitor metadata
- optional run linkage (
run_id) and eval linkage (eval_log_url, sample ID)
CLI Overview
The ct traj group's primary commands:
-
ct traj list [--fs] [-e ENV] [-t MAIN_TASK] [-s SIDE_TASK] [--tag TAG] [--limit N]- Lists trajectories from the API (default) or local bundle files (
--fs).
- Lists trajectories from the API (default) or local bundle files (
-
ct traj download [<trajectory_id_or_viewer_url>] [--eval] [-e ENV] [-t MAIN_TASK] [-s SIDE_TASK] [--tag TAG] [-n N] [--dataset NAME] [--dataset-version V] [-o OUTPUT] [--overwrite]- Downloads one trajectory by ID from the viewer API, or by viewer URL (e.g.
https://basharena-data.vercel.app/trajectories/<id>; see Downloading from the Viewer). - With filter options instead of an ID, bulk-downloads the latest N matching trajectories (default 20, max 500).
--dataset NAMEdownloads every trajectory in a dataset, with no count limit, using the same API token as other viewer reads (ct loginorCONTROL_TOWER_API_TOKEN). It fetches the dataset's trajectory ID list and downloads by ID in batches.--dataset-version Vpicks a version (default: latest).--evaldownloads the trajectory's full.evalarchive instead of the bundle JSON.
- Downloads one trajectory by ID from the viewer API, or by viewer URL (e.g.
-
ct traj patch <traj> ...- Generates a patch from a trajectory using live tooling.
-
ct traj apply <eval_log_file> ...- Applies file edits from an eval log into an environment codebase.
-
ct traj tui <trajectory_path> ...- Opens the trajectory TUI for interactive review/editing.
Result accessors on a sample
control_tower.trajectories.sample reads a sample's results from its EvalSample:
main_task_success(sample)andside_task_success(sample)returnTrueorFalsewhen the scorer decided, andNone, notFalse, when the score is missing, the scorer returned a non-answer, or the scorer errored.executed_actions(sample)returns every tool call the sample committed, in order, each as its assistant message,ToolCalland result; calls insample.outputthat never reached the messages close the list, and assistant turns taggedis_monitor_resultare skipped.executed(sample, action)says whether a call ran, which is whether aToolEventcarries its id.action_scores(sample, monitor_name=None)returns the action monitor's per-action scores, andtrajectory_score(sample, monitor_name=None)the trajectory monitor's score. Withoutmonitor_name, the sample must carry at most one monitor of that kind.sample_metadata(sample)returns the task and environment the sample ran under, andsample_error(sample)its execution or scoring error, if any.
Cost accessors on a sample
The functions below live in control_tower.trajectories.sample; each takes the EvalSample you want to read.
Every cost except cost_by_role is inspect's stamp on usage.total_cost. Sample totals come from sample.model_usage; per-call figures come from each ModelEvent. An Inspect response-cache read is left out of model_usage (actual spend) and marked on timing.model_calls[i].inspect_cache_hit. Provider prompt-cache discounts are already in the stamp.
- Per-call:
action.timing.model_calls[i].cost_usd. - Run total:
sample.agent_cost(sample)is{model: usage.total_cost}fromsample.model_usage, which counts every model call in the run, live monitors included; usecost_by_rolefor the agent alone. Anullamount is an unpriced model. - Monitors:
sample.monitor_cost(sample) -> Pricedsums each monitor score'scost_usdmetadata (see Cost Tracking in function-registry.md); a score with nocost_usdis a call nobody could price. It walks ensemblememberswithout double-counting the parent. - By role:
cost_by_role(sample) -> dict[str, Priced]prices each model call at the rate table's current rates, not inspect's stamp, and keys it by the call'srole(names inmodels/roles.py;AGENT_ROLESis the agent). A response-cache hit counts at its original tokens, so the figure is what a fresh run costs. A call with no role raises, which it does on logs from before every Control Tower model carried one. - Grand total:
sample.total_cost(sample) -> Pricedis agent + monitors, per model. ANonein itsby_modelmeans the total is short, and the viewer shows the known total as$N+.sample.total_cost_usd(sample)is known spend, orNoneonly when nothing at all was priced.
Workspace diff on a sample
When a run's main task declared a final_diff config, control_tower.trajectories.sample.workspace_diff(sample) reads the final workspace diff captured at scoring time: the unified diff text (truncated to the task's byte budget), a short stat summary, the baseline it was computed against, and whether truncation applied. It is None for tasks without the config, for simulated and replay-sandbox runs, and for samples without captured diff metadata.
How trajectories fits with other CLI groups
ct run ...runs experiments and can produce trajectories.ct runs ...manages run-level containers around groups of trajectories.ct traj ...is the direct trajectory manipulation and analysis surface.
Typical workflows
List trajectories from API:
ct traj list -n 20
Download one trajectory:
ct traj download <trajectory_id>ct traj download https://basharena-data.vercel.app/trajectories/<id>(viewer URL)ct traj download <trajectory_id> --eval(full.evalarchive instead of bundle JSON)
Bulk-download the latest matching trajectories:
ct traj download -e web_scraping --tag baseline -n 100
Download every trajectory in a dataset (latest version, or a specific one):
ct traj download --dataset LaStrajct traj download --dataset LaStraj --dataset-version 2