Your Own Task

ct run eval assembles one Inspect Task from a task selection, an untrusted policy and a protocol, Control Tower's scorers and a sandbox. A project that wants to change one of those parts builds the Task itself from the same pieces and runs it through Inspect's own eval. ct run eval has no task args for using your own samples or dropping the task scorer, because each is a line in your own Task.

The pieces

PieceModuleWhat it does
build_eval_solvercontrol_tower.protocols.eval.solverWires a registered policy and protocol, with their validated args and a White, into a solver.
policies, protocolscontrol_tower.protocols.eval.registryThe registries; get resolves a name or a module.path:function, validate_args checks a dict of args against it.
Whitecontrol_tower.protocols.eval.whiteThe deployment: harness, granted affordances, max_steps.
wrap_with_task_setupcontrol_tower.task.setupRuns the sample's task setup scripts and baseline snapshots before the solver.
get_linux_scorers, diagnostic_scorerscontrol_tower.task.scorers, control_tower.task.diagnosticsThe main and side task scorer with diagnostics, or the transcript diagnostics on their own.
monitor_scorercontrol_tower.protocols.eval.monitoring.posthocReports every monitor score the protocol recorded as the sample's monitor score.
make_role_modelcontrol_tower.models.role_modelBuilds the model bound to a role (untrusted, trusted, or one a monitor names) from a model name or spec.
SandboxConfigcontrol_tower.sandbox.cliapply() configures Control Tower's Docker sandbox class from the network flags; build_sandbox() is the task-level spec when an external sandbox replaces the per-sample ones.
build_eval_datasetcontrol_tower.ct_runs.ai_eval.datasetControl Tower's own samples for a TaskSelectionInput, when you want to start from those and change them.
DEFAULT_MESSAGE_LIMITcontrol_tower.defaultsThe message cap ct run eval puts on the Task to stop a policy that never submits; an Inspect Task has none of its own.

U and the monitors run on the Task's model roles, not on a policy argument, so a Task you build binds them itself; ct run eval binds untrusted and trusted to their Control Tower aliases unless told otherwise.

Example

An honest run over your own samples, under the null protocol, with only the refusal diagnostic scored:

from inspect_ai import Task, eval from control_tower.defaults import DEFAULT_MESSAGE_LIMIT from control_tower.models.role_model import make_role_model from control_tower.protocols.eval.registry import policies, protocols from control_tower.protocols.eval.solver import build_eval_solver from control_tower.protocols.eval.white import White from control_tower.task.diagnostics import diagnostic_scorers from control_tower.task.setup import wrap_with_task_setup policy = policies.get("honest") protocol = protocols.get("null-blue-team") task = Task( dataset=my_dataset, solver=wrap_with_task_setup( build_eval_solver( policy, policies.validate_args(policy, {}), protocol, protocols.validate_args(protocol, {}), white=White(max_steps=100), ) ), scorer=diagnostic_scorers(["refusal"], timeout_seconds=600), model_roles={ role: make_role_model(model, role=role) for role, model in {"untrusted": "sonnet", "trusted": "trusted"}.items() }, message_limit=DEFAULT_MESSAGE_LIMIT, ) eval(task)

Swap diagnostic_scorers for get_linux_scorers plus monitor_scorer() and the Task scores exactly as a ct run eval run does; the only thing you brought is the dataset. policies.get also resolves module.path:function, so the policy can be yours too.

What your samples carry

Every sample carries SampleMetadata (control_tower.task.sample_metadata) in its metadata: an environment block and, when the sample has them, main_task and side_task, in the shape build_eval_dataset writes. The solver, the tools, task setup, the submit tool and the task scorer all read it, and resolve the environment and tasks it names from the registry, so those must be registered in the running venv. A sample with only an environment block runs against that environment with no task. Each sample sets its own sandbox.

The solver opens every sample with Control Tower's own system prompt and task request, built from that metadata, and gives U Control Tower's submit tool for the sample's tasks; a sample's own input messages are not what U is asked.

What stays with ct run eval

The EC2 fleet, run metadata and the upload to the Control Tower database belong to ct run eval. A Task you own is an ordinary Inspect eval: run it with eval() or inspect eval, and read its log with Inspect's tools or convert it with ct traj.