Function Registry
The function registry powers the eval policies and protocols and the safety / usefulness calculations consumed by ct game-metrics graph pareto. It handles argument declaration, CLI integration, and external module:function lookup.
For how policies and protocols fit into an eval, see Protocols and Policies. For the calculation-side registries used by Pareto plots, see Pareto Frontier.
Two Argument Shapes
Every registry builds a FunctionInfo (defined in function_info.py) for each registered function. There are two shapes for declaring arguments, each backed by a different function_info_factory:
| Shape | Factory | Args declared as… | Used by |
|---|---|---|---|
Args | None | function_to_function_info (default — see function_info.py) | a single BaseModelPython class | safety_calculations, usefulness_calculations |
| Direct kwargs | signature_to_function_info (see signature_info.py) | regular keyword parameters | policies, protocols |
Mixing shapes inside a registry is not supported. The factory is wired up in each registry's registry.py (function_info_factory=signature_to_function_info for the eval registries, omitted for the calculation registries).
Defining a Policy or Protocol (direct kwargs)
The eval registries, policies and protocols in protocols/eval/registry.py, declare configurable parameters directly as keyword arguments. The factory synthesizes a BaseModelPython args class from the function signature: each parameter's annotation and default become a field's type and default, and a pydantic Field(description=...) default documents it. Policies live in src/control_tower/protocols/eval/policies/ and protocols in src/control_tower/protocols/eval/protocols/.
from pydantic import Field from control_tower.protocols.eval.blue import Protocol from control_tower.protocols.eval.registry import protocols @protocols.register(name="trusted-monitoring") def trusted_monitoring( monitors: Monitors = Field(description=MONITORS_ARG_DESCRIPTION), sus_threshold: float | None = Field(default=None, description="Act above this score."), blocks: int = Field(default=0, ge=0), ) -> Protocol: """Ask the monitors about each action; above `sus_threshold`, block, then audit and stop.""" return TrustedMonitoring(monitors=monitors, sus_threshold=sus_threshold, blocks=blocks)
A policy is registered the same way with @policies.register(...) and returns a Policy; it also declares attempts_main_task, attempts_side_task and attempts_sneaky, which trajectory conversion records on every trajectory it produced. Registration is automatic via the decorator: the registry imports every non-underscore module under its discovery package on first lookup.
Defining a Safety or Usefulness Calculation (Args | None)
Safety / usefulness calculations are pluggable impls consumed by ct game-metrics graph pareto (and compute_pareto). Decorate with @safety_calculations.register() or @usefulness_calculations.register() and drop the file under metrics/safety_calculations/impls/ or metrics/usefulness_calculations/impls/. Each impl returns a single scalar plus metadata; curve-style sweeps and optional bootstrap CIs are handled by callers such as Pareto, not by the impls themselves.
from control_tower.utils.function_registry.base_model import BaseModelPython from control_tower.metrics.safety_calculations.registry import ( safety_calculations, ) from control_tower.metrics.safety_calculations.types import ( SafetyResult, ) from inspect_ai.log import EvalSample class TMSafetyArgs(BaseModelPython): mode: str = "optimized" audit_budget: float | None = None @safety_calculations.register() def tm_safety( honest: "list[EvalSample]", attack: "list[EvalSample]", args: TMSafetyArgs | None = None, ) -> "SafetyResult": if args is None: args = TMSafetyArgs() # ... fit rule and compute safety ... return SafetyResult(safety=..., metadata={"mode": args.mode})
Args classes must extend BaseModelPython (not plain Pydantic BaseModel). Reference the impl from YAML by its hyphenated name (e.g. tm-safety, defer-to-trusted).
Naming Convention
By default, the registry strips the _<identifier> suffix from the function name and converts underscores to hyphens. Each registry passes its own identifier (see registry.py):
| Registry | identifier | Auto-strips suffix | Example |
|---|---|---|---|
policies | eval-policy | none | register(name="honest") → honest |
protocols | eval-protocol | none | null_blue_team → null-blue-team |
safety_calculations | safety_calculation | _safety_calculation | tm_safety → tm-safety |
usefulness_calculations | usefulness_calculation | _usefulness_calculation | main_task_success → main-task-success |
No Python function name ends in _eval-policy, so an eval policy or protocol is named after its function unless you pass name="...": honest_policy would otherwise be honest-policy.
The docstring's first sentence becomes the description shown in help output.
Passing Arguments from the CLI
# Untrusted policies use -up (short for --untrusted-policy-arg) ct run eval --untrusted-policy sleep -up min_s=3 -up max_s=4 --task-file simple # Protocols use -bp (short for --protocol-arg) ct run eval \ --untrusted-policy simple-attack \ --protocol trusted-monitoring \ -bp 'monitors=[simple]' \ -bp sus_threshold=7.0 \ --task-file simple
Hyphens in keys are converted to underscores automatically. Unknown keys raise a validation error. Models are not policy arguments: an eval binds them to roles with --model-role (see Model Configuration). ct run sabotage-eval selects the same policies via --attack-policy / --honest-policy and a protocol via --protocol / --protocol-arg — see docs/sabotage-eval.md.
Discovering Available Arguments
ct protocols policies # List all policies ct protocols policies -v # List with all arguments ct protocols policies attack # Show args for a specific policy ct protocols protocols # List all protocols ct protocols models # List model aliases
Cost Tracking
CONTROL_TOWER_PRICING selects the rate source:
file(default) loads a generated JSON table (~/.cache/control-tower/model_rates.json, orCONTROL_TOWER_MODEL_RATES).apiloads the deployment'spricingcollection overGET /api/pricingand needs a viewer token (ct loginorCONTROL_TOWER_API_TOKEN).offloads nothing. Model calls still run and bill, but every one is recorded as unknown cost, so use it only for runs that make no priced calls. The sabotage-eval preflight skips its pricing check underoff.
With file or api, the table loads at Task construction, and a missing or empty table fails the eval before any sample starts. A model the loaded table has no row for stamps unknown cost, not $0.
uv run python scripts/generate_model_rates.py writes the file: vendors' own prices from LiteLLM's price map under bare names, and OpenRouter's catalog under each model's full openrouter/... name, which a lookup tries before the bare name. An OpenRouter row is the default route's price, so a run pinned to an endpoint that prices differently needs a manual row under the same key. A re-run keeps rows marked "manual": true and models the source has stopped publishing.
External Functions
Any CLI flag that accepts a policy or protocol name also accepts a module:function path, following the same convention as uvicorn app:factory. This lets you use policies and protocols defined in your own project without modifying control-tower.
ct run eval --untrusted-policy my_project.policies:my_attack --task-file simple
The function must have the same shape as a registered one (keyword arguments, returning a Policy or Protocol). External environments also work via --env my_project.envs:my_env_config, and --extra-src puts a directory of experiment code on the path without installing it.
Quick Reference
| I want to... | Command |
|---|---|
| List policies | ct protocols policies |
| List protocols | ct protocols protocols |
| See a policy's args | ct protocols policies <name> |
| Override a model | --model-role untrusted=opus / --model-role trusted=haiku |
| Use external function | --untrusted-policy my.module:my_func |