Execution & Agents
How AXIS executes scenarios, manages agent processes, and isolates workspaces.
Execution Model
When you run axis run, AXIS loads your config, discovers scenarios, and executes
each scenario/agent combination as an independent job. Jobs run in parallel up to the configured
concurrency limit (default: 15).
Scenarios are selected in a fixed order, narrowing at each step. AXIS walks the
scenarios source, applies the active suite's
include then exclude, narrows to
each agent's own scenarios list, then applies the --scenario filter.
Because each step only narrows, an agent filter or a -s flag cannot reach a
scenario the suite excluded.
Each job follows the same lifecycle:
- Setup -Run setup actions defined in the scenario (if any).
- Spawn -Start the agent process in an isolated workspace.
- Capture -Stream and record the full interaction transcript.
- Score -Evaluate the transcript against the judge and interaction signals (unless
--no-scoreis set). - Teardown -Run teardown actions (if any).
- Save -Write the result to the report.
Each scenario has a 15-minute timeout by default (configurable via limits.scenario.time_minutes
or per-scenario limits). If the agent does not finish in time, AXIS sends SIGTERM,
waits briefly, then SIGKILL. Timed-out runs are marked as failed with a timeout error.
Supported Agents
AXIS ships native adapters for Claude Code, Codex, and Gemini, plus every major AI coding agent that speaks the Agent Client Protocol (ACP). See Built-in Agents for the full list with required environment variables.
Custom Agents
You can test any agent by creating a custom agent module. Use the createAgentAdapter()
factory for agents that produce NDJSON or plain text streams, or
createAcpBasedAdapter() for agents that speak the Agent Client Protocol (any CLI
with an --acp mode). Register the module in your config.
// adapters/my-agent.ts
import { createAgentAdapter } from "@netlify/axis";
export default createAgentAdapter<{ stdout: string }>({
name: "my-agent",
resolveCommand: () => ({ command: "my-cli", prefixArgs: [] }),
buildArgs: (input) => [input.prompt],
initialState: () => ({ stdout: "" }),
streamConfig: {
mode: "aggregate",
onChunk: (chunk, ctx) => {
ctx.state.stdout += chunk;
},
},
getResult: (ctx) => ({
result: ctx.state.stdout.trim() || null,
}),
});
The example above uses the default promptVia: "argv" behavior: buildArgs
places input.prompt on the command line. Set promptVia: "stdin" and AXIS
writes the prompt to the process's stdin and closes it instead -in that mode
buildArgs must omit the prompt. The built-in claude-code and
codex adapters use "stdin": argv rejects null bytes and caps argument
length, and agent transcripts can contain either.
promptVia: "stdin",
buildArgs: () => [],
Register it in axis.config.json:
{
"adapters": {
"my-agent": "./adapters/my-agent.ts"
},
"agents": ["my-agent"]
} Stream Modes
Custom agents support two modes for processing output:
- Lines mode -For agents that emit NDJSON (one JSON object per line). AXIS
parses each line and passes the parsed object to your
onLinehandler. The nativeclaude-codeandcodexadapters use this mode. ACP-based adapters bypassstreamConfigentirely; the ACP SDK handles framing. - Aggregate mode -For agents that emit plain text or non-JSON output. Raw
chunks are passed to
onChunkand accumulated in state. Use this for agents with custom output formats or simple stdout capture.
The module must export an AgentAdapter as the default export or as a named
adapter export.
Workspace Isolation
Each agent run gets a fresh temporary directory as its workspace. AXIS isolates the following to prevent configuration leakage and cross-run interference:
- HOME directory: Set to a per-job sibling directory of the workspace (not the workspace itself), so each agent's own config files (
.claude/,.codex/, etc.) live alongside the workspace and never appear in the directory the agent is scanning. - Agent-specific dirs:
CLAUDE_CONFIG_DIR,CODEX_HOME,GEMINI_CLI_HOMEare all set to isolated paths inside the per-job HOME. - Environment variables: Only explicitly listed vars and system essentials (
PATH,USER,SHELL,LANG,TERM,TMPDIR) are passed through. AXIS_CONFIG_DIR: Set to the absolute path of the directory containingaxis.config.{json,ts,js,mjs}. Lifecycle scripts and the agent process can use it to reference versioned fixtures or helper scripts. See Lifecycle environment variables.
For NDJSON-style agents, MCP server configuration files are written into each agent's isolated
config directory (under the per-job HOME) before spawn, in the format native to each CLI. Claude
Code additionally runs with --strict-mcp-config, so it loads only the servers a
scenario declares and ignores any MCP config discovered on the host (e.g. a local
~/.claude.json). ACP-based adapters pass MCP servers through the ACP
session/new call instead. See
MCP Servers in the configuration reference.
Multiple Runs
Agents are stochastic. Run the same scenario twice and the score moves, even when nothing
about your system changed. Set runs above 1 to sample a scenario/agent pair
several times, and AXIS reports a single representative result plus the spread around it.
The repeat count resolves from the most specific source that sets it:
--runs on the CLI, then a scenario's own runs, then
settings.runs, then 1. It must be odd; see
why below.
{
"settings": { "runs": 3 },
"agents": ["claude-code"]
} # or per invocation
axis run --runs 3
Repeats are off by default because they are not free. Each extra run is a full agent
execution plus its four judge calls, so runs: 3 costs roughly three times the
tokens and three times the wall clock of a single pass. Turn it on where score stability
matters (baselines and CI gates) rather than everywhere.
Run counts must be odd
AXIS rejects an even runs value. The reason is that both statistics a repeated
pair reports stop behaving without a middle run.
The median of the composites falls between two runs, so the number you would report
belongs to no run at all. That is precisely what representative selection exists to prevent.
And at runs: 2 the selection degenerates completely: with two samples the
per-dimension median is their midpoint, both runs sit at identical distance from it, every
comparison ties, and the tie-break hands back run 1 no matter which run scored better. Swap
the order of the two runs and the "representative" follows position rather than merit.
So runs: 2 as a cheap smoke check is not the bargain it looks like: it buys you
spread and reliability, but nothing at all for the central estimate, while presenting a
positional pick as a considered one. Use 3.
Even values are rejected outright rather than rounded. Turning a 4 into a 3 or a 5 silently would change both the sample size and the bill, so the error names the nearest legal counts and lets you choose.
Picking the representative run
AXIS does not average the composites. An average is a number no run actually earned, with no transcript behind it to explain it. Instead it selects a representative run: the real run whose composite sits nearest the median of the pair's composites.
Because run counts are odd, the median is itself one of the composites, so the representative's score is the median rather than merely being close to it. Two properties follow, and both are the point:
-
The headline equals the median of the runs listed directly beneath it in the report. A
reader who eyeballs
#1 81, #2 92, #3 87and expects 87 gets 87. - The live CLI can compute the same pick from the composite alone, so the number in your terminal and the number in the report are the same run's score by construction, not two statistics that happen to agree.
Ties go to the lowest run index, so selection is deterministic and re-aggregating the same data never moves the headline. Only successful runs are candidates, so a crash can neither become the headline nor drag the median.
Medians rather than means throughout, which matches the log-normal calibration AXIS already uses for its category scores: these distributions are skewed, so one outlier run should not drag the number everyone reads.
Everything in the report that describes a single run (the transcript, the interaction audits, the waterfall, the sparse index) belongs to the representative. The other runs are recorded alongside it with their own composite, their own four dimension scores, and their own spend. Keeping dimensions per run rather than summarising them is deliberate: a pair whose composite barely moves can still hide a goal-achievement score swinging 30 points against an agent score that compensates, and only the per-run columns show that.
Failures and reliability
A failed run is excluded from the score aggregate and counted against the pair's
reliability instead, reported as a
succeeded / total fraction. Folding a crash into the score as a zero would make
flakiness look like low quality, and the two call for different fixes.
A run whose score was withheld is treated differently again. If a judge dies or
returns something unparseable, that is a measurement failure, not an agent failure, so the run
leaves the reliability denominator entirely rather than counting against the agent. A pair with
three runs where one judge call failed reports 2/2 scored (1 withheld).
Pair-level status follows from this: a pair is completed when at least one of
its runs produced a score, and failed only when none did. So a single flaky run
no longer fails your suite on its own, while a pair that never once succeeded still does. The
run-level counts appear separately in the summary as runsTotal and
runsFailed.
Scheduling and interference
Repeats are queued run-major: every pair's run 1 is scheduled before any pair's run 2. Left in discovery order, a pair's three runs would sit adjacent in the queue and start together, which maximizes contention on the same provider rate limit and turns correlated 429s into what looks like variance in the agent.
Each run still gets its own isolated workspace and HOME, so repeats cannot interfere through the
filesystem. They can interfere through shared external state: a scenario whose setup
provisions a real site or service will have several runs collide on the same resource. Setup and
teardown scripts receive AXIS_RUN_INDEX and AXIS_RUN_COUNT for exactly
this reason, so you can namespace per run:
{
"setup": [
{ "action": "run_script", "command": "netlify sites:create --name axis-$AXIS_SCENARIO-$AXIS_RUN_INDEX" }
]
}
Both variables are always set, reading 1 and 1 for an unrepeated pair,
so a script never has to branch on whether repeats are configured.
One thing to adjust when turning repeats on:
settings.limits.run is a budget for the whole
run, not per pair, so runs: 3 exhausts it about three times sooner. AXIS prints a
note at startup when a run-level limit is set alongside repeats. Per-scenario limits are per
run and need no change.
Reading a repeated run
The live CLI keeps one row per pair, with a chip per run and a running median, so a suite does not triple in height when repeats are on:
cms/create-post
◐ claude-code|opus 2/3 · med 86 [87✓ 84✓ ●] 1:42 ~12.4k
In the report, a repeated pair is still a single row, tagged with a
3 runs · σ 5.1 pill and expanding to a per-run table that marks the
representative. The terminal equivalent appears under
axis reports <id>:
cms/create-post claude-code|opus 87 88 92 84 80
↳ 3 runs · representative #2 · median 87 · range 81-92 · σ 5.5
runs: #1 81, #2 87*, #3 92 (* representative) See Reports for where each run's files land on disk.
Baselines and regression detection
This is the main payoff. A baseline built from repeats records the standard deviation it
measured, and --compare-baseline then sizes its noise tolerance from that data
instead of a fixed constant: a drop counts as a regression only when it exceeds two sigma of
the baseline's own observed spread, with a floor of 1 point. A single-run baseline has no
measured sigma and falls back to the flat floor, which is why a flaky scenario reports a
regression every other run.
Reliability is compared too. A pair that scores as well as ever on the runs it completes, but now completes fewer of them, is reported as a regression on its own.
Multi-variant Scenarios
A single scenario file can produce multiple jobs by defining variants. Each variant runs as an independent job with its own key, inheriting the base scenario's fields and applying any overrides. This is useful for testing the same task under different tool configurations, prompts, or agent restrictions without duplicating scenario files.
For example, a scenario with two variants and two agents produces four jobs (2 variants × 2
agents). Each variant appears as a separate row in the CLI output and a separate entry in
reports, identified by its @-suffixed key (e.g., create-post@with-mcp).
Report manifest entries include a failed boolean computed from the full run output before
transcripts are stripped from report.json. This preserves the correct status for agents that
return a final result even when their process exits non-zero during cleanup.
See Writing Scenarios → Variants for the full field reference and examples.