CLI Reference
Complete reference for the axis command-line interface. All commands can be run
directly or via npx @netlify/axis.
axis init
axis.config.json and a sample scenario in ./scenarios. When run in
a TTY without flags, prompts interactively for the scenarios directory and agents.
-s, --scenarios <path> Path to scenarios directory (default: ./scenarios). -a, --agent <names>
Agent(s) to include — comma-separated (e.g. claude-code,codex,gemini). Default:
claude-code. --agents is accepted as an alias.
-f, --force Overwrite existing files. --format <format>
Config file format: json, js, or ts (default: json).
JS and TS scaffolds let you compose the config programmatically.
--no-skills
Skip the automatic install of AXIS skills. By default, init runs
npx skills add netlify/axis --all after writing the config so the
configure-axis skill is wired into every detected agent. Pass
--no-skills to opt out.
axis run
-c, --config <path> Config file (default: axis.config.json). -s, --scenario <keys>
Run specific scenarios. Accepts a comma-separated list and glob patterns, e.g.
cms/*, hello-world,cms/create-post@*.
--scenarios is accepted as an alias, since the option takes a list.
-a, --agent <names>
Run with specific agents. Accepts a comma-separated list and glob patterns, e.g.
claude-code|*, codex|gpt,claude-code|opus.
--agents is accepted as an alias.
-p, --profile <name>
Apply a named profile from the config's profiles map before the run. An
unknown name is an error rather than a fallback to the default suite. See
Profiles.
--json JSON output to stdout (no live terminal display). -v, --verbose Detailed per-step logging. -o, --output-dir <dir> Also write report files to this directory. --concurrency <n> Max parallel jobs (default: 15). --runs <n>
Run each scenario/agent pair n times and report the
representative run (default: 1, or settings.runs). Overrides
both settings.runs and any per-scenario runs. Must be
odd (1, 3, 5, …) and at most 19; an even value is an
error rather than being rounded. Each extra run costs a full agent execution plus its
judge calls, so the spend scales linearly. See
multiple runs.
--debug Capture raw agent stdout for debugging. --failed [reportId]
Re-run only the scenario/agent pairs marked failed in a previous report
(default: latest). Pairs no longer in the config are silently
dropped. Mutually exclusive with --scenario and
--agent, but combinable with --profile: retry the
report under the same profile that produced it, since pairs outside the
selected suite (or named by another suite's agent matrix) match nothing and
the run exits non-zero with no jobs discovered. A repeated pair is retried in
full rather than run by run: its score is a property of the whole sample, so
refilling only the failed slots would leave the retry report holding a partial
sample that no single run can represent.
--no-score Skip scoring (raw results only). --refresh-skills Force re-clone remote skills. --refresh-repos
Force re-clone of repositories cached under .axis/repos/ for
copy actions pointed at a git URL.
--compare-baseline [name]
Compare against a baseline after scoring (default name: default). Exits with
code 1 if any regressions are detected.
Because --scenario and --agent both take a list, their plural
spellings (--scenarios, --agents) are accepted as aliases wherever
they appear. The one exception is axis init --scenarios <path>, which is a
real and different option: it sets where the scenarios directory is written.
Exit codes
axis run exits non-zero in three cases. The last two are about coverage rather
than scores: a suite that quietly stops loading a scenario would otherwise keep reporting a
clean average over whatever is left.
| Condition | Exit |
|---|---|
| Every job ran and passed | 0 |
| One or more jobs failed | 1 |
A file in the scenarios tree failed to load. The files and reasons are printed, stored
in the report manifest under loadFailures, and shown in the HTML report.
| 1 |
No jobs were discovered at all, e.g. a --scenario or --agent
filter that matches nothing.
| 1 |
Scenarios opted out with skip: true do not affect the exit code. They are
counted separately from files that failed to load, in both the terminal summary and the
report.
axis reports
[reportId] Report ID or latest. Omit to list all reports. [scenarioKey] Drill into a specific scenario detail. -a, --agent <name...> Filter by agent(s), repeatable. --agents is accepted as an alias. --json Output as JSON. --html Open report in browser. -n, --limit <count> Max reports to list (default: 10). axis baseline
Every axis baseline subcommand also accepts -c, --config <path> to
point at a non-default config file.
--from <reportId> Use a specific report (default: latest). --json Output as JSON. --report <reportId> Specific report to compare (default: latest). --json Output as JSON.