Reports & Baselines
Every AXIS run produces a report with full scoring breakdowns and interaction transcripts. Baselines let you snapshot scores and detect regressions over time.
Understanding Reports
Every run automatically saves a report to .axis/reports/. Each report is a
directory containing a manifest and per-scenario result files.
.axis/reports/{reportId}/
report.json # Manifest with summary + metadata
report.html # Visual report (after scoring)
scenarios/{key}/{agent}.json # Full result with transcript + scores
scenarios/{key}/{agent}.raw.ndjson # Raw agent stdout
scenarios/{key}/{agent}.sparse-index.txt # Compressed transcript for scoring
scenarios/{key}/{agent}/artifacts/ # Captured files when artifacts is configured The manifest
report.json contains the run metadata and a summary of every scenario/agent result:
the composite AXIS Result, per-dimension scores, token usage, duration, and any error messages.
This is the file you read when scripting against AXIS output. A run started with
--profile also records the profile name in a
top-level profile field, so a report is attributable to the suite that produced it;
the field is absent when no profile was selected.
Scenario files
Each {agent}.json file under scenarios/ contains the full result
for one scenario/agent combination: the complete interaction transcript, judge evaluations,
per-interaction signal scores, and the judge assessment.
Repeated pairs
When a pair runs more than once (see Multiple Runs), each run gets its own directory instead of sharing one file, so nothing is overwritten and every run keeps its own transcript, raw stdout, and artifacts:
.axis/reports/{reportId}/
report.json
scenarios/{key}/{agent}/run-1/result.json
scenarios/{key}/{agent}/run-1/raw.ndjson
scenarios/{key}/{agent}/run-1/sparse-index.txt
scenarios/{key}/{agent}/run-1/artifacts/
scenarios/{key}/{agent}/run-2/result.json
... A pair that runs once keeps the flat layout above, unchanged, so reports written before repeats existed stay readable and nothing needs migrating.
The manifest still carries one entry per scenario/agent pair regardless of the
repeat count, so anything scripting against report.json keeps working. For a
repeated pair, the entry's score, durationMs, tokenUsage,
and file all describe the
representative run, which keeps them mutually
consistent and means file always explains the headline score. Three fields are
added:
| Field | Description |
|---|---|
runCount | Total runs configured for the pair. Absent when it ran once. |
runs |
One entry per run, ordered by runIndex: its composite, its four
dimensionScores, its complete score, duration, tokens, cost,
result file, and whether it failed or was withheld. Exactly
one is flagged representative, and its score is the same
object as the pair's top-level one. A withheld run publishes no score at all, since its
zeros stand in for "unknown" rather than a grade; a failed run publishes its zeros,
which it earned.
|
spread |
Statistics across the pair's successful runs: axisScore median, min, max,
mean, and sample standard deviation, plus representativeRunIndex. Failed
and withheld runs are excluded, so the band describes the agent's variance when it
works rather than mixing in zeros. Per-dimension variance is not summarised here; each
entry in runs carries its own dimensionScores, which shows
the whole distribution instead of four medians that hide it.
|
reliability | { succeeded, total, withheld }. total counts only the runs
that could be measured, so a judge outage shrinks the denominator instead of looking like
an agent crash.
|
The run-level counts appear in summary as runsTotal and
runsFailed. summary.total, completed, and
failed always count pairs, so the headline numbers keep their meaning when
runs changes.
Switching between runs
In the HTML report, a repeated pair's detail panel opens on the
representative run and lists every run in a table
above the breakdown. Selecting a row switches the whole breakdown (the waterfall, the
per-dimension cards, the interaction audits, the transcript) to that run. A
Representative column marks which run the pair reports, and its header
carries a tooltip explaining what that means.
Each run's full score is embedded in the page, so switching fetches nothing and works with
the report opened straight off disk over file://, where a browser would refuse
to load a sibling JSON file. The cost is report size: every run carries its own score and
sparse index, so a 20-scenario suite at runs: 3 produces roughly a 3 MB
report.html against 0.9 MB at runs: 1. The per-run
result.json files are still on disk if you want them, at the paths shown above.
Viewing Reports
# List all reports
npx @netlify/axis reports
# View the latest report summary
npx @netlify/axis reports latest
# View a specific scenario detail
npx @netlify/axis reports latest hello-world
# Filter by agent
npx @netlify/axis reports latest --agent claude-code
HTML reports
Open the visual report in your browser for the richest view:
npx @netlify/axis reports latest --html
The HTML report includes:
- Composite and per-dimension score breakdowns with visual indicators.
- The full interaction transcript with tool calls and results.
- Judge evaluations for each judge check and interaction signal.
- Score insights identifying the weakest signals for low-scoring dimensions.
-
A captured-file tree per run when scenarios configure
artifacts — preview text/images in a
modal, download individual files, or grab everything as a .zip.
-
Markdown "Setup notes" and "Teardown notes" panels per run when lifecycle scripts write to
$AXIS_OUTPUT — useful for
capturing workspace state, external probes, or diagnostic context alongside each scored run.
JSON output
For scripting and CI integration, use --json to get machine-readable output:
npx @netlify/axis reports latest --json
Baselines
Baselines snapshot your scores at a point in time. You compare future runs against a baseline
to detect regressions -scores that dropped by more than the noise tolerance (1 point).
Setting a baseline
# Save from the latest report
npx @netlify/axis baseline set
# Save with a name (for multiple baselines)
npx @netlify/axis baseline set v1.0
# Save from a specific report (report IDs use YYYY-MM-DD-HHMMSS format)
npx @netlify/axis baseline set --from 2026-04-15-143022
Comparing against a baseline
# Compare during a run (automatic)
npx @netlify/axis run --compare-baseline
# Compare explicitly after a run
npx @netlify/axis baseline compare
# Compare against a named baseline
npx @netlify/axis baseline compare v1.0
The comparison shows deltas for each score. Changes within the noise tolerance are reported as
unchanged. Regressions are highlighted and the command exits with code 1 if any are detected.
The tolerance is 1 point by default. A baseline captured from a pair that ran
several times also records the standard deviation it measured, and the tolerance for that row
widens to two sigma of that spread (the 1-point floor still applies). A noisy scenario
therefore has to move further before it counts as a regression, which is the main reason to
build baselines from repeated runs: a single-run baseline
cannot know its own variance and will report ordinary noise as a regression.
Reliability is compared alongside the score. A pair that still scores as well on the runs it
completes, but now completes fewer of them, is reported as a regression on its own, so
increasing flakiness cannot pass as unchanged.
Statistical significance
When both the baseline and the current report ran a pair more than once, every metric also
gets a Welch's two-sample t-test over the two distributions, reported
alongside the tolerance band rather than replacing it. The two answer different questions:
the band asks whether the value moved further than this scenario usually wobbles, and the test
asks whether the difference could be chance. They can legitimately disagree, so both are
shown and you decide which to act on.
cms/create-post claude-code|opus 87 74 -13 ▼
↳ tolerance ±8.0 from the baseline's measured run spread
AXIS Result 87.0 → 74.0 d=-3.29 p=0.016 ● large regression
Goal Achievement 88.0 → 70.0 d=-6.00 p=0.002 ● large regression
Environment 92.0 → 91.0 d=-1.00 p=0.288 ○ not significant
Duration 62.0s → 91.0s d=+9.67 p<0.001 ● large regression
Tokens 150,000 → 152,000 d=+0.35 p=0.708 ○ not significant
d is Cohen's effect size and p is the two-sided p-value. A filled
marker means a significant move with a large effect, half-filled means significant but
smaller, and open means the test could not separate the two samples. The verdict is
direction-aware: Duration and Tokens are lower-is-better, so
a drop in either reads as an improvement while the same drop in a score dimension reads as a
regression.
Significance testing covers more than the band does. Alongside the four score dimensions and
the composite, duration and tokens are compared too, each
with its own measured spread.
Worth knowing before reading too much into it: at three runs a side, a t-test needs about
2.27 standard deviations of separation where the band needs 2.00, so at the usual
runs: 3 the test is marginally less sensitive and the two mostly agree.
The test earns its place at higher run counts, where it tightens and the fixed band does not:
nine runs a side detects a 1.00 sigma move, eighteen detects 0.68 sigma, while the band stays
at 2.00 sigma however much you sample.
The exit code still follows the tolerance band, so turning repeats on cannot
silently change which runs fail CI. The significance tally is reported separately in
summary.significant, and every metric's full test output is in
entries[].metrics[].significance for --json consumers.
When to set baselines
- After establishing a good score: Run your scenarios, review the results, and
if you are satisfied, save the baseline. This becomes your quality floor.
- After intentional changes: If you change your project structure, APIs, or
agent configuration and scores change as expected, update the baseline to reflect the new
normal.
- Named baselines for releases: Use named baselines (
baseline set v2.0)
to track scores across major versions.
Managing baselines
# List all baselines
npx @netlify/axis baseline list
# View baseline contents
npx @netlify/axis baseline show
# Delete a baseline
npx @netlify/axis baseline delete v1.0
CI Integration
AXIS is designed to run in CI environments. The key patterns:
-
--json -Machine-readable output to stdout. No live terminal
display, no color codes. Suitable for piping to other tools or saving as artifacts.
-
--compare-baseline -Exits with code 1 if regressions are
detected. Use this as a CI gate: the build fails if agent experience degrades.
-
--concurrency -Control resource usage in constrained CI
environments.
-
--profile -Pick which suite a job runs. A dispatch input maps
straight onto a profile name, so one workflow can offer several matrices without maintaining
a config file per mode. An unknown name fails the job instead of falling back to the default
suite. See Profiles.
- API keys via environment -Pass
ANTHROPIC_API_KEY,
CODEX_API_KEY, GEMINI_API_KEY, or META_API_KEY as CI
secrets. In CI you should always set explicit API keys: claude-code,
codex, and muse have a local-login fallback for laptop use, but that
path is unsuitable for CI because it bills against an individual subscription rather than a
service account.
GitHub Actions example
# GitHub Actions example
- name: Run AXIS tests
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
run: npx @netlify/axis run --json --compare-baseline
Report Storage
Baselines are stored in .axis/baselines/ and designed to be checked into version
control so your team shares the same regression thresholds.
Reports and cached clones should not be committed:
# .gitignore
.axis/reports/
.axis/remotes/
.axis/repos/