Skip to content
Youness Aamiri
Go back

BenchmarkGate, Part 4: benchmark-gate compare, and Why Order Matters

benchmark-gate compare, and Why Order Matters

In part three, v0.3.0-alpha.1 added validate — catching a broken policy or baseline file before check ever saw it. That release didn’t touch what check actually does, though: match benchmarks, match metrics, calculate deltas, apply thresholds, decide pass or fail — all inside one evaluator, one pass, no way to get “what changed” without also asking “is that acceptable.”

v0.4.0-alpha.1 splits that in two. benchmark-gate compare answers the first question. check still answers the second — but it answers it by applying policy on top of the same comparison compare produces, instead of computing its own.

A real diff

Here’s compare against a real baseline from this project’s own benchmark suite:

benchmark-gate compare --baseline ..\baseline.json --results .\BenchmarkDotNet.Artifacts\results\CedarRecon.Tests.Performance.ClassifierPhaseBenchmark-report-full-compressed.json
Suite: CedarRecon  Comparable: 21  Added: 0  Removed: 6

Benchmark                                Metric               Reference    Candidate    Abs Delta      % Delta    Direction      Status
ClassifierPhaseBenchmark.BuildDictionar. allocatedBytesPerOp. 5.465 MB     5.465 MB     0 B            +0.00%     Unchanged      Comparable
                                         meanNanoseconds      4.998 ms     4.998 ms     0.000 ns       +0.00%     Unchanged      Comparable
[... 19 more Comparable benchmarks, all Unchanged ...]

Removed benchmarks:
  ExceptionClassifierBenchmark.DictionaryClassifier|job=Job-SNYTAA|N=10000
  ExceptionClassifierBenchmark.DictionaryClassifier|job=Job-SNYTAA|N=100000
  ExceptionClassifierBenchmark.DictionaryClassifier|job=Job-SNYTAA|N=1000000
  ExceptionClassifierBenchmark.IndexedClassifier|job=Job-SNYTAA|N=10000
  ExceptionClassifierBenchmark.IndexedClassifier|job=Job-SNYTAA|N=100000
  ExceptionClassifierBenchmark.IndexedClassifier|job=Job-SNYTAA|N=1000000

Nothing in the 21 Comparable benchmarks moved. But an entire benchmark class — ExceptionClassifierBenchmark, six parameterized entries — is gone from the current run. No policy was involved in producing that output. compare doesn’t know whether a benchmark disappearing is a deliberate rename, a benchmark someone deleted on purpose, or a sign the project no longer builds that code path. It just reports the fact: this existed in the baseline, it doesn’t exist now. Whether that’s fine is a judgment call for a human, or for check with a policy attached — not for compare.

The same fact, as data

--format json writes the identical comparison as a versioned, machine-readable document — same command, --format json --output compare.json instead of printing a table. This isn’t just another report format: it’s the same comparison facts, serialized as a versioned artifact that another tool — or a future BenchmarkGate command — can consume without rerunning the comparison. The JSON preserves the comparison engine’s canonical ordering exactly; reporters don’t re-sort the results.

{
  "schemaVersion": 1,
  "suite": "CedarRecon",
  "comparable": 21,
  "added": 0,
  "removed": 6,
  "benchmarks": [
    {
      "identity": "ExceptionClassifierBenchmark.DictionaryClassifier|job=Job-SNYTAA|N=10000",
      "status": "Removed",
      "metrics": [
        {
          "metricName": "allocatedBytesPerOperation",
          "status": "MissingCandidateMetric",
          "reference": { "value": 5087065, "unit": "bytes" }
        },
        {
          "metricName": "meanNanoseconds",
          "status": "MissingCandidateMetric",
          "reference": { "value": 4945967.65625, "unit": "ns" }
        }
      ]
    }
  ]
}

Removed means “present in the reference baseline but absent from the candidate run.” It does not imply an error.

Two more things worth noticing here. First, there’s no candidateStability field on this benchmark at all — not null, genuinely absent. A removed benchmark has no current-run observation, so there’s nothing to source stability facts from; the field is omitted rather than filled with a placeholder. Second, each metric’s reference value is still the exact number from the baseline, full precision, un-rounded — compare’s JSON output never rounds a value for display the way the console table does.

Why order matters

The title isn’t about the order benchmarks appear in the output. It’s about the order two things happen in.

Before this release, the pipeline was:

baseline + current + policy → evaluation

One evaluator performed benchmark matching, metric matching, delta calculation, and threshold evaluation in a single pass. There was no way to ask “what changed” without also asking “is that acceptable” — the two questions were the same function call.

Now it’s:

baseline + current → comparison → (+ policy) → evaluation

BenchmarkComparisonEngine is the only place benchmark matching, metric matching, and delta calculation happen. It doesn’t know what a policy is. compare calls it directly and reports the result. check calls the exact same engine, then hands its output to RegressionEvaluator, which applies policy on top — never recomputing a delta, never re-matching a benchmark, just interpreting facts that already exist. In other words: arithmetic belongs to comparison; judgment belongs to evaluation.

That ordering is the whole point of this release. history (v0.5.0) and anything after it can depend on “what changed” being a real, callable thing on its own — not something baked one level deep inside an evaluator that also happens to know about warning thresholds.

What v0.4.0 deliberately didn’t do

MetricComparisonStatus has a UnitMismatch value. compare never produces it.

That’s not a bug — it’s a gap in the data, not the logic. Neither a baseline entry nor a current observation carries a per-value unit today; a metric’s unit comes entirely from MetricCatalog, keyed by metric name. Both sides currently derive their unit from the same catalog entry, so a mismatch can never be observed. The status stays in the schema, documented as reserved, because it’s part of the intended shape; it just needs baseline and candidate values to carry real source-unit metadata before it can mean anything. That’s separate, future work, not something to fake with a check that can never fire.

The bigger picture

The important part isn’t the new command. It’s that “what changed” is now a first-class concept. Once comparison exists independently of policy, other features — history, trend analysis, dashboards, and explanation — can all build on the same artifact instead of reimplementing comparison logic.

What’s next

v0.5.0 — filesystem history. check --history <path> as an alternative to --baseline: a plain directory of snapshots instead of one file, aggregated (median of the last N, by default) into a baseline on the fly. That only becomes possible because comparison is now a standalone artifact instead of an internal evaluator step — history needs something stable to compare against, and now it has one. Full sequencing is in ROADMAP.md.

dotnet tool install --global Bijecta.BenchmarkGate.Tool --version 0.4.0-alpha.1

Repo, README, and the full roadmap: github.com/Bijecta/BenchmarkGate.


Share this post:

Next Post
BenchmarkGate, Part 3: benchmark-gate validate