Metrics that can be defined from observed CI evidence
A benchmark must start with a precise source contract. Candidate operational metrics include jobs per run, completed job duration, observed queue wait where timestamps support it, retry incidence, runner-label distribution, workflow-event distribution, job fan-out, history coverage, and evidence quality.
A metric name must not imply stronger semantics than its source fields support. Missing timestamps or bounded history remain null/unknown rather than being filled with synthetic values.
Why there are no customer-derived cohort pages yet
Training eligibility is not public publication authority. Before an aggregate benchmark can become indexable, PipelineSpend requires a separate publication policy, privacy filtering, deterministic cohort classification, minimum independent organization/repository counts, contribution-dominance limits, and enough materially distinct metrics to make the page useful.
If those thresholds are not met, the correct publication state is blocked—not a thin page populated with generated filler.
Freshness means material change
Recomputing the same metrics tomorrow should not make a page look newly updated. Future benchmark snapshots separate computation time from the last material content change, and sitemap freshness should change only when public data actually changes.