Benchmark methodology
Every metric on a benchmark report page carries a source badge — Measured by VARTA, Vendor published, or Third-party benchmark — so a claim about latency, accuracy or cost is never presented without saying where the number came from.
First-party measurements
Numbers labelled “Measured by VARTA” were measured by VARTA directly against the provider's production API, using the same call harness and network path VARTA workflows run in production — not a synthetic lab benchmark. Each metric row on a report records the exact date it was tested (“testedOn”), since provider performance changes over time and a stale number is a wrong number.
Vendor-published and third-party figures
Where a first-party re-run isn't yet available for a given provider or metric, reports fall back to figures the vendor has published themselves (“Vendor published”) or to independent third-party benchmarks (“Third-party benchmark”). These are never silently blended with first-party numbers in the same row — the source badge makes the provenance of every single figure explicit.
Caveats and re-testing
Each report's own Caveats section documents anything specific to that comparison — test conditions, known limitations, or why a provider was excluded. Reports are re-run periodically as providers ship new models; the “Last updated” date at the top of each report reflects the most recent pass.
Raw data
Every report offers its underlying metric rows as a CSV download, so the numbers behind any chart or table can be independently verified rather than taken on faith.