Third-party record & evaluation disclosure

Ox Alpha benchmarks: what the LiveBench listing shows — and its limits.

This page records one public LiveBench listing for the model label ox-alpha-max, then explains what that listing does and does not establish. It is not presented as a verified result for Tokenra’s stealth/ox-alpha route.

Record: third-party public listingIdentity: label-level onlyPage reviewed: 2026-08-23
01 / Third-party record

LiveBench-reported values for ox-alpha-max.

LiveBench publicly lists the label ox-alpha-max in its LiveBench-2026-06-25 snapshot. The values below are transcribed from that listing as a third-party observation; they are not independently reproduced measurements by Ox Alpha.

LiveBench-reported values for the label “ox-alpha-max”
MetricLiveBench-reported value
Overall69.2
Reasoning76.6
Coding75.8
Agentic coding52.6
Mathematics77.5
Data analysis75.8
Language66.1
Instruction following60.3
Cost per successful task$0.000

Source and scope: LiveBench public leaderboard, snapshot “LiveBench-2026-06-25”, model label “ox-alpha-max”; accessed 23 August 2026. LiveBench defines Overall as the mean of category averages. Its displayed $0.000 is the leaderboard’s reported cost per successful task for this snapshot and configuration; it is not a claim that the model is free or that Tokenra’s current route has zero cost.

This record supports

What the listing displays

The label, snapshot name, Overall value, seven category values, and LiveBench-reported cost field shown above.

This record does not support

Route attribution

It does not establish that ox-alpha-max maps to Tokenra’s stealth/ox-alpha route.

Reproducibility status

Incomplete evidence

The public listing does not supply a model snapshot, full request setup, provider route, or raw execution artifacts.

Do not turn this listing into a route claim. The available LiveBench record cannot verify a provider, model snapshot, deployment configuration, or direct comparison for Tokenra’s documented route.
02 / Publication status

No route-verified, reproducible Ox Alpha scores are published here.

This page publishes one clearly labeled third-party observation, not a verified result for a Tokenra route. A trend spike, a community post, a model description, or an isolated demo does not establish how a route performs on a named evaluation. Until the supporting record is complete, this site does not rank Ox Alpha against other models or make claims such as “outperforms.”

Required

Exact model record

Model ID, provider route, date, and any available provider-side snapshot or revision identifier.

Required

Runnable setup

Dataset version and split, prompts, generation settings, tools, retries, and scoring code.

Required

Inspectable evidence

Raw outputs or permitted artifacts, aggregate calculations, and source links for readers to audit.

03 / Disclosure standard

What an Ox Alpha benchmark result must disclose.

Each published row will identify the evaluation rather than relying on a model name alone. A clear record lets a reader distinguish a difference in model behavior from a difference in route, prompt, tool access, or scoring procedure.

Disclosure itemWhy it matters
Model and routeState stealth/ox-alpha, the Tokenra route, collection date, and any available model/version identifier.
Benchmark definitionName the task, version, split, scoring implementation, license, and whether prompts or fixtures were changed.
Generation configurationRecord temperature, token limits, reasoning or tool settings, system prompt, number of samples, retries, and timeout policy.
Execution environmentDescribe sandboxing, network and tool access, dependency versions, hardware where relevant, and test-time restrictions.
Raw evidenceLink to permitted raw outputs, aggregate calculations, failures, exclusions, and the code used to derive a published metric.
Comparison parityFor comparison claims, disclose the same conditions, dates, prompts, budgets, and scoring protocol for every model.
04 / Evaluation method

Measure a task, then report the boundaries.

Evaluation begins by selecting a task that represents a decision a reader actually needs to make. Coding, long-context retrieval, tool use, and general reasoning are distinct workloads; a result from one cannot stand in for all others. The benchmark page will separate task categories rather than presenting a blended score as a universal capability claim.

For each run, the record should preserve a fixed dataset split and prompt template, then document any model-specific adapter required by the provider API. If a route exposes optional settings, their enabled or disabled state belongs in the record. If a test relies on tools, the available tools, permissions, result format, and failure behavior matter as much as the model output.

When variance is possible, results should report the sampling policy and repeat count. A single success can be illustrative, but it is not enough to imply a stable rate. Exclusions must be listed rather than silently removed from a total.

05 / Interpretation limits

What a benchmark cannot tell you.

Benchmarks simplify production work. They may miss proprietary context, live tools, latency requirements, output-format constraints, safety controls, prompt changes, and the cost of human review. A score can help narrow an experiment; it cannot validate an integration by itself.

Do not compare results across unmatched setups. A different provider route, date, prompt, token budget, tool policy, or benchmark version can change the result. Comparable labels require comparable conditions.

For a production decision, run a small, representative evaluation through the same Tokenra route, application controls, and user inputs you plan to deploy. Review error responses and failure modes alongside successful outputs.

06 / Reproduction

How a route-verified result becomes publishable.

  1. Provide the model ID, Tokenra route, date, and applicable model snapshot or provider listing evidence.
  2. Provide the benchmark repository or specification, dataset version and split, license context, and unmodified scoring procedure.
  3. Provide prompts, generation settings, tool policy, retry behavior, resource budgets, and environment details.
  4. Provide raw or appropriately redacted run artifacts plus the calculation that produces each published metric.
  5. For a comparison, run every listed model under documented parity conditions or state the material differences directly.

The LiveBench record above does not currently meet these route-verification and reproducibility requirements. Once the missing mapping and run evidence are available, this page can add separately labeled route-verified results. Until then, use the guide to reading LLM benchmarks to evaluate third-party claims, or review the Ox Alpha API reference before testing the Tokenra route yourself.

Validate the route you will actually use.

Start with the API reference, then build an evaluation that matches your own application constraints.

Read API docs