LiveBench-reported values for ox-alpha-max.
LiveBench publicly lists the label ox-alpha-max in its LiveBench-2026-06-25 snapshot. The values below are transcribed from that listing as a third-party observation; they are not independently reproduced measurements by Ox Alpha.
| Metric | LiveBench-reported value |
|---|---|
| Overall | 69.2 |
| Reasoning | 76.6 |
| Coding | 75.8 |
| Agentic coding | 52.6 |
| Mathematics | 77.5 |
| Data analysis | 75.8 |
| Language | 66.1 |
| Instruction following | 60.3 |
| Cost per successful task | $0.000 |
Source and scope: LiveBench public leaderboard, snapshot “LiveBench-2026-06-25”, model label “ox-alpha-max”; accessed 23 August 2026. LiveBench defines Overall as the mean of category averages. Its displayed $0.000 is the leaderboard’s reported cost per successful task for this snapshot and configuration; it is not a claim that the model is free or that Tokenra’s current route has zero cost.
What the listing displays
The label, snapshot name, Overall value, seven category values, and LiveBench-reported cost field shown above.
Route attribution
It does not establish that ox-alpha-max maps to Tokenra’s stealth/ox-alpha route.
Incomplete evidence
The public listing does not supply a model snapshot, full request setup, provider route, or raw execution artifacts.
No route-verified, reproducible Ox Alpha scores are published here.
This page publishes one clearly labeled third-party observation, not a verified result for a Tokenra route. A trend spike, a community post, a model description, or an isolated demo does not establish how a route performs on a named evaluation. Until the supporting record is complete, this site does not rank Ox Alpha against other models or make claims such as “outperforms.”
Exact model record
Model ID, provider route, date, and any available provider-side snapshot or revision identifier.
Runnable setup
Dataset version and split, prompts, generation settings, tools, retries, and scoring code.
Inspectable evidence
Raw outputs or permitted artifacts, aggregate calculations, and source links for readers to audit.
What an Ox Alpha benchmark result must disclose.
Each published row will identify the evaluation rather than relying on a model name alone. A clear record lets a reader distinguish a difference in model behavior from a difference in route, prompt, tool access, or scoring procedure.
| Disclosure item | Why it matters |
|---|---|
| Model and route | State stealth/ox-alpha, the Tokenra route, collection date, and any available model/version identifier. |
| Benchmark definition | Name the task, version, split, scoring implementation, license, and whether prompts or fixtures were changed. |
| Generation configuration | Record temperature, token limits, reasoning or tool settings, system prompt, number of samples, retries, and timeout policy. |
| Execution environment | Describe sandboxing, network and tool access, dependency versions, hardware where relevant, and test-time restrictions. |
| Raw evidence | Link to permitted raw outputs, aggregate calculations, failures, exclusions, and the code used to derive a published metric. |
| Comparison parity | For comparison claims, disclose the same conditions, dates, prompts, budgets, and scoring protocol for every model. |
Measure a task, then report the boundaries.
Evaluation begins by selecting a task that represents a decision a reader actually needs to make. Coding, long-context retrieval, tool use, and general reasoning are distinct workloads; a result from one cannot stand in for all others. The benchmark page will separate task categories rather than presenting a blended score as a universal capability claim.
For each run, the record should preserve a fixed dataset split and prompt template, then document any model-specific adapter required by the provider API. If a route exposes optional settings, their enabled or disabled state belongs in the record. If a test relies on tools, the available tools, permissions, result format, and failure behavior matter as much as the model output.
When variance is possible, results should report the sampling policy and repeat count. A single success can be illustrative, but it is not enough to imply a stable rate. Exclusions must be listed rather than silently removed from a total.
What a benchmark cannot tell you.
Benchmarks simplify production work. They may miss proprietary context, live tools, latency requirements, output-format constraints, safety controls, prompt changes, and the cost of human review. A score can help narrow an experiment; it cannot validate an integration by itself.
For a production decision, run a small, representative evaluation through the same Tokenra route, application controls, and user inputs you plan to deploy. Review error responses and failure modes alongside successful outputs.
How a route-verified result becomes publishable.
- Provide the model ID, Tokenra route, date, and applicable model snapshot or provider listing evidence.
- Provide the benchmark repository or specification, dataset version and split, license context, and unmodified scoring procedure.
- Provide prompts, generation settings, tool policy, retry behavior, resource budgets, and environment details.
- Provide raw or appropriately redacted run artifacts plus the calculation that produces each published metric.
- For a comparison, run every listed model under documented parity conditions or state the material differences directly.
The LiveBench record above does not currently meet these route-verification and reproducibility requirements. Once the missing mapping and run evidence are available, this page can add separately labeled route-verified results. Until then, use the guide to reading LLM benchmarks to evaluate third-party claims, or review the Ox Alpha API reference before testing the Tokenra route yourself.