The live model exchange. Every model from every lab β graded hourly on tasks a program can verify, priced by the real cost of a done task, tracked the moment it launches. Free, for everyone.
π’ β measured liveπ β models trackedπ β launched this monthπ± priced by CPAT β cost per verified taskπ refreshed hourly
Real list prices, context and modalities for every model on the market (public feed, refreshed hourly), with our measured grade and CPAT wherever we have run the battery. Search, filter, sort, and tick up to three to compare side by side.
default: measured models by grade, then the newest launches Β· 50 at a time
Comparing:
How a grade is earned
We give every model the same real, checkable tasks and score only what actually works β not opinions.
β
π€
We send a real taskβ
π€
The model answersβ
π
We check it's correctβ
What Meerada sells (everything above stays free)
ποΈ LLManager β one app for every model, on your keys; import your Claude/ChatGPT history and switch models mid-conversation. Free tier Β· Pro $99/yr.
Grades from our own hourly runs on the verifiable battery (hidden unit tests, real SQL, exact answers, faithful summaries). Coverage grows with every free-tier key we add; paid frontier models are graded when a lab sponsors the run.
#
Model
Grade
Quality (95% CI)
n
CPAT
TTAT
trend
Status
updated
π Frontier Watch β how fast, how autonomous, how safe
No doom numbers β three things we can actually measure, every hour, on the same battery: capability (verified grade), autonomy (multi-step agentic plans executed correctly, in policy order) and safety behaviour (refusing clear harm, keeping secrets and scope under pressure). The velocity line is how fast the frontier's best measured grade is moving.
Frontier grade (best measured)
β
Velocity (last 10 cycles)
β
Autonomy (avg, measured)
β
Safety (avg, measured)
β
#
Model
Capability
Autonomy
Safety
Trend
Autonomy and safety fill in as the hourly grader re-runs each model on the extended battery (agentic + safety clusters added 9 Sep 2026).
ποΈ The labs β who makes the models, and what they're worth
The exchange's "issuers": every lab behind the models above, with its parent company or listing, last reported valuation (private) or ticker (public), how many models it fields, its best measured grade and its cheapest verified task. Valuations are last-reported figures with a date and a source β not live prices.
Grades attribute to the lab that made the model, wherever it is served β e.g. OpenAI's open-weights gpt-oss models measured on Groq's free tier count for OpenAI.
Lab
Parent / listing
Ticker
Valuation
Models
Best grade
Cheapest CPAT
Latest launch
π New on the market β found on the web, queued for grading
Every hour we scan the public model feeds for launches (stealth and alpha models included). Anything with a free tier is graded automatically on the next run; paid-only models are listed at their public price until a sponsor key covers them.
launched
Model
$ in / out per M
context
Bourse status
π± Outcome Exchange β buy a done task, not tokens
This is what makes the arena a real market. Post the job and the quality bar; the exchange routes to the cheapest model that clears it. The price is the live CPAT β cost per verified task. As grades move, the clearing price moves. prototype
π§ The Bourse in 30 seconds
1 Β· Every hour we run the same verifiable tasks (schema, regex, unit-checkable) on every model we can reach, and check the outputs by program β not by opinion. 2 Β· Each model gets a grade (verified success rate, with its confidence interval) and a price: what one verified task cost at the provider's public list price β CPAT β plus how long it took (TTAT). 3 Β· You post a job and a quality bar. The exchange clears at the cheapest model over the bar. That clearing price is the market price of a done task β and it moves as models change. Why only a middleman can run it: pricing a "done task" needs independent verification across every vendor. A vendor can't grade itself; a token reseller doesn't verify. The interface people work through (LLManager) is where that measurement comes from and where the routing gets applied.
π Order book β models that clear the bar, by price
#
Model
Grade (this task)
CPAT
Monthly at your volume
The clearing price is the lowest CPAT among models over your bar. Only we can run this market β because pricing a "done task" needs verified measurement, which is exactly the grade. A production venue locks that price for prepaid capacity (futures) and takes the routing spread as margin.
π§ͺ The test battery β what actually moves a grade
No opinions, no vibes. Every model runs the same programmatically-checkable tasks. Light tests run constantly so the grade stays fresh; heavy tests run less often but weigh more. Hundreds of these micro-checks roll up into the category scores, and the categories into the single grade β that's how small, checkable tasks extrapolate to a macro verdict you can trust.
β‘ LIGHT β fast, high-frequency, keep the grade current
π HEAVY β complex, weighted, the real separators
Live pass-rates update as tests run above. A test is scored only when its output is machine-verifiable β code that must pass hidden asserts, JSON whose totals must add up, a retrieval whose cited fact must be exact. That's what separates a Meerada grade from a leaderboard vote.
ποΈ Meet LLManager β it manages every model for you
Not a translator β a manager for the communication and tasks between you and your models. Tell it what you want in plain words; it shapes the lean instruction, sends it, checks the answer held up, and keeps the thread. Same result, far fewer tokens β on any model you use. See the full product β
STEP 1 Β· CONNECT
Your keys, your models
Add your provider keys once. LLManager runs locally and drives Claude, GPT, Gemini, DeepSeek on your own account β nothing routed through us.
STEP 2 Β· IT READS YOUR WORKLOAD
Finds where spend leaks
It sees how you actually call models and flags the expensive mistakes: re-sent context that should be cached, flagship models doing work a cheap one clears, runaway reasoning, no output ceiling.
STEP 3 Β· IT RUNS THEM RIGHT
Same result, fraction of the bill
Caches stable context, routes each task to the cheapest model that still passes your quality bar, caps reasoning and output β verified, not guessed.
Analyze your workload illustrative Β· LLManager measures your real numbers
ποΈ Session Console β one cockpit for every model
LLManager is the go-to launcher for any model. Say "open Claude and draft the release notes" β it opens Claude, shapes your ask into a lean instruction, and runs it. Fire several at once, across different models, and manage them all from one place instead of juggling tabs. Full product β
preview Β· SIMULATION
How it works: LLManager holds your keys locally and drives each provider's API (or app) for you. One plain request β the router picks the cheapest model that clears the quality bar, shapes your words into an efficient instruction, and streams the result back. Parallel sessions share one context so you don't re-paste β and one bill's worth of tokens does the work of many.
β¬οΈ Get LLManager β any device
DESKTOP Β· mac / win / linux
CLI + tray app
pip install meerada then meerada up. Runs locally, holds your keys, opens as a menu-bar cockpit. Your prompts never leave your machine.
PHONE Β· iOS / Android
Share-sheet + PWA
Install the web app, or hit Share β Meerada from any app. Speak or paste an intent; it routes to your model and returns the lean answer.
EVERYWHERE Β· bring-your-keys
Your models, your bill
Claude, GPT, Gemini, DeepSeek, local Ollama β add a key once. LLManager sits in front of all of them as one interface.
Pick the model you run today and a candidate. We estimate the savings and quality delta β and guarantee, via output equivalence, that you don't lose quality silently.
FROM (today)
β
TO (candidate)
Run the real Handshake βReplays your golden set on the candidate, returns a measured per-cluster gap report. Content never leaves your machine.
π Why our grade beats the leaderboards
measuresβ¦
LMArena
Artificial Analysis
$/token
Meerada
the real cost of a done task
β
β
β
β
counts retries & reasoning burn
β
β
β
β
verified success, not preference
β
~
β
β
updates within the hour
weeks
~
static
β
on YOUR workload
β
β
β
β
A model can win on $/token yet lose on cost per done task β cheap tokens don't help if it needs three tries. That gap is exactly what we measure.