Which AI delivers more, for less?
Here we hand the same job, at the same moment, to several AI models — and measure live what each one delivers: speed, cost and accuracy on the very same task. Pick a case to start.
News
The same headlines from an RSS feed go out to several AI models at once. Each label shows up the moment it arrives: whoever crosses the line first shows up first. We compare time per request, cost and how much the models agree.
Churn
Customers from a real telecom dataset (the classic Telco one, from IBM/Kaggle) are drawn at random and sent to the models — without the Churn column. Each estimates the chance of cancelling, and the score is simple: right or wrong, against the real answer. A logistic regression trained on the same data joins as the classic-method reference.
Hangman: uncovering words, letter by letter
Here the task changes nature: instead of classifying or estimating, the model has to decide step by step. Each one gets the same hidden word and keeps guessing, letter by letter, until it finds it. We compare how much time, how many rounds and how much money each solution took.
The running leaderboard
Every race run here adds to this tally. It's the average of everything that has gone through the lab — how long each model takes and how much it costs per item, and, for churn, how often it gets it right.
Rows marked as reference don't compete for the podium: logistic regression is the classic statistical method, and the frontier models — Claude Opus 5, Kimi K3 and GPT-5.6 Sol — don't race live (it would be slow and pricey). We measured each frontier model on 50 requests per case, with the same prompts, on September 20, 2026. They're a yardstick.
Adding up the history…
So, what is JEV?
Most AI models were built to chat: you ask, they write. JEV, from TypeSafe AI, was born for something else — deciding. Instead of handing back text for us to interpret, it returns the decision ready to use, with a confidence level: "this headline is Economy, 92% sure". TypeSafe calls this a System One model: fast, direct, no small talk.
And that's exactly what this lab — jev.welljose.net — puts to the test. Is a model built to decide really faster, more accurate and cheaper than chat models adapted to the same task? We don't assume here: we measure.
Every model runs through OpenRouter, with the same yardstick and the same material — and JEV joins as typesafe/jev-1.13. Want the official version, straight from the source? TypeSafe's announcement is at typesafe.ai/blog.