JEV Lab · test bench

Which AI delivers more, for less?

Here we hand the same job, at the same moment, to several AI models — and measure live what each one delivers: speed, cost and accuracy on the very same task. Pick a case to start.

Available

News

The same headlines from an RSS feed go out to several AI models at once. Each label shows up the moment it arrives: whoever crosses the line first shows up first. We compare time per request, cost and how much the models agree.

Live RSS feed · paste any feed you likeOpen case
Available

Churn

Customers from a real telecom dataset (the classic Telco one, from IBM/Kaggle) are drawn at random and sent to the models — without the Churn column. Each estimates the chance of cancelling, and the score is simple: right or wrong, against the real answer. A logistic regression trained on the same data joins as the classic-method reference.

Real data with answers · LLMs vs. the classic methodOpen case
New

Hangman: uncovering words, letter by letter

Here the task changes nature: instead of classifying or estimating, the model has to decide step by step. Each one gets the same hidden word and keeps guessing, letter by letter, until it finds it. We compare how much time, how many rounds and how much money each solution took.

Sequential decisions · cost and time to solveOpen case

The running leaderboard

Every race run here adds to this tally. It's the average of everything that has gone through the lab — how long each model takes and how much it costs per item, and, for churn, how often it gets it right.

Rows marked as reference don't compete for the podium: logistic regression is the classic statistical method, and the frontier models — Claude Opus 5, Kimi K3 and GPT-5.6 Sol — don't race live (it would be slow and pricey). We measured each frontier model on 50 requests per case, with the same prompts, on September 20, 2026. They're a yardstick.

Adding up the history…

So, what is JEV?

Most AI models were built to chat: you ask, they write. JEV, from TypeSafe AI, was born for something else — deciding. Instead of handing back text for us to interpret, it returns the decision ready to use, with a confidence level: "this headline is Economy, 92% sure". TypeSafe calls this a System One model: fast, direct, no small talk.

And that's exactly what this lab — jev.welljose.net — puts to the test. Is a model built to decide really faster, more accurate and cheaper than chat models adapted to the same task? We don't assume here: we measure.

Every model runs through OpenRouter, with the same yardstick and the same material — and JEV joins as typesafe/jev-1.13. Want the official version, straight from the source? TypeSafe's announcement is at typesafe.ai/blog.

Every case uses the same yardstick: the models get exactly the same material at the same moment, and results show up in order of arrival — no editing, no imposed answers. When models disagree, the disagreement is highlighted for you to judge.