OpenRouter · test bench · with answers

Who's going to cancel?

We draw customers from a real telecom dataset. Each model gets the profile — without knowing what happened — and estimates the chance of cancelling. Then we compare with reality. The race is between AI models; a logistic regression trained on this data joins as the classic-method reference.

Heads-up: the tasks themselves run in Portuguese (Brazilian news feed, Portuguese words, same prompts for every model). That keeps every run comparable — and the leaderboard in one piece.

1

Start here: let's draw the customers

This test runs on a real telecom dataset. We draw customers at random — you get to see who's in the sample before any model weighs in.

customers picked at random

Each customer becomes a separate request per model, run in sequence — the fairest way to measure. Models see the profile (tenure, contract, services, charges), never the ID or the Churn column. If a model takes longer than 20 seconds, we pull it from the race and the report moves on with whatever it already delivered.

2

Now just kick off the race

Draw a sample of customers to get started.
About the dataset

Telco Customer Churn, from IBM — one of the classic machine learning datasets, with 7,043 customers from a fictional California telecom company. Publicly available on Kaggle. Here we only use the Churn column as the answer key: it's never sent to the models — click any model's logo to see the exact request. The reference logistic regression is retrained on this same dataset every session; the LLMs never see any data from the population. Click the Σ to see what it learned.