Evaluation · Model Selection · Retrieval · Benchmarks
We tested it against 10 other models on two jobs your product already has: routing support tickets and reranking search results. Jev was the fastest and cheapest model in both tests, by margins no other model in the table approaches. 7 of those 10 beat it on accuracy. It also has an input ceiling that existing write-ups leave out. The trade is below, with the code to check it.
What we found, in plain words
Jev takes a situation and a list of options, then picks one. It returns no prose. In this test, it was fast and cheap. 7 of the 10 other models we tested beat it on accuracy.
The short version
Use Jev when the answer is one of a short list you already wrote and you can absorb being wrong 8% of the time. In our ticket-routing test, 92% of tickets went to the right queue. The strongest model reached 97%. Jev answered 4 times faster and charged $0.018 per thousand tickets against $0.026.
Speed and price do not erase an accuracy gap. Your workload decides how much that gap matters. Route a thousand tickets a day and the gap between 92% and 97% drops 50 tickets a day into the wrong queue. An agent may spot each one in ten seconds. A ticket that sits in the wrong queue for two days can cost more than the model saved.
Fig. 01 A fit check for Jev, ordered so the first no rules it out. both benchmarks in this post
Three jobs where we would pick it
These jobs have the same shape. You already wrote the answer list, and speed or cost matters before the last points of accuracy do.
- Sorting things into buckets you defined. Put tickets in queues, email in folders, and expenses in categories. Jev was built for this, and this is where it scored best.
- Filtering a firehose. Give Jev ten thousand items when the expensive model can read only a hundred. Jev chooses the hundred. A miss costs one unnecessary read by the expensive model.
- Anything you ship where 385 milliseconds beats 1,732. Use it for a dropdown that reorders while someone types or a check that runs on every keystroke. The user feels the speed and will not see the accuracy gap.
Three jobs where teams use it and should not
The public index of Jev projects has several hundred entries. Most are sound. Three recurring categories would stop us from shipping it. Each fails for a different reason.
The confidence number does not mean what it looks like
Do not treat Jev's confidence as a probability. Trading signals and risk scores are the dangerous cases. One independent evaluation measured a confidence of 1.00 on 86% of Jev's answers. It found excellent calibration on support tickets and useless calibration on logic puzzles, with the model and settings held fixed. The number does not tell you which case you have.
The input does not fit
Do not hand Jev a long document and ask one question about all of it. We made this mistake. We asked it to pick the best 5 out of 100 passages in a single request. It came back on 265 of our 350 questions and returned nothing at all on the other 85, because the provider rejected requests that size. We then split the same work into batches of a few passages each, and it came back on all 350. Nothing here is about right or wrong answers. It is about whether an answer arrives.
Someone on the other side wants a particular answer
Keep Jev out of decisions where someone has an incentive to game the result. CV screening and content moderation both appear on the public list. In the same independent evaluation, one line claiming authority flipped Jev's answer on 147 of 200 support tickets. A sentence of text should not move a control you rely on.
One place it beat our expectations
Search reranking works well when you stop asking about everything at once. A handful of passages per request put Jev fifth of 7 on quality. It answered every question we gave it, cost 2 times less than the highest-scoring model, and came back 4 times sooner.
How we measured it
We ran two end-to-end benchmarks against live APIs. The repository contains the code and raw model output, so you can rebuild every number below. Only the paid runs need an API key.
Benchmark 1: ticket routing
2,400 calls, 20 arms. Every arm sees the same 100 tickets in the same order and chooses one of five queues. Jev gets one Choice call with the five queue descriptions as criteria. The chat arms get the same descriptions in a prompt and return a queue name. We negotiate reasoning effort before each run and record the provider's response. That keeps an ignored setting from creating a different arm in the table.
The diagram traces the path from the ticket file to the published table. It shows the two passes, throttle, validity gate, provider-cost reconciliation, and offline scoring from stored records.
Source: report/classify.architecture.json
Fig. 02 Ticket-routing accuracy against price. Jev is the cheapest point on the chart and ranks 12th of the 19 arms plotted on accuracy. results.json, 19 arms that answered, 100 tickets each
| Model | Answered | Correct | 95% interval | $ / 1,000 | p50 ms |
|---|---|---|---|---|---|
| DeepSeek v4.1 Flash · no reasoning | 100 | 97% | 92–99 | $0.026 | 1,732 |
| Gemini 3.8 Flash · no reasoning | 100 | 96% | 90–98 | $0.357 | 2,194 |
| Claude Sonnet 5 | 100 | 96% | 90–98 | $0.542 | 2,156 |
| Gemini 3.8 Flash | 100 | 95% | 89–98 | $0.691 | 2,660 |
| GLM 5.3 Flash | 98 | 95% | 89–98 | $0.045 | 1,691 |
| Kimi K3 · no reasoning | 100 | 95% | 89–98 | $0.443 | 901 |
| GLM 5.3 Flash · no reasoning | 100 | 94% | 88–97 | $0.021 | 922 |
| Claude Opus 5 | 100 | 94% | 88–97 | $2.441 | 3,253 |
| DeepSeek v4.1 Flash | 97 | 93% | 86–97 | $0.059 | 1,427 |
| Grok 4.6 | 100 | 93% | 86–97 | $3.162 | 6,336 |
| Grok 4.6 · no reasoning | 100 | 93% | 86–97 | $2.391 | 4,295 |
| Jev · Choice | 100 | 92% | 85–96 | $0.018 | 385 |
| Kimi K3 | 98 | 92% | 85–96 | $1.771 | 1,944 |
| GPT-5.3 Codex · no reasoning | 100 | 91% | 84–95 | $0.408 | 1,256 |
| GPT-5.3 Codex | 100 | 90% | 83–94 | $1.041 | 1,544 |
| Qwen 3.8 Flash | 98 | 90% | 83–94 | $0.084 | 1,989 |
| Qwen 3.8 Flash · no reasoning | 100 | 90% | 83–94 | $0.026 | 671 |
| Mistral Small | 96 | 86% | 78–91 | $0.025 | 565 |
| Mistral Small · no reasoning | 84 | 79% | 70–86 | $0.022 | 600 |
| Llama 4 Scout | 0 | no score | — | — | — |
Where a model supports them, we ran two reasoning settings. That gives 12 models and 20 arms. 19 of those arms, covering 11 models, returned a usable answer. Jev finishes 12th of those 19 arms, at 92% with a 95% interval of 85–96. That interval overlaps every arm from 86% upward. With a hundred tickets, accuracy does not establish a settled twelfth place. Cost and latency separate much more cleanly: Jev is the cheapest and the fastest arm in the table, by a margin much larger than the accuracy interval.
Fig. 03 Median time to route one ticket, showing the eight fastest arms. results.json, p50 over 100 calls per arm
Benchmark 2: reranking retrieved passages
350 questions, 9 arms, depth 100. BM25 searches 261,046 passages from three public corpora, keeps the top 100 for each question, and gives every arm the same candidates in the same order. Each arm returns its best 10. At report time, we score the stored rankings against the datasets' gold labels. The results file stores no metric, and nobody on this project wrote a question or judged an answer.
Source: report/pipeline.architecture.json
Coverage comes before quality. Coverage is a result in its own right. An arm that answers easy questions and skips the rest can look excellent on its chosen subset. The first table is therefore not the leaderboard. It counts how many questions each arm answered at all.
Fig. 04 Questions answered by each arm. Coverage comes before quality: a score over the subset an arm picked for itself will not compare with a score over all 350 questions. data/rag/results-d100-merged.json
Fig. 05 Top-five recall for reranking. We withhold any arm that answered under 90% of the questions. data/rag/results-d100-merged.json, 350 questions, depth 100
| Model | Answered | recall@5 | nDCG@10 | Calls/q | p50 ms | $ / 1,000 |
|---|---|---|---|---|---|---|
| GLM 5.3 Flash · no reasoning | 95.1% | 0.369 | 0.400 | 1 | 3,156 | $3.061 |
| Claude Sonnet 5 | 90.0% | 0.368 | 0.405 | 1 | 3,963 | $67.585 |
| Qwen 3.8 Flash · no reasoning | 99.4% | 0.361 | 0.391 | 1 | 4,513 | $3.394 |
| DeepSeek v4.1 Flash · no reasoning | 98.9% | 0.359 | 0.393 | 1 | 3,047 | $4.796 |
| Jev · Score | 100.0% | 0.349 | 0.369 | 2 | 833 | $1.545 |
| Jev · Noul | 100.0% | 0.337 | 0.364 | 2 | 594 | $1.419 |
| BM25 (no model) | 100.0% | 0.259 | 0.259 | 1 | 0 | $0.000 |
| Gemini 3.8 Flash · no reasoning | 82.0% | withheld | withheld | 1 | 3,708 | $16.130 |
| Jev · Choice | 75.7% | withheld | withheld | 1 | 582 | $0.885 |
The top four arms differ in quality by 1.1 points. Across the whole table, price differs 44-fold. Their scores sit between 0.369 and 0.359 at 350 questions. That spread gives you no stable ranking to act on. Over the same questions, the priciest arm in the table, Claude Sonnet 5, charged $67.59 per thousand against $1.55 for the cheapest arm we scored, and answered 90.0% of the questions against 100.0%.
One ceiling limits all of it: 46.5% of judged-correct passages were ever put in front of the models. We average that figure per question over 1,531 judged passages. A reranker cannot surface a passage its retriever never fetched. Every table value stays below that ceiling. A better retriever is the cheapest improvement here.
The optimisation: fewer questions, more passages in each
Our first version sent one question per passage through Jev's per-pair primitives. Ranking one question's candidates cost 100 HTTP calls. The cookbook shows this pattern, and it made an earlier draft call Jev slow and expensive at reranking. That conclusion was wrong. Our harness was the bottleneck.
The planner in jevdemo/rag/jev_pair.py now packs passages into one request until the token budget fills. It packs greedily in retrieval order, so the grouping stays deterministic across runs. Three measured constants control it:
CHARS_PER_TOKEN = 2.89 # densest packing seen across the three corpora
BATCH_TOKEN_BUDGET = 24_000 # under the ~32,768 ceiling the Choice arm measured
BATCH_RETRY_SPLITS = 1 # a batch that returns nothing is halved once
Using the densest observed packing over-estimates token count for the other two corpora. That is the safe error. A batch that is too small costs one extra round trip. A rejected batch costs the whole question.
Fig. 06 The same arm and the same 350 questions, with 50x fewer requests. Jev · Noul is plotted here and Jev · Score moved the same way. data/rag/results-d100.json vs results-d100-batched.json
The arm, 350 questions, depth, and candidate lists stayed the same. 100 calls became 2, 6,841 ms became 594 ms, and $2.70 per thousand became $1.42. Quality held: recall@5 moved from 0.334 to 0.337, well inside the interval. The savings come from the request envelope. We used to send the query text and instructions 100 times per question. Now we send them 2 times.
Fig. 07 The request sets the ceiling. An encoding that can split the work across requests clears it. A single-request encoding loses the question. jevdemo/rag/jev_pair.py, BATCH_TOKEN_BUDGET and the measured 32,850-token largest success
The thing worth taking away
Jev's ceiling belongs to the request rather than to the model. The Choice encoding must fit the query and every candidate into one request. If it does not, there is no smaller fallback request and the question is lost. A per-pair encoding splits the same work across two requests. The arm we expected to be strongest at depth 100 is the one whose score this post withholds.
Reproducing it
The scoring path runs offline, stays deterministic, and costs nothing. The raw model output is committed, so you can rebuild the tables without spending a cent.
git clone https://github.com/rachit-srivastava-devx/jev-classification-benchmark
cd jev-classification-benchmark
python3 -m venv .venv && .venv/bin/pip install aiohttp requests pyarrow pypdf
python3 -m pytest -q # offline, no key, no spend
# rebuild every number in this post from the committed runs
.venv/bin/python scripts/build_blog.py # -> blog.html
# the paid runs, if you want to re-measure rather than re-check
.venv/bin/python -m jevdemo.runner # benchmark 1
.venv/bin/python scripts/run_rag.py --depth 100 \
--cap-micro 60000000 # benchmark 2
--cap-micro is a hard spend cap. We check it per query against the provider's reported cost, rather than an estimate. We added it after an earlier run hit an account spend limit mid-flight.
The reranking run combines three sessions. We re-ran the two batched Jev arms and Claude Sonnet 5 separately, then merged them with scripts/merge_rag_runs.py, which refuses to splice runs differing in depth, top-k, task set, or scoring cutoff, and records each arm's source file in the output's provenance block.
All 85 failures return that string. None is a timeout, rate limit, or malformed answer. The arm never ranked those 85 questions, so we withhold its score instead of printing a low one.
Open questions
What we still do not know
Four open questions could still change the conclusions above. The repository does not answer them.
- Where the Choice input ceiling sits. Our largest successful request measured 32,850 tokens, so the wall is somewhere above it. We never bisected the exact boundary, and it may not be constant.
- Whether the probabilities are calibrated. Noul returns a number between 0 and 1. We used it for sorting, never as a threshold. This repository does not check whether 0.8 is right eight times in ten, and we found nobody else who has checked it.
- How it behaves outside English support text and product search. We tested three corpora in one language.
- Whether the batching win survives a harder token budget. We tuned it to one ceiling on one endpoint.
If a rerun differs, you can locate the disagreement. The runs are committed and scoring happens offline.



