Jev is cheap.Is it good enough?

Independent evidence for Jev.

Benchmarks · Cost · Accuracy · Fallback

INPUTQUESTIONCONFIDENCEPASSFALLBACK
Faster
Cheaper
Structured
But
Accurate
enough?
Where Jev wins
0 ms
Median latency
vs 687 ms · Claude Haiku 4.5
$0.000
Measured workload cost, per 1,000
vs $0.462
Where Jev loses
0.0%
Accuracy
vs 81.3% · Claude Haiku 4.5

Different tasks.
Different outcomes.

Across the runs recorded here Jev ranges from 62.6% to 95.4%. There is no universal Jev accuracy, and our schema has nowhere to store one.

PhishNChips v5.2 · Label an email as phishing or legitimate. · n=2,000 · evidence

Same model, same task, different prompt.

85.7%
Phishing recall, text-only prompt
98.4%
Same set, same model, enriched evidence
43.2%
A different set, text-only protocol

All three are Jev detecting phishing. The gap between the first two is nothing but the prompt, on the same 5,733-email set. A recall figure quoted without its protocol is not a fact about the model — which is the whole reason every number here travels with the run that produced it. See the run.

The author is explicit that the enriched criteria were informed by labelled errors, so part of that gain is supervision a cold-start deployment would not have.

Cost × accuracy

The trade-off, on one dataset where both were measured.

Scoped to a single run deliberately. A cost–accuracy plane implies the points are alternatives to one another, and that is only true when they were measured on the same data under the same protocol.

$0.1$1.0051%61%72%83%93%COST PER 1,000 CALLS → (log)ACCURACY ↑Jev62.6% · $0.038Claude Haiku 4.581.3% · $0.462

Hover or tap a point for its full record.

Recent benchmark runs

Verified runs. No marketing claims.

Evidence stream
  • 09.19Jev pricing verified at OpenRouter
  • 09.19Claude Haiku 4.5 pricing verified at Anthropic (first-party API)
  • 09.19Source verified · OpenRouter
  • 09.19Source verified · Anthropic
  • 09.19Source verified · Latent Space (AINews)
  • 09.19Source verified · anisselbd
  • 09.19Source verified · ickma2311
  • 09.19Source verified · themsquared
Fallback economics

How much quality can you recover before the saving is gone?

A Jev call on every request, escalating to Claude Haiku 4.5 on the ones it is not confident about. The break-even fallback rate is where that cascade stops being cheaper than running the baseline on everything.

Model your own workload
0.0%
Break-even fallback rate

Jev can fall back on up to 99.8% of requests before this hybrid path costs more than baseline-only.

Jev only$12.60/mo
Hybrid at 20%$1,023/mo
Baseline only$5,050/mo

1M requests/month · Jev 300 input tokens · Claude Haiku 4.5 500 input + 910 output · verified list pricing both sides

Extraordinary claims need independent evidence.

10
Verified benchmarks, plus 2 routing evaluations — 4 of them pre-registered
7
Independent sources behind them
0
Independently reproduced — a public repository is not a reproduction, and we check
Read the methodology
What this site will not do
No universal Jev score, no overall ranking, and no comparison page without a direct head-to-head. Where a number is missing it stays missing — 12 sources are recorded, and the ones we have not read yet sit in an intake queue rather than on a page.