Short answer
For a bounded coding decision, pixel classify asks a model you configure and returns one probability per label. With deepseek-v4.1-flash it answered all 14 public JevBench coding items right; Jev publishes 0.839 on its 56. The samples differ, so read it as a direction. The rest of Pixel needs no decision model.
Side by side
Published figures, different samples. Pixel's own measurement set beside Jev's published figure: not a head-to-head, and the samples differ.
| Measure | Pixel | Jev |
|---|---|---|
| JevBench coding items | 14 of 14 with deepseek-v4.1-flash | 83.9% |
| Model | The one you choose (13 of 14 with gpt-oss) | Jev's own |
| Items scored | 14, the public ones | 56 |
Every figure above is on the benchmarks page, with its sample size, its method and the cases where Pixel loses.
Where Jev wins
Its score covers all 56 of its coding items, where Pixel's run covers only the 14 public ones.
How they differ
Jev is a model built for one job: settling a coding question with a bounded set of answers. pixel classify puts the same kind of question to a model you configure (OpenRouter, Ollama Cloud or a local server) and returns one probability per label. It is an optional add-on: nothing else in Pixel calls a model.
Choosing
Pick Jev for a dedicated model scored on its whole benchmark. Pick pixel classify to keep the choice of model, the bill and the data path in your hands, next to an index your agent already queries.