When to use it
pixel classify chooses between labels you supply: task routing, severity or a review gate. It does not retrieve code and is not needed to try the local index. Unlike index queries, it uses a model and its answer is not deterministic.
Use --criterion to define each label and --context to explain the decision. The output gives a probability per label and predicted: names the highest score. Remote engines send the question to the configured provider; Ollaya runs a local model. Model scores are a decision aid, not evidence that the choice is correct.
Example and published evidence
The example below uses illustrative routing labels. Its scores are from one local run; they are not a comparison of the named coding models. Accuracy and latency figures are upstream measurements with different samples, not a Pixel reproduction.
Ollaya reports 0.722 typed-decisions accuracy for winnow:e4b and 0.738 for TypeSafe Jev. The separate latency measurements below use different hardware and requests; they do not establish a speed winner.
These are upstream measurements, not a Pixel reproduction. The benchmark page describes the sample, scoring and source.
- 89ms winnow:e4b on Ollaya RTX 4090, five questions · 0.722 accuracy
- 236–276ms TypeSafe Jev Hosted API, median request · 0.738 accuracy
$ pixel classify "Review this Rust PR, find correctness bugs, and propose a safe patch." \ --engine ollaya \ --context "Choose the cheapest model that can reliably handle the request." \ --label 'claude-haiku-4.5_(fast)' \ --label 'claude-opus-5.5_(strong)' \ --label 'claude-fable-5.1_(reasoning)' \ --criterion 'claude-haiku-4.5_(fast)=Simple rewriting, extraction, or classification; no deep reasoning.' \ --criterion 'claude-opus-5.5_(strong)=Complex coding, multi-file review, or tool use; accuracy matters.' \ --criterion 'claude-fable-5.1_(reasoning)=Multi-step analysis, difficult debugging, or high uncertainty.'- claude-fable-5.1_(reasoning) 0.192
- claude-haiku-4.5_(fast) 0.098
- claude-opus-5.5_(strong) 0.710
predicted: claude-opus-5.5_(strong)
predicted: is the highest-scoring choice. Remote engines are also available and report their own scores, renormalized to sum to 1 — a decision aid, not a measurement.Benchmark source and method · typed-decisions accuracy, higher is better; latency, lower is better
Further reading
Source, sample and scoring · Jev comparison · Agent protocol