Optional classification

A model-backed decision aid, separate from Pixel’s local code index.

Updated

When to use it

pixel classify chooses between labels you supply: task routing, severity or a review gate. It does not retrieve code and is not needed to try the local index. Unlike index queries, it uses a model and its answer is not deterministic.

Use --criterion to define each label and --context to explain the decision. The output gives a probability per label and predicted: names the highest score. Remote engines send the question to the configured provider; Ollaya runs a local model. Model scores are a decision aid, not evidence that the choice is correct.

Example and published evidence

The example below uses illustrative routing labels. Its scores are from one local run; they are not a comparison of the named coding models. Accuracy and latency figures are upstream measurements with different samples, not a Pixel reproduction.

Ollaya reports 0.722 typed-decisions accuracy for winnow:e4b and 0.738 for TypeSafe Jev. The separate latency measurements below use different hardware and requests; they do not establish a speed winner.

These are upstream measurements, not a Pixel reproduction. The benchmark page describes the sample, scoring and source.

  • 89ms winnow:e4b on Ollaya RTX 4090, five questions · 0.722 accuracy
  • 236–276ms TypeSafe Jev Hosted API, median request · 0.738 accuracy
$ pixel classify "Review this Rust PR, find correctness bugs, and propose a safe patch." \    --engine ollaya \    --context "Choose the cheapest model that can reliably handle the request." \    --label 'claude-haiku-4.5_(fast)' \    --label 'claude-opus-5.5_(strong)' \    --label 'claude-fable-5.1_(reasoning)' \    --criterion 'claude-haiku-4.5_(fast)=Simple rewriting, extraction, or classification; no deep reasoning.' \    --criterion 'claude-opus-5.5_(strong)=Complex coding, multi-file review, or tool use; accuracy matters.' \    --criterion 'claude-fable-5.1_(reasoning)=Multi-step analysis, difficult debugging, or high uncertainty.'
  1. claude-fable-5.1_(reasoning) 0.192
  2. claude-haiku-4.5_(fast) 0.098
  3. claude-opus-5.5_(strong) 0.710

predicted: claude-opus-5.5_(strong)

One local run using Ollaya: each score is the local model's probability for that label; predicted: is the highest-scoring choice. Remote engines are also available and report their own scores, renormalized to sum to 1 — a decision aid, not a measurement.

Benchmark source and method · typed-decisions accuracy, higher is better; latency, lower is better

Further reading

Source, sample and scoring · Jev comparison · Agent protocol