ollanet docs

Bench

Fixed suite for speed plus a light quality check on one host, model list, suite, and run count. Those checks are not a general model-quality ranking. Results do not compare across hardware without the same conditions.

A common use is a Finetuna before and after on one host. Finetuna is optional. You can bench any named model already on the machine.

ollanet bench <machine> [model...]
ollanet bench studio --all
ollanet bench studio gemma3:12b --hot --runs 5
ollanet bench studio --suite full

Suites

Suite Speed Quality
quick (default) 256-token count, --runs times → tok/s (med) ping, math, haiku
full that, plus one 1024-token prose shot → tok/s long + json, reason

The 256-token case is enumerative on purpose so models hit num_predict instead of early EOS. It is peak pinned decode — not “how a long chat feels.” Use --suite full for the chat-shaped column.

Useful flags

Flag Meaning
--runs <n> Counted 256-token repeats (default 3)
--hot Discard first 256-token run; keep models loaded
--num-predict <n> Pin for the 256-token case only (not the long case)
--cold-load Measure cold load via unload + /api/ps
--json / --save Machine-readable / persist under benchmarks/

See docs/bench-spec.md for measurement rules.