Token Speed Test

Pick a model, hit run, and watch a simulated tokens-per-second benchmark for your hardware.

Detected Hardware

Detecting...

Detected via your browser. Used only to calibrate the simulation below.

108 models across 20 families

0.0tokens / sec

Estimated for Mistral 7B on your hardware

This is a simulation. No model actually runs on your device — results are estimated from your reported hardware specs and typical local-inference throughput for each model size.

Known Bugs & Limitations

  • GPU often shows as "Hidden by browser" on Brave, Firefox (resistFingerprinting), and Safari — these browsers block the real GPU name for privacy. Pick your GPU manually from the dropdown that appears when this happens.
  • Memory can read lower than your real RAM, or "Unknown." Chrome-based browsers cap the reported value at 8 GB regardless of actual RAM; Firefox and Safari don't expose it at all.
  • CPU core count includes hyperthreads/efficiency cores, not just physical performance cores, so the estimate can skew on hybrid CPUs (e.g. recent Intel) and heavily multithreaded mobile chips.
  • Mobile GPU detection is generally unreliable — integrated mobile GPUs report generic or inconsistent renderer strings across browsers.
  • Results are simulated estimates for illustration only, not a real inference benchmark. Actual local throughput depends on quantization, context length, batch size, thermal throttling, and the inference engine used.