Token Speed Test
Pick a model, hit run, and watch a simulated tokens-per-second benchmark for your hardware.
Detected Hardware
Detecting...
Detected via your browser. Used only to calibrate the simulation below.
108 models across 20 families
Estimated for Mistral 7B on your hardware
This is a simulation. No model actually runs on your device — results are estimated from your reported hardware specs and typical local-inference throughput for each model size.
Known Bugs & Limitations
- GPU often shows as "Hidden by browser" on Brave, Firefox (resistFingerprinting), and Safari — these browsers block the real GPU name for privacy. Pick your GPU manually from the dropdown that appears when this happens.
- Memory can read lower than your real RAM, or "Unknown." Chrome-based browsers cap the reported value at 8 GB regardless of actual RAM; Firefox and Safari don't expose it at all.
- CPU core count includes hyperthreads/efficiency cores, not just physical performance cores, so the estimate can skew on hybrid CPUs (e.g. recent Intel) and heavily multithreaded mobile chips.
- Mobile GPU detection is generally unreliable — integrated mobile GPUs report generic or inconsistent renderer strings across browsers.
- Results are simulated estimates for illustration only, not a real inference benchmark. Actual local throughput depends on quantization, context length, batch size, thermal throttling, and the inference engine used.