Wally Labs
Every claim comes with numbers.
We publish our benchmark results and engineering deep-dives openly. On-device inference: fast, private, hardware-native.
QHexRT
NPU inference for all Qualcomm Hexagon devices. First benchmarks on Snapdragon silicon.
MetalRT
Custom kernel inference engine for Apple Silicon. Record-setting LLM, speech, vision, and speech-to-speech performance.
MetalRT · Speech-to-SpeechMar 15, 2026
MetalRT Now Does Speech-to-Speech. 1.52x Faster Than mlx-audio.
Read the benchmarks123 tok/s
S2S throughput
MetalRT · VisionMar 13, 2026
MetalRT Now Runs Vision Language Models. Fastest on Apple Silicon.
Read the benchmarks287 tok/s
vision decode
MetalRT · SpeechMar 9, 2026
The First Complete AI Inference Engine for Apple Silicon. Now with Speech.
Read the benchmarks101ms
STT latency
MetalRT · LLMMar 3, 2026
We Built the Fastest LLM Decode Engine for Apple Silicon.
Read the benchmarks658 tok/s
LLM decode
FastVoice
End-to-end on-device voice AI. Co-scheduled inference for sub-100ms first-audio latency.