Wally
Home
ResearchNews10.3K
Talk to usLogin
Y Combinator

Backed by Y Combinator

Engineering notes

From every layer of the stack.

AllQHexRT2MetalRT4SDKs3Agents2Voice2
We Put a 27B Model in Your Pocket: PrismML Bonsai 1-bit on iPhone, Android, and Mac
QHexRT·July 19, 2026

We Put a 27B Model in Your Pocket: PrismML Bonsai 1-bit on iPhone, Android, and Mac

Run a 27 billion parameter LLM on your phone. PrismML's 1-bit Bonsai reasoning models are live in the Wally apps on iOS, Android, and macOS, including the first true 1-bit model ever to run on any NPU.

QHexRT Is Live: Full-Stack NPU Inference for Qualcomm Hexagon
QHexRT·June 25, 2026

QHexRT Is Live: Full-Stack NPU Inference for Qualcomm Hexagon

QHexRT is officially live — the first inference engine built to run LLM, VLM, STT, TTS, and embeddings 100% on Qualcomm Hexagon NPUs. First model: LFM 2.5 230M at 12,540 tok/s prefill and 36ms flat TTFT on v81.

Wally

Wally hosted inference, local neural accelerators, and open-source SDKs from one inference lab.

Wally

  • Overview
  • Documentation
  • Console login

Developers

  • NeuRT
  • QHexRT
  • Local SDKs
  • GitHub

Research

  • Publications
  • Engineering blog
  • Inference Radar

Company

  • About
  • Press
  • Talk to us

© 2026 Wally, Inc.

  • Terms
  • Privacy
  • Refunds
  • Acceptable use