All issues
Inference Radar·2026-W24·Jun 11 — Jun 17, 2026·15 min read

SGLang Drags DFlash Into Serving

“The week brought few new base models, but a lot of work that decides whether those models can run at scale. vLLM, SGLang, FlashInfer, llama.cpp, MLX, OpenVINO, ExecuTorch, and ROCm all pushed on the same limit: KV cache, memory movement, and hardware-specific kernels.”

Cover for SGLang Drags DFlash Into Serving
3,893 commits
2,147 PRs
925 issues
76 releases
79 active repos
Weekly activity by organization

Weekly briefing

Get the next issue in your inbox.

One email, every week. Every link cited. No fluff, no crypto analogies.

Subscribe on Inference Radar