llama.cpp 0.3.0 · dots3-note multimodal, DeepSeek 4 tensor-split, ggml 0.22 — from issue No. 017nofeed.dev
Issue No. 017 · 26 August 2026
Use it

llama.cpp 0.3.0 · dots3-note multimodal, DeepSeek 4 tensor-split, ggml 0.22

v0.3.0 (25 Aug): dots3-note with DSA-ISWA KV cache and vision/audio in mtmd; DeepSeek 4 gets -sm tensor and multi-sequence rollback fixes; GLM-4.5-Air MTP; ggml → 0.22.0 (meta-backend tensor split, per-op Metal kernels, non-in-place ggml_clamp).5 mtmd: WebP via ffmpeg, Pillow-accurate resize. Server: LLAMA_SERVER_SLOTS_N_DIFF. Web UI: tabbed chat. Nightly pointer b10621.5

MATTERS TO · local llama.cpp / server / multimodal runners, especially DeepSeek 4 multi-GPU
WHAT THE VERDICT MEANS
Available, works, worth your time today.

Also in this issue

read the whole thing →
Use itClaude Code 2.1.246 · MCP honesty, background sessions, gateway edge cases

Get it by email.

One issue every weekday. The whole thing, not a teaser.

Or RSS, if you would rather we never had your address.