Published 18:27 UTC 14 Sep. llama_sampler_chain_n() returns int32_t rather than int. GDN normalisation moves from max to rsqrt for affected Qwen, Kimi and GLM models, and MTP context KV cache allocation is fixed for DeepSeek2 and GLM-MoE — bad output from those served locally may have been this.5
| Use it | Claude Code v2.1.271 · per-command allowed_domains, plugin commands pinned by hash |
One issue every weekday. The whole thing, not a teaser.