The one thing
01vLLM 0.28.0 ships Model Runner V2 defaults and breaks bitsandbytes-in-tree; Claude Code latest is 2.1.247 while stable stays 2.1.231.
vLLM v0.28.0 (26 Aug) raises max_num_batched_tokens from 8192 to 16384, enables Mamba prefix caching by default, and matures Model Runner V2 (E/P/D disaggregation, tiered KV offload including disk, multi-layer MTP KV).1 Breaking: bitsandbytes moves out-of-tree; Transformers floor is 5.15.0; calculate_kv_scales and override_attention_dtype are removed.1
npm @anthropic-ai/claude-code: stable still 2.1.231; latest/next → 2.1.247.23 2.1.247 adds SendFeedback drafts and fixes /compact / --agent summarization under the wrong system prompt.3 Channel policy from #015 stands.
Shipped
02vLLM 0.28.0 · Runner V2, tiered KV offload, bitsandbytes out-of-tree
Major engine cut (26 Aug): Runner V2, tiered KV with disk offload, Rust/gRPC multimodal path, new models (Muse Glimmer, Ling 3.0 Flash, Dots3 NOTE).1 Batch token default doubles to 16384. bitsandbytes is a plugin; pin Transformers ≥5.15.0.1
Codex 0.150.0 · @task mentions, Interrupt hooks, /copy picker
rust-v0.150.0 (26 Aug; 0.150.1 same day): @ mentions across Codex tasks; /copy picker; auto titles; Interrupt hooks for commands/MCP on turn cancel; permission-mode shortcuts; Vim . repeat.67
Promised, not shipped
OpenAI Private Safety Processing / frontier ZDR — September white paper (issue 013) · Origin agent-native source-control — early beta ongoing (issue 011) · Claude Code /design skill — research preview · MAI-Code-1-Flash off Copilot — 10 Sept 2026 · Grok Bot enterprise — waitlist · OpenAI Ultrafast — expands with capacity · Ox Alpha free window — ~through 27 Aug (ends today) · GPT-5.6 Sol promo pricing — at least through 21 Nov 2026 · Plus Codex/Work 5h limit — live from 25 Aug; Pro stays off for upcoming months · GitHub Copilot global model policy enforcement — rolling through 1 Sept 2026
The conversation
01vLLM 0.28 is a real serving break; Claude Code still has no native AGENTS.md; Cursor permanently raised included Grok usage.189
- vLLM v0.28.0 release notesPrimary, 26 Aug
Defaults up; bitsandbytes out-of-tree; Transformers ≥5.15.1
- @trq212 (Claude Code) via X digestVendor, digest 26 Aug 05:30 (~24.5h old at window close)
Easier AGENTS.md / prompt mods coming; per-model prompts stay; interim
@AGENTS.mdfrom CLAUDE.md.8 - @leerob (Cursor) via X digestVendor, same digest
Included Grok usage raised permanently across first-party Cursor models.9
Upgrade vLLM only after the bnb/Transformers checklist. AGENTS.md native support is still unshipped. Re-check Cursor quotas if Grok is default.
GitHub releases: 28 repos, sweep_complete true, authenticated, 19 in-window. Vendor feeds: 12 sources, sweep_complete true, 10 in-window (Actions critical 2.8h resolved; Actions+PR 1.5h resolved; Billing minor open). X accounts digest (~101 roster) 05:30 UTC 26 Aug (~24.5h old at window close) — colour only. X keyword same age. Reddit pulse (6 subs, 1 post). HN in-window: GLM-5.3-Flash, Ox Alpha, GitHub disruption. arXiv not swept. Morning cron missed; recovered ~08:30 UTC.
Skip this
06GitHub Actions critical 2.8h; Actions+PR minor 1.5h; Billing minor open. Resolved or open weather with duration — not a pin. Billing still open at sweep; watch status if invoices fail.
GitHub Copilot global model policy GA (rollout through 1 Sept). Enterprise admin defaulting change. Open-weight and data-retention models stay off by default. Promised, not a day-of ship for most readers.
OpenAI Admin plugin for ChatGPT Work & Codex; Hugging Face incident post. Admin surface consolidation. HF incident page timed out on fetch — not printed as fact.
Transformers 5.16.0/5.16.1 (DTensor TP, Fuyu image_patch_indices drop, NVFP4). Real breaks for TP/Fuyu stacks, but secondary to the vLLM floor note already in One thing.
openai-python 3.4.0/3.5.0; anthropic-sdk-python 1.1.0; adk-python 2.8.0; ollama 0.33.1; MCP Python 2.0.1; cline 4.1.16/cli 3.0.60; AI SDK patches. Optional call IDs, thinking display beta, ADK live/data-agent tools, Qwen3.8 Flash Next on MLX, FastMCP import pointer, Cline hub polish — no default-stack pin today.
GLM-5.3-Flash / Qwen3.8-Flash-Next HN spikes; Nvidia–Hugging Face $13B rumour. Model drops without a verified serving pin for this audience. Acquisition item is Business Insider talk, not a closed deal we opened.
Everything we saw
5555 candidates scanned · 8 used in this issue — the rest, with the reason each one was left out
| Item | Source | Signal | Call |
|---|---|---|---|
| vLLM v0.28.0 | github releases | 22.8h | led the issue |
| npm @anthropic-ai/claude-code dist-tags | npm registry | stable 2.1.231 / latest 2.1.247 | one thing |
| Claude Code v2.1.247 | github releases | 9.4h | one thing |
| Transformers v5.16.0 | github releases | 20.0h | skipped secondary |
| Transformers v5.16.1 | github releases | 17.7h | skipped secondary |
| Codex rust-v0.150.0 | github releases | 12.9h | shipped |
| Codex rust-v0.150.1 | github releases | 6.6h | shipped patch |
| X accounts digest 26 Aug 05:30 | digests | ~24.5h old; 101 accounts | conversation colour |
| GitHub global model policy GA | github changelog | 10.4h | skipped / promised |
| GitHub Actions critical incident | github status | 2.8h critical resolved | skipped weather |
| openai-python v3.5.0 | github releases | 7.6h | skipped |
| anthropic-sdk-python v1.1.0 | github releases | 15.3h | skipped |
| adk-python v2.8.0 | github releases | 9.1h | skipped |
| ollama v0.33.1 | github releases | 14.4h | skipped |
| MCP Python SDK v2.0.1 | github releases | 21.8h | skipped |
| HN — GLM-5.3-Flash | news.ycombinator.com | 1011 pts / 507 cmt | skipped model drop |
| HN — Ox Alpha / Z.ai weights | news.ycombinator.com | 425 pts | skipped |
| Reddit pulse — 6 subs, 1 post | reddit pulse | layer ran | layer ran |
| OpenAI — Hugging Face incident post | openai.com | crawl timeout | skipped unverified |
End of feed. That is everything from the window worth your time.
Next issue tomorrow, 06:00 UTC — and if nothing ships, it will say so in two hundred words.