Saturday, October 3, 2026
DeepSeek V4.1 Flash (Sept): 552B-parameter MoE, only ~8B active per token, new architecture cuts per-token memory to ~1/4; pricing $0.30/$1.20 per M tokens, halved off-peak; since mid-Sept DeepSeek routes V4 Pro requests to Flash at Flash rates — undercutting its own flagship Inception Mercury 2.5 (Sept): diffusion-based language model — generates many tokens at once then refines; 1,107 tokens/sec on standard Nvidia GPUs; 260K context; 80% launch discount ($0.04/$0.15 per M tokens); targets voice agents/search/RAG latency New Muse abilities: none — no new Muse launch since Sept 29 Small Busine
Transcript
The full episode script, for reading or reference: