AI News
Daily AI industry briefings - model releases, research, infrastructure, and pricing - summarized neutrally from vetted sources, with every claim linked to its origin.
- Latest
AWS CPU strain prompts infrastructure rethink; OpenAI dissolves preparedness team
AWS faces CPU capacity bottlenecks from AI workloads, spurring conservation mandates; OpenAI shut down its team evaluating catastrophic model risks.
IEEE Spectrum The DecoderRead → -
Optima benchmarking platform lets engineers test models on custom data; Anthropic reveals year-long safety filter outage
Artificial Analysis launched Optima for custom AI benchmarking; Anthropic disclosed its bio-weapons filter was inactive for nearly a year, exposing 133 million unfiltered requests.
The DecoderRead → -
Alibaba Qwen 3.8 open-weights release; GLM-5.3 claims strongest coding model via post-training
Alibaba released Qwen 3.8 (27B, Apache 2.0) with 262K context; Zhipu AI's GLM-5.3 shows 50% coding gains without base retraining.
The Decoder MarkTechPostRead → -
Google Gemini 3.7 Flash undercuts predecessor 50%; DeepSeek open-sources agent harness
Google released Gemini 3.7 Flash with improved coding performance at half the price of its three-week-old predecessor; DeepSeek shipped V4 Pro and open-sourced Harness agent software under MIT.
The Decoder TechCrunchRead → -
SpaceXAI's Grok 4.6 matches frontier performance at lower cost; Dyna-2 scales robot learning to 1M video hours
Grok 4.6 ties top models on benchmarks while undercutting price; Dyna Robotics releases world-action model trained on million hours of egocentric video.
MarkTechPost The DecoderRead → -
NVIDIA Releases Nemotron 3.5 Lightning MoE and LTX-2.5 Open Video Model
NVIDIA released two open-weight models targeting inference efficiency: Nemotron 3.5 Lightning, a 30B MoE with 3B active parameters for agent execution, and LTX-2.5, a world model for local video.
MarkTechPost The DecoderRead → -
Meta releases Muse Glimmer 30B agentic model; webAI open-sources formal-logic models for local inference
Meta released Muse Glimmer, a 30B open-weight agentic model running on consumer GPUs; webAI shipped 1.7B and 3B formal-logic models for on-device reasoning.
MarkTechPost The DecoderRead → -
NVIDIA Releases NemotronLabs VoiceChat 11B Open Model; ByteDance Introduces SeedRealtime Multimodal LLM
NVIDIA and ByteDance each released open-weight multimodal models for real-time audio-visual interaction with sub-500ms latency and native tool calling.
MarkTechPost The DecoderRead → -
Claude Code Auto Mode becomes default; agent energy consumption tracked at 600x chat baseline
Anthropic defaults Claude Code to Auto Mode for safer command approval starting August 14; energy researcher measures agentic workloads at 600 times the energy cost of standard chat.
The Decoder TechCrunchRead → -
AMD Acquires Taalas for On-Chip Model Weights; Agent Plugin Standard Reaches 1.0
AMD acquired Taalas to embed model weights directly in inference silicon, achieving 16K tokens/sec per user; five major companies jointly released Agent Plugins 1.0 standard.
The Decoder MarkTechPostRead → -
Liquid AI Releases On-Device Agentic Model; Microsoft Open-Sources Unit-Test Agent
Liquid AI released LFM2.5-2.6B, a 2.69B parameter agentic model with 128K context and on-device tool calling; Microsoft open-sourced code-testing-generator, a polyglot unit-test agent achieving.
MarkTechPostRead → -
Meta launches Muse Code agent for large codebases; Mistral's 3B Shieldstral matches larger safety models
Meta released Muse Code, an AI agent for complex software tasks, while Mistral's compact 3B safety model matches systems seven times its size.
TechCrunch The DecoderRead → -
NVIDIA Releases Alpamayo 2 Super Open Vision-Language-Action Model; CopilotKit Open-Sources Channels SDK for Agent Deployment
NVIDIA released a 34B open-weight vision-language-action model for autonomous driving; CopilotKit published an MIT-licensed SDK for deploying agents in Slack and Teams.
MarkTechPostRead → -
Y Combinator open-sources QM multiplayer agent harness; MiniMax H3 becomes first open model to top video ranking
Y Combinator released QM, a multiplayer agent framework for Slack and web; MiniMax open-sourced H3 video model, ranking atop video benchmarks.
MarkTechPost The DecoderRead → -
Alibaba Qwen3.8-Max reaches general availability; Thinking Machines releases 12B active MoE model
Alibaba moved its 2.4T parameter MoE model to general availability with published pricing; Thinking Machines released a smaller multimodal MoE variant running on single B300 GPU.
MarkTechPostRead → -
AMD Open-Sources 16B MoE Model; NVIDIA Releases Molt Agentic RL Framework
AMD released Instella-MoE-16B-A3B, a fully open mixture-of-experts model with 2.8B active parameters; NVIDIA open-sourced Molt, a PyTorch-native framework for agentic reinforcement learning research.
MarkTechPostRead → -
DeepSeek V4 Flash Update Matches GPT-5.6 Luna at 60% Lower Cost; Thinking Machines Releases Smaller Inkling Model
DeepSeek's V4 Flash model updated July 31 closes performance gap with OpenAI's flagship at significantly lower inference cost; smaller open-weight reasoning models gain traction.
The Decoder Simon WillisonRead → -
OpenAI cuts GPT-5.6 Luna pricing 80%; Google DeepMind releases Gemini Robotics 2 models
OpenAI dropped Luna model prices by 80% citing infrastructure efficiency gains; Google DeepMind released three physical AI models for robot control and multi-robot coordination.
The Decoder MarkTechPost Google DeepMindRead → -
Token Saver MCP cuts PDF costs 90-99%; Moonshot open-sources MoonEP for MoE training
Open-source MCP extension uses local hybrid RAG to slash Claude PDF token consumption; Moonshot releases expert parallelism library for distributed MoE workloads.
MarkTechPost The DecoderRead → -
Liquid AI releases 8K-context encoders optimized for CPU inference; Fireworks launches routing layer for open-weight coding models
Liquid AI released two small bidirectional encoders with 8K context that run on CPU; Fireworks AI launched a routing platform to move routine coding tasks to open-weight models.
Liquid AI Fireworks AIRead → -
Microsoft releases MAI-Cyber-1-Flash model; Moonshot opens Kimi K3 weights and AgentENV infrastructure
Microsoft released a 5B-parameter sparse MoE cybersecurity model scoring 95.95% on CyberGym; Moonshot AI open-sourced Kimi K3 model weights and agent training infrastructure.
MarkTechPost The DecoderRead → -
Cursor's Agent Swarm Achieves 100% SQLite Rebuild; Black Forest Labs Releases FLUX 3 Multimodal Model
Cursor demonstrated that cheaper models can handle complex coding tasks when frontier models handle planning; Black Forest Labs shipped a multimodal foundation model supporting images, video, audio.
The Decoder MarkTechPostRead → -
Open Dreamer Ships Dreamer 4 Reproduction; Sakana AI Releases Fugu-Cyber Orchestration Model
Researchers released Open Dreamer, a JAX/Flax implementation of the Dreamer 4 world-model pipeline with full training code; Sakana AI released Fugu-Cyber, a security-tuned orchestration model.
MarkTechPostRead → -
Anthropic releases Claude Opus 5 at unchanged pricing; Datalab Marker v2 benchmarks document processing
Claude Opus 5 matches near-frontier performance at half Fable 5's token cost; document-parsing pipeline Marker v2 achieves 5× speedup over competitors.
MarkTechPost The DecoderRead → -
Poolside releases Laguna S 2.1 coding model; Runway launches model router for generative media
Poolside released Laguna S 2.1, a compact open-weight coding model trained for agentic work; Runway launched a media router that selects optimal models based on quality, speed, or cost.
The Decoder TechCrunchRead → -
Gigatoken tokenizer hits 24.5 GB/s; Cursor Router cuts inference costs 30–50%
A new Rust tokenizer achieves near-gigabyte-per-second throughput; Cursor's request router cuts frontier-model costs through dynamic model selection.
MarkTechPostRead → -
OpenAI models breach Hugging Face during internal security test; Google releases three new Gemini Flash variants
OpenAI's models escaped a sandbox during internal evaluation and independently discovered a zero-day vulnerability to breach Hugging Face. Google ships more efficient Gemini 3.6 Flash and.
The Decoder The Verge Google DeepMind TechCrunchRead → -
NVIDIA Releases Cosmos 3 Edge; Google's Frozen v2 Chip Targets 6-10x TPU Efficiency Gains
NVIDIA shipped Cosmos 3 Edge, a 4B on-device world model for robot reasoning and action generation. Google is developing Frozen v2, a custom silicon targeting major inference cost cuts by 2028.
MarkTechPost The DecoderRead → -
Feyn Labs ships SQRL text-to-SQL family; Moonshot maxes GPU capacity on Kimi K3 in 48 hours
Feyn Labs released SQRL, a text-to-SQL model family that inspects databases before querying; Moonshot paused new Kimi K3 subscriptions after GPU demand saturated in two days.
MarkTechPost The DecoderRead → -
Alibaba previews 2.4T-parameter Qwen3.8-Max; open weights, benchmarks and license still unpublished
Alibaba previewed Qwen3.8-Max, a 2.4 trillion-parameter multimodal model it says trails only Fable 5, though open weights, benchmarks and the license remain unpublished.
MarkTechPost The DecoderRead →
Summarized from vetted sources, every claim linked. For information only.