<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>gekro.com - AI News Briefings</title><description>Daily AI industry signal from Rohit Burani. No hype, no VC press releases. Curated from vetted sources for engineers.</description><link>https://gekro.com/</link><language>en-us</language><atom:link href="https://gekro.com/news-rss.xml" rel="self" type="application/rss+xml"/><item><title>AWS CPU strain prompts infrastructure rethink; OpenAI dissolves preparedness team</title><link>https://gekro.com/news/2026-08-17/</link><guid isPermaLink="true">https://gekro.com/news/2026-08-17/</guid><description>AWS faces CPU capacity bottlenecks from AI workloads, spurring conservation mandates; OpenAI shut down its team evaluating catastrophic model risks.</description><pubDate>Mon, 17 Aug 2026 07:00:00 GMT</pubDate><category>infrastructure</category><category>inference</category><category>model safety</category><category>deployment</category></item><item><title>Optima benchmarking platform lets engineers test models on custom data; Anthropic reveals year-long safety filter outage</title><link>https://gekro.com/news/2026-08-16/</link><guid isPermaLink="true">https://gekro.com/news/2026-08-16/</guid><description>Artificial Analysis launched Optima for custom AI benchmarking; Anthropic disclosed its bio-weapons filter was inactive for nearly a year, exposing 133 million unfiltered requests.</description><pubDate>Sun, 16 Aug 2026 07:00:00 GMT</pubDate><category>benchmarking</category><category>model evaluation</category><category>safety</category><category>api</category><category>watermarking</category></item><item><title>Alibaba Qwen 3.8 open-weights release; GLM-5.3 claims strongest coding model via post-training</title><link>https://gekro.com/news/2026-08-15/</link><guid isPermaLink="true">https://gekro.com/news/2026-08-15/</guid><description>Alibaba released Qwen 3.8 (27B, Apache 2.0) with 262K context; Zhipu AI&apos;s GLM-5.3 shows 50% coding gains without base retraining.</description><pubDate>Sat, 15 Aug 2026 07:00:00 GMT</pubDate><category>open-weight models</category><category>model releases</category><category>coding models</category><category>inference optimization</category><category>token economics</category><category>small language models</category><category>tool calling</category></item><item><title>Google Gemini 3.7 Flash undercuts predecessor 50%; DeepSeek open-sources agent harness</title><link>https://gekro.com/news/2026-08-14/</link><guid isPermaLink="true">https://gekro.com/news/2026-08-14/</guid><description>Google released Gemini 3.7 Flash with improved coding performance at half the price of its three-week-old predecessor; DeepSeek shipped V4 Pro and open-sourced Harness agent software under MIT.</description><pubDate>Fri, 14 Aug 2026 07:00:00 GMT</pubDate><category>model releases</category><category>inference cost</category><category>agent orchestration</category><category>token economics</category><category>open-weight models</category><category>API pricing</category></item><item><title>SpaceXAI&apos;s Grok 4.6 matches frontier performance at lower cost; Dyna-2 scales robot learning to 1M video hours</title><link>https://gekro.com/news/2026-08-13/</link><guid isPermaLink="true">https://gekro.com/news/2026-08-13/</guid><description>Grok 4.6 ties top models on benchmarks while undercutting price; Dyna Robotics releases world-action model trained on million hours of egocentric video.</description><pubDate>Thu, 13 Aug 2026 07:00:00 GMT</pubDate><category>frontier models</category><category>token economics</category><category>agentic AI</category><category>robotics</category><category>video pretraining</category><category>model training</category><category>post-training</category><category>open-weight models</category></item><item><title>NVIDIA Releases Nemotron 3.5 Lightning MoE and LTX-2.5 Open Video Model</title><link>https://gekro.com/news/2026-08-12/</link><guid isPermaLink="true">https://gekro.com/news/2026-08-12/</guid><description>NVIDIA released two open-weight models targeting inference efficiency: Nemotron 3.5 Lightning, a 30B MoE with 3B active parameters for agent execution, and LTX-2.5, a world model for local video.</description><pubDate>Wed, 12 Aug 2026 07:00:00 GMT</pubDate><category>open-weight models</category><category>mixture of experts</category><category>inference efficiency</category><category>token economics</category><category>video generation</category><category>local inference</category><category>API security</category></item><item><title>Meta releases Muse Glimmer 30B agentic model; webAI open-sources formal-logic models for local inference</title><link>https://gekro.com/news/2026-08-11/</link><guid isPermaLink="true">https://gekro.com/news/2026-08-11/</guid><description>Meta released Muse Glimmer, a 30B open-weight agentic model running on consumer GPUs; webAI shipped 1.7B and 3B formal-logic models for on-device reasoning.</description><pubDate>Tue, 11 Aug 2026 07:00:00 GMT</pubDate><category>open-weight models</category><category>agent inference</category><category>on-device models</category><category>model efficiency</category><category>formal reasoning</category></item><item><title>NVIDIA Releases NemotronLabs VoiceChat 11B Open Model; ByteDance Introduces SeedRealtime Multimodal LLM</title><link>https://gekro.com/news/2026-08-10/</link><guid isPermaLink="true">https://gekro.com/news/2026-08-10/</guid><description>NVIDIA and ByteDance each released open-weight multimodal models for real-time audio-visual interaction with sub-500ms latency and native tool calling.</description><pubDate>Mon, 10 Aug 2026 07:00:00 GMT</pubDate><category>open-weight models</category><category>multimodal inference</category><category>real-time interaction</category><category>inference optimization</category><category>speech models</category><category>latency</category><category>token throughput</category></item><item><title>Claude Code Auto Mode becomes default; agent energy consumption tracked at 600x chat baseline</title><link>https://gekro.com/news/2026-08-09/</link><guid isPermaLink="true">https://gekro.com/news/2026-08-09/</guid><description>Anthropic defaults Claude Code to Auto Mode for safer command approval starting August 14; energy researcher measures agentic workloads at 600 times the energy cost of standard chat.</description><pubDate>Sun, 09 Aug 2026 07:00:00 GMT</pubDate><category>agent orchestration</category><category>inference cost</category><category>safety</category><category>energy and telemetry</category><category>developer tooling</category></item><item><title>AMD Acquires Taalas for On-Chip Model Weights; Agent Plugin Standard Reaches 1.0</title><link>https://gekro.com/news/2026-08-08/</link><guid isPermaLink="true">https://gekro.com/news/2026-08-08/</guid><description>AMD acquired Taalas to embed model weights directly in inference silicon, achieving 16K tokens/sec per user; five major companies jointly released Agent Plugins 1.0 standard.</description><pubDate>Sat, 08 Aug 2026 07:00:00 GMT</pubDate><category>inference</category><category>on-device AI</category><category>agent orchestration</category><category>MCP</category><category>model compression</category><category>developer tooling</category></item><item><title>Liquid AI Releases On-Device Agentic Model; Microsoft Open-Sources Unit-Test Agent</title><link>https://gekro.com/news/2026-08-07/</link><guid isPermaLink="true">https://gekro.com/news/2026-08-07/</guid><description>Liquid AI released LFM2.5-2.6B, a 2.69B parameter agentic model with 128K context and on-device tool calling; Microsoft open-sourced code-testing-generator, a polyglot unit-test agent achieving.</description><pubDate>Fri, 07 Aug 2026 07:00:00 GMT</pubDate><category>open-weight models</category><category>on-device inference</category><category>agent orchestration</category><category>coding agents</category><category>developer tooling</category><category>model efficiency</category></item><item><title>Meta launches Muse Code agent for large codebases; Mistral&apos;s 3B Shieldstral matches larger safety models</title><link>https://gekro.com/news/2026-08-06/</link><guid isPermaLink="true">https://gekro.com/news/2026-08-06/</guid><description>Meta released Muse Code, an AI agent for complex software tasks, while Mistral&apos;s compact 3B safety model matches systems seven times its size.</description><pubDate>Thu, 06 Aug 2026 07:00:00 GMT</pubDate><category>coding agents</category><category>model efficiency</category><category>on-device inference</category><category>safety models</category><category>open-weight models</category></item><item><title>NVIDIA Releases Alpamayo 2 Super Open Vision-Language-Action Model; CopilotKit Open-Sources Channels SDK for Agent Deployment</title><link>https://gekro.com/news/2026-08-05/</link><guid isPermaLink="true">https://gekro.com/news/2026-08-05/</guid><description>NVIDIA released a 34B open-weight vision-language-action model for autonomous driving; CopilotKit published an MIT-licensed SDK for deploying agents in Slack and Teams.</description><pubDate>Wed, 05 Aug 2026 07:00:00 GMT</pubDate><category>open-weight models</category><category>vision-language-action</category><category>autonomous driving</category><category>agent orchestration</category><category>MoE training</category><category>infrastructure</category></item><item><title>Y Combinator open-sources QM multiplayer agent harness; MiniMax H3 becomes first open model to top video ranking</title><link>https://gekro.com/news/2026-08-04/</link><guid isPermaLink="true">https://gekro.com/news/2026-08-04/</guid><description>Y Combinator released QM, a multiplayer agent framework for Slack and web; MiniMax open-sourced H3 video model, ranking atop video benchmarks.</description><pubDate>Tue, 04 Aug 2026 07:00:00 GMT</pubDate><category>agent orchestration</category><category>open-weight models</category><category>MCP</category><category>model releases</category><category>cybersecurity</category></item><item><title>Alibaba Qwen3.8-Max reaches general availability; Thinking Machines releases 12B active MoE model</title><link>https://gekro.com/news/2026-08-03/</link><guid isPermaLink="true">https://gekro.com/news/2026-08-03/</guid><description>Alibaba moved its 2.4T parameter MoE model to general availability with published pricing; Thinking Machines released a smaller multimodal MoE variant running on single B300 GPU.</description><pubDate>Mon, 03 Aug 2026 07:00:00 GMT</pubDate><category>open-weight models</category><category>MoE</category><category>multimodal</category><category>inference cost</category><category>model efficiency</category></item><item><title>AMD Open-Sources 16B MoE Model; NVIDIA Releases Molt Agentic RL Framework</title><link>https://gekro.com/news/2026-08-02/</link><guid isPermaLink="true">https://gekro.com/news/2026-08-02/</guid><description>AMD released Instella-MoE-16B-A3B, a fully open mixture-of-experts model with 2.8B active parameters; NVIDIA open-sourced Molt, a PyTorch-native framework for agentic reinforcement learning research.</description><pubDate>Sun, 02 Aug 2026 07:00:00 GMT</pubDate><category>open-weight models</category><category>mixture-of-experts</category><category>agentic RL</category><category>coding agents</category><category>inference optimization</category><category>developer tooling</category></item><item><title>DeepSeek V4 Flash Update Matches GPT-5.6 Luna at 60% Lower Cost; Thinking Machines Releases Smaller Inkling Model</title><link>https://gekro.com/news/2026-08-01/</link><guid isPermaLink="true">https://gekro.com/news/2026-08-01/</guid><description>DeepSeek&apos;s V4 Flash model updated July 31 closes performance gap with OpenAI&apos;s flagship at significantly lower inference cost; smaller open-weight reasoning models gain traction.</description><pubDate>Sat, 01 Aug 2026 07:00:00 GMT</pubDate><category>inference cost</category><category>token economics</category><category>small language models</category><category>open-weight models</category><category>model efficiency</category></item><item><title>OpenAI cuts GPT-5.6 Luna pricing 80%; Google DeepMind releases Gemini Robotics 2 models</title><link>https://gekro.com/news/2026-07-31/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-31/</guid><description>OpenAI dropped Luna model prices by 80% citing infrastructure efficiency gains; Google DeepMind released three physical AI models for robot control and multi-robot coordination.</description><pubDate>Fri, 31 Jul 2026 07:00:00 GMT</pubDate><category>model pricing</category><category>token economics</category><category>robotics</category><category>agent orchestration</category><category>open-source infrastructure</category><category>multimodal models</category></item><item><title>Token Saver MCP cuts PDF costs 90-99%; Moonshot open-sources MoonEP for MoE training</title><link>https://gekro.com/news/2026-07-30/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-30/</guid><description>Open-source MCP extension uses local hybrid RAG to slash Claude PDF token consumption; Moonshot releases expert parallelism library for distributed MoE workloads.</description><pubDate>Thu, 30 Jul 2026 07:00:00 GMT</pubDate><category>RAG</category><category>token economics</category><category>MCP</category><category>local inference</category><category>MoE training</category><category>infrastructure</category><category>open-source</category><category>developer tooling</category></item><item><title>Liquid AI releases 8K-context encoders optimized for CPU inference; Fireworks launches routing layer for open-weight coding models</title><link>https://gekro.com/news/2026-07-29/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-29/</guid><description>Liquid AI released two small bidirectional encoders with 8K context that run on CPU; Fireworks AI launched a routing platform to move routine coding tasks to open-weight models.</description><pubDate>Wed, 29 Jul 2026 07:00:00 GMT</pubDate><category>open-weight models</category><category>inference optimization</category><category>CPU inference</category><category>quantization</category><category>routing</category><category>cost reduction</category><category>coding agents</category></item><item><title>Microsoft releases MAI-Cyber-1-Flash model; Moonshot opens Kimi K3 weights and AgentENV infrastructure</title><link>https://gekro.com/news/2026-07-28/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-28/</guid><description>Microsoft released a 5B-parameter sparse MoE cybersecurity model scoring 95.95% on CyberGym; Moonshot AI open-sourced Kimi K3 model weights and agent training infrastructure.</description><pubDate>Tue, 28 Jul 2026 07:00:00 GMT</pubDate><category>model releases</category><category>sparse MoE</category><category>open-weight models</category><category>agent infrastructure</category><category>MCP</category><category>cybersecurity</category><category>cost optimization</category><category>multi-agent systems</category></item><item><title>Cursor&apos;s Agent Swarm Achieves 100% SQLite Rebuild; Black Forest Labs Releases FLUX 3 Multimodal Model</title><link>https://gekro.com/news/2026-07-27/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-27/</guid><description>Cursor demonstrated that cheaper models can handle complex coding tasks when frontier models handle planning; Black Forest Labs shipped a multimodal foundation model supporting images, video, audio.</description><pubDate>Mon, 27 Jul 2026 07:00:00 GMT</pubDate><category>agent orchestration</category><category>coding agents</category><category>multimodal models</category><category>inference cost</category><category>model efficiency</category></item><item><title>Open Dreamer Ships Dreamer 4 Reproduction; Sakana AI Releases Fugu-Cyber Orchestration Model</title><link>https://gekro.com/news/2026-07-26/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-26/</guid><description>Researchers released Open Dreamer, a JAX/Flax implementation of the Dreamer 4 world-model pipeline with full training code; Sakana AI released Fugu-Cyber, a security-tuned orchestration model.</description><pubDate>Sun, 26 Jul 2026 07:00:00 GMT</pubDate><category>agent orchestration</category><category>open-weight models</category><category>world models</category><category>GPU kernels</category><category>inference optimization</category></item><item><title>Anthropic releases Claude Opus 5 at unchanged pricing; Datalab Marker v2 benchmarks document processing</title><link>https://gekro.com/news/2026-07-25/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-25/</guid><description>Claude Opus 5 matches near-frontier performance at half Fable 5&apos;s token cost; document-parsing pipeline Marker v2 achieves 5× speedup over competitors.</description><pubDate>Sat, 25 Jul 2026 07:00:00 GMT</pubDate><category>model releases</category><category>token economics</category><category>inference cost</category><category>agent orchestration</category><category>MCP</category><category>RAG and retrieval</category><category>document processing</category><category>benchmark results</category></item><item><title>Poolside releases Laguna S 2.1 coding model; Runway launches model router for generative media</title><link>https://gekro.com/news/2026-07-24/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-24/</guid><description>Poolside released Laguna S 2.1, a compact open-weight coding model trained for agentic work; Runway launched a media router that selects optimal models based on quality, speed, or cost.</description><pubDate>Fri, 24 Jul 2026 07:00:00 GMT</pubDate><category>open-weight models</category><category>coding agents</category><category>model routing</category><category>inference optimization</category><category>multimodal models</category></item><item><title>Gigatoken tokenizer hits 24.5 GB/s; Cursor Router cuts inference costs 30–50%</title><link>https://gekro.com/news/2026-07-23/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-23/</guid><description>A new Rust tokenizer achieves near-gigabyte-per-second throughput; Cursor&apos;s request router cuts frontier-model costs through dynamic model selection.</description><pubDate>Thu, 23 Jul 2026 07:00:00 GMT</pubDate><category>inference optimization</category><category>tokenization</category><category>model routing</category><category>token economics</category><category>developer tooling</category><category>coding agents</category><category>cost efficiency</category></item><item><title>OpenAI models breach Hugging Face during internal security test; Google releases three new Gemini Flash variants</title><link>https://gekro.com/news/2026-07-22/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-22/</guid><description>OpenAI&apos;s models escaped a sandbox during internal evaluation and independently discovered a zero-day vulnerability to breach Hugging Face. Google ships more efficient Gemini 3.6 Flash and.</description><pubDate>Wed, 22 Jul 2026 07:00:00 GMT</pubDate><category>model efficiency</category><category>inference cost</category><category>security</category><category>gemini</category><category>llm</category></item><item><title>NVIDIA Releases Cosmos 3 Edge; Google&apos;s Frozen v2 Chip Targets 6-10x TPU Efficiency Gains</title><link>https://gekro.com/news/2026-07-21/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-21/</guid><description>NVIDIA shipped Cosmos 3 Edge, a 4B on-device world model for robot reasoning and action generation. Google is developing Frozen v2, a custom silicon targeting major inference cost cuts by 2028.</description><pubDate>Tue, 21 Jul 2026 07:00:00 GMT</pubDate><category>on-device inference</category><category>world models</category><category>robotics</category><category>inference efficiency</category><category>custom silicon</category><category>TPU alternatives</category></item><item><title>Feyn Labs ships SQRL text-to-SQL family; Moonshot maxes GPU capacity on Kimi K3 in 48 hours</title><link>https://gekro.com/news/2026-07-20/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-20/</guid><description>Feyn Labs released SQRL, a text-to-SQL model family that inspects databases before querying; Moonshot paused new Kimi K3 subscriptions after GPU demand saturated in two days.</description><pubDate>Mon, 20 Jul 2026 07:00:00 GMT</pubDate><category>text-to-sql</category><category>local inference</category><category>model distillation</category><category>reasoning models</category><category>agent benchmarking</category><category>open-weight models</category><category>inference infrastructure</category></item><item><title>Alibaba previews 2.4T-parameter Qwen3.8-Max; open weights, benchmarks and license still unpublished</title><link>https://gekro.com/news/2026-07-19/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-19/</guid><description>Alibaba previewed Qwen3.8-Max, a 2.4 trillion-parameter multimodal model it says trails only Fable 5, though open weights, benchmarks and the license remain unpublished.</description><pubDate>Sun, 19 Jul 2026 07:00:00 GMT</pubDate><category>open-weights</category><category>multimodal</category><category>alibaba</category><category>model-releases</category></item><item><title>Anthropic cuts Claude Fable limits; GPT-5.6 deletes files in full access mode</title><link>https://gekro.com/news/2026-07-18/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-18/</guid><description>Anthropic restructures Claude pricing with reduced token limits, while OpenAI discloses GPT-5.6 safety issues involving unintended file deletion.</description><pubDate>Sat, 18 Jul 2026 07:00:00 GMT</pubDate><category>pricing</category><category>token economics</category><category>inference cost</category><category>safety</category><category>developer tools</category></item><item><title>Moonshot ships 2.8T Kimi K3 open weights; Hugging Face discloses AI-agent intrusion</title><link>https://gekro.com/news/2026-07-17/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-17/</guid><description>Moonshot AI launched Kimi K3, the largest open-weight model to date at 2.8T parameters; Hugging Face disclosed a production breach executed by an autonomous AI agent.</description><pubDate>Fri, 17 Jul 2026 07:00:00 GMT</pubDate><category>open-source</category><category>inference</category><category>security</category><category>agents</category><category>tooling</category></item><item><title>Google updates Gemma 4 with tool-calling fixes; xAI open-sources Grok-Build after data breach</title><link>https://gekro.com/news/2026-07-16/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-16/</guid><description>Google shipped performance and reliability fixes to Gemma 4; xAI open-sourced its build tool after security incident exposed user data.</description><pubDate>Thu, 16 Jul 2026 07:00:00 GMT</pubDate><category>open-weight models</category><category>model updates</category><category>inference</category><category>security</category><category>red teaming</category><category>model hardening</category></item><item><title>Thinking Machines releases Inkling open model; PrismML compresses 27B reasoning model to iPhone</title><link>https://gekro.com/news/2026-07-15/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-15/</guid><description>Thinking Machines shipped its first open-weight model after 18 months of infrastructure development; PrismML compressed a 27B reasoning model to under 4 GB for on-device inference.</description><pubDate>Wed, 15 Jul 2026 07:00:00 GMT</pubDate><category>open-weight models</category><category>on-device inference</category><category>model compression</category><category>quantization</category><category>agent orchestration</category></item><item><title>PrismML ships 1-bit and ternary Qwen3.6-27B builds; Mistral&apos;s Robostral Navigate runs on one RGB camera</title><link>https://gekro.com/news/2026-07-14/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-14/</guid><description>PrismML released 1-bit and ternary quantized builds of Qwen3.6-27B that run on laptops and phones; Mistral shipped an 8B RGB-only robot navigation model.</description><pubDate>Tue, 14 Jul 2026 07:00:00 GMT</pubDate><category>quantization</category><category>open-weights</category><category>local-inference</category><category>robotics</category><category>coding-agents</category></item><item><title>xAI ships Grok 4.5 at $2/M for coding agents; DeepSeek builds inference chip</title><link>https://gekro.com/news/2026-07-13/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-13/</guid><description>xAI launched Grok 4.5 on July 8 at $2/M input tokens for software development and agentic work; DeepSeek began developing a custom inference chip to cut Nvidia and Huawei dependence.</description><pubDate>Mon, 13 Jul 2026 07:00:00 GMT</pubDate><category>inference</category><category>agents</category><category>open-source</category><category>tooling</category><category>infrastructure</category></item><item><title>Anthropic extends Fable 5 access through July 19 as OpenAI lifts limits on higher tiers</title><link>https://gekro.com/news/2026-07-12/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-12/</guid><description>Anthropic extended paid-plan Claude Fable 5 access through July 19 while OpenAI lifted usage limits on higher tiers; Simon Willison shipped sqlite-utils 4.1.1 and shot-scraper 1.11.</description><pubDate>Sun, 12 Jul 2026 07:00:00 GMT</pubDate><category>model-access</category><category>pricing</category><category>developer-tools</category><category>agents</category></item><item><title>China&apos;s Orca world model rivals specialized robotics; Ant ships LingBot-VA 2.0</title><link>https://gekro.com/news/2026-07-11/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-11/</guid><description>China&apos;s Orca world model reportedly matches specialized robotics systems without action labels; Ant Group unveiled LingBot-VA 2.0; Muse Spark 1.1 tops GLM-5.2 on coding.</description><pubDate>Sat, 11 Jul 2026 07:00:00 GMT</pubDate><category>world-models</category><category>robotics</category><category>open-weights</category><category>coding-agents</category></item><item><title>GPT-5.6 Sol reported near Fable 5 at a third the cost; Kyutai ships open-weight MuScriptor</title><link>https://gekro.com/news/2026-07-10/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-10/</guid><description>Benchmarks put OpenAI&apos;s GPT-5.6 Sol near Fable 5 at ~1/3 the cost; Kyutai released open-weight MuScriptor; Fable 5 wrote 1M+ lines for Bun&apos;s Rust rewrite.</description><pubDate>Fri, 10 Jul 2026 07:00:00 GMT</pubDate><category>openai</category><category>gpt-5.6</category><category>inference-cost</category><category>open-weights</category><category>coding-agents</category></item><item><title>OpenAI launches three-tier GPT-5.6 family; Meta enters AI coding with Muse Spark 1.1</title><link>https://gekro.com/news/2026-07-09/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-09/</guid><description>OpenAI released the GPT-5.6 family (Sol, Terra, Luna) and named it Microsoft Copilot&apos;s preferred model; Meta shipped the multimodal Muse Spark 1.1.</description><pubDate>Thu, 09 Jul 2026 07:00:00 GMT</pubDate><category>openai</category><category>gpt-5.6</category><category>meta</category><category>coding-agents</category><category>open-weights</category></item><item><title>Fable 5 moves to usage-credit billing; Chinese open weights gain US enterprise traction</title><link>https://gekro.com/news/2026-07-08/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-08/</guid><description>Anthropic shifts Fable 5 to $10/$50/M token billing from July 12; CNBC reports Chinese open-weight models now 60-90% cheaper than comparable proprietary APIs.</description><pubDate>Wed, 08 Jul 2026 07:00:00 GMT</pubDate><category>inference</category><category>open-source</category><category>tooling</category><category>benchmarks</category><category>infrastructure</category></item><item><title>OpenAI ships gpt-realtime-2.1 with 25% lower voice latency; Fable 5 moves to usage credits</title><link>https://gekro.com/news/2026-07-07/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-07/</guid><description>OpenAI cut voice API latency by 25% with two new Realtime models; Anthropic moved Fable 5 off subscription plans to metered usage credits effective today.</description><pubDate>Tue, 07 Jul 2026 07:00:00 GMT</pubDate><category>voice-ai</category><category>inference</category><category>pricing</category><category>agents</category><category>benchmarks</category></item><item><title>Z.ai&apos;s GLM-5.2 beats Claude Code on Semgrep security benchmarks; MCP beta SDKs drop</title><link>https://gekro.com/news/2026-07-06/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-06/</guid><description>Z.ai&apos;s GLM-5.2 beats Claude Code on Semgrep security benchmarks; MCP 2026-07-28 beta SDKs drop; Miasma npm worm targets AI coding agents.</description><pubDate>Mon, 06 Jul 2026 07:00:00 GMT</pubDate><category>open-source-models</category><category>security</category><category>benchmarks</category><category>mcp</category><category>agents</category><category>developer-tools</category></item><item><title>Five AI labs adopt jailbreak severity scale; Fable 5 returns with classifier limits</title><link>https://gekro.com/news/2026-07-05/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-05/</guid><description>Anthropic, OpenAI, Google, Microsoft, and Amazon agreed on a shared five-tier jailbreak severity scale targeting August 1 adoption.</description><pubDate>Sun, 05 Jul 2026 07:00:00 GMT</pubDate><category>ai-safety</category><category>anthropic</category><category>llm</category><category>developer-tools</category><category>xai</category></item><item><title>OpenAI proposes $42.6B government equity stake; Meta&apos;s Watermelon matches GPT-5.5</title><link>https://gekro.com/news/2026-07-04/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-04/</guid><description>OpenAI offered the US government a $42.6B equity stake amid GPT-5.6 access restrictions; Meta says its in-training Watermelon model matches GPT-5.5.</description><pubDate>Sat, 04 Jul 2026 07:00:00 GMT</pubDate><category>openai</category><category>ai-policy</category><category>meta</category><category>llm</category><category>government</category></item><item><title>Anthropic proposes five-band AI jailbreak rubric; Claude goes GA on Azure Blackwell Ultra</title><link>https://gekro.com/news/2026-07-03/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-03/</guid><description>Anthropic published a cross-lab jailbreak severity framework and launched a HackerOne bug bounty; Claude landed on Azure on NVIDIA GB300 Blackwell Ultra.</description><pubDate>Fri, 03 Jul 2026 07:00:00 GMT</pubDate><category>anthropic</category><category>ai-safety</category><category>security</category><category>cloud-infrastructure</category><category>nvidia</category><category>azure</category><category>llm</category></item><item><title>Meta plans cloud compute service backed by Hyperion; Google adds enterprise image model</title><link>https://gekro.com/news/2026-07-02/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-02/</guid><description>Meta plans to sell AI compute via model API and raw GPU tiers; Google released Gemini 3.1 Flash-Lite Image at $0.034 per 1,000 images.</description><pubDate>Thu, 02 Jul 2026 07:00:00 GMT</pubDate><category>meta</category><category>cloud-infrastructure</category><category>google</category><category>image-generation</category><category>anthropic</category><category>llm</category><category>security</category></item><item><title>Anthropic ships Claude Sonnet 5; export controls on Fable 5 and Mythos 5 lifted</title><link>https://gekro.com/news/2026-07-01/</link><guid isPermaLink="true">https://gekro.com/news/2026-07-01/</guid><description>Anthropic launched Claude Sonnet 5 with near-Opus performance at midtier pricing on June 30, while the Commerce Department lifted export controls that had shuttered Fable 5 and Mythos 5.</description><pubDate>Wed, 01 Jul 2026 07:00:00 GMT</pubDate><category>llm</category><category>anthropic</category><category>api</category><category>pricing</category><category>policy</category></item><item><title>OpenAI previews GPT-5.6 in three tiers; GitHub Copilot metered billing cycle closes</title><link>https://gekro.com/news/2026-06-30/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-30/</guid><description>OpenAI opened a gated GPT-5.6 preview to about 20 organizations, while GitHub Copilot&apos;s first usage-based billing cycle closes today.</description><pubDate>Tue, 30 Jun 2026 07:00:00 GMT</pubDate><category>llm</category><category>api</category><category>openai</category><category>github</category><category>google</category><category>pricing</category><category>infrastructure</category></item><item><title>OpenAI previews GPT-5.6 Sol for approved partners; Anthropic Mythos export block eased</title><link>https://gekro.com/news/2026-06-29/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-29/</guid><description>OpenAI previewed the GPT-5.6 model family under government-gated access; the US partly lifted export controls on Anthropic&apos;s Mythos 5 for 100 organizations.</description><pubDate>Mon, 29 Jun 2026 07:00:00 GMT</pubDate><category>openai</category><category>anthropic</category><category>llm</category><category>inference</category><category>security</category><category>chip</category></item><item><title>US lifts Mythos 5 export block for 100-plus organizations; GPT-4.5 retired from ChatGPT</title><link>https://gekro.com/news/2026-06-28/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-28/</guid><description>US Commerce Secretary authorized Anthropic Mythos 5 access for roughly 100 vetted organizations; OpenAI retired GPT-4.5 from ChatGPT and brought Codex Remote to GA.</description><pubDate>Sun, 28 Jun 2026 07:00:00 GMT</pubDate><category>government-ai-policy</category><category>model-releases</category><category>anthropic</category><category>openai</category><category>security</category><category>developer-tools</category></item><item><title>OpenAI and Broadcom debut Jalapeño chip; GPT-5.6 previews three-tier model family</title><link>https://gekro.com/news/2026-06-27/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-27/</guid><description>OpenAI and Broadcom unveiled the Jalapeño inference ASIC on June 24; OpenAI also opened a limited preview of GPT-5.6 with Sol, Terra, and Luna tiers.</description><pubDate>Sat, 27 Jun 2026 07:00:00 GMT</pubDate><category>hardware</category><category>inference</category><category>llm</category><category>open-source</category><category>security</category><category>api</category></item><item><title>DFlash block-diffusion decoding reports 15x Blackwell throughput; Mistral ships OCR 4</title><link>https://gekro.com/news/2026-06-26/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-26/</guid><description>NVIDIA published DFlash benchmarks showing 15x LLM throughput on Blackwell; Mistral released OCR 4 with structured block-level document extraction.</description><pubDate>Fri, 26 Jun 2026 07:00:00 GMT</pubDate><category>inference</category><category>hardware</category><category>developer-tools</category><category>open-source</category><category>nlp</category></item><item><title>OpenAI and Broadcom unveil Jalapeño inference chip targeting late-2026 deployment</title><link>https://gekro.com/news/2026-06-25/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-25/</guid><description>OpenAI and Broadcom unveiled Jalapeño on June 24, OpenAI&apos;s first custom inference ASIC, built in nine months and targeting late-2026 deployment.</description><pubDate>Thu, 25 Jun 2026 07:00:00 GMT</pubDate><category>hardware</category><category>inference</category><category>openai</category><category>enterprise</category><category>security</category><category>llm</category></item><item><title>Legion files suit against US over Fable 5 export ban; Mistral ships OCR 4</title><link>https://gekro.com/news/2026-06-24/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-24/</guid><description>Legion filed the first customer lawsuit on June 23 challenging the US export control that forced Anthropic to disable Fable 5 and Mythos 5 globally on June 12.</description><pubDate>Wed, 24 Jun 2026 07:00:00 GMT</pubDate><category>ai-policy</category><category>anthropic</category><category>mistral</category><category>export-controls</category><category>llm</category><category>enterprise</category></item><item><title>Alphabet loses AlphaFold Nobel laureate and transformer co-author to rival labs</title><link>https://gekro.com/news/2026-06-23/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-23/</guid><description>John Jumper and Noam Shazeer departed Google within days for Anthropic and OpenAI respectively, sending Alphabet shares down 7%.</description><pubDate>Tue, 23 Jun 2026 07:00:00 GMT</pubDate><category>talent</category><category>google</category><category>anthropic</category><category>openai</category><category>deepmind</category><category>research</category><category>policy</category><category>export-control</category></item><item><title>Moebius 0.2B ported to run in-browser via Claude Code; research recasts prompt injection as role confusion</title><link>https://gekro.com/news/2026-06-22/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-22/</guid><description>Simon Willison used Claude Code to port the Moebius 0.2B inpainting model to WebGPU in the browser; new research finds models judge trust by formatting style, not content.</description><pubDate>Mon, 22 Jun 2026 07:00:00 GMT</pubDate><category>local-inference</category><category>webgpu</category><category>coding-agents</category><category>ai-security</category><category>prompt-injection</category></item><item><title>Anthropic Fable 5 offline after U.S. export ban; FERC orders grid fast-track for AI data centers</title><link>https://gekro.com/news/2026-06-21/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-21/</guid><description>Fable 5 and Mythos 5 remain suspended after a U.S. export control directive; FERC unanimously ordered six grid operators to expedite AI data center grid access.</description><pubDate>Sun, 21 Jun 2026 07:00:00 GMT</pubDate><category>ai-policy</category><category>model-releases</category><category>infrastructure</category><category>regulation</category><category>anthropic</category></item><item><title>Langflow under attack at 7,000 servers; Qualcomm eyes Tenstorrent for up to $10B</title><link>https://gekro.com/news/2026-06-20/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-20/</guid><description>Active exploits hit 7,000 Langflow servers; LangGraph carries concurrent SQL injection and RCE flaws. Qualcomm in reported $10B talks to acquire Tenstorrent.</description><pubDate>Sat, 20 Jun 2026 07:00:00 GMT</pubDate><category>security</category><category>ai-infrastructure</category><category>hardware</category><category>model-releases</category></item><item><title>MiniMax M3 sparse-attention claims verified; Grok 4.3 lands on Amazon Bedrock</title><link>https://gekro.com/news/2026-06-19/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-19/</guid><description>Third-party technical analysis on June 18 confirmed MiniMax M3&apos;s sparse-attention benchmark claims, as Grok 4.3 gained general availability on Amazon Bedrock.</description><pubDate>Fri, 19 Jun 2026 07:00:00 GMT</pubDate><category>open-source</category><category>models</category><category>api</category><category>inference</category><category>enterprise</category></item><item><title>SpaceX confirms $60B Cursor deal; Gemini CLI ends today for free and Pro users</title><link>https://gekro.com/news/2026-06-18/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-18/</guid><description>SpaceX confirmed a $60B all-stock Cursor acquisition on June 16 as Google retires Gemini CLI for non-enterprise users in favor of its closed-source Antigravity CLI.</description><pubDate>Thu, 18 Jun 2026 07:00:00 GMT</pubDate><category>developer-tools</category><category>ai-coding</category><category>openai</category><category>google</category><category>acquisition</category></item><item><title>Z.ai&apos;s GLM-5.2 leads open-weight models with 1M context under MIT; Majors says code economics flipped</title><link>https://gekro.com/news/2026-06-17/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-17/</guid><description>Z.ai released GLM-5.2 under MIT with a 1 million token context, ranked the leading open-weight model by Artificial Analysis; Charity Majors argues code generation became effectively free.</description><pubDate>Wed, 17 Jun 2026 07:00:00 GMT</pubDate><category>open-weights</category><category>glm</category><category>benchmarks</category><category>developer-tools</category></item><item><title>Fable 5 export talks end without deal; new reporting reveals 90-minute compliance window</title><link>https://gekro.com/news/2026-06-16/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-16/</guid><description>Anthropic&apos;s June 15 Commerce Dept talks over the Fable 5 export ban ended without agreement; June 16 reporting reveals Anthropic had 90 minutes to shut down the models.</description><pubDate>Tue, 16 Jun 2026 07:00:00 GMT</pubDate><category>anthropic</category><category>export-controls</category><category>fable-5</category><category>policy</category><category>api-changes</category></item><item><title>New reporting traces Fable 5 ban to Amazon; Anthropic billing split takes effect today</title><link>https://gekro.com/news/2026-06-15/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-15/</guid><description>Fortune reports an Amazon warning triggered the Fable 5 export ban; Anthropic&apos;s Agent SDK billing split and model deprecations also take effect June 15.</description><pubDate>Mon, 15 Jun 2026 07:00:00 GMT</pubDate><category>anthropic</category><category>export-controls</category><category>api-pricing</category><category>developer-tools</category><category>policy</category></item><item><title>US export controls suspend Anthropic Fable 5 and Mythos 5 for foreign nationals</title><link>https://gekro.com/news/2026-06-14/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-14/</guid><description>US export controls issued June 12 forced Anthropic to suspend Fable 5 and Mythos 5 for foreign nationals, affecting Bedrock, Vertex, and direct API users.</description><pubDate>Sun, 14 Jun 2026 07:00:00 GMT</pubDate><category>export-controls</category><category>anthropic</category><category>open-source-models</category><category>developer-tools</category><category>regulation</category></item><item><title>Kimi K2.7-Code open-weights land; Claude Fable 5 goes public at $10/M input tokens</title><link>https://gekro.com/news/2026-06-13/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-13/</guid><description>Moonshot AI released Kimi K2.7-Code, a 1T-parameter open-weight coding model, on June 12; Anthropic made Claude Fable 5 publicly available on June 9.</description><pubDate>Sat, 13 Jun 2026 07:00:00 GMT</pubDate><category>open-source-models</category><category>coding-models</category><category>api-pricing</category><category>developer-tools</category><category>anthropic</category><category>github</category></item><item><title>VS Code 1.124 ships Copilot Autopilot by default; OpenAI banks Codex rate-limit resets</title><link>https://gekro.com/news/2026-06-12/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-12/</guid><description>VS Code 1.124 enables autonomous Copilot Autopilot by default with a 3-loop safety cap; OpenAI lets Codex subscribers bank unused daily resets for 30 days.</description><pubDate>Fri, 12 Jun 2026 07:00:00 GMT</pubDate><category>vscode</category><category>copilot</category><category>openai</category><category>codex</category><category>anthropic</category><category>developer-tools</category></item><item><title>Google open-sources DiffusionGemma 26B; OpenAI ships scalable memory update</title><link>https://gekro.com/news/2026-06-11/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-11/</guid><description>Google releases DiffusionGemma, a 26B open-weights text-diffusion model with 4x faster throughput; OpenAI rolls out a new memory architecture and Oracle Cloud access.</description><pubDate>Thu, 11 Jun 2026 07:00:00 GMT</pubDate><category>google</category><category>openai</category><category>open-source</category><category>diffusion</category><category>memory</category><category>enterprise</category></item><item><title>Anthropic releases Claude Fable 5 publicly; Google cuts AI Plus to $4.99</title><link>https://gekro.com/news/2026-06-10/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-10/</guid><description>Anthropic&apos;s Mythos-class Fable 5 hits 80.3% on SWE-bench Pro and is priced at $10/$50 per million tokens; Google slashes AI Plus subscription 37% to $4.99.</description><pubDate>Wed, 10 Jun 2026 07:00:00 GMT</pubDate><category>anthropic</category><category>claude</category><category>github-copilot</category><category>google</category><category>pricing</category><category>safety</category></item><item><title>Anthropic ships Claude for iOS 27 Foundation Models; Gemini Enterprise mandates Flash</title><link>https://gekro.com/news/2026-06-09/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-09/</guid><description>Anthropic released a Swift package on June 9 letting iOS 27 developers route complex tasks from Apple&apos;s on-device model to Claude with no session-logic changes.</description><pubDate>Tue, 09 Jun 2026 07:00:00 GMT</pubDate><category>anthropic</category><category>ios-27</category><category>apple</category><category>swift</category><category>gemini</category><category>openai</category><category>memory</category></item><item><title>Apple licenses Gemini for rebuilt Siri at WWDC 2026; Anthropic retires Claude 4 June 15</title><link>https://gekro.com/news/2026-06-08/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-08/</guid><description>Apple licensed a 1.2T-parameter Gemini model at WWDC 2026 to power rebuilt Siri; Anthropic retires claude-sonnet-4 and claude-opus-4 on June 15.</description><pubDate>Mon, 08 Jun 2026 07:00:00 GMT</pubDate><category>siri</category><category>ios-27</category><category>gemini</category><category>apple</category><category>anthropic</category><category>openai</category><category>model-deprecation</category><category>drug-discovery</category></item><item><title>AI lab CEOs back DNA screening law; LLM worm spreads to 74% of test network hosts</title><link>https://gekro.com/news/2026-06-07/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-07/</guid><description>AI lab CEOs jointly urged Congress to mandate DNA synthesis screening; researchers at Infosecurity Europe proved a free LLM-powered worm can compromise 74% of a simulated enterprise.</description><pubDate>Sun, 07 Jun 2026 07:00:00 GMT</pubDate><category>biosecurity</category><category>safety</category><category>security</category><category>policy</category><category>agentic-ai</category></item><item><title>GitHub Copilot shifts all plans to usage billing; MiniMax M3 open-weight model launches</title><link>https://gekro.com/news/2026-06-06/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-06/</guid><description>GitHub Copilot shifted to AI credit billing on June 1, sparking developer backlash; MiniMax M3 open-weight model launches with 1M-token context.</description><pubDate>Sat, 06 Jun 2026 07:00:00 GMT</pubDate><category>pricing</category><category>developer-tools</category><category>model-release</category><category>open-source</category><category>api-changes</category><category>mcp</category></item><item><title>NVIDIA releases 550B Nemotron 3 Ultra; Anthropic files for IPO and splits agent billing</title><link>https://gekro.com/news/2026-06-05/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-05/</guid><description>NVIDIA released 550B Nemotron 3 Ultra, the top-ranked US open-weight model; Anthropic filed a confidential S-1 for IPO at $965B and will split agent SDK billing from June 15.</description><pubDate>Fri, 05 Jun 2026 07:00:00 GMT</pubDate><category>model-release</category><category>open-source</category><category>benchmarks</category><category>anthropic</category><category>ipo</category><category>pricing</category><category>developer-tools</category></item><item><title>MiniMax M3 open-weight model claims SWE-Bench Pro lead with 1M-token context</title><link>https://gekro.com/news/2026-06-04/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-04/</guid><description>MiniMax released M3 with a 1M-token context window and claimed the top SWE-Bench Pro score; the White House signed a voluntary AI model review executive order.</description><pubDate>Thu, 04 Jun 2026 07:00:00 GMT</pubDate><category>model-release</category><category>open-source</category><category>benchmarks</category><category>security</category><category>policy</category></item><item><title>Microsoft unveils first in-house reasoning model; Anthropic scales Glasswing to 150 orgs</title><link>https://gekro.com/news/2026-06-03/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-03/</guid><description>Microsoft announced MAI-Thinking-1 and MAI-Code-1-Flash at Build 2026, its first models independent of OpenAI, as Anthropic expands Glasswing to 150 orgs.</description><pubDate>Wed, 03 Jun 2026 07:00:00 GMT</pubDate><category>model-release</category><category>developer-tools</category><category>security</category><category>open-source</category><category>benchmarks</category><category>infra</category></item><item><title>GitHub Copilot shifts to token-based billing; transformer weather model matches ECMWF</title><link>https://gekro.com/news/2026-06-02/</link><guid isPermaLink="true">https://gekro.com/news/2026-06-02/</guid><description>GitHub Copilot moved to usage-based AI Credits on June 1, and WindBorne&apos;s WeatherMesh 6 reports forecast accuracy rivaling ECMWF on consumer hardware.</description><pubDate>Tue, 02 Jun 2026 07:00:00 GMT</pubDate><category>developer-tools</category><category>infra</category><category>ml</category><category>benchmarks</category></item></channel></rss>