01
Alibaba's Qwen3.8-Max Goes Open-Weight: 2.4T Parameters, Beats GPT-5.6 and Claude Fable 5 on Agentic Benchmarks
Alibaba released Qwen3.8-Max, a 2.4 trillion parameter Mixture-of-Experts model with 95B active parameters and 1M token context, scoring 86.1 on OSWorld-Verified to beat GPT-5.6 Sol Max (83.2) and Claude Fable 5 (85.0). The Artificial Analysis Intelligence Index rates it at 56, level with Claude Opus 4.8 and ahead of all US companies except Anthropic and OpenAI. Open weights are promised within days of launch.
02
Tencent Opens Hy3 Globally Under Apache 2.0: 295B-Parameter MoE Tops OpenRouter Leaderboard
Tencent launched its Hy3 model globally with 295B total / 21B active parameters, available under the commercially permissive Apache 2.0 license on Hugging Face, ModelScope, and OpenRouter. Hy3 ranked #1 on OpenRouter's global LLM usage leaderboard within a week of launch and is free on WorkBuddy until August 31, with API pricing at $0.1288/$0.5336 per million tokens.
03
DeepSeek V4 Flash Becomes World's Cheapest Major AI Model, Beats Its Own Flagship Pro on Agent Benchmarks
DeepSeek's V4 Flash scored 50 on the Artificial Analysis Intelligence Index (a 10-point jump) and beats V4-Pro-Preview on all nine published agent benchmarks, including 82.7 on Terminal Bench 2.1. At $0.14/$0.28 per million input/output tokens, it costs roughly 3 cents per benchmark run, compared to $1.86 for GPT-5.6 Sol and $3.15 for Claude Fable 5. It now processes 8 trillion daily tokens.
04
Prime Intellect's Open-Source Prime Agent Beats Human Experts on ARC-AGI-3
Prime Intellect released Prime Agent, an MIT-licensed self-improving coding harness built on a Recursive Language Model (RLM) that treats context as a variable and subagent delegation as function calls. With Claude Opus 5, it scored 95.5% on ARC-AGI-3, surpassing the human expert baseline of 95.4%. The startup is backed by $150M in total funding.
05
White House AI Security Framework Exempts Open-Weight Models From Vetting
The Trump administration's voluntary AI oversight framework, finalized this week, will only apply to "closed" proprietary models, explicitly excluding open-weight models from pre-release security review. The framework was shared with OpenAI, Anthropic, and other AI labs on Tuesday but remains undisclosed to the public, raising concerns about transparency and the exclusion of open-source developers.
06
Anthropic Confirms In-House Custom Chip Design Team for Claude
Anthropic confirmed it is building an in-house chip design team to co-design custom silicon for its Claude models, with salaries up to $485,000 for senior engineers. The move aims to cut inference costs and reduce dependence on Nvidia, joining Google and OpenAI in the push for AI hardware independence amid persistent chip shortages.