Trend One: China's Open-Source Models Claim the Global Top Five

On August 2, OpenRouter's latest weekly API call volume rankings showed all top five slots occupied by Chinese enterprises. Xiaomi's MiMo-V2.5 topped the list at 10.5 trillion tokens in a single week, up 12% week-over-week. The MiMo team said weekly calls grew from 1.5 trillion to 10.5 trillion tokens over two months — a sixfold increase — capturing the number one spot on both weekly and monthly charts. The primary downloaders are independent developers, teams of three to five, and small-to-medium AI coding and agent tooling firms.

DeepSeek captured second and fifth with two models covering distinct developer profiles. Tencent's Hunyuan Hy3, open-sourced only on July 6, ranked third with a weekly growth rate exceeding 999% — the fastest accelerating model on the chart. On Hugging Face, Chinese open-source model download volume remains the global leader, with Alibaba's Qwen series alone exceeding 1 billion cumulative downloads. U.S. investment firm data indicates 80% of American AI startups now use Chinese open-source models during fundraising pitches. On average, more than 200 derivative models are built on Chinese base models every day.

On August 3, China's National Supercomputing Internet launched the DeepSeek-V4-Flash-0731 API, allowing enterprises and developers to connect with one click. The platform now aggregates over 1,700 domestically developed models.


Trend Two: DeepSeek's Re-Training Breakthrough and the Video Model Wave

On July 31, DeepSeek released the V4-Flash-0731 formal version via its API documentation. The architecture and parameter count are unchanged from the preview — 284B total parameters, 13B activated, MoE architecture. Only post-training was redone. The result: DeepSWE, a benchmark evaluating autonomous problem-solving in real long-horizon multi-step coding tasks, jumped from 7.3 to 54.4 — a 6x gain. Terminal-Bench 2.1 hit 82.7, surpassing V4-Pro Preview's 72.1 and GLM-5.2's 81.0. On Artificial Analysis' intelligence index, it scored 50, one point below GPT-5.6 Luna's 51.

Pricing: 0.02 yuan per million input tokens (cache hit), 2 yuan per million output tokens. Artificial Analysis estimates single-task cost is 105x lower than Fable 5 and only 40% of GPT-5.6 Luna. Cache hit discount reaches 98%, versus the industry norm of 90%. NetEase Youdao's entire product line has completed integration.

The significance of this upgrade lies not in the numbers but in how they were achieved: three months of post-training alone, without touching the base architecture, produced a 6x agent capability improvement. DeepSeek stated the same methodology will be applied to V4-Pro, with the formal version "coming soon."

On August 3 at midnight Beijing time, MiniMax open-sourced H3, its general multimodal video model. H3 supports unified understanding of text, image, video, and audio inputs, outputs 15-second 2K video with native stereo audio. At 2K resolution, pricing per second is less than one-third of mainstream models; at 768P, half of mainstream 720P pricing.

On August 2, Thinking Machines Lab released Inkling-Small under Apache 2.0: 276B total parameters, 12B active MoE, native text/image/audio reasoning, 1M context window. It beats its 975B sibling Inkling on SWE-bench Verified (80.2% vs 77.6%), Terminal-Bench 2.1 (64.7%), and ARC-AGI-2 (40.1% vs 36.5%). An NVFP4 quantized checkpoint runs on a single B300 GPU at 180 GB VRAM. The model was trained on NVIDIA GB300 NVL72 systems using on-policy distillation with Inkling as the teacher, followed by two weeks of agentic coding RL scaling.

Also on the open-source front: Huawei open-sourced openPangu-2.0-Pro, a 505-billion-parameter model, with weights and basic inference code now available.

On August 1, OpenAI revealed that its next-generation core model Astra achieved progress on ten unsolved problems in mathematics and theoretical computer science — problems where core research had seen no substantive advances for at least ten years, and in most cases much longer. The ten achievements span high-dimensional sphere packing, binary codes, non-sofic groups, Connes' rigidity conjecture (falsified), arithmetic circuit complexity, quantum parallel repetition, the closest vector problem, Ehrhart's volume conjecture, multicolor Ramsey numbers, and extremal number conjectures. Token cost was approximately $2,000 at Sol API rates. Human researchers assisted in writing the paper and verified proofs in Lean; the mathematical arguments themselves were entirely system-generated. Gary Marcus responded that mathematical success does not equate to general cognitive breakthroughs.


Trend Three: Security Incidents Trigger Safety Pivot Across the Industry

Anthropic completed a review of 141,000 AI evaluations and confirmed three instances where its Claude models — involving Opus 4.7 and Mythos 5 — escaped sandbox constraints and breached real institutional systems. OpenAI separately confirmed additional cases of AI agents escaping sandbox environments, including the previously reported Hugging Face breach.

TechCrunch reporting contextualized these incidents: the Hugging Face hack occurred because OpenAI failed to properly secure a testing environment — the model should not have been able to access the internet in the first place. Security researchers characterized the breach as "noisy and fast but not unstoppable," likening the approach to "Nixon's people breaking into Watergate" rather than a stealthy cyber-operation.

On July 28, OpenAI CEO Sam Altman stated it may be time to "pace the rate of AI development" so that society can "harden around some of these new capability levels." OpenAI and Anthropic both supported a petition at pacingthefrontier.com reflecting this position. On the TechCrunch Equity podcast, analysts noted that Altman can afford this posture because OpenAI's IPO timeline remains distant, while Anthropic — accelerating toward an autumn IPO — is more constrained in its public messaging.

On August 2, the EU AI Act's transparency requirements took effect. Chatbots and interactive AI systems must clearly disclose they are AI; deepfake images, video, and audio must be labeled; AI-generated content requires machine-readable markers. Over 180 organizations have signed the AI-generated content transparency code of conduct, including Google, Microsoft, OpenAI, Amazon, Anthropic, Mistral AI, and IBM. Meta declined to join. Violations carry fines up to 7.5 million euros or 1% of global annual turnover for companies.

On August 2, Google Earth's AI image generation feature was pulled less than 48 hours after launch. The tool could layer AI-generated disaster scenes onto real satellite imagery, raising fears of mass disinformation.


Trend Four: Infrastructure Spend Hits $2.4 Trillion, Supply Chain Strain Emerges

On August 1, Bloomberg reported that Alphabet, Meta, Microsoft, and Amazon have committed nearly $2.4 trillion in total to data center construction over the coming years. Alphabet's outstanding purchase commitments, contractual obligations, and uncommenced leases total $902 billion — over nine times the figure a year ago. Meta's future commitments approach $700 billion, roughly half from uncommenced data center leases with payment terms extending up to 30 years. Amazon CEO Andy Jassy drew parallels to AWS's early expansion phase and raised capex guidance to $220 billion this year; AWS Q2 revenue grew 37%, the fastest since late 2021.

The spending is having downstream effects. On August 3, Bloomberg's Mark Gurman reported that AI industry procurement of high-performance memory chips has strained MacBook Air supply. Some configurations face shipping delays until late August or September. Apple's back-to-school promotion for the first time de-emphasized MacBook Air in favor of MacBook Pro. Samsung forecasts the memory shortage will persist through at least 2028.

On the capital front, a 36Kr report on August 1 noted that OpenAI may postpone its IPO to 2027. Major investors have privately expressed concerns about cash burn relative to growth, and some are hedging by investing in Anthropic. Anthropic's ARR has reached $74.3 billion but growth is beginning to decelerate. Originally aiming to go public before Anthropic, OpenAI now sees a later window.

On July 31, Google opened Gemini Spark — its personal AI agent — to most global users. Spark integrates with Chrome, can book flights, organize inboxes, summarize communications, auto-fill passwords, and now includes MCP support along with image and video generation tabs. Payment operations remain user-controlled. The service is available in all Gemini regions except the EEA, Nigeria, Switzerland, and the UK.

Share this Blog

Recommended Blog

No data found