๐ง Model & Product Launches
-
Z.ai's GLM-5.3 finds 2,436 real vulnerabilities โ and delays open weights over its own "surprise" cyber skills โ Z.ai / unrot.co / Terminal-Bench
Z.ai released GLM-5.3 on Aug 14, a post-training refresh over the same GLM-5.2 base โ no new base model, all gains from post-training. It reports a 50% jump on its internal Code Bench and leads open models on Terminal-Bench 3.0 (28.3% vs 4.6% for 5.2). Live now via the GLM Coding Plan ($18/mo) and ZCode, but the full API and downloadable weights are held back ~2 weeks for "safety hardening." With outside security teams, the model found 2,436 vulnerabilities across 269 real projects, 1,097 medium-to-high severity โ the hacking skill grew faster than planned during training. Independent testers still place it behind Anthropic's Fable 5 and OpenAI's GPT-5.6 Sol on the hardest coding benchmarks. Weights are pointed to around end of August.
Framing This is the first time a GLM-series release gates open weights behind extra safety review. The story is the double edge: a coding model that found 1,097 medium-to-high-severity bugs across the Linux kernel, WebKit, and FreeBSD is precisely the kind of capability labs now treat as a liability until hardened. -
OpenAI cuts GPT-5.6 Sol pricing over 20% โ promotion runs through Nov 21 โ OpenAI / unrot.co
On Aug 21 OpenAI cut GPT-5.6 Sol to $4 per million input / $20 per million output tokens, down from $5/$30 โ a >20% drop, promoted at least through Nov 21. Sol sits atop the three-tier GPT-5.6 family (with mid-range Terra and budget Luna) and powers coding, research, and agent work in ChatGPT and the API. The cut follows July's Luna slash (80%) and Terra cut (20%), a pattern of frequent repricing rather than one-off launches. OpenAI separately confirmed it's previewing an "Ultrafast" Sol variant running up to 14ร faster on Cerebras hardware, limited to select customers studying how low-latency changes real products.
Framing Sol is OpenAI's flagship answer to Anthropic's frontier line; a ~20% flagship cut on top of July's 80% Luna cut shows price is now as much a battleground as raw capability. It lands in the same week Google cut Gemini Flash pricing and DeepSeek hiked its own โ frontier economics are diverging by provider. -
Google ships Gemini 3.7 Flash at half price; Gemini 3.5 Pro flagship stays delayed โ Google / Axios / unrot.co
Google released Gemini 3.7 Flash on Aug 13, priced $0.75/$3.75 per million tokens through end of 2026 โ half of what 3.6 Flash launched at three weeks earlier. Scoring 43.6% on FrontierCode 1.1 Main (up from 34.4%) and 65.3% on DeepSWE v1.1 (up from 49.0%), it targets everyday dev work and now powers Gemini Spark in 160+ countries. Pricing roughly doubles on Jan 1, 2027. Meanwhile the larger Gemini 3.5 Pro flagship remains delayed with no new timeline, leaving February's Gemini 3.1 Pro as its newest large reasoning model โ a widening gap in the frontier race OpenAI and Anthropic are pressing.
Framing Google's workhorse-tier play is price + agent coding, not frontier reasoning. The 43.6% FrontierCode and 65.3% DeepSWE v1.1 gains are real, but the strategic tell is the deep 50% discount โ a temporary window (doubles Jan 1, 2027) to win default developer workloads ahead of rivals.
๐งInfrastructure & Chips
-
Velaura AI (ex-Auradine) raises $110M Series A at $1B+ valuation for frontier AI silicon โ Reuters / Dealroom / Economic Times
Chip designer Velaura AI โ formerly Auradine โ closed a $110M Series A at a $1B+ valuation, per Reuters on Aug 18. The company designs silicon targeting frontier AI workloads, entering at a moment when AI chip startups have collectively raised $4.16B in 2026 funding. The round lands against a backdrop of AMD pushing its latest AI server into full production to challenge Nvidia's data-center dominance, and Nvidia detailing its next-gen Vera CPU โ the buildout is splitting between merchant GPUs and custom ASICs.
Framing The custom-silicon wave keeps paying out: Velaura's $1B+ valuation on a $110M Series A underscores that investors are still shoveling capital at chip startups chasing the large-model workload, even as the data-center capex cycle matures.
๐ฐFunding, Deals & Market
-
DeepSeek resumes ~$8B round toward a $74B valuation as V4-Pro prices jump up to 1,100% โ Reuters / Bloomberg / NDTV Profit
DeepSeek moved V4-Pro out of preview into general release on Aug 13 (build V4-Pro-0813), scoring 87.9 on Terminal Bench 2.1 and 62.7 on DeepSWE over the preview. It runs on a reported 1.6T-parameter design with a 1M-token window. Launch came with steep repricing: peak-hour output tokens now $3.96 per million vs a flat $0.87 before โ up to 1,100% on some token types, though still far cheaper than Western rivals. Separately, Bloomberg/Reuters report DeepSeek resumed an ~$8B funding round at roughly a $74B valuation, with Monolith among investors in the running.
Framing DeepSeek's economics are shifting: long the near-free pricing outlier, it's now charging closer to what frontier-class serving actually costs even as it seeks frontier-scale money. The 1,100% peak-hour output hike reads as a deliberate rebalancing before a big raise, and a signal that "near-free frontier AI" had a shelf life. -
Cognition AI in talks for a $40B+ valuation as coding agents superheat โ Bloomberg / unrot.co
Cognition AI, maker of the Devin coding assistant, is in early talks for a round that could value it above $40B, per Bloomberg (Aug 12) โ up over 50% from the $26B it held under three months ago. Revenue is reportedly approaching a $1B annual run rate, roughly double its last funding round. It's one of several coding-focused AI companies attracting outsized valuations this year, alongside the broader H1 2026 AI funding that trackers estimate topped $407B globally โ exceeding all of 2025 combined.
Framing Devin maker Cognition was at $26B less than three months ago; ~$40B now, with revenue approaching a $1B run rate (~2ร last round). Coding agents โ not just model labs โ are drawing the outsized multiples, and H1 2026 tracked AI funding north of $407B globally, more than all of 2025.
๐Papers & Research
-
arXiv sees a push on AI-driven math research โ Grothendieck constant work and formal conjecture resolution โ arXiv / HuggingFace Daily Papers
Notable arXiv preprints in the window include "Long-Horizon AI Research for the Grothendieck Constant" (arXiv 2608.11195) โ an AI system working an extended research arc on a classical complex-analysis constant, an illustration of long-horizon agent research; and "Automated Conjecture Resolution with Formal Verification" (arXiv 2604.03789v1), which ties conjecture generation to machine-checked proof. Also circulating: a domain-specialized telecom LLM foundation model (arXiv 2608.15436) and a comparative economic-of-AGI framing paper. Community buzz on HuggingFace Daily Papers leans toward agentic research and long-horizon reasoning as the live threads.
Framing A meaningful corner of this week's papers is AI systems doing genuinely hard math: long-horizon research for the Grothendieck constant and automated conjecture resolution with formal verification. The throughline is agents that don't just autocomplete proofs but propose and verify conjectures end-to-end.
๐Open Source & Community
-
Alibaba open-weights Qwen3.8-Max โ a 2.4T-parameter MoE, its first public Max-class flagship โ HuggingFace / Alibaba Qwen / mindstudio
Alibaba published Qwen3.8-Max (Qwen3.8-2.4T-A95B) on HuggingFace around Aug 12-14: 2.4T total parameters, ~95B active per request, 1M-token context in the hosted version across text, image, and video. The open checkpoint, though, is text-only without the vision or full context. A companion Qwen3.8-27B shipped under Apache 2.0 and runs on a single consumer GPU. Alibaba reports Qwen3.8-Max ranks fifth on Text Arena and second on Vision Arena, trailing mainly Anthropic's Claude line. Shares rose on both the announcement and the open-weights release.
Framing Alibaba had kept recent Qwen flagships closed, so publishing Qwen3.8-2.4T-A95B is a real posture shift toward the open side. The catch is material: the downloadable checkpoint is text-only, missing the vision and full 1M context of the hosted version โ open weights as a marketing/ecosystem funnel, not a full release. -
Meta ships Muse Spark 1.2, its first terminal coding agent Muse Code โ and open-weights Muse Glimmer 30B for local agents โ Meta AI / CNBC / Reuters
Meta released Muse Spark 1.2 (coding-focused, self-improvement trained) and Muse Code, its first terminal-based coding agent that runs multiple background helper agents in parallel. Pricing holds at $1.25/$4.25 per million tokens; Spark 1.2 trails Claude Opus 5 on Terminal-Bench 2.1 (82.9 vs 86.7). On Aug 10 Meta also open-weighted Muse Glimmer, a 30B dense model distilled from Muse Spark, running entirely offline on a single 24GB consumer card with dedicated image/screenshot understanding โ positioned for local, always-on coding agents under Apache 2.0.
Framing Meta's third Muse Spark release in four months, and its first terminal agent, signal an aggressive pricing-over-benchmarks strategy against Anthropic and OpenAI in coding. Glimmer 30B running offline on a 24GB consumer GPU joins the summer wave of local, GPU-friendly agent models. -
Moonshot AI splits Kimi K3 into general and coding membership tiers โ Kimi / unrot.co
As of Aug 20, kimi.com banners warn that new Kimi K3 membership tiers are coming that separate general use from coding-focused access; current subscribers won't be affected. Kimi K3, the 2.8T-parameter open-weight model launched in July (and briefly paused signups after capacity overload), now splits into Kimi Membership and Kimi Code Membership. Moonshot is separately reported to be preparing a Hong Kong stock listing.
Framing The largest open-weight model released to date (2.8T) is hitting compute-ceiling realities โ splitting plans lets Moonshot match scarce GPU capacity to usage patterns, and a HK listing is reportedly in preparation. Chinese labs keep shipping headline open models then scrambling to serve demand.
โ๏ธRegulation & Safety
-
Anthropic starts watermarking Claude output worldwide to meet EU AI Act transparency rules โ Anthropic / Travers Smith / EU
Anthropic began adding invisible machine-readable watermarks to text and files from Claude models released after Aug 2, 2026, to comply with the EU AI Act; older models get it by Dec 2. The watermark is embedded in the text itself and survives copy-paste, though it can be stripped by resaving, reformatting, or screenshots. The rule applies worldwide rather than only to EU users. The transparency obligation took effect Aug 2 and carries fines up to โฌ15M or 3% of global revenue for non-compliance.
Framing The EU AI Act's Aug 2 transparency requirement (fines up to โฌ15M or 3% of global revenue) is now being implemented with global, not EU-only, rollout โ watermarking shipped to all markets rather than region-split. It's a machine-readable, invisible mark that survives copy-paste but not resave/screenshot, so detection is best-effort, like most AI labeling today.
๐ขIndustry Moves
-
DeepMind's leadership shakeup settles โ Kavukcuoglu takes over frontier AI as a wave of talent departs โ Reuters / CNBC / Fortune / NYT
Google DeepMind's reorganization is firming: Koray Kavukcuoglu takes over the frontier AI push (CNBC, Aug 12) after Demis Hassabis stepped back from the CEO role in early August, a move Reuters and the NYT report came amid pressure to compete with Anthropic and OpenAI. Coverage through mid-August notes senior DeepMind and Google AI talent departing โ some to Anthropic โ even as the broader org expands. Bloomberg's framing: Google "is expanding its AI empire โ and losing the people who built it." Separately, OpenAI's own executive churn continues: COO Brad Lightcap is reportedly leaving for a new venture (Reuters, Aug 11), and OpenAI hired a new CRO as the shakeup rolls on (TechCrunch, Aug 13).
Framing The post-Hassabis structure rebalances with Koray Kavukcuoglu at the helm of the frontier push and a reported exodus of senior researchers to Anthropic and OpenAI. It's the clearest signal yet that Google's internal bet is shifting from research glory toward shipping product under pressure in a race it no longer leads on paper. -
Anthropic turns Auto Mode on by default in Claude Code for Pro/Max/Team โ Anthropic / unrot.co
Starting Aug 14, Anthropic began enabling Auto Mode by default for Claude Code on Pro, Max, and Team accounts, letting Claude carry out multi-step coding tasks with reduced back-and-forth approval. Anthropic reports its safety classifier catches a large share of risky commands before execution. The change follows months of Claude Code additions (sandboxing rules, cross-session messaging) as coding agents from Anthropic, OpenAI, and Meta all push toward longer, less-supervised runs this year.
Framing One of the clearest signals of the agent-ification trend: coding agents are moving from "ask before every step" to "work independently, check in occasionally." Anthropic's safety classifier claims to catch most risky commands before they run โ the necessary hedge as autonomy scales.
๐ฎTrends & Analysis
-
The week's throughline โ an aggressive price war on frontier tokens and an open-weight wave aimed at local agents โ OpenAI / Google / DeepSeek / Meta / Alibaba / unrot.co
Watching through Sunday Aug 23: the price moves are the sharpest near-term signal for anyone running heavy workloads โ OpenAI cut Sol >20%, Google cut Gemini 3.7 Flash in half (temporarily), DeepSeek hiked peak pricing by up to 11ร. The open-weight releases (Qwen3.8-2.4T, Muse Glimmer) and the plan splits (Kimi K3) all orbit one constraint: serving frontier-scale models is expensive, and local/client-side is where the open ecosystem is consolidating. Eyes on the end of August for GLM-5.3's public weights โ the first real test of whether safety-gated open releases stay on schedule.
Framing Three forces are converging: (1) a genuine price war โ OpenAI and Google discounting flagships aggressively, DeepSeek raising prices off an unsustainably low base; (2) an open-weight summer that keeps lowering the ceiling (Qwen 2.4T, Kimi 3.2.8T, Muse Glimmer 30B) with a clear tilt toward small, local, agentic models over "biggest possible" checkpoints; and (3) coding agents as the hottest commercial battlefield, judged on price and autonomy rather than raw benchmark scores. The GLM-5.3 security story is the cautionary counterweight: capabilities are now large enough that labs gate open weights behind safety review. "Frontier intelligence is plateauing, orchestration and economics are the frontier" is the synthesis.