๐ง Model & Product Launches
-
GPT-5 sighted in the wild, launch expected early August with mini and nano variants โ Ars Technica / The Verge / TechCrunch
References to "gpt-5-reasoning-alpha-2025-07-13" have been spotted in code, showing "reasoning_effort: high" in model configuration. Sam Altman confirmed GPT-5 on Theo Von's podcast, saying the model answered a question he couldn't himself โ "it was a weird feeling" feeling useless relative to the AI. The Verge reports Microsoft engineers have been preparing server capacity for GPT-5 since late May. GPT-5 will unify the GPT series' multimodality with o-series reasoning capabilities, shipping with mini and nano versions through the API. Before GPT-5, OpenAI still plans to release its first open-weights model since GPT-2 โ though that's been delayed indefinitely for safety testing.
Framing The unification of OpenAI's GPT and o-series reasoning lines marks a structural shift in their product strategy โ one model to rule them all, with configurable reasoning effort. -
OpenAI launches ChatGPT Agent โ a general-purpose browser-based agent โ TechCrunch / OpenAI
The ChatGPT agent combines Operator's ability to click around websites with Deep Research's multi-source synthesis. It can navigate calendars, generate presentations, run code, and complete multi-step computer tasks. Rolling out to Pro, Plus, and Team subscribers. Users activate it via the "agent mode" dropdown. Early AI agents have struggled with complex tasks, but OpenAI claims this is far more capable than previous offerings.
Framing This is OpenAI's boldest attempt to pivot ChatGPT from Q&A to autonomous action โ combining Operator's web navigation with Deep Research's synthesis capabilities in one agentic mode. -
Thinking Machines Lab releases Inkling โ open-weight 975B MoE model โ TechCrunch / Thinking Machines Lab
Former OpenAI CTO Mira Murati's startup released Inkling, a mixture-of-experts model with 975B total parameters but only ~41B active per task. Trained on 45 trillion tokens of text, image, audio, and video. Open-weight, meaning developers can download and modify it directly. Features calibrated uncertainty flagging and adjustable thinking effort. On one benchmark, Inkling uses 1/3 the tokens of Nvidia's Nemotron 3 Ultra for the same coding performance. The company explicitly states Inkling is "not the strongest overall model available" โ targeting well-rounded enterprise adaptability instead.
Framing Mira Murati's startup is betting the enterprise will prefer adaptable, calibratable AI over one-size-fits-all frontier models โ Inkling is their proof point. -
Google delays Gemini 3.5 Pro for full architectural rebuild, scraps Gemini 2.5 Pro base โ BigGo Finance / Reuters / Geeky Gadgets
Google DeepMind delayed Gemini 3.5 Pro to July 17, 2026, abandoning the existing 2.5 Pro architecture for a ground-up rebuild. The overhaul targets mathematical reasoning, SVG scene generation, and image quality improvements. The new model introduces a 2 million token context window, Deep Think Reasoning Layer, and autonomous workflow capabilities. Meanwhile, Reuters reports Google updated lightweight Gemini models but the flagship remains delayed. Google is also developing Nano Banana Pro for image generation and Gemini 4 Flash for speed.
Framing Google scrapped months of work and started a new pre-training cycle โ a sign DeepMind is struggling to keep pace with OpenAI and Anthropic on the frontier. -
Vibe-coding platform Base44 launches own LLM (Base1) for app generation โ TechCrunch
Base44, acquired by Wix for $80M in 2025, has started rolling out its own custom LLM trained on tens of millions of user interactions. Founder Maor Shlomo says owning the model allows optimizations on latency, cost, and efficiency โ a defensive moat against competitors reliant on external frontier models. The first iteration, Base1, is trained on platform interaction data.
๐งInfrastructure & Chips
-
Meta to put custom AI chip into production in September, aims to double compute capacity โ Reuters
Meta's custom AI chip is slated for production in September 2026, as the company looks to double its computing capacity. The chip is designed for inference workloads specific to Meta's social media and recommendation systems. The move follows similar custom silicon initiatives at Google (TPU), Amazon (Trainium/Inferentia), and Microsoft. Meta's chip strategy is a direct play to reduce dependence on Nvidia GPUs for inference, though training still relies on Nvidia.
Framing Meta is following the hyperscaler playbook โ custom silicon for inference means escaping Nvidia margins and optimizing for Meta-specific workloads at planetary scale. -
Intel targets Nvidia and AMD with "Crescent Island" AI data center chip, sampling H2 2026 โ Barchart / Financial Times
Intel plans to launch a new AI data center GPU code-named Crescent Island by year-end. The chip, led by Eric Demers (ex-AMD), started sampling in H2 2026. Intel CEO Lip-Bu Tan has bet the company on a fresh GPU effort. The stock is up 190% YTD and 442% over 52 weeks, making it one of the strongest large-cap chip plays. Nvidia CEO Jensen Huang has projected $1 trillion in data center buildouts by 2028.
Framing Intel is attempting a comeback in AI silicon โ 442% stock gain over 52 weeks suggests the market is buying the turnaround story before the product ships. -
AMD transforms into a systems company with Helios rack-scale platform โ SiliconANGLE / AMD Advancing AI event
AMD is repositioning beyond chip specs into full rack-scale systems with its Helios platform, unifying GPUs, CPUs, networking, and software. The company's acquisition of ZT Systems accelerated this transition. AMD reported Q1 data center revenue of $5.8B (up 57% YoY) and landed a 6-gigawatt OpenAI deal. Meanwhile, AMD is splitting next-generation data center GPU lines to handle specialized AI demand, with HBM4 memory requirements driving design decisions.
-
China's Z.AI completes 1-gigawatt AI data center using only Chinese-made chips โ Bloomberg / Yahoo Finance
Z.AI has completed a 1-gigawatt AI data center powered entirely by Chinese-manufactured chips. The facility is capable of training frontier AI models without reliance on Western semiconductor supply chains. This represents a major milestone for China's chip independence efforts, particularly important given US export controls on advanced Nvidia and AMD GPUs to China.
Framing A geopolitical inflection point โ China can now build AI infrastructure at scale without Nvidia or AMD. The implications for export controls and the global chip market are significant. -
AI data centers squeezing memory supply, cross-sector coalition warns โ Data Center Knowledge
A cross-sector coalition warns that AI data center buildout is tightening global memory supply. Beyond GPUs, vendors are investing across networking, memory, CPUs, and orchestration software to improve utilization and remove bottlenecks. HPE's Discover 2026 announcements focused on cutting network-induced latency so GPU clusters stay busy and efficient.
๐ฐFunding, Deals & Market
-
Global startup investment hits record $510B in H1 2026 โ OpenAI and Anthropic alone account for 43% โ Crunchbase
Crunchbase data shows global venture funding reached $510B in H1 2026, surpassing the $440B invested in all of 2025. OpenAI and Anthropic alone accounted for $217B (43%). Q2 2026 was the second-largest quarter on record, with $205B across 5,000+ startups. IPOs and acquisitions returned in force, with Q2 being one of the strongest quarters for venture-backed exits in years. Late-stage funding dominated, but the seed and early-stage ecosystems remain active.
Framing Venture capital is concentrating at an unprecedented rate โ two companies consumed 43% of all startup funding in H1. This is either the new normal or peak froth. -
AI agent startups raised $1.8B in July across 12+ deals, valuations up 40% QoQ โ AI Funding Tracker
AI agent startup funding in July reached $1.8B across 12+ deals, up 35% from June. Enterprise automation agents captured 58% of capital. Developer tooling agents raised $420M. Sequoia led four deals including two $100M+ rounds. Average deal size hit $150M. 42% of deals were outside Silicon Valley, with London, Tel Aviv, and Paris emerging as agent innovation hubs.
-
Microsoft cuts 4,800 roles citing AI automation; 120,000 tech layoffs in 2026 โ TechCrunch / Layoffs.fyi
Microsoft eliminated ~4,800 roles (2.1% of workforce), saying AI is "changing how work gets done" though claiming the roles aren't being replaced by AI directly. Oracle cut 21,000 over 12 months. GitLab cut 14% of staff. Roughly 120,000 tech roles have been cut in 2026. Companies citing AI as a factor include Microsoft, Oracle, GitLab, Salesforce, Google, and others. May was the highest single month for tech layoffs in years.
Framing The paradox of AI-era tech: record revenues + record layoffs. Companies are automating their way to efficiency while shedding the pandemic-era hiring overhang.
๐Papers & Research
-
Dataset Distillation by Influence Matching โ outcome-centric synthetic data โ arXiv / HuggingFace Daily Papers
A new paper (Inf-Match) revisits dataset distillation by aligning the final training outcome rather than intermediate process surrogates. The method introduces a differentiable, sample-level influence estimator that quantifies parameter shifts from adding or removing data. On Tiny-ImageNet (IPC=10), Inf-Match achieves 31.5%, a +4.7% improvement over prior art. Scales to vision-language distillation on Flickr30K. Code released on GitHub. @url: https://arxiv.org/abs/2607.16859
-
GraphVid โ interactive graph-controlled video generation โ HuggingFace Daily Papers / UIUC
Researchers from UIUC introduce GraphVid, a graph-conditioned image-to-video model enabling interactive control through structured interaction graphs. Includes GraphVid-Bench, a large-scale interaction-centric video dataset. Versus Motion-I2V, GraphVid reduces FID by 39.9% and FVD by 37.6%, with substantial PSNR and SSIM improvements. Demonstrates that structured semantic interfaces can outperform pixel-level trajectory control. @url: https://arxiv.org/abs/2607.21580
-
Kronos โ a foundation model for the language of financial markets โ GitHub Trending / arXiv
Kronos is a new foundation model specifically designed for financial market language โ trained on market data, financial reports, and trading signals. Trending on GitHub with significant community interest. Represents the growing trend of domain-specific foundation models rather than general-purpose LLMs adapted for finance. @url: https://github.com/shiyu-coder/Kronos
๐Open Source & Community
-
GitHub trending highlights: browser agents, AI skills frameworks, and code review โ GitHub Trending
Several notable repos trending today: ego-lite (3K stars, fastest browser for AI agents for web automation), alibaba/open-code-review (12.5K stars, hybrid architecture code review with deterministic pipelines + LLM Agent), and obra/superpowers (agentic skills framework and software development methodology). Also trending: anthropics/claude-cookbooks, mattpocock/skills (agent skills for engineers), and Automattic/harper (Rust-powered offline grammar checker).
-
Thinking Machines Lab releases Inkling as open-weight model โ TechCrunch / Thinking Machines Lab
Inkling joins the growing ecosystem of open-weight models available for local deployment. At 975B params (41B active), it's a significant addition โ especially notable for enterprise developers who want the ability to fine-tune and modify rather than consume via API. The calibrated uncertainty feature is novel for open models. Available for download from Thinking Machines.
โ๏ธRegulation & Safety
-
EU AI Act โ transparency obligations kick in August 2, high-risk rules delayed โ Technology Org / EU Council
On August 2, 2026, the EU AI Act's Article 50 transparency obligations become enforceable โ requiring chatbot disclosure, synthetic content marking, and deepfake labeling. However, high-risk obligations for standalone Annex III systems have been pushed to December 2, 2027. AI embedded in regulated products under Annex I gets until August 2, 2028. The changes came through the "Digital Omnibus" simplification package approved by the EU Council in June 2026. This is the first enforceable deadline of the Act.
Framing The EU's phased approach means transparency (chatbot disclosure, deepfake labeling) hits first while the contentious high-risk rules get years more runway โ a pragmatic split that may define how the Act is remembered. -
2026 tech layoffs near 150,000 as AI reshapes workforce โ TechCrunch / Challenger, Gray & Christmas
AI was the most-cited reason for tech layoffs in 2026, according to outplacement firm Challenger, Gray & Christmas. Roughly 150,000 roles cut across tech, per Layoffs.fyi. TechCrunch notes the paradox: companies reporting record revenues while cutting headcount, pointing to AI as both growth engine and reason for reductions. Microsoft, Oracle, GitLab, and Salesforce were among the largest.
๐ขIndustry Moves
-
Microsoft and OpenAI restructure partnership, loosening exclusive ties โ Microsoft Blog / The Verge / THE Journal
April 2026 saw a major restructuring of the Microsoft-OpenAI partnership, loosening exclusive arrangements that had been core to the relationship. The new deal allows OpenAI more flexibility with other cloud providers while Microsoft retains significant commercial rights. The AGI clause โ which would force Microsoft to relinquish revenue rights if OpenAI achieves AGI โ remains a critical unresolved tension, which is one reason GPT-5's potential AGI declaration matters so much. @url: https://blogs.microsoft.com/blog/2026/04/27/the-next-phase-of-the-microsoft-openai-partnership/
-
Oracle cuts 21,000 jobs over 12 months, cites AI adoption โ TechCrunch / SEC Filing
Oracle disclosed a 21,000-person workforce reduction (13% of staff) in its annual SEC filing, explicitly stating "the adoption and deployment of AI technologies across our operations have resulted, and may continue to result, in reductions to our workforce." The cuts were larger than previously known.
๐ฎTrends & Analysis
-
The open-weight model landscape is diversifying โ niche over general
Three separate releases this week (Inkling for enterprise, Base1 for vibe-coding, Kronos for financial markets) illustrate a broader shift. The market is fragmenting: general frontier supremacy still matters for the lab PR wars, but actual value creation is moving to domain-optimized, adaptable, and transparent models. Thinking Machines Lab's explicit decision to prioritize calibration over benchmark supremacy is particularly notable as a counter-signal to the "bigger is better" frontier race.
Framing Inkling, Base1, Kronos, and Nemotron represent a shift: the most interesting AI releases this week aren't trying to beat GPT-5 on general benchmarks. They're optimizing for specific verticals, calibratable trust, or enterprise adaptability. -
The EU AI Act goes live โ enforcement begins August 2
The combination of EU enforcement and US state-level activity (Colorado AI Act) is creating a patchwork compliance landscape. The EU's decision to push high-risk rules to 2027-2028 while enforcing transparency now is pragmatic โ it gets the easy wins first while giving the industry and lawmakers time to figure out the hard questions around systemic risk and liability.
Framing The first enforceable deadline of the world's first comprehensive AI regulation is days away. How companies respond to the transparency obligations will set the compliance precedent for everything that follows.