SCALAC.AI

Last Month in AI – June 2026

June 2026 was arguably the most chaotic and consequential month in AI history. We witnessed the launch and immediate government suspension of Anthropic’s most capable model, a massive wave of 27 major model releases, the emergence of trillion-dollar IPO filings, and the hardware industry’s pivot toward local AI at Computex. Here is everything that mattered in June.

Models

Claude Fable 5

Anthropic Statement

On June 9, Anthropic launched Claude Fable 5 and Mythos 5, only to have the US government issue an export control directive suspending all access just 72 hours later. The government cited national security concerns, claiming to have found a jailbreak method. Anthropic publicly disputed this, stating the jailbreak only found „previously known, minor vulnerabilities” — essentially asking the model to read a codebase and fix software flaws — that other models like GPT-5.5 can also find. Anthropic warned that applying this standard industry-wide would „essentially halt all new model deployments.” Despite the suspension, the community managed to distill Fable 5 before it was taken down, releasing „Qwable-v1” on Hugging Face.

Released on June 18 — the same day Fable access disappeared — Zhipu AI’s GLM-5.2 arrived as a timely open-source alternative. It is a 753B-parameter Mixture-of-Experts model with a 1M context window, released under an MIT license. It quickly became the third-best model overall on the Artificial Analysis Intelligence Index and the top open-weight model for creative writing and agentic coding tasks.

DiffusionGemma

DiffusionGemma Blog

Google introduced a radical architectural shift with DiffusionGemma, a 26B MoE model that uses text diffusion rather than autoregressive generation. This approach yields 4x faster text generation, running at over 2,000 tokens per second, though it trades off some accuracy at the margins. The release signals Google’s willingness to explore fundamentally different generation paradigms beyond the transformer autoregressive standard.

Google also released Gemma 4 12B on June 4, an encoder-free multimodal model under Apache 2.0 that targets 16GB VRAM local setups. Rather than bolting separate vision or audio encoders onto a language model, it uses one unified network — making smaller multimodal models cheaper, cleaner, and easier to run locally. It is the first Gemma model to natively understand images and audio without a separate encoder stack.

Microsoft MAI-Thinking-1

Microsoft Blog | Technical Report

Microsoft stepped out of OpenAI’s shadow at Build 2026, launching MAI-Thinking-1, a 1T total (35B active) MoE reasoning model trained from scratch on 33T tokens without distillation. The launch signals Microsoft’s ambition to become a frontier model lab in its own right, rather than solely an OpenAI distribution channel. The model is available via Azure AI Foundry and ships directly into GitHub Copilot alongside the companion MAI-Code-1-Flash.

NVIDIA released Nemotron 3 Ultra on June 4, a 550B-parameter sparse MoE with 55B active parameters built for long-running agentic harnesses. It features a hybrid Mamba/Transformer architecture and an unusually complete open release: weights, training data, recipes, a GenRM reward model, and an NVFP4 quantized checkpoint. It is designed to run natively in agentic frameworks like OpenCode and Hermes.

H Company released Holo 3.1 on June 4, a family of local computer-use agent models ranging from 0.8B to 35B parameters with new quantized checkpoints. The lineup targets running screen-driving agents on local hardware rather than in the cloud, making autonomous desktop control accessible without a cloud dependency.

Ideogram released Ideogram 4.0 on June 4, a 9.3B-parameter text-to-image model with open weights under a non-commercial license. It leads open-weight image models on typography and layout, with bounding-box-style prompting that trades casual generation ease for precise structured control. It is the first open-weight model to seriously challenge Flux and Stable Diffusion on layout accuracy.

JetBrains released Mellum 2 on June 4, a 12B Mixture-of-Experts coding model with only 2.5B active parameters, trained from scratch by a small team using a three-stage curriculum over 10 trillion tokens. The release demonstrates how IDE companies can convert years of developer-workflow context into model advantage, and it is also available on CoreWeave Inference for cloud deployments.

MiniMax announced M3 on June 4, a natively multimodal coding and agentic model with a one-million-token sparse attention context window and open weights promised soon. It scored 59 on SWE-bench Pro and builds on MiniMax’s reputation for cheap agentic tool calling, making it a strong contender for cost-sensitive enterprise deployments.

xAI launched the full release of Grok Imagine Video 1.5 on June 18, featuring nearly 2x faster generation, native synchronized audio, and a claimed #1 position on the video generation leaderboard. The model supports image-to-video with audio, placing xAI firmly in the race alongside Google’s Veo 3 and Sora for the video generation crown.

GPT-5.6 Sol, Terra, and Luna

OpenAI Blog | System Card | TechCrunch Coverage

OpenAI unveiled GPT-5.6 on June 26 as a family of three models: Sol, the flagship; Terra, a balanced everyday option with competitive performance to GPT-5.5 at 2x lower cost; and Luna, the fastest and most affordable. However, following the Fable 5 precedent set two weeks earlier, the White House asked OpenAI to limit the rollout to a „small group of trusted partners whose participation has been shared with the government” before any wider release. OpenAI complied, but made its displeasure clear: „We don’t believe this kind of government access process should become the long-term default. It keeps the best tools from users, developers, enterprises, cyber defenders, and global partners who need them.” The company cited its ongoing work with the administration on a cyber Executive Order framework as the reason for cooperating in the short term. The episode is fast establishing a pattern: the US government now effectively holds veto power over the public launch of frontier AI models, with no clearly defined safety standards to guide those decisions.

Hardware

NVIDIA RTX Spark

The Verge Coverage

At Computex 2026, NVIDIA unveiled RTX Spark, its first family of consumer PC chips combining an Arm CPU with a Blackwell GPU. Featuring 128GB of unified memory and roughly one petaflop of local AI compute, the platform targets local AI agents and 120B-class local inference. Arriving in laptops this fall from partners including Asus, Dell, and Microsoft Surface, NVIDIA claims it is „the most efficient PC chip ever built.” The platform directly competes with Apple Silicon for the local AI workload market.

Intel Arc G3 Extreme

Intel Newsroom

Intel announced the Arc G3 Extreme at Computex, a $1,699 custom chip with 32GB of RAM designed for handheld gaming devices such as the Acer Predator Atlas 8. The chip targets the emerging category of portable AI-capable gaming hardware, bringing meaningful VRAM to a form factor that previously topped out at 16GB.

AMD at Computex

AMD Blog

AMD focused on longevity and value at Computex, announcing the Radeon RX 9070 GRE and confirming AM5 socket support through 2029. The promise of a decade-long platform lifecycle offers a stable upgrade path for builders navigating the rapid hardware churn of the AI era, and the RX 9070 GRE brings competitive GPU performance to the mid-range market.

Other

The Trillion-Dollar IPO Race

Reuters Coverage

The financial scale of the AI industry reached unprecedented heights in June. Both OpenAI and Anthropic filed for US IPOs, with each company targeting valuations up to $1 trillion. This massive capitalization push comes as the companies engage in a brutal price war, with OpenAI exploring drastic token pricing cuts to defend its enterprise turf against the soon-to-be-public Anthropic. The dual filings mark the formal arrival of AI as a public-market asset class.

SpaceX/xAI Acquires Cursor

Reuters Coverage

In a massive consolidation of the AI coding space, SpaceX/xAI reportedly acquired Cursor for $60 billion. The acquisition signals a fundamental shift in how frontier labs view developer tooling — moving from providing APIs to owning the entire developer surface and workflow. Combined with xAI’s Grok models and Kimi K2.7 Code powering Cursor Composer 2, the deal creates a vertically integrated AI coding stack.

Midjourney Medical

Midjourney Announcement

Midjourney shocked the industry by announcing Midjourney Medical, a full-body ultrasound scanner concept capable of capturing 806TB per scan in under 60 seconds. This marks a striking pivot for the AI-native company, moving beyond image generation into hardware, imaging, and healthcare infrastructure. It signals a broader trend of AI-native companies expanding into physical products.

OpenRouter launched Fusion API in June, routing and ensembling a panel of lower-cost models to reach near-frontier results. According to episode notes, it lands within roughly 1% of Claude Fable 5’s performance at half the price, beating GPT-5.5 and Claude Opus 4.8 in some comparisons. For developers who lost access to Fable 5, Fusion API emerged as an immediate practical alternative.

Dario Amodei’s Local AI Controversy

Bloomberg Interview | Axios Coverage | Yahoo Finance Coverage

Anthropic CEO Dario Amodei sparked a fierce backlash in June after a series of public statements in which he argued that open-source AI models pose a national security threat comparable to bioweapons, and called on the US government to have the legal authority to block dangerous AI deployments. In a Bloomberg interview, he warned of what he called „China’s open-source threat” and argued that open-weight models represent a „Mythos-class” cyber risk. He also made several technical claims that drew immediate criticism from the AI community — including that „with open source software you can see the source, but here you cannot see inside the model,” that the collaborative benefits of open source „don’t work the same way” for AI, and that „ultimately you have to host it on the cloud.” Critics were swift and pointed: open-weight models like GLM-5.2 and Nemotron 3 Ultra are fully inspectable, actively improved through community fine-tuning and LoRAs, and routinely run on consumer hardware without any cloud dependency. Many observers noted the irony that Amodei’s statements came precisely as Anthropic’s own Fable 5 was suspended by the same government he was urging to take a harder line — and that the open-source community was already distributing a distilled version of Fable 5 as a direct response to that suspension. The controversy crystallised a growing tension in the AI industry between safety-focused closed-source labs lobbying for regulatory barriers and the open-source community that views such moves as anticompetitive gatekeeping dressed up as safety advocacy.

Fun