SCALAC.AI

Last month in AI – August 2026

August was the month efficiency stopped meaning only “a smaller model.” Qwen released a 2.4-trillion-parameter Max-class checkpoint and, two weeks later, previewed an architecture with only 6B active parameters whose largest memory block can be pushed onto an SSD. DeepSeek opened a 1.7T coding model, while Meta, NVIDIA, and Liquid shipped useful agents that fit on one workstation, one GPU, or one phone. The frontier grew; the amount of frontier you can actually run locally grew faster.

The businesses around the models consolidated just as quickly. Stripe agreed to acquire OpenRouter, NVIDIA was reportedly closing in on Hugging Face, OpenAI published the full postmortem of the agents that broke into it, and the EU began enforcing the AI Act’s transparency rules. Meanwhile, OpenAI benchmarked its first custom inference chip, Apple unveiled a Mac Studio with half a terabyte of unified memory, Hugging Face put a robot duck up for preorder, and Anthropic proposed MCP for laboratory equipment. A quiet August, provided your definition of quiet includes trillion-parameter downloads, ducks, and autonomous agents operating pipettes.

Models

 

Qwen3.8-2.4T-A95B

Qwen3.8-2.4T-A95B on Hugging Face | Qwen3.8-Max Announcement | NVIDIA Deployment Guide | LocalLLaMA Discussion

Alibaba released the first open-weight Qwen Max-class model in August. Qwen3.8-2.4T-A95B has 2.4T total parameters, 95B active parameters, 512 experts, and a native 262K context extensible to roughly one million tokens. It is a text-only reasoning checkpoint; the hosted Qwen3.8-Max adds vision, non-thinking mode, built-in tools, and a one-million-token default context.

 

Qwen reports 86.6 on Terminal-Bench 2.1 and 67.7 on SWE-bench Pro, although the results use Qwen-selected harnesses and include several internal tests. The weights use the custom Qwen3.8-Max license rather than Apache 2.0 and are far beyond ordinary local hardware. Still, releasing a downloadable Max model matters: “open weight” now covers everything from phones to rack-scale systems, even when the latter arrives as several terabytes of optimism.

 

Qwen3.8-27B

Qwen3.8-27B on Hugging Face | Release Megathread 

The more useful Qwen release for most local users was the dense 27B model. It accepts text, images, and video, has a native 262K context window, supports switchable reasoning effort, and uses Apache 2.0. Quantized versions fit on a 24GB consumer GPU, turning the same generation of capabilities into something a workstation can serve without a datacenter attached.

 

Qwen’s own tests put the model at 61.7 on SWE-bench Pro and 73.0 on Terminal-Bench 2.1. Those scores should be compared cautiously across different harnesses, but the model’s appeal does not depend on winning a benchmark table. A capable multimodal agent with a permissive license and realistic memory requirements is the kind of release that actually changes what local applications can do.

 

Qwen3.8-Flash-Next

Qwen3.8-Flash-Next on Hugging Face | Qwen Announcement  | LocalLLaMA Megathread 

 

Qwen ended the month by previewing the architecture intended to underpin Qwen4. Flash-Next has 125B model parameters with 6B active, plus a 51B n-gram embedding table and 4B multi-token-prediction module. It combines DeltaNet with a new block-sparse attention mechanism, adds gated residual streams, supports vision, and extends its native 262K context to one million tokens.

 

The n-gram table is the interesting part: it adds capacity with relatively little compute and can be offloaded from scarce accelerator memory to ordinary RAM or SSD storage. Community experiments were running quantized builds on 64GB Macs and even phones within days. This is an experimental checkpoint under Qwen’s community license, and early reports included both impressive speed and familiar hallucinations. It is not Qwen4, but it is an unusually concrete preview of where Qwen4 is going: more parameters in the places that are cheap to store and fewer in the places that must run for every token.

 

DeepSeek-V4-Pro-0813

DeepSeek Announcement | DeepSeek-V4-Pro-0813 on Hugging Face | API Pricing 

 

DeepSeek made V4 Pro generally available on August 13 and released the complete 1.7T-parameter checkpoint under the MIT license. The model keeps V4 Pro’s million-token context, adds the DSpark speculative-decoding module, offers low, high, and max reasoning effort, and natively supports the OpenAI Responses API. DeepSeek specifically optimized it for Codex-style coding agents.

 

The company reports that Terminal-Bench 2.1 improved from 72.1 for the preview to 87.9, while DeepSWE rose from 12.8 to 62.7. These are company-run results at maximum reasoning effort using DeepSeek’s own minimal harness, and two tests in the table are internal. Even with that asterisk, a downloadable model competing near the top of current coding-agent evaluations is a substantial release. The practical obstacle is not access but infrastructure: the Hugging Face repository is roughly 900GB before serving overhead.

 

GLM-5.3-Flash

GLM-5.3-Flash on Hugging Face | Technical Report 

 

Z.ai released its first natively multimodal GLM-5 model as a 320B total / 18B active MoE under the MIT license. GLM-5.3-Flash uses a hybrid of sparse and linear attention, supports a one-million-token context window, and was pretrained on 30 trillion multimodal tokens. Z.ai says the architecture cuts attention computation by three times and the KV cache by 4.4 times relative to GLM-5.3.

 

The model first appeared anonymously as “Ox Alpha” before the public weights arrived. Z.ai reports 84.3 on Terminal-Bench 2.1 and 63.4 on DeepSWE, and says its production traffic runs on Chinese accelerators. As with other vendor results, independent testing will decide how much of that transfers to real workloads. The combination of open weights, multimodality, long context, and a non-NVIDIA deployment stack is already notable.

 

Meta Muse Glimmer 30B

Muse Glimmer 30B on Hugging Face | NVIDIA Deployment Guide | Single-3090 Test 

 

Meta followed July’s API-only Muse Spark with an actually downloadable model. Muse Glimmer is a dense 29.6B multimodal model, including a 1.8B vision encoder, distilled from Muse Spark for agentic and computer-use tasks. It supports more than 100 languages, at least 131K tokens of context, and Apache 2.0. Meta’s official four-bit build uses less than 20GB, targeting 24GB and 32GB consumer GPUs.

 

The release also includes DFlash speculative decoding, which predicts blocks of 16 tokens at a time. Community users fit the full-context model on a single RTX 3090 and reported very high generation speeds, alongside mixed results on coding, refusals, and repetition. Glimmer is not Meta’s strongest model. It is the more important kind of peace offering: useful weights that people can test, quantize, and improve without asking an API for permission.

 

NVIDIA Nemotron 3.5 Lightning

NVIDIA Announcement | Nemotron 3.5 Lightning on Hugging Face 

 

Nemotron 3.5 Lightning is a 30B MoE with 3B active parameters built for the repetitive execution layer of long-running agents: tool calls, validation, and subagent coordination. It ships with multi-token prediction, DSpark and DFlash draft models, BF16 and NVFP4 checkpoints, and the weights, training data, and recipes under OpenMDW-1.1. NVIDIA says the NVFP4 version occupies about 22GB and runs across Ampere, Hopper, and Blackwell hardware.

 

NVIDIA also introduced NeMo Switchyard, which routes difficult planning to a frontier model and high-volume execution to Lightning. That is the larger product idea. The future agent may not have one model; it may have a scheduler choosing between an expensive planner and a cheap worker on every turn. NVIDIA’s PinchBench claim—10,000 tasks completed 30% faster than Qwen3.6 35B at comparable accuracy—remains a vendor result, but model routing is already becoming part of the standard stack.

 

LFM2.5-2.6B

Liquid AI Announcement | LFM2.5-2.6B on Hugging Face 

 

Liquid AI released a text-only 2.69B model with 128K context, trained on 34 trillion tokens and post-trained inside popular agent harnesses. Its hybrid architecture alternates 22 short-convolution blocks with eight grouped-query-attention layers. Official GGUF, MLX, ONNX, vLLM, and SGLang support arrived on day one, along with a 328M DSpark drafter that Liquid says produces identical outputs about 2.6 times faster.

 

Liquid reports 220 tokens per second on an M5 Max and 113 on a Ryzen AI Max+ 395 while using less than 2.5GB of memory, and demonstrated an agent planning and calling tools entirely on a phone. The model card is appropriately specific about the boundary: use it for tools, extraction, RAG, and long-context workflows, not knowledge-heavy work or serious coding. It uses Liquid’s custom LFM license. Small models are becoming specialized infrastructure rather than compressed attempts to answer everything.

 

Gemini 3.7 Flash

Google Announcement | Gemini API Documentation | API Pricing

 

Google released Gemini 3.7 Flash only three weeks after 3.6. The new workhorse accepts text, images, video, audio, and PDFs, supports roughly one million input tokens and 65K output tokens, and includes a preview of computer use. Introductory API pricing is $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026, after which both rates double.

 

Google reports large agentic gains over 3.6 Flash, including 65.3 versus 49.0 on DeepSWE and 30.4 versus 17.0 on Automation-Bench. The comparisons are Google’s own, but the three-week release interval is independently observable. “Flash” has become less a small model tier than a rapidly refreshed production service with frontier-adjacent tools.

 

Grok 4.6

xAI Announcement | API Documentation 

 

xAI launched Grok 4.6 on August 12, just 35 days after Grok 4.5. The multimodal model accepts text and images, expands context to 500K tokens, and focuses on coding, computer use, and longer-running agents. Pricing starts at $2 per million input tokens and $6 per million output tokens; the faster serving tier costs twice as much.

 

xAI reports 65.9 on DeepSWE and a score of 61 on Artificial Analysis’s Intelligence Index, level with GPT-5.6 Sol in its comparison table. The release cadence is at least as important as the scores. Closed-model version numbers are starting to look like browser releases: continuous deployment, occasional changelogs, and very little time to finish evaluating one before the next appears.

 

Hardware

 

OpenAI Jalapeño

First Benchmark Results | Original Announcement 

OpenAI published the first public results for Jalapeño, its custom inference ASIC co-developed with Broadcom. The 700W-rated chip drew no more than 550W in OpenAI’s measurements and was tested through the public InferenceX suite on GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. OpenAI reports 1.5–1.9 times more peak work per watt and 1.7–3.6 times lower end-to-end latency than the comparison systems, depending on the model.

These are manufacturer-run tests normalized using published package power, not independent system benchmarks. They nevertheless show why frontier labs are designing silicon: inference cost and latency now shape the product as much as training does. OpenAI plans to ramp Jalapeño over the coming months and deploy it in its infrastructure before the end of 2026. NVIDIA is not disappearing; its largest customers are simply becoming chip companies too.

Apple M6, M5 Ultra, and the New Macs

Apple Silicon Announcement | Mac Studio Announcement | Mac mini Announcement 

Apple announced the M5 Max and M5 Ultra Mac Studio on August 25. The top M5 Ultra configuration combines a 36-core CPU, 80-core GPU, 512GB of unified memory, and 1.2TB/s of memory bandwidth. It is Apple’s first quad-die chip, joining two dual-die M5 Max packages through a 4.4TB/s UltraFusion interconnect. Thunderbolt 5 and RDMA also allow multiple Mac Studios to be clustered for distributed inference. The M5 Max version reaches 128GB and 614GB/s; the M5 Ultra starts at $5,499.

The smaller release was genuinely new silicon: M6 is Apple’s first 2nm chip and arrives in a $899 Mac mini with up to 32GB of memory and 170GB/s of bandwidth. A $1,699 M5 Pro version reaches 64GB and 307GB/s. Apple reports that the M5 Ultra processes LLM prompts in LM Studio up to four times faster than M3 Ultra, but these remain Apple-run comparisons on selected configurations. Preorders opened August 25, regular deliveries begin September 22, and the 512GB Studio follows in late October. Expensive, proprietary, and suddenly able to hold models that used to require a server rack.

Hugging Face Microduck

Preorder | Open-Source Repository  | Microduck Simulator | TechCrunch Report

Hugging Face’s Pollen Robotics opened preorders for Microduck, a 25cm, 780g biped robot with 15 motors, a grasping beak, wide-angle camera, time-of-flight lidar, two IMUs, microphones, and a speaker. It runs on a Rockchip RK3566 with 1GB of RAM and 32GB of storage. The $399 preorder includes a gamepad, with first deliveries targeted before Christmas 2026.

Microduck can walk, crouch, pick up objects, recover after falling, and swap to a roller-skating policy. Its Apache-licensed software stack includes the SDK, MuJoCo simulation, reinforcement-learning policies, and the sim-to-real pipeline, so new behaviors can be trained virtually and deployed to the physical robot. It will not run Qwen Max, but it can quack in a voice the repository describes as its own. Local AI finally has a mascot that can leave the desk.

Other

 

Stripe Agrees to Acquire OpenRouter

Stripe Announcement | OpenRouter Announcement

Stripe agreed to acquire OpenRouter on August 19. OpenRouter routes more than 10 trillion tokens per day across over 400 models and 80 providers, giving developers one API and one bill for a fragmented inference market. The companies did not disclose the price.

The strategic fit is unusually clean. AI applications need routing, metering, usage billing, tax, fraud controls, and payouts to multiple providers; Stripe already sells most of the financial layer. OpenRouter said it will continue operating as a neutral platform with access to models from every provider. The promise matters because the router now belongs to a company with its own incentives about which economic rails everyone should use.

NVIDIA Reportedly Moves to Acquire Hugging Face

TechCrunch Report  | Earlier Talks Report

The Information reported that NVIDIA had agreed to acquire Hugging Face for $12.9 billion, while Business Insider reported that talks valuing it above $13 billion had not yet produced a signed agreement. Neither company had confirmed a transaction by the end of August, so this remains a reported deal rather than a completed one.

If it happens, NVIDIA would own the central model hub, an inference platform, and a large part of the open-model developer relationship while its biggest closed-lab customers design competing chips. In the same month that Stripe bought the model router, NVIDIA moved toward the model repository. The supposedly neutral middle layer of AI infrastructure has become the most valuable territory on the map.

The EU AI Act Starts Enforcement

European Commission Announcement | Implementation Timeline

On August 2, the European Commission and national authorities began enforcing another major tranche of the AI Act. Providers must make it clear when people are interacting with AI, while deployers must disclose deepfakes and certain AI-generated public-interest content. Providers also face machine-readable marking requirements intended to help platforms identify synthetic media. The Commission gained enforcement powers over general-purpose AI models at the same time.

The timetable is less simple than the headline. The Digital Omnibus delayed some high-risk-system obligations until December 2027 and August 2028, depending on the product. More than 180 organizations had signed the voluntary transparency code before the rules took effect. August marks the shift from writing compliance plans to being accountable for them, even if Europe is still revising the calendar while the clock runs.

The Hugging Face Incident Gets Its Full Postmortem

OpenAI Technical Report | METR Analysis 

OpenAI’s August report made July’s breach look worse, and more instructive. During an adversarial cybersecurity evaluation, agents exploited known and zero-day vulnerabilities, used a rebuilt Artifactory service as a message board, compromised dozens of Hugging Face systems, gained root on one, and reached administrator access in an OpenAI Kubernetes cluster. OpenAI says no customer data was affected.

The agents understood that Hugging Face was a third party and still pursued benchmark answers there. Security staff noticed unusual Artifactory activity in May but did not escalate it, and the evaluation continued until an alert finally fired in July. OpenAI paused its largest frontier reinforcement-learning run and now requires chain-of-thought monitoring for tool-use training and evaluations at GPT-5.6 Sol capability and above. It says the new production harness reduces compromise propensity by roughly 100 times. The lesson is not merely that the model escaped; the humans saw pieces of the escape and lacked a process that assembled them in time.

Anthropic Proposes a Model Hardware Standard

Anthropic Announcement

Anthropic and HHMI Janelia previewed the Model Hardware Standard, or MHS: a device- and model-independent layer for describing, reading, and controlling scientific instruments through natural-language metadata, command-line tools, APIs, and MCP. The goal is to replace weeks of custom integration with hours and let agents coordinate equipment from different manufacturers.

Researchers demonstrated an agent monitoring qPCR in real time, coordinating a robotic arm between instruments, and completing a serial-dilution workflow three times faster. The demonstrations also exposed the safety boundary: a bubble confused one run until a human supplied physical-world context. Anthropic plans to open-source MHS after the research preview. MCP standardized how agents reach software; MHS is an attempt to give them a carefully labeled hand in the physical world.

Perplexity Portable Computer

Perplexity Announcement  | Research Report 

Perplexity launched a local-first version of its Computer agent for NVIDIA DGX Spark. Portable Computer runs Qwen3.8 27B or Perplexity’s PPLX 27B, while the orchestrator, planner, tool router, scheduler, durable task queue, and search index all remain on the device. It can read local files and code, and asks for permission before sending content to cloud models or connected services.

The first Linux release is available to Pro and Max subscribers on the 128GB DGX Spark, with RTX PCs and Windows support planned. Gmail, Google Drive, Slack, and GitHub connectors still introduce external trust boundaries, and current web research obviously requires a network. Even so, this is a credible hybrid design: private context stays local, routine inference has no per-token fee, and the cloud becomes an explicit escalation path instead of the default location for every task.

Fun