
Grok 4.6, Gemini 3.7 Flash, GLM-5.3, and DeepSeek V4 Pro 0813 all shipped inside the same two weeks. xAI's Grok Bot and Block's Buzz/Berd stack both push agents from "tool you call" to "colleague with a computer." Plus open weights landing at every size from 2.6B to 2.4T, a fresh Arabic voice comparison, and the observability stack you need before any of this ships to a client.
The last round closed on a curve: capability commoditizing and capability becoming dangerous are the same line, not two separate stories. This round is what that curve looks like when it's moving fast in every direction at once. In sixteen days, four labs — xAI, Google, Zhipu, and DeepSeek — refreshed their flagship or near-flagship model. In the same window, two entirely different companies shipped their version of "the agent has a desk now": xAI's Grok Bot gives every agent its own persistent cloud computer, and Block open-sourced both halves of a workspace where humans and agents share the same identity system.
Underneath the headline launches, the pieces that actually change how a team like Conneqt builds kept moving too — a fresh round of open weights small enough to self-host on a laptop, a genuine second option for Arabic speech recognition, and the observability tooling every one of these agent deployments needs before it touches a client account. Here's the full read on August 4–20, and what to actually do with it.
SpaceXAI shipped Grok 4.6 with a sharper focus on long-running agents and "more ambitious interactive and visual work." It scored 61 on the Artificial Analysis Intelligence Index — up 5 points from Grok 4.5, matching GPT-5.6 Sol, two points behind Claude's newest flagship — while holding pricing flat at $2/$6 per million tokens, still over 60% below GPT-5.6 Sol. The catch: that "matches Sol" claim holds only on the nine-benchmark composite. On xAI's own ten-row table, Claude's flagship wins the most rows outright, and Grok 4.6 loses Terminal-Bench 3.0 by roughly 8.6 points to GPT-5.6 Sol — a gap the launch post's prose didn't mention. Grok 4.6 is strongest on knowledge work and legal reasoning, weakest on terminal use. No independent third-party replication existed as of a week after launch.
Google's fastest release cadence of the year: Gemini 3.7 Flash landed just 23 days after 3.6 Flash, with no architecture change — Google says the gains came entirely from algorithmic post-training improvements. DeepSWE v1.1 jumped from 49.0% to 65.3%; FrontierCode 1.1 Main from 34.4% to 43.6%; WebDev Arena Elo from 1538 to 1588. Introductory pricing is half of 3.6 Flash's launch price. The context that matters: Google still hasn't shipped Gemini 3.5 Pro, promised for June, and industry chatter attributes the rapid Flash-tier iteration partly to internal AI talent departures. Live now in AI Studio, Android Studio, Google Antigravity, the Gemini Enterprise Agent Platform, and Spark (Google's always-on personal agent).
Z.ai's thesis for this release is explicit: no new base model, no bigger architecture — "scaling post-training is all we did." Every claimed gain, including a reported 50% jump over GLM-5.2, comes from training method alone. DeepSWE rose more than twenty points to 66.9%; GDPval-AA v2 hit 1769, at launch the best published figure from any open-model lab. The real novelty is cybersecurity: 84.5% on CyberGym topped Z.ai's own launch chart outright, and 54.4% on ExploitBench set a launch-time high among open-model labs (though closed Anthropic and OpenAI flagships still lead there). One departure from GLM-5.2's pattern worth flagging: open weights and API access are staged behind a safety review rather than shipped same-day — available today only through the GLM Coding Plan and ZCode.
No blog post, no press release — DeepSeek simply updated its pricing page's version string from preview to DeepSeek-V4-Pro-0813, four months after the April 24 preview debut. On DeepSeek's own harness it scored 87.9 on Terminal-Bench 2.1 (versus 72.1 for the April preview), 62.7 on DeepSWE, 61.5 on NL2Repo. Alongside the GA, DeepSeek open-sourced its own agent software, DeepSeek Harness v0.1, under MIT — a continuous session log, resumable/branchable/replayable runs, and a minimal shell-plus-file-editor mode (the exact configuration DeepSeek used for its own benchmark numbers). The less-covered footnote: prices jumped substantially on August 16 with new peak/off-peak billing — off-peak output landing at roughly 2.28× the old rate, peak at 4.55×.
| Model | Released | Input/Output ($/M) | Headline number |
|---|---|---|---|
| Grok 4.6 | Aug 12 | $2 / $6 | 61 on AA Intelligence Index |
| Gemini 3.7 Flash | Aug 13 | $0.75 / $3.75* | 65.3% DeepSWE v1.1 |
| GLM-5.3 | Aug 14 | Staged rollout | 84.5% CyberGym |
| DeepSeek V4 Pro 0813 | Aug 12 | $0.44 / $0.87→up 16th | 87.9% Terminal-Bench 2.1 |
*Introductory price through Dec 31, 2026
Meta Superintelligence Labs distilled Muse Glimmer down to ~30B parameters (29.6B dense + 1.8B vision encoder) from the larger Muse Spark, built specifically for always-on local agents — function calling, coding, screen/document reading, LLM-as-judge — on a single Mac or PC GPU. 131K context, 100+ languages, works with OpenClaw out of the box. Benchmarks are genuinely mixed: it leads Qwen3.6-27B and Gemma4-31B on MCP Atlas (75.5) and reasoning (AIME 2026: 94.7), but trails on OSWorld-Verified (65.9 vs Qwen's 75.6) and Terminal-Bench 2.1 (60.7). CEO Zuckerberg used the launch to press for lighter US open-source policy, framing it against Chinese open-weight competition.
Four days after Muse Glimmer, Alibaba open-weighted both the 2.4-trillion-parameter Qwen3.8-Max and, more usefully for most teams, Qwen3.8-27B — a 27.78B dense multimodal model (text, image, video) under Apache 2.0 with a 262K context window. On Qwen's own numbers it beats Muse Glimmer across all 8 direct comparisons and surpasses Claude Opus 4.6 on 15 of 19 overlapping tests: Terminal-Bench 2.1 up from 63.4 to 73.0, DeepSWE 1.1 from 13.3 to 42.2, OSWorld-Verified from 63.9 to 84.3. Runs on 24GB VRAM. As with every self-reported number this round, the standardized-eval caveat applies — but this is currently the strongest local-deployable candidate for teams that need vision plus agentic coding on a single workstation.
Liquid AI rounded out an on-device family across three weeks: LFM2.5-2.6B (Aug 4) decodes at 220 tokens/s on an Apple M5 Max and 113 tok/s on a Ryzen CPU, under 2.5GB of memory, beating models up to 4× its size on ToolSandbox and instruction-following. LFM2.5-VL-3B (Aug 12) pairs the same backbone with a SigLIP2 vision encoder for phone/laptop-class vision-language work — RefCOCO grounding jumped roughly 30 points over the prior generation. Both ship under the LFM Open License with day-one llama.cpp, MLX, and ONNX support.
xAI's framing is direct: each Bot gets its own always-on cloud computer, signs into your existing tools with your own credentials, and works through multi-step jobs unsupervised, surfacing only when it needs approval. Internal xAI use cases (now public examples): an outbound-sales Bot that researches accounts overnight and queues personalized drafts, a demo-readiness Bot that fixes stale seed data before morning calls, bug-reproduction Bots that hand off to debugging Bots. Bundled into SuperGrok Heavy, Cursor Ultra ($200/mo), and Cursor Teams Premium ($120/seat/mo); live on desktop and iOS, enterprise on a waitlist.
Anthropic shipped three governance-focused additions in the same window. Session budgets put a hard spend cap on any Managed Agents session — hit the ceiling and it pauses with a budget_reached stop reason instead of silently burning more; adjustable or removable to resume. Advisor models let a session's primary thread consult a separate, at-least-as-capable model mid-turn — a built-in second opinion without restructuring the agent loop. Inference geo pinning and GitHub-hosted skills round out the release, giving teams real data-residency and version-control options for the first time. Separately: Sonnet 5's introductory pricing ($2/$10 per Mtok) became the permanent standard price on August 10 — the scheduled increase to $3/$15 will not happen.
Block — Jack Dorsey's company, behind Square, Cash App, and Tidal — released two complementary open-source products this stretch, and Nous Research shipped official Hermes Agent integration for both.
Buzz is an open-source, self-hostable workspace built on Nostr, the decentralized protocol that gives every participant — human or agent — their own cryptographic keypair instead of a bot token. It looks like Slack (channels, DMs, threads, reactions) but agents are full workspace members, not sidebar bots. Every message, reaction, and approval is a signed, auditable event on a relay you control. Model-agnostic: ships harnesses for Claude Code, Codex, and Block's own Goose, all speaking the Agent Client Protocol (ACP) — meaning any ACP-compatible agent can join the same channel.
Where Buzz is collaborative, Berd is what Block's own teams use to work with agents privately before work becomes shared. A local-first desktop app (Tauri 2 + React 19, Apache 2.0, macOS/Windows/Linux) that gives one builder a single environment across models, tools, and projects — talking to Goose underneath via ACP, with conversation history stored on-device. Block's stated arc: "Agent work often begins privately, but useful work rarely stays solo. What we learn in Berd will help us shape Buzz." Reached v0.6.2 with 91 contributors within days of open-sourcing.
Nous Research shipped official support connecting Hermes Agent to Buzz via three paths — a zero-config Buzz Desktop runtime, a relay bridge for a hosted identity, or a full Hermes Gateway connection preserving native skills, long-term memory, approvals, and cron scheduling. A documented multi-profile pattern lets one team run several specialized Hermes identities in parallel: a #research channel running a web-search profile, a #content channel running a writing profile, a #infra channel running a DevOps profile — each with its own model, skills, memory, and Nostr keypair, all inside one Buzz workspace alongside Claude Code and Codex on the same channel.
MiniMax open-sourced the core of H3 — its full-modal generation system understanding text, image, video, and audio context and outputting native dual-channel audio-video at up to 2K, 15 seconds. It ranks #1 globally on Artificial Analysis's video editing leaderboard, with pricing cut to roughly a third of comparable flagships. What's actually open: H3-Base (the 768p-class core generator) is on Hugging Face with day-0 ComfyUI, Diffusers, and WanGP support — runs on 12GB VRAM at 480p with the pruned int8 checkpoint. The preprocessing layer (Context-IR) and the 2K upscale pass (Regenerate-2K) remain API-only for now.
Three months after Sonic-3.5, Cartesia's state-space-model TTS now leads both Artificial Analysis speech arenas — 1,283 Elo on the open Provider Voice board and, more tellingly, 1,123 on the Controlled Voice board, which clones every competing model onto the same eight reference voices to isolate the synthesis engine itself from voice-catalog quality. Sub-90ms time-to-first-audio, inline expression tags ([laughter]), 10-second instant voice cloning, and IPA pronunciation overrides. Available in beta now.
Last round covered Cohere's Transcribe Arabic (25.87 WER, Apache 2.0, Gulf/Egyptian/Levantine dialect coverage). Worth noting alongside it: Speechmatics' Arabic-English bilingual model, live since earlier this year, reports 4.5% WER on Arabic-only transcription and a 35%-lower error rate than Google specifically on Arabic-English code-switching — the exact bilingual pattern common in Gulf business and clinical settings — plus a dedicated bilingual medical variant for regulated healthcare deployments. Speechmatics moved to unified credit-based billing across its whole product line on August 1, with no price change.
Unsloth's first desktop app moves well past "download and chat" — it runs and trains local models (LoRA workflows, 2× training speedup, 70% less VRAM in vendor testing), generates images and video locally including day-0 MiniMax-H3 support (a B200 dropped a 960×544, 124-frame generation from 70+ seconds to ~13), and connects Claude Code or Codex directly to a local model via unsloth start --as-subagent. Tool calling is reported 50% more accurate with self-healing, auto-retried malformed calls. Supports Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, DeepSeek-V4, GLM, and FLUX out of the box.
Made possible by OpenAI cutting Luna's cost 80% on July 30, Replit now runs chat, ideation, and simple agent tasks on GPT-5.6 Luna at no credit cost for Core and Pro subscribers — Replit's own agent automatically escalates to Sol when a task actually needs frontier reasoning, then returns to Free Mode once finished. Core subscribers reportedly get roughly 30× more usable output than before. Part of a deepening Replit-OpenAI partnership with more joint launches signaled.
A widely-shared list this round is worth keeping on hand for exactly that reason — as agent deployments multiply across Hermes, Claude Managed Agents, Buzz, and Grok Bot, tracing what they actually did stops being optional.