// conneqtme.com · Saudi Arabia
KSA AI NEWS ↗AI SHOPPING ↗AI DASHBOARD ↗|ENعربي
AI News August 2026 Round 1
● AUGUST 2026 · ROUND 1 · WEEKS 2–3

Four Frontier Labs Refreshed Their Flagship in Ten Days — and Every AI Just Got an Employee ID.

Grok 4.6, Gemini 3.7 Flash, GLM-5.3, and DeepSeek V4 Pro 0813 all shipped inside the same two weeks. xAI's Grok Bot and Block's Buzz/Berd stack both push agents from "tool you call" to "colleague with a computer." Plus open weights landing at every size from 2.6B to 2.4T, a fresh Arabic voice comparison, and the observability stack you need before any of this ships to a client.

PUBLISHED · AUGUST 20, 2026
WINDOW · AUGUST 4–20
COVERAGE · FRONTIER · OPEN WEIGHTS · AGENTS · ARABIC AI

The last round closed on a curve: capability commoditizing and capability becoming dangerous are the same line, not two separate stories. This round is what that curve looks like when it's moving fast in every direction at once. In sixteen days, four labs — xAI, Google, Zhipu, and DeepSeek — refreshed their flagship or near-flagship model. In the same window, two entirely different companies shipped their version of "the agent has a desk now": xAI's Grok Bot gives every agent its own persistent cloud computer, and Block open-sourced both halves of a workspace where humans and agents share the same identity system.

Underneath the headline launches, the pieces that actually change how a team like Conneqt builds kept moving too — a fresh round of open weights small enough to self-host on a laptop, a genuine second option for Arabic speech recognition, and the observability tooling every one of these agent deployments needs before it touches a client account. Here's the full read on August 4–20, and what to actually do with it.

01 · FRONTIER

Four flagships, sixteen days: Grok 4.6, Gemini 3.7 Flash, GLM-5.3, DeepSeek V4 Pro.

AUG 12$2/$6 PER M · FLAT PRICING

SpaceXAI shipped Grok 4.6 with a sharper focus on long-running agents and "more ambitious interactive and visual work." It scored 61 on the Artificial Analysis Intelligence Index — up 5 points from Grok 4.5, matching GPT-5.6 Sol, two points behind Claude's newest flagship — while holding pricing flat at $2/$6 per million tokens, still over 60% below GPT-5.6 Sol. The catch: that "matches Sol" claim holds only on the nine-benchmark composite. On xAI's own ten-row table, Claude's flagship wins the most rows outright, and Grok 4.6 loses Terminal-Bench 3.0 by roughly 8.6 points to GPT-5.6 Sol — a gap the launch post's prose didn't mention. Grok 4.6 is strongest on knowledge work and legal reasoning, weakest on terminal use. No independent third-party replication existed as of a week after launch.

Grok 4.6 — Opus-adjacent, five weeks after Grok 4.5

Google's fastest release cadence of the year: Gemini 3.7 Flash landed just 23 days after 3.6 Flash, with no architecture change — Google says the gains came entirely from algorithmic post-training improvements. DeepSWE v1.1 jumped from 49.0% to 65.3%; FrontierCode 1.1 Main from 34.4% to 43.6%; WebDev Arena Elo from 1538 to 1588. Introductory pricing is half of 3.6 Flash's launch price. The context that matters: Google still hasn't shipped Gemini 3.5 Pro, promised for June, and industry chatter attributes the rapid Flash-tier iteration partly to internal AI talent departures. Live now in AI Studio, Android Studio, Google Antigravity, the Gemini Enterprise Agent Platform, and Spark (Google's always-on personal agent).

Gemini 3.7 Flash — three weeks after 3.6, half the price, sixteen points higher

Z.ai's thesis for this release is explicit: no new base model, no bigger architecture — "scaling post-training is all we did." Every claimed gain, including a reported 50% jump over GLM-5.2, comes from training method alone. DeepSWE rose more than twenty points to 66.9%; GDPval-AA v2 hit 1769, at launch the best published figure from any open-model lab. The real novelty is cybersecurity: 84.5% on CyberGym topped Z.ai's own launch chart outright, and 54.4% on ExploitBench set a launch-time high among open-model labs (though closed Anthropic and OpenAI flagships still lead there). One departure from GLM-5.2's pattern worth flagging: open weights and API access are staged behind a safety review rather than shipped same-day — available today only through the GLM Coding Plan and ZCode.

GLM-5.3 — "Built to Code. Ready for Cyber Defense."

No blog post, no press release — DeepSeek simply updated its pricing page's version string from preview to DeepSeek-V4-Pro-0813, four months after the April 24 preview debut. On DeepSeek's own harness it scored 87.9 on Terminal-Bench 2.1 (versus 72.1 for the April preview), 62.7 on DeepSWE, 61.5 on NL2Repo. Alongside the GA, DeepSeek open-sourced its own agent software, DeepSeek Harness v0.1, under MIT — a continuous session log, resumable/branchable/replayable runs, and a minimal shell-plus-file-editor mode (the exact configuration DeepSeek used for its own benchmark numbers). The less-covered footnote: prices jumped substantially on August 16 with new peak/off-peak billing — off-peak output landing at roughly 2.28× the old rate, peak at 4.55×.

DeepSeek V4 Pro 0813 — the quietest flagship launch of the year

AUG 13$0.75/$3.75 PER M · INTRO THRU DEC 31
AUG 14743B PARAMS · SAME BASE AS 5.2
AUG 121.6T MoE · 49B ACTIVEPRICE HIKE AUG 16
ModelReleasedInput/Output ($/M)Headline number
Grok 4.6Aug 12$2 / $661 on AA Intelligence Index
Gemini 3.7 FlashAug 13$0.75 / $3.75*65.3% DeepSWE v1.1
GLM-5.3Aug 14Staged rollout84.5% CyberGym
DeepSeek V4 Pro 0813Aug 12$0.44 / $0.87→up 16th87.9% Terminal-Bench 2.1

*Introductory price through Dec 31, 2026

For model selection this quarter
All four numbers above are vendor-reported and largely unaudited a week out — treat every one as a starting hypothesis, not a procurement decision. The practical move: if DeepSeek V4 Pro is in your routing table, act before the August 16 price change compounds; if evaluating GLM-5.3, note that weights aren't public yet, so it isn't self-hostable this week regardless of the benchmark chart.
02 · OPEN WEIGHTS

Open weights landed at every size this round — 2.6B to 2.4T.

AUG 10APACHE 2.0DISTILLED FROM MUSE SPARK

Meta Superintelligence Labs distilled Muse Glimmer down to ~30B parameters (29.6B dense + 1.8B vision encoder) from the larger Muse Spark, built specifically for always-on local agents — function calling, coding, screen/document reading, LLM-as-judge — on a single Mac or PC GPU. 131K context, 100+ languages, works with OpenClaw out of the box. Benchmarks are genuinely mixed: it leads Qwen3.6-27B and Gemma4-31B on MCP Atlas (75.5) and reasoning (AIME 2026: 94.7), but trails on OSWorld-Verified (65.9 vs Qwen's 75.6) and Terminal-Bench 2.1 (60.7). CEO Zuckerberg used the launch to press for lighter US open-source policy, framing it against Chinese open-weight competition.

Meta's Muse Glimmer — 30B, runs on one consumer GPU

Four days after Muse Glimmer, Alibaba open-weighted both the 2.4-trillion-parameter Qwen3.8-Max and, more usefully for most teams, Qwen3.8-27B — a 27.78B dense multimodal model (text, image, video) under Apache 2.0 with a 262K context window. On Qwen's own numbers it beats Muse Glimmer across all 8 direct comparisons and surpasses Claude Opus 4.6 on 15 of 19 overlapping tests: Terminal-Bench 2.1 up from 63.4 to 73.0, DeepSWE 1.1 from 13.3 to 42.2, OSWorld-Verified from 63.9 to 84.3. Runs on 24GB VRAM. As with every self-reported number this round, the standardized-eval caveat applies — but this is currently the strongest local-deployable candidate for teams that need vision plus agentic coding on a single workstation.

Qwen3.8 goes open — both the flagship and the practical size

Liquid AI rounded out an on-device family across three weeks: LFM2.5-2.6B (Aug 4) decodes at 220 tokens/s on an Apple M5 Max and 113 tok/s on a Ryzen CPU, under 2.5GB of memory, beating models up to 4× its size on ToolSandbox and instruction-following. LFM2.5-VL-3B (Aug 12) pairs the same backbone with a SigLIP2 vision encoder for phone/laptop-class vision-language work — RefCOCO grounding jumped roughly 30 points over the prior generation. Both ship under the LFM Open License with day-one llama.cpp, MLX, and ONNX support.

Liquid AI's LFM2.5 family — the edge tier

AUG 142.4T MAX + 27B DENSE
AUG 4–122.6B TEXT + 3.1B VISION
Reading self-reported benchmarks this round
Nearly every number above comes from the releasing lab, not an independent evaluator, and several outlets flagged inconsistent prompting or in-house-modified benchmark variants. The pattern worth internalizing: open-weight releases are now shipping faster than anyone can independently verify them. Build your own small eval set from real Conneqt tasks (as recommended in prior rounds) rather than ranking models by launch-day charts.
03 · AGENTS

The agent gets a desk: Grok Bot, and new Claude Managed Agents production controls.

AUG 11 · BETAPERSISTENT CLOUD COMPUTER

xAI's framing is direct: each Bot gets its own always-on cloud computer, signs into your existing tools with your own credentials, and works through multi-step jobs unsupervised, surfacing only when it needs approval. Internal xAI use cases (now public examples): an outbound-sales Bot that researches accounts overnight and queues personalized drafts, a demo-readiness Bot that fixes stale seed data before morning calls, bug-reproduction Bots that hand off to debugging Bots. Bundled into SuperGrok Heavy, Cursor Ultra ($200/mo), and Cursor Teams Premium ($120/seat/mo); live on desktop and iOS, enterprise on a waitlist.

Grok Bot — "AI teammates you can give real work to"

Anthropic shipped three governance-focused additions in the same window. Session budgets put a hard spend cap on any Managed Agents session — hit the ceiling and it pauses with a budget_reached stop reason instead of silently burning more; adjustable or removable to resume. Advisor models let a session's primary thread consult a separate, at-least-as-capable model mid-turn — a built-in second opinion without restructuring the agent loop. Inference geo pinning and GitHub-hosted skills round out the release, giving teams real data-residency and version-control options for the first time. Separately: Sonnet 5's introductory pricing ($2/$10 per Mtok) became the permanent standard price on August 10 — the scheduled increase to $3/$15 will not happen.

Claude Managed Agents — session budgets, advisor models, geo pinning

ROLLING OUTBUDGETS · ADVISOR · GEO PINNING
For Conneqt's Dawaa Brain and client deployments
Session budgets are the one to wire in immediately for any client-facing Managed Agents deployment — a hard spend ceiling per session removes the single biggest billing-surprise risk in autonomous agent work. Advisor models are worth pairing with the highest-stakes deliverables (client pitch decks, DCOM proposals) as a built-in second opinion, similar in spirit to the Hermes MoA pattern from July.
A gap worth knowing before you deploy
xAI's marketing page says each Bot has "their own computer." Its own technical documentation, published the same day, says the opposite: every Bot on an account shares one persistent cloud computer, and separate Bots should never be treated as a security boundary. For anyone considering giving a sales Bot and a finance Bot separate credentials to contain blast radius — that separation doesn't currently exist at the infrastructure level. Read the docs, not the launch page, before connecting anything sensitive.
04 · WORKSPACES

Block open-sources both halves of "the agent has a desk": Buzz and Berd.

Block — Jack Dorsey's company, behind Square, Cash App, and Tidal — released two complementary open-source products this stretch, and Nous Research shipped official Hermes Agent integration for both.

Buzz — the shared workspace, released July 31

Buzz is an open-source, self-hostable workspace built on Nostr, the decentralized protocol that gives every participant — human or agent — their own cryptographic keypair instead of a bot token. It looks like Slack (channels, DMs, threads, reactions) but agents are full workspace members, not sidebar bots. Every message, reaction, and approval is a signed, auditable event on a relay you control. Model-agnostic: ships harnesses for Claude Code, Codex, and Block's own Goose, all speaking the Agent Client Protocol (ACP) — meaning any ACP-compatible agent can join the same channel.

Berd — the solo-builder desktop app, released August 18

Where Buzz is collaborative, Berd is what Block's own teams use to work with agents privately before work becomes shared. A local-first desktop app (Tauri 2 + React 19, Apache 2.0, macOS/Windows/Linux) that gives one builder a single environment across models, tools, and projects — talking to Goose underneath via ACP, with conversation history stored on-device. Block's stated arc: "Agent work often begins privately, but useful work rarely stays solo. What we learn in Berd will help us shape Buzz." Reached v0.6.2 with 91 contributors within days of open-sourcing.

Hermes Agent connects to both

Nous Research shipped official support connecting Hermes Agent to Buzz via three paths — a zero-config Buzz Desktop runtime, a relay bridge for a hosted identity, or a full Hermes Gateway connection preserving native skills, long-term memory, approvals, and cron scheduling. A documented multi-profile pattern lets one team run several specialized Hermes identities in parallel: a #research channel running a web-search profile, a #content channel running a writing profile, a #infra channel running a DevOps profile — each with its own model, skills, memory, and Nostr keypair, all inside one Buzz workspace alongside Claude Code and Codex on the same channel.

JUL 31THREE INTEGRATION PATHS
For Conneqt's internal ops
This is the most directly applicable infrastructure release of the round. A self-hosted Buzz workspace with three or four Hermes profiles (research, content, DevOps, chief-of-staff) — plus Claude Code and Codex joining the same channels via ACP — replicates the multi-agent Slack-like setup teams have been assembling piecemeal since June, but now as one open, self-hostable, audit-trailed system with no per-seat SaaS bill. Worth a real pilot this month, starting with Berd for solo experimentation before scaling to a shared Buzz workspace.
05 · VOICE + ARABIC AI

MiniMax opens its video model, Cartesia takes the voice crown, and Arabic STT gets a genuine second option.

JUL 31 · WEIGHTS AUG 22K · 15-SECOND AUDIO-VIDEO

MiniMax open-sourced the core of H3 — its full-modal generation system understanding text, image, video, and audio context and outputting native dual-channel audio-video at up to 2K, 15 seconds. It ranks #1 globally on Artificial Analysis's video editing leaderboard, with pricing cut to roughly a third of comparable flagships. What's actually open: H3-Base (the 768p-class core generator) is on Hugging Face with day-0 ComfyUI, Diffusers, and WanGP support — runs on 12GB VRAM at 480p with the pruned int8 checkpoint. The preprocessing layer (Context-IR) and the 2K upscale pass (Regenerate-2K) remain API-only for now.

MiniMax H3 — open-source omni-modal video, ranked #1 on video editing

Three months after Sonic-3.5, Cartesia's state-space-model TTS now leads both Artificial Analysis speech arenas — 1,283 Elo on the open Provider Voice board and, more tellingly, 1,123 on the Controlled Voice board, which clones every competing model onto the same eight reference voices to isolate the synthesis engine itself from voice-catalog quality. Sub-90ms time-to-first-audio, inline expression tags ([laughter]), 10-second instant voice cloning, and IPA pronunciation overrides. Available in beta now.

Cartesia Sonic-3.6 — #1 on both speech leaderboards

Last round covered Cohere's Transcribe Arabic (25.87 WER, Apache 2.0, Gulf/Egyptian/Levantine dialect coverage). Worth noting alongside it: Speechmatics' Arabic-English bilingual model, live since earlier this year, reports 4.5% WER on Arabic-only transcription and a 35%-lower error rate than Google specifically on Arabic-English code-switching — the exact bilingual pattern common in Gulf business and clinical settings — plus a dedicated bilingual medical variant for regulated healthcare deployments. Speechmatics moved to unified credit-based billing across its whole product line on August 1, with no price change.

Arabic speech recognition: now a genuine two-horse race

AUG 1844 LANGUAGES · SUB-90MS
For Conneqt voice pipelines
Two credible, production-ready options now exist for Gulf-dialect Arabic STT with genuinely different strengths: Cohere for open, self-hostable, Apache 2.0 deployment inside PDPL-bound infrastructure; Speechmatics for hosted deployment with the strongest published code-switching numbers and a purpose-built medical variant. Worth benchmarking both directly against a real Hermes or Dawaa Brain voice sample set before standardizing on either.
06 · TOOLS + RESOURCES

A local AI workstation, a genuinely free coding tier, and the observability stack for everything above.

AUG 11OPEN SOURCE · MAC/WIN/LINUX

Unsloth's first desktop app moves well past "download and chat" — it runs and trains local models (LoRA workflows, 2× training speedup, 70% less VRAM in vendor testing), generates images and video locally including day-0 MiniMax-H3 support (a B200 dropped a 960×544, 124-frame generation from 70+ seconds to ~13), and connects Claude Code or Codex directly to a local model via unsloth start --as-subagent. Tool calling is reported 50% more accurate with self-healing, auto-retried malformed calls. Supports Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, DeepSeek-V4, GLM, and FLUX out of the box.

Unsloth Desktop — run and train models locally, one app

Made possible by OpenAI cutting Luna's cost 80% on July 30, Replit now runs chat, ideation, and simple agent tasks on GPT-5.6 Luna at no credit cost for Core and Pro subscribers — Replit's own agent automatically escalates to Sol when a task actually needs frontier reasoning, then returns to Free Mode once finished. Core subscribers reportedly get roughly 30× more usable output than before. Part of a deepening Replit-OpenAI partnership with more joint launches signaled.

Replit Free Mode — GPT-5.6 Luna, no credit cost, default for paid users

A widely-shared list this round is worth keeping on hand for exactly that reason — as agent deployments multiply across Hermes, Claude Managed Agents, Buzz, and Grok Bot, tracing what they actually did stops being optional.

The observability stack you need before any of the above ships to a client

AUG 19$20/MO CORE · $100/MO PRO
  • Langfuse and Arize Phoenix lead for self-hosted LLM/RAG tracing with strong eval support
  • OpenLLMetry and OpenLIT are OpenTelemetry-native instrumentation standards that pipe traces into infrastructure you already run
  • MLflow covers the full GenAI lifecycle including prompt versioning
  • Prometheus + Grafana remain the standard metrics/dashboard pairing underneath all of it
  • AI Fairness 360 is the odd one out — a bias/fairness auditing toolkit worth a look for any client-facing scoring or ranking agent
For Conneqt Brain and Hermes deployments
If nothing is currently instrumented, Langfuse is the fastest credible starting point — MIT-licensed, self-hostable, and it already understands OpenTelemetry spans, so switching or adding Phoenix later doesn't mean re-instrumenting. Wire it in before the next client-facing Managed Agents or Hermes MoA deployment, not after something goes wrong in production.
07 · BY THE NUMBERS

Sixteen days, in numbers.

4
Flagship Models Refreshed
2.4T
Qwen3.8-Max · Now Open
2.6B
LFM2.5 · Runs Under 2.5GB
23
Days Between Gemini Flash Releases
80%
GPT-5.6 Luna Price Cut
12
Open-Source Observability Tools Tracked
08 · BOTTOM LINE

What to do this week.

FOR MODEL ROUTING
Re-run your own eval set, not the launch charts
Four vendor-reported benchmark tables landed this round with no independent replication. Test Grok 4.6, Gemini 3.7 Flash, GLM-5.3, and DeepSeek V4 Pro against 10–20 real Conneqt tasks before any routing change.
FOR LOCAL DEPLOYMENT
Pilot Qwen3.8-27B on one workstation
Open today, 24GB VRAM, vision plus agentic coding in one dense model. The strongest self-hostable candidate this round for PDPL-bound work.
FOR AGENT INFRASTRUCTURE
Pilot Berd, then a shared Buzz workspace
Start solo with Berd this week. If it earns its keep, stand up a self-hosted Buzz workspace with two or three Hermes profiles plus Claude Code — no per-seat SaaS bill, full audit trail.
FOR CLIENT-FACING AGENTS
Wire in session budgets today
Any Claude Managed Agents deployment touching a client account should have a hard session budget set before its next run — this closes the single biggest billing-surprise risk in autonomous work.
FOR VOICE WORKFLOWS
Bench Cohere vs. Speechmatics on real Gulf audio
Two credible Arabic STT options now exist with different deployment models. Test both against an actual Hermes or Dawaa Brain sample set before standardizing.
FOR OBSERVABILITY
Instrument with Langfuse before the next deployment
MIT-licensed, self-hostable, OpenTelemetry-native. Wire it into the next Hermes MoA or Managed Agents rollout rather than retrofitting after an incident.
The pattern this round
Two curves crossed in the same sixteen days. The model curve: four labs proved a flagship refresh now ships in as little as three weeks, at a fraction of the previous cost, with benchmark gains coming from post-training technique as often as raw scale. The agent curve: two unrelated companies — a builder of rockets and payment terminals, and a social network turned AI lab — both concluded independently that the next unit of AI product isn't a smarter reply, it's a persistent identity with its own computer, its own credentials, and its own audit trail. Put together, the operational reality for any team building on this stack is the same one from every round this summer, just sharper: the model you route to and the agent identity you grant access to are both now things that change faster than any single deployment decision should assume. Build the routing and governance layer first. Everything above it is going to keep moving.