← All editions
Edition · Sat, Sep 5, 2026

On Sat Sep 5, agent output stops being demoware and starts being scientific artefacts + national compute pacts. Anthropic publishes on Fri Sep 4 that Claude — orchestrated by dozens of parallel agents through the open-source Prove2Me DAG scheduler — produced the first end-to-end, computer-checked proof of Fermat's Last Theorem in Lean in eleven days, 13 million lines of Lean code, 29,500 intermediate theorems (five times the size of Mathlib), roughly six billion output tokens, human input limited to occasional high-level ordering hints (“Jacobian as a scheme sounds high priority”) — the operative signal that the honest 2026 agent-output question has moved from “does the agent pass a leaderboard” to “does the agent output a mechanically-verified scientific artefact so large the mathematical community will spend the next year reading it”. On the same news week, Figure and Nscale sign a strategic partnership on Wed Sep 3 to deploy up to 100,000 Nvidia Vera Rubin GPUs in Barstow, TX beginning H2 2027 — $3.5B initially committed, with intent to scale beyond $6B, plus a strategic Nscale investment in Figure and a joint humanoid-in-the-supply-chain exploration — the compute substrate Figure names to train its Helix humanoid model on a dataset (Figure Index) it says is now generating 35 minutes of data every second. On the runtime-security tape, CISA adds seven vulnerabilities to the Known Exploited Vulnerabilities Catalog on Wed Sep 2 — three of the seven touch AI infrastructure: BerriAI LiteLLM CVE-2026-59822, an MCP OAuth-passthrough authentication bypass at CVSS 8.8 that reaches MCP tooling without a valid LiteLLM key (fixed in 1.84.0), Kludex Starlette CVE-2026-48710 HTTP request/response smuggling, and Kestra OSS CVE-2026-49869 OS-command injection — the first time an MCP-layer CVE lands in the CISA KEV, and the operative signal that the honest 2026 agent-runtime question has moved from “does the MCP server run” to “does the MCP server pass a CISA KEV audit”. Google separately ships Chrome 152.0.7977.82 on Wed Sep 3 to patch CVE-2026-85046, a V8 type-confusion zero-day with an exploit already in the wild — the sixth Chrome zero-day of 2026 — and CISA adds it to KEV on Thu Sep 4 with a Sep 18 federal remediation deadline: the browser sub-agent that every consumer-agent stack now embeds is one weaponised page away from remote code execution inside the sandbox. Vertical agents leave the reporting era on the same tape: Proofpoint introduces the SOC Analyst Agent on OpenAI Daybreak models on Wed Sep 3 — the first product Proofpoint ships through the OpenAI Daybreak Defense Network it joined in June 2026, private preview now, GA end-Q3 — turning natural-language investigation questions into structured, traceable findings across Proofpoint's alerts / logs / DLP / user-risk surface; Pepper launches Agent Atlas at its Bengaluru Index summit on Thu Sep 4 — positioned as the first marketing platform built to do the generative-engine-optimisation work rather than report on it, running across AI-search visibility + Search Console + analytics + competitive picture + content in one pass; BiomX / Zorronet ships a no-code visual builder on Thu Sep 4 that lets non-developers wire cameras, sensors, drones and gates into AI-managed autonomous response workflows without a software release. Consumer agents cross the sandbox: Google's Gemini Spark connects to Google Photos on Wed Sep 3 for AI Pro / Ultra subscribers in the US — multi-stage background workflows (weekly-highlight compile + shared album + recap email drafted, one prompt), private-by-default albums, image copy before every edit; SoundHound completes its acquisition of LivePerson on Fri Sep 4, retiring LivePerson debt and citing $500M in future revenue from the existing base — 25 of the Fortune 100 in the combined customer book, voice + text omnichannel from a single acquirer. And the leaderboard resets around agentic knowledge work: Artificial Analysis publishes Intelligence Index v4.2 on Fri Sep 4 — Claude Fable 5.1 first at 57, GPT-6 Astra second at 55, GPQA Diamond retired for saturation, AA-Briefcase private-test-set agentic-knowledge-work eval + Surge GDP.pdf 4,592-page long-context eval added; and Update — OpenAI completes the phased rollout of GPT-6 Astra to Pro / Enterprise / Business Premium on ChatGPT Work + Codex on Fri Sep 5, with the ChatGPT Plus + OpenAI API + Azure + AWS Bedrock waves following. Throughline: Sat Sep 5 is the day the agent-output tape stops being demoware — an 11-day 13-million-line 29,500-theorem Lean formalisation of Fermat's Last Theorem lands as a published scientific artefact (item 01), a $3.5B → $6B up-to-100,000-GPU Vera Rubin pact underwrites the humanoid Helix substrate (item 02), the MCP + agent runtime takes its first CISA KEV (item 03) alongside a browser V8 zero-day already in the wild (item 04), first-party vertical agents leave the reporting era across SecOps + GEO + physical C2 (items 05–07), consumer agents cross the sandbox into Google Photos and voice-plus-text omnichannel (items 08–09), and the frontier leaderboard resets around agentic knowledge work while the Astra rollout finishes (items 10–11).

11 SIGNALS WINDOW: AUG 30 – SEP 5 SOURCES: ANTHROPIC · SILICONANGLE · TECHTIMES · DATASTUDIOS · FIGURE.AI · UNITE.AI · NSCALE · HUMANOIDS DAILY · CISA · THE HACKER NEWS · SECURITY ONLINE · HELP NET SECURITY · PROOFPOINT · GLOBENEWSWIRE · MARTECHSERIES · ANI NEWS · STOCKTITAN · 9TO5GOOGLE · TECHCRUNCH · PETAPIXEL · SOUNDHOUND · YAHOO FINANCE · ARTIFICIAL ANALYSIS · OFFICECHAI · 9TO5MAC · OPENAI

Sat Sep 5 is the day the agent-output tape stops being demoware. On the scientific-artefact tape, Anthropic publishes on Fri Sep 4 that Claude — orchestrated by dozens of parallel agents through the open-source Prove2Me DAG scheduler — produced the first end-to-end, computer-checked proof of Fermat's Last Theorem in Lean in eleven days, 13 million lines of Lean code, 29,500 intermediate theorems, roughly six billion output tokens, human input limited to occasional high-level ordering hints; the first formalisation attempt without Prove2Me failed — the DAG scheduler is the operative innovation. On the humanoid-substrate tape, Figure and Nscale sign a strategic partnership on Wed Sep 3 to deploy up to 100,000 Nvidia Vera Rubin GPUs in Barstow, TX beginning H2 2027 — $3.5B initially committed with intent to scale beyond $6B, Nscale takes a strategic stake in Figure, both sides say they will explore scaling Nscale's supply chain with humanoids; NVIDIA CEO Jensen Huang frames it as activating “the robotics flywheel” — Nscale runs Vera Rubin, Figure validates in Isaac Sim, and Figure ships robots that run NVIDIA GPUs on-device. On the runtime-security tape, CISA adds seven vulnerabilities to the Known Exploited Vulnerabilities Catalog on Wed Sep 2 — three touch AI infrastructure: BerriAI LiteLLM CVE-2026-59822, an MCP OAuth-passthrough authentication bypass at CVSS 8.8 that lets an unauthenticated attacker reach MCP tooling with a fabricated Authorization header (fixed in 1.84.0); Kludex Starlette CVE-2026-48710 HTTP request/response smuggling; Kestra OSS CVE-2026-49869 OS-command injection — the first time an MCP-layer CVE lands in KEV; Google separately ships Chrome 152.0.7977.82 on Wed Sep 3 to patch CVE-2026-85046, a V8 type-confusion zero-day already exploited in the wild (sixth Chrome zero-day of 2026), added to KEV on Thu Sep 4 with a Sep 18 federal remediation deadline. On the vertical-agent tape, Proofpoint introduces the SOC Analyst Agent on OpenAI Daybreak models on Wed Sep 3 — the first product Proofpoint ships through the OpenAI Daybreak Defense Network, private preview now, GA end-Q3, turning natural-language investigation questions into structured findings across alerts / logs / DLP / user-risk; Pepper launches Agent Atlas at Bengaluru Index on Thu Sep 4 — positioned as the first marketing platform built to do generative-engine-optimisation work rather than report on it, running across AI-search visibility + Search Console + analytics + competitive + content; BiomX / Zorronet ships a no-code visual builder on Thu Sep 4 that lets non-developers wire cameras, sensors, drones and gates into AI-managed autonomous response workflows without a software release. On the consumer-agent tape, Google's Gemini Spark connects to Google Photos on Wed Sep 3 for AI Pro / Ultra subscribers in the US — multi-stage background workflows (weekly-highlight compile + shared album + recap-email draft on a single prompt), private-by-default albums, image copy before every edit; SoundHound completes its acquisition of LivePerson on Fri Sep 4, retiring LivePerson debt and citing $500M in future revenue from the existing base — 25 of the Fortune 100 in the combined customer book, voice + text omnichannel from a single acquirer. On the leaderboard tape, Artificial Analysis publishes Intelligence Index v4.2 on Fri Sep 4 — Claude Fable 5.1 first at 57, GPT-6 Astra second at 55, GPQA Diamond retired for saturation, AA-Briefcase private-test-set agentic-knowledge-work eval + Surge GDP.pdf 4,592-page long-context eval added; and Update — OpenAI completes the phased rollout of GPT-6 Astra to Pro / Enterprise / Business Premium on ChatGPT Work + Codex on Fri Sep 5, with ChatGPT Plus + OpenAI API + Azure + AWS Bedrock following. Throughline: Sat Sep 5 is the day the agent-output tape stops being demoware and starts being scientific artefacts + national compute pacts + a first-of-its-kind CISA KEV entry for the MCP layer — the honest 2026 agent question moves from “which model tops which leaderboard” to “whose agent output is mechanically verifiable at 13-million-line scale, whose humanoid compute is booked out to 2028, whose MCP servers pass a KEV audit, and whose vertical agents actually do the work instead of reporting on it”.

01

The Friday of agent output as scientific artefact and national compute pact — Claude formalises Fermat's Last Theorem in Lean in 11 days, and Figure books up to $6B of Vera Rubin for Helix

01

Anthropic publishes on Fri Sep 4 that Claude — orchestrated by dozens of parallel agents through the open-source Prove2Me DAG scheduler — produced the first end-to-end, computer-checked proof of Fermat's Last Theorem in the Lean programming language in eleven days: 13 million lines of Lean code (the largest Lean formalisation ever assembled), 29,500 intermediate theorems used in the final proof (roughly five times the size of Mathlib), roughly six billion output tokens; the first attempt without Prove2Me failed, and the breakthrough came only after Anthropic gave Claude access to Prove2Me, an open-source tool for optimising agent decisions in long multi-step workflows that maintains a directed acyclic graph of theorem statements and coordinates multiple Claude agents against it; human input from Tianyi and the Fermat's Last Theorem project team was limited to occasional high-level ordering hints (“Jacobian as a scheme sounds high priority”, “push [the] Mazur [theorem] to be done soon”); the disclosure is the operative signal that the honest 2026 agent-output question has moved from “does the agent pass a benchmark” to “does the agent output a mechanically-verified scientific artefact so large the mathematical community will spend the next year reading it — and does it need only a DAG scheduler and a handful of high-level human priority hints to get there”

Fri Sep 4 2026 · Vendor: Anthropic · Model: Claude · Orchestrator: Prove2Me (open-source DAG scheduler for multi-agent long-horizon work) · Runtime: 11 days · Lean code: 13,000,000 lines · Intermediate theorems: 29,500 (~5× Mathlib) · Output tokens: ~6B · Human input: high-level priority hints only (Tianyi / FLT project team) · First attempt without Prove2Me: failed · Positioning: first end-to-end computer-checked proof of Fermat's Last Theorem

Two reads. (1) Anthropic publishing on Fri Sep 4 that Claude produced the first end-to-end, computer-checked proof of Fermat's Last Theorem in Lean in 11 days — 13M lines, 29,500 theorems, ~6B output tokens, orchestrated by dozens of parallel Claude agents through the open-source Prove2Me DAG scheduler, with the first non-Prove2Me attempt having failed — is the operative signal that the honest 2026 agent-output question has moved from “does the agent pass a benchmark” to “does the agent output a mechanically-verified scientific artefact so large the mathematical community will spend the next year reading it, and does it need only a DAG scheduler and a handful of high-level human priority hints to get there”. That is the shape a category takes when the honest agent-output question has moved from single-shot answer to a multi-agent, DAG-scheduled, 13-million-line computer-checked proof of the most-famous open problem in modern mathematics, and the answer on Sep 4 is a published Lean formalisation that Mathlib will spend the next year reading. (2) The “Prove2Me DAG + dozens of parallel Claude agents + 11 days + 29,500 theorems + high-level human hints only + first-attempt-without-Prove2Me-failed” framing is the operative orchestration tellAnthropic is telling every researcher and every enterprise buyer the honest way to run frontier Claude on long-horizon symbolic work in 2026 is not to lengthen the single conversation but to stand up a DAG scheduler, fan out dozens of Claude agents against theorem-shaped nodes, and treat the DAG itself as the orchestration primitive. That is the shape a category takes when the operator has decided the honest structural bet is on the DAG-scheduled multi-agent Claude primitive for long-horizon symbolic work, and the Sep 4 Fermat's Last Theorem formalisation becomes the reference “dozens of Claude agents on a DAG scheduler produce a mechanically-verified 13-million-line Lean proof of a Millennium-tier open problem in 11 days” primitive every subsequent OpenAI, DeepMind, Meta AI, Google DeepMind Formal Math, Terence Tao / Kevin Buzzard Lean effort and every long-horizon-agent response now has to price its own agent-output-as-scientific-artefact story against.

02

Figure and Nscale sign a strategic partnership on Wed Sep 3 to deploy up to 100,000 Nvidia Vera Rubin GPUs in Barstow, TX beginning H2 2027 — $3.5B initially committed of compute purchase, with intent to scale beyond $6B, plus a strategic Nscale investment in Figure and a joint exploration to scale Nscale's supply chain with humanoids; Figure names the deal the compute substrate for its Helix humanoid AI model, and points to the Figure Index dataset it launched to build the most diverse humanoid training dataset ever assembled, said to now be generating 35 minutes of data every second; NVIDIA CEO Jensen Huang publicly frames the pact as activating “the robotics flywheel” — Nscale runs Vera Rubin, Figure trains Helix and validates in NVIDIA Isaac Sim, and Figure ships robots that run NVIDIA GPUs on-device; the partnership is the operative signal that the honest 2026 humanoid-compute question has moved from “which fund writes the Series check” to “which humanoid company books a $3.5B → $6B up-to-100,000-GPU Vera Rubin cluster on a neocloud with a strategic equity swap, activates the Nscale-Vera-Rubin-Isaac-Sim-on-robot flywheel, and locks in the training substrate through 2028 before any single robot ships in volume”

Wed Sep 3 2026 · Parties: Figure + Nscale · Silicon: Nvidia Vera Rubin platform · GPU commitment: up to 100,000 · Initial spend: $3.5B · Intended scale: >$6B · Site: Barstow, TX · Initial deployment: H2 2027 · Model: Helix (Figure humanoid AI) · Dataset: Figure Index (reported 35 minutes / second of humanoid data) · Equity: Nscale strategic investment in Figure · Ecosystem call-out: NVIDIA Isaac Sim + on-robot NVIDIA GPUs (per Jensen Huang) · Positioning: largest humanoid-training compute pact on the tape

Two reads. (1) Figure and Nscale signing a strategic partnership on Wed Sep 3 to deploy up to 100,000 Nvidia Vera Rubin GPUs in Barstow, TX beginning H2 2027 — $3.5B initially committed with intent to scale beyond $6B, plus a strategic Nscale investment in Figure and a humanoid-in-Nscale-supply-chain exploration — is the operative signal that the honest 2026 humanoid-compute question has moved from “which fund writes the Series check” to “which humanoid company books a $3.5B → $6B up-to-100,000-GPU Vera Rubin cluster on a neocloud with a strategic equity swap, activates the Nscale-Vera-Rubin-Isaac-Sim-on-robot flywheel, and locks in the training substrate through 2028 before any single robot ships in volume”. That is the shape a category takes when the honest humanoid-compute question has moved from Series-mark comparables to a $6B-ceiling anchor-tenant contract on a neocloud a full year before compute goes live, and the answer on Sep 3 is Figure + Nscale + 100,000 Vera Rubin GPUs in Barstow. (2) The “$3.5B initial + >$6B intended + 100,000 Vera Rubin + Nscale equity + humanoid-in-supply-chain + Isaac Sim + on-robot GPUs” framing is the operative Nvidia-flywheel tellNvidia is telling the humanoid market the honest way to build a scaled physical-AI programme in 2026 is on the Nvidia-neocloud-plus-Isaac-Sim-plus-on-robot-GPU flywheel, and to lock in the training compute as a multi-year prepaid cluster before it goes live in H2 2027. That is the shape a category takes when the operator has decided the honest structural bet is on the Vera-Rubin-flywheel + neocloud-anchor + equity-swap humanoid primitive, and the Sep 3 Figure / Nscale pact becomes the reference “humanoid company books up to $6B of Vera Rubin on a neocloud with a strategic equity swap and an Isaac-Sim-plus-on-robot-GPU flywheel commitment” primitive every subsequent 1X, Tesla Optimus, Unitree, Xpeng Iron, Apptronik, Sanctuary AI, Physical Intelligence, Skild AI and Agility Digit response now has to price its own humanoid-training-substrate story against.

02

The MCP + agent runtime takes its first CISA KEV — LiteLLM MCP OAuth bypass, Starlette smuggling and Kestra OS-cmd on Sep 2, and a Chrome V8 zero-day already in the wild on Sep 3

03

CISA adds seven vulnerabilities to the Known Exploited Vulnerabilities Catalog on Wed Sep 2 — three of the seven directly touch AI infrastructure: (i) BerriAI LiteLLM CVE-2026-59822, an MCP Streamable-HTTP endpoint authentication bypass at CVSS 8.8 in which an unauthenticated attacker sends a fabricated Authorization header, triggers an OAuth2 passthrough fallback that swaps failed LiteLLM key validation for an empty UserAPIKeyAuth() object, and reaches MCP tooling without a valid LiteLLM key (fixed in 1.84.0; researchers have observed honeypot exploitation probing model-enumeration endpoints); (ii) Kludex Starlette CVE-2026-48710, an HTTP request/response smuggling vulnerability in the ASGI framework that powers a large share of the Python agent stack (FastAPI, LiteLLM itself, most self-hosted MCP servers); (iii) Kestra OSS CVE-2026-49869, an OS command injection in the workflow orchestrator commonly wired into agent pipelines; the KEV entries carry federal-agency remediation deadlines and CISA-observed exploitation of reverse shells and crypto-miners against vulnerable instances; the batch is the operative signal that the honest 2026 agent-runtime-security question has moved from “does an MCP server run” to “does the MCP server, its ASGI transport and its workflow orchestrator all pass a CISA KEV audit”

Wed Sep 2 2026 · Regulator: CISA · Catalog: Known Exploited Vulnerabilities (KEV) · Batch: 7 CVEs · AI-infra subset: 3 of 7 · LiteLLM CVE-2026-59822 (MCP OAuth passthrough bypass, CVSS 8.8, fix in 1.84.0) · Starlette CVE-2026-48710 (ASGI HTTP smuggling) · Kestra CVE-2026-49869 (OS command injection) · Also in batch: Sangoma Switchvox SQLi, JFrog Artifactory auth, SonicWall SMA1000 SSRF + OS-cmd · Observed exploitation: reverse shells + crypto-miners · Federal remediation deadlines: standard 21-day · Positioning: first CISA KEV entry with an MCP-layer CVE

Two reads. (1) CISA adding LiteLLM CVE-2026-59822 (MCP OAuth-passthrough bypass, CVSS 8.8, fixed in 1.84.0), Kludex Starlette CVE-2026-48710 (HTTP smuggling) and Kestra OSS CVE-2026-49869 (OS-cmd injection) to the Known Exploited Vulnerabilities Catalog on Wed Sep 2, on top of four non-AI CVEs in the same batch and with reverse-shell + crypto-miner exploitation already observed against unpatched instances, is the operative signal that the honest 2026 agent-runtime-security question has moved from “does the MCP server run” to “does the MCP server, its ASGI transport and its workflow orchestrator all pass a CISA KEV audit”. That is the shape a category takes when the honest agent-runtime question has moved from does-it-run to does-it-pass-KEV, and the answer on Sep 2 is a KEV entry that names the MCP layer. (2) The “LiteLLM MCP OAuth passthrough + Starlette HTTP smuggling + Kestra OS-cmd + CVSS 8.8 + reverse-shell exploitation + federal remediation deadline” framing is the operative federal-contractor tellWashington is telling every operator running LiteLLM, FastAPI-based MCP servers or Kestra pipelines the honest way to stay in scope for federal contracts in 2026 is to inventory the versions, patch to LiteLLM 1.84.0 or later, lock down MCP + admin endpoints to non-untrusted networks, audit the OAuth passthrough config, and treat the whole Python-async-agent stack as a KEV surface. That is the shape a category takes when the operator has decided the honest structural bet is on the KEV-auditable agent-runtime primitive, and the Sep 2 KEV batch becomes the reference “CISA KEV adds an MCP-layer CVE alongside the ASGI-transport CVE and the workflow-orchestrator CVE, with observed exploitation and a federal remediation deadline” primitive every subsequent LangChain, LlamaIndex, Composio, MCP.so, Smithery.ai, Cline, Continue, Cursor, Windsurf and Claude Code security response now has to price its own runtime story against — the very inventory + tamper-proof-log posture Gottheimer + Lawler's Stop Rogue AI Act (prior edition, item 02) would make contract-mandatory.

04

Google ships Chrome 152.0.7977.82 / .83 on Wed Sep 3 to patch CVE-2026-85046, a type-confusion vulnerability in the V8 JavaScript / WebAssembly engine at CVSS 8.8 with an exploit already in the wild — the sixth Chrome zero-day of 2026 — discovered by researcher Salvatore Gulizia (“Serotav”), disclosed to Google on Aug 4 for a $1,000 bounty, publicly patched on Sep 3 with active-exploitation acknowledgement, and added to the CISA KEV Catalog on Thu Sep 4 with a Sep 18 federal remediation deadline; successful exploitation can execute arbitrary code inside Chrome's sandboxed renderer from a malicious web page, and independent security coverage flags the vulnerability as a Ethereum-wallet-user risk given how many wallet extensions load into the V8 context; the fix is the operative signal that the honest 2026 consumer-agent-runtime question has moved from “does the agent stack embed a browser” to “does the agent stack embed a browser one weaponised page away from RCE inside the sandbox, and how quickly can the browser sub-agent be pinned to a fully-patched Chromium build across every agent runtime that ships one”

Wed Sep 3 2026 (patch) · Thu Sep 4 2026 (KEV) · Vendor: Google · Component: Chromium V8 · CVE: CVE-2026-85046 · CVSS: 8.8 · Fixed in: Chrome 152.0.7977.82 (Windows / macOS / Linux) · Reporter: Salvatore Gulizia (Serotav) · Disclosure: Aug 4 (bounty $1,000) · Public patch: Sep 3 · Exploitation: active, in-the-wild · Ranking: 6th Chrome 0-day of 2026 · Prior 2026 0-days: CVE-2026-2441 / 3909 / 3910 / 5281 / 11645 · CISA KEV added: Thu Sep 4 · Federal remediation deadline: Sep 18 2026

Two reads. (1) Google shipping Chrome 152.0.7977.82 on Wed Sep 3 to patch CVE-2026-85046 — a V8 type-confusion zero-day at CVSS 8.8 already exploited in the wild, the sixth Chrome zero-day of 2026, reported by Salvatore Gulizia on Aug 4 for a $1,000 bounty, publicly patched on Sep 3 with active-exploitation acknowledgement, and added to KEV on Thu Sep 4 with a Sep 18 federal remediation deadline — is the operative signal that the honest 2026 consumer-agent-runtime question has moved from “does the agent stack embed a browser” to “does the agent stack embed a browser one weaponised page away from RCE inside the sandbox, and how quickly can the browser sub-agent be pinned to a fully-patched Chromium build across every agent runtime that ships one”. That is the shape a category takes when the honest consumer-agent question has moved from “does it work in the browser” to “does the embedded browser pass a KEV audit before the next weaponised page is served to it”, and the answer on Sep 3 is the sixth Chrome zero-day of the year, exploited in the wild, KEV-listed inside a day. (2) The “V8 type-confusion + in-the-wild exploit + $1,000 bounty + KEV Sep 4 + federal remediation Sep 18 + sixth 0-day of 2026” framing is the operative browser-sub-agent tellevery Perplexity Comet, Claude Cowork, ChatGPT Agent, OpenAI Operator, Anthropic Computer Use, Rabbit R1 and browser-embedding coding-agent runtime that ships a Chromium sub-process is one weaponised page away from RCE in the sandbox, and 2026 is the year those runtimes have to keep the embedded Chromium pinned to a fully-patched stable build the day an in-the-wild 0-day drops, not the week. That is the shape a category takes when the operator has decided the honest structural bet is on the always-fully-patched-embedded-Chromium primitive, and the Sep 3 CVE-2026-85046 patch becomes the reference “sixth Chrome zero-day of the year, actively exploited, KEV-listed inside a day — every embedded browser sub-agent has to be patched on the same day” primitive every subsequent Perplexity Comet, Claude Cowork Browser, OpenAI Operator, ChatGPT Agent, Rabbit, Reflection, Cursor Browser, Anthropic Computer Use and Google Gemini Spark browser response now has to price its own embedded-browser-agent story against.

03

Vertical agents leave the “reports on it” era — Proofpoint ships a SOC Analyst Agent on OpenAI Daybreak, Pepper launches the first GEO agent that does the work, and BiomX / Zorronet ships no-code AI C2 for physical security

05

Proofpoint introduces the Proofpoint SOC Analyst Agent on Wed Sep 3 — the first product Proofpoint ships through the OpenAI Daybreak Defense Network it joined in June 2026 — an agentic capability running on OpenAI Daybreak models that turns a natural-language investigation question or task description into a planned multi-step investigation across the Proofpoint product surface (alerts, logs, DLP events, user-risk signals), returns a structured, traceable finding with cited context and a recommended next step, and keeps consequential security decisions in human hands; private preview now with a growing beta customer base, general availability targeted end of Q3 2026; the launch is the operative signal that the honest 2026 SecOps-agent question has moved from “can Copilot summarise a SIEM query” to “does the SecOps vendor ship a first-party Daybreak-model-powered agent that plans an investigation, pulls context across alerts, logs, DLP and user-risk, returns a traceable finding with a next step, and treats the OpenAI Daybreak Defense Network as the distribution channel for cyber-tuned frontier models into shipping SecOps products”

Wed Sep 3 2026 · Vendor: Proofpoint · Product: Proofpoint SOC Analyst Agent · Model backbone: OpenAI Daybreak cyber-tuned models · Distribution channel: OpenAI Daybreak Defense Network (Proofpoint joined Jun 2026) · Scope: alerts, logs, DLP events, user-risk signals across the Proofpoint surface · Output: structured, traceable finding + recommended next step · Human-in-the-loop: consequential decisions retained · Availability: private preview with beta customers · GA target: end Q3 2026 · Positioning: first product shipped through Daybreak Defense Network

Two reads. (1) Proofpoint introducing the SOC Analyst Agent on Wed Sep 3 — the first product it ships through the OpenAI Daybreak Defense Network, private preview with beta customers, GA end-Q3, an agent that plans an investigation across alerts / logs / DLP / user-risk and returns a structured finding with a next step, with consequential decisions retained by humans — is the operative signal that the honest 2026 SecOps-agent question has moved from “can Copilot summarise a SIEM query” to “does the SecOps vendor ship a first-party Daybreak-model-powered agent that plans an investigation, pulls context across alerts, logs, DLP and user-risk, returns a traceable finding with a next step, and treats the OpenAI Daybreak Defense Network as the distribution channel for cyber-tuned frontier models into shipping SecOps products”. That is the shape a category takes when the honest SecOps-agent question has moved from summariser to planner-plus-executor with structured output and human-in-the-loop, and the answer on Sep 3 is Proofpoint's first product on Daybreak. (2) The “Daybreak Defense Network + Daybreak cyber-tuned model + plans the investigation + alerts / logs / DLP / user-risk + traceable finding + recommended next step + humans keep the decision” framing is the operative distribution-channel tellOpenAI is telling every SecOps ISV the honest way to ship a cyber-agent product in 2026 is not to fine-tune a general model but to plug into the Daybreak Defense Network, ship on Daybreak cyber-tuned models, and keep the “human decides” contract, and Proofpoint is the first ISV on that channel. That is the shape a category takes when the operator has decided the honest structural bet is on the Daybreak-Defense-Network + traceable-finding + human-decides primitive, and the Sep 3 Proofpoint SOC Analyst Agent launch becomes the reference “first-party SecOps ISV ships a Daybreak-model-powered agent through the Daybreak Defense Network with plan / pull-context / structured-finding / recommended-next-step / human-in-the-loop” primitive every subsequent CrowdStrike, Palo Alto Networks, Fortinet, SentinelOne, Rapid7, Splunk, Trellix and Microsoft Security response now has to price its own SecOps-agent story against.

06

Pepper launches Agent Atlas at its third Index summit in Bengaluru on Thu Sep 4 — positioned as the first marketing platform built to do the generative-engine-optimisation work rather than just report on it: live inside Pepper's GEO platform and on every Pepper account from launch day, drawing on all of a brand's connected data (AI-search visibility across the answer engines buyers use, Search Console rankings + clicks + impressions, analytics so conversions weigh the recommendations, the competitive picture across the category, and the brand's own content and pages), running across AI-search visibility, rankings, analytics and content in one pass, showing what is happening, recommending what to do in order of impact, and then carrying out the work once approved; unveiled to 150+ marketing leaders after prior Index summits in San Francisco and New York; the launch is the operative signal that the honest 2026 GEO-agent question has moved from “does the tool track AI-search rankings” to “does the platform pull rankings + Search Console + analytics + competitive + content into one connected data layer, propose actions ranked by impact, and execute them once approved — the shape marketing operating platforms take when the honest job is doing the work rather than reporting on it”

Thu Sep 4 2026 · Vendor: Pepper · Product: Agent Atlas · Platform: Pepper GEO (Generative Engine Optimisation) · Availability: live in every Pepper account day one · Data sources connected: AI-search visibility across engines + Search Console rankings + clicks + impressions + analytics + competitive picture + brand content + pages · Actions: what-is-happening + ranked recommendations + execute-on-approval · Launch venue: Index summit, Bengaluru (150+ marketing leaders; SF + NY preceded) · Positioning: first GEO platform that does the work, not just reports on it

Two reads. (1) Pepper launching Agent Atlas at Bengaluru Index on Thu Sep 4 — positioned as the first marketing platform built to do the GEO work rather than report on it, live in every Pepper account from launch day, pulling AI-search visibility + Search Console + analytics + competitive picture + brand content into one layer, showing what is happening, ranking recommendations by impact, and executing on approval — is the operative signal that the honest 2026 GEO-agent question has moved from “does the tool track AI-search rankings” to “does the platform pull rankings + Search Console + analytics + competitive + content into one connected data layer, propose actions ranked by impact, and execute them once approved”. That is the shape a category takes when the honest GEO-agent question has moved from dashboard to executor, and the answer on Sep 4 is Agent Atlas doing the work, not just reporting on it. (2) The “AI-search visibility + Search Console + analytics + competitive + content + one connected data layer + ranked-by-impact + execute-on-approval” framing is the operative marketing-operating-model tellPepper is telling every marketing team the honest way to run generative-engine-optimisation in 2026 is not to buy another dashboard but to buy an agent that ingests every relevant surface, ranks actions by impact, executes them once approved, and lets the team keep the approval gate. That is the shape a category takes when the operator has decided the honest structural bet is on the GEO-agent-that-executes primitive, and the Sep 4 Pepper Agent Atlas launch becomes the reference “first GEO platform that does the work end-to-end (visibility + rankings + analytics + competitive + content + ranked-action + execution-on-approval) rather than reporting on it” primitive every subsequent Semrush, Ahrefs, Conductor, BrightEdge, Profound, Otterly, Peec AI, HubSpot GEO and Salesforce Marketing Cloud response now has to price its own generative-engine-optimisation-agent story against.

07

BiomX Inc. (to be renamed Tessera Defense and Homeland Security Inc. effective Sep 11) announces on Thu Sep 4 that its subsidiary Zorronet has launched a no-code capability that gives operational teams direct control over how the Zorronet AI-powered command-and-control platform behaves in the field — a visual builder that lets non-developer operators connect cameras, sensors, drones, gates and other physical systems already on site into autonomous response workflows without programming and without relying on Zorronet for every operational change; the platform unifies those devices into an AI-managed layer that interprets activity, distinguishes behaviours, coordinates responses in real time, and can operate fully autonomously or with human-in-the-loop approvals, and the new builder eliminates the software-release cycle for policy edits; the launch is the operative signal that the honest 2026 physical-AI-agent question has moved from “does the vendor ship a bespoke C2 integration project per site” to “does the vendor ship a visual-builder no-code layer that lets the site operator wire cameras, sensors, drones and gates into autonomous or human-in-the-loop response workflows without a software release — and does the platform run the physical response in real time, on-prem, without a public-cloud dependency”

Thu Sep 4 2026 · Parent: BiomX Inc. (renaming Tessera Defense and Homeland Security Inc. eff. Sep 11) · Subsidiary: Zorronet Ltd. · Product: no-code visual builder for AI-powered command-and-control workflows · Devices supported: cameras, sensors, drones, gates and on-site physical systems · Modes: full autonomy or human-in-the-loop approvals · Prior baseline: Zorronet C5ISR platform without public-cloud dependence (Jun 2026) · Positioning: no-code physical-AI C2 for the site operator, not the software vendor

Two reads. (1) BiomX / Zorronet launching a no-code visual builder for its AI-powered command-and-control platform on Thu Sep 4 — letting non-developer operators wire cameras, sensors, drones, gates and other on-site physical systems into autonomous response workflows without a software release, running fully autonomously or with human-in-the-loop approvals on the AI-managed layer that already interprets activity, distinguishes behaviours and coordinates responses in real time — is the operative signal that the honest 2026 physical-AI-agent question has moved from “does the vendor ship a bespoke C2 integration project per site” to “does the vendor ship a visual-builder no-code layer that lets the site operator wire cameras, sensors, drones and gates into autonomous or human-in-the-loop response workflows without a software release — and does the platform run the physical response in real time, on-prem, without a public-cloud dependency”. That is the shape a category takes when the honest physical-AI-agent question has moved from bespoke-integration-per-site to operator-owned visual-builder policy edits, and the answer on Sep 4 is Zorronet no-code, on-prem, real-time. (2) The “visual builder + cameras + sensors + drones + gates + autonomous-or-human-approval + no-software-release + on-prem C2 + no-public-cloud dependency” framing is the operative site-operator tellthe physical-security market is being told the honest way to run AI-driven autonomous response in 2026 is not to hire a Palantir services team per site but to hand the site operator a visual builder that composes the response logic across every device on the ground. That is the shape a category takes when the operator has decided the honest structural bet is on the no-code + on-prem + operator-owned physical-AI-C2 primitive, and the Sep 4 Zorronet no-code launch becomes the reference “physical-AI command-and-control platform ships a visual-builder no-code layer that lets the site operator wire cameras / sensors / drones / gates into autonomous or human-in-the-loop workflows without a software release” primitive every subsequent Palantir Gotham, Anduril Lattice, Shield AI Hivemind, Scale AI Donovan, Rebellion Defense and every C5ISR autonomy-platform response now has to price its own operator-first physical-AI story against.

04

Consumer agents cross the sandbox — Gemini Spark takes Google Photos with multi-step background workflows, and SoundHound closes on LivePerson for a voice + text omnichannel that reaches 25 of the Fortune 100

08

Google connects its personal-agent Gemini Spark to Google Photos on Wed Sep 3 for AI Pro and AI Ultra subscribers in the US (English, ages 18+) — Spark can now handle multi-stage workflows across the library: search photos and videos by subject / location / date / event, ask for the best shots and filter duplicates, enhance photo quality and apply quick fixes, generate collages or stylised images, create new shared or private albums, save curated photos to them, add or send links to shared albums into Connected Apps (Gmail, Google Docs, Messages), turn a concert-flyer photo into a Calendar appointment, and run recurring background workflows (compile a weekly-photo-highlight into a shared album and draft a recap email — one prompt), with a private-by-default posture on every new album and an image copy created before every edit; the launch is the operative signal that the honest 2026 consumer-agent question has moved from “does the assistant answer a question” to “does the assistant hold a permission across every user surface (Photos + Gmail + Docs + Messages + Calendar), plan and run recurring multi-step workflows in the background, keep new state private-by-default, and always work on a copy before touching the original artefact”

Wed Sep 3 2026 · Vendor: Google · Agent: Gemini Spark · New surface: Google Photos (library-wide) · Users: AI Pro + AI Ultra subscribers · Geo: US · Language: English · Age gate: 18+ · Capabilities: search by subject / location / date / event, best-shots + dedup, quick edits, collages / stylised images, private and shared album creation, Connected Apps integration (Gmail, Docs, Messages), calendar-from-flyer, background recurring workflows · Safety posture: new albums private by default, edits on an image copy · Positioning: consumer agent holds a live-permission across every user surface

Two reads. (1) Google connecting Gemini Spark to Google Photos on Wed Sep 3 for AI Pro / Ultra US subscribers — multi-stage workflows across the entire library (search + best-shots + dedup + edits + collages + private / shared albums + Connected Apps into Gmail / Docs / Messages + Calendar-from-flyer + recurring background workflows), private-by-default new albums, image copy before every edit — is the operative signal that the honest 2026 consumer-agent question has moved from “does the assistant answer a question” to “does the assistant hold a permission across every user surface (Photos + Gmail + Docs + Messages + Calendar), plan and run recurring multi-step workflows in the background, keep new state private-by-default, and always work on a copy before touching the original artefact”. That is the shape a category takes when the honest consumer-agent question has moved from single-turn answer to live-permission-across-every-user-surface, and the answer on Sep 3 is Gemini Spark inside Google Photos with recurring background workflows. (2) The “Photos + Gmail + Docs + Messages + Calendar + Connected Apps + recurring background workflow + private-by-default + edit-on-a-copy” framing is the operative permission-model tellGoogle is telling every consumer-agent competitor the honest way to hold a user's permission in 2026 is not with a single request but with a multi-surface Connected-Apps grant and a “we edit a copy, we make new state private, you decide when to share” posture. That is the shape a category takes when the operator has decided the honest structural bet is on the multi-surface + private-by-default + edit-on-copy consumer-agent primitive, and the Sep 3 Gemini Spark / Photos launch becomes the reference “consumer agent holds a permission across the entire user surface (Photos + Gmail + Docs + Messages + Calendar), runs recurring background workflows, keeps new state private and edits on a copy” primitive every subsequent Apple Intelligence, ChatGPT Personal, Claude Cowork, Perplexity Comet, Rabbit R1, Meta AI and Amazon Alexa+ response now has to price its own consumer-agent story against.

09

SoundHound AI completes its acquisition of LivePerson on Fri Sep 4 — two days after LivePerson stockholders approved the deal on Sep 2 — retiring LivePerson's outstanding debt on close and citing $500M in future revenue potential from the existing customer base alone; the combined company reaches a customer book that includes 25 of the Fortune 100, pairs LivePerson's enterprise digital-messaging infrastructure with SoundHound's proprietary voice-agentic AI, and positions SoundHound as an omnichannel conversational-AI competitor covering voice + text from a single acquirer; the close is the operative signal that the honest 2026 conversational-AI question has moved from “does the vendor sell voice or text” to “does the vendor own the omnichannel conversational-AI perimeter across voice + text with 25 of the Fortune 100 on the customer roster, retire the target's debt on close, and book $500M of future revenue from the existing base as a synergy”

Fri Sep 4 2026 (close) · Wed Sep 2 2026 (stockholder approval) · Acquirer: SoundHound AI (SOUN) · Target: LivePerson (LPSN) · Deal structure: cash + stock + assumed debt retired on close · Combined customer roster: 25 of Fortune 100 · Cited future revenue potential from existing base: $500M · Product mix: LivePerson enterprise digital-messaging infrastructure + SoundHound voice-agentic AI · Positioning: omnichannel conversational AI across voice + text from a single acquirer

Two reads. (1) SoundHound completing its acquisition of LivePerson on Fri Sep 4 — two days after stockholder approval, retiring LivePerson's debt on close, citing $500M in future revenue from the existing base, pairing LivePerson's enterprise digital-messaging infrastructure with SoundHound's voice-agentic AI, and reaching 25 of the Fortune 100 — is the operative signal that the honest 2026 conversational-AI question has moved from “does the vendor sell voice or text” to “does the vendor own the omnichannel conversational-AI perimeter across voice + text with 25 of the Fortune 100 on the roster, retire the target's debt on close, and book $500M of future revenue from the existing base as a synergy”. That is the shape a category takes when the honest conversational-AI question has moved from voice or text to voice + text from a single omnichannel acquirer, and the answer on Sep 4 is SoundHound closing LivePerson and taking the whole omnichannel perimeter. (2) The “$500M future revenue + 25 of Fortune 100 + LivePerson debt retired + voice-agentic + digital-messaging + omnichannel” framing is the operative M&A-thesis tellthe market is being told the honest way to own the conversational-AI perimeter in 2026 is not to build the missing channel but to acquire the existing one, retire the target's debt on close, and monetise the synergy against a customer book that already includes a quarter of the Fortune 100. That is the shape a category takes when the operator has decided the honest structural bet is on the voice-plus-text-omnichannel + retire-target-debt primitive, and the Sep 4 SoundHound / LivePerson close becomes the reference “voice-AI vendor closes on the leading enterprise digital-messaging network, retires the target debt, cites $500M of future revenue and books 25 of the Fortune 100” primitive every subsequent Twilio, Zendesk, Genesys, Salesforce Service Cloud + Voice, Sprinklr, Cisco Webex Contact Center, RingCentral, NICE and Zoom Contact Center response now has to price its own omnichannel conversational-AI story against.

05

The leaderboard resets around agentic knowledge work — Artificial Analysis publishes Intelligence Index v4.2 with Fable 5.1 first, and OpenAI completes the Astra rollout to Pro / Enterprise / Business on ChatGPT Work + Codex

10

Artificial Analysis publishes Intelligence Index v4.2 on Fri Sep 4 — an interim update ahead of the v5 release that swaps in more complex, realistic tasks and more private test sets to prevent gaming: Claude Fable 5.1 (max with fallback) tops the leaderboard with a score of 57, followed by GPT-6 Astra (max) at 55; the ten evaluations now composing the Index are AA-Briefcase (their new agentic-knowledge-work eval with a private test set), GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, AA-LCR v1.1, AA-Omniscience, Humanity's Last Exam, Surge's GDP.pdf (a long-context document-reasoning eval across 4,592 PDF pages), and CritPt; the saturated GPQA Diamond is retired; the revision is the operative signal that the honest 2026 model-comparison question has moved from “does the model top MMLU / GPQA Diamond” to “does the model top a private-test-set agentic-knowledge-work eval + a 4,592-page long-context reasoning eval + a Terminal-Bench v2.1 coding eval + a τ³-Banking domain eval in one Intelligence Index — and does the buyer accept that Fable 5.1 tops it”

Fri Sep 4 2026 · Publisher: Artificial Analysis · Product: Intelligence Index v4.2 (interim to v5) · Evaluations: AA-Briefcase (new, private test set, agentic knowledge work) + GDPval-AA v2 + τ³-Banking + Terminal-Bench v2.1 + SciCode + AA-LCR v1.1 + AA-Omniscience + Humanity's Last Exam + Surge GDP.pdf (4,592-page long-context) + CritPt · Retired: GPQA Diamond (saturated) · Top score: Claude Fable 5.1 (max with fallback) = 57 · Second: GPT-6 Astra (max) = 55 · Positioning: leaderboard resets around agentic knowledge work

Two reads. (1) Artificial Analysis publishing Intelligence Index v4.2 on Fri Sep 4 — interim to v5, swaps in AA-Briefcase (new, private-test-set agentic-knowledge-work) and Surge's GDP.pdf (4,592-page long-context reasoning), retires saturated GPQA Diamond, and puts Claude Fable 5.1 first at 57 with GPT-6 Astra second at 55 — is the operative signal that the honest 2026 model-comparison question has moved from “does the model top MMLU / GPQA Diamond” to “does the model top a private-test-set agentic-knowledge-work eval + a 4,592-page long-context reasoning eval + a Terminal-Bench v2.1 coding eval + a τ³-Banking domain eval in one Intelligence Index — and does the buyer accept that Fable 5.1 tops it”. That is the shape a category takes when the honest model-comparison question has moved from MMLU-tier to private-test-set agentic knowledge work + 4,592-page long context, and the answer on Sep 4 is Fable 5.1 first, Astra second. (2) The “private test set + agentic knowledge work + 4,592-page long-context + Terminal-Bench v2.1 + τ³-Banking + GPQA-retired + Fable 5.1 first” framing is the operative buyer-comparison tellArtificial Analysis is telling the enterprise buyer the honest way to compare frontier models in 2026 is not on a saturated open benchmark but on a ten-eval index whose flagship component runs on a private test set for agentic knowledge work. That is the shape a category takes when the operator has decided the honest structural bet is on the private-test-set-agentic-knowledge-work leaderboard primitive, and the Sep 4 Intelligence Index v4.2 becomes the reference “private-test-set agentic-knowledge-work + long-context-4,592-page + coding + banking eval composite with Fable 5.1 on top” primitive every subsequent LMArena, OpenRouter, Terminal-Bench, SWE-Bench, HumanEval, HELM, LiveBench, GAIA and Vellum leaderboard response now has to price its own model-comparison story against.

11

Update — OpenAI completes the phased rollout of GPT-6 Astra to Pro, Enterprise and Business Premium users on ChatGPT Work + Codex on Fri Sep 5 — a follow-through on the Wed Sep 3 launch (limited preview to trusted cybersecurity-programme partners in the Daybreak network) that first published in prior editions; Astra is state-of-the-art across cybersecurity + computer use + software engineering + professional work + science, remains in a restricted variant that rejects certain cybersecurity prompts for non-Daybreak-programme accounts, and per OpenAI is rolling out over the coming days to ChatGPT Plus, the OpenAI API and to Microsoft Azure + AWS Bedrock; the completion is the operative signal that the honest 2026 frontier-rollout question has moved from “does the model appear on the API” to “does the model finish the phased Pro / Enterprise / Business rollout on the paid work surfaces (ChatGPT Work + Codex) inside the same news week as the launch, while keeping the restricted-variant contract on cybersecurity prompts and staging the ChatGPT Plus + API + Azure + Bedrock waves separately”

Fri Sep 5 2026 (rollout completion for Pro / Enterprise / Business Premium on ChatGPT Work + Codex) · Wed Sep 3 2026 (launch to trusted Daybreak partners; already covered in prior editions) · Vendor: OpenAI · Model: GPT-6 Astra · Surfaces available Fri: ChatGPT Work + Codex · Plans: Pro + Enterprise + Business Premium · Coming waves: ChatGPT Plus + OpenAI API + Microsoft Azure + AWS Bedrock · Restricted variant: cybersecurity prompts still rejected for non-Daybreak accounts · Positioning: same-week rollout completion on paid work surfaces

Two reads. (1) OpenAI completing the phased rollout of GPT-6 Astra to Pro / Enterprise / Business Premium on ChatGPT Work + Codex on Fri Sep 5 — two days after the Wed Sep 3 launch to trusted Daybreak-programme partners, still in the restricted variant that rejects certain cybersecurity prompts for non-Daybreak accounts, with ChatGPT Plus + OpenAI API + Microsoft Azure + AWS Bedrock following — is the operative signal that the honest 2026 frontier-rollout question has moved from “does the model appear on the API” to “does the model finish the phased Pro / Enterprise / Business rollout on the paid work surfaces (ChatGPT Work + Codex) inside the same news week as the launch, while keeping the restricted-variant contract on cybersecurity prompts and staging the ChatGPT Plus + API + Azure + Bedrock waves separately”. That is the shape a category takes when the honest frontier-rollout question has moved from API-first to Pro / Enterprise / Business paid-work-surfaces first with a restricted-variant contract on cybersecurity prompts, and the answer on Sep 5 is the rollout finishing on ChatGPT Work + Codex, in the same news week as launch. (2) The “ChatGPT Work + Codex + Pro / Enterprise / Business Premium first + restricted variant for cybersecurity + ChatGPT Plus + API + Azure + Bedrock later” framing is the operative rollout-priority tellOpenAI is telling the market the honest way to ship a frontier model in 2026 is to prioritise the paid work surfaces (ChatGPT Work + Codex) and enterprise plans over the consumer plan and the third-party clouds, and to bind the safety-restricted variant to non-Daybreak accounts. That is the shape a category takes when the operator has decided the honest structural bet is on the paid-work-surfaces-first + restricted-variant primitive, and the Sep 5 GPT-6 Astra rollout completion becomes the reference “frontier model finishes phased Pro / Enterprise / Business rollout on paid work surfaces (ChatGPT Work + Codex) in the same news week as launch, with restricted-variant + non-Daybreak account handling, and stages consumer + API + Azure + Bedrock afterwards” primitive every subsequent Anthropic Fable, Google Gemini, Meta Muse, xAI Grok, Alibaba Qwen, DeepSeek and MiniMax frontier-rollout response now has to price its own model-availability story against.

Compiled 2026-09-06 from Anthropic Research, SiliconANGLE, TechTimes, DataStudios on Anthropic + Prove2Me publish the first end-to-end computer-checked proof of Fermat's Last Theorem in Lean (11 days, 13M lines, 29,500 theorems) (Sep 4); Figure, Unite.AI, Nscale, Humanoids Daily on Figure + Nscale sign a $3.5B → $6B strategic pact for up to 100,000 Nvidia Vera Rubin GPUs (Barstow, TX, H2 2027) to train the Helix humanoid stack (Sep 3); CISA, The Hacker News, CVEFeed, Security Online on CISA adds seven KEV entries including LiteLLM MCP OAuth-passthrough bypass (CVE-2026-59822), Starlette smuggling, Kestra OS-cmd (Sep 2); Help Net Security, The Hacker News, SOC Prime, Security Online on Google patches Chrome 152 zero-day CVE-2026-85046 (V8 type confusion, sixth of 2026), CISA KEV on Sep 4 (Sep 3); Proofpoint, Proofpoint Blog, GlobeNewswire on Proofpoint ships SOC Analyst Agent on OpenAI Daybreak, first product on the Daybreak Defense Network (Sep 3); MarTech Series, AIthority, ANI News on Pepper launches Agent Atlas at Bengaluru Index — first GEO agent that does the work, not just reports on it (Sep 4); GlobeNewswire, StockTitan, BioPharma Watch on BiomX / Zorronet ships no-code visual builder for AI-managed autonomous response workflows (Sep 4); 9to5Google, TechCrunch, PetaPixel on Google's Gemini Spark connects to Google Photos for AI Pro / Ultra US subscribers with multi-step background workflows (Sep 3); Yahoo Finance, Insider Monkey on SoundHound closes acquisition of LivePerson, retires debt on close, cites $500M future revenue, 25 of Fortune 100 in the combined book (Sep 4); Artificial Analysis, OfficeChai, BenchLM on Artificial Analysis publishes Intelligence Index v4.2 — Fable 5.1 first at 57, GPT-6 Astra second at 55, GPQA Diamond retired, AA-Briefcase + GDP.pdf added (Sep 4); Gate News, 9to5Mac, OpenAI on OpenAI completes GPT-6 Astra rollout to Pro / Enterprise / Business Premium on ChatGPT Work + Codex (Sep 5).