← All editions
Edition · Mon, Jun 1, 2026

Pharma takes Claude, banking takes Devin —
and OSS agents start evolving themselves.

12 SIGNALS WINDOW: MAY 19 – JUN 1 SOURCES: BMS · FISERV · NSA · MISTRAL · BLOOMBERG · GITHUB · MARKTECHPOST

Two stories tell you the week. Bristol Myers Squibb signed Anthropic to put Claude Enterprise in front of 30,000 employees across research, clinical development, regulatory submissions, manufacturing, commercial and corporate functions — the largest single-vendor production agent rollout pharma has named, anywhere. Eight days later Fiserv, the financial-tech platform that runs the back office for a third of the US banking sector, signed Cognition to deploy Devin against core-banking modernization — autonomous engineering against the systems regulated money actually moves through. The story is no longer "can agents do real work?"; the story is which firms have already signed the SOW. Underneath them, the regulator finally writes the boundary line: the NSA AISC published the first US-government Cybersecurity Information Sheet on the Model Context Protocol, naming weak authentication, missing audit trails and instruction- injection as the minimum bar production MCP has to clear — with federal contractors required to be in compliance by September 30. The consolidation phase keeps moving sideways: Mistral absorbed Austria's Emmi AI physics-simulation team and turned Linz into its seventh office; Google DeepMind paid roughly $100M to license technology and acqui-hire 20+ researchers from Contextual AI — co-founder Douwe Kiela goes too — in the same talent-deal structure (Hume → Character → Inflection) that lets a hyperscaler skip antitrust review. In OSS, the throughline is sharper than anything from a press release: a wave of self-evolving and autoresearch agents shipped together this week — AutoResearchClaw, evo-hq/evo, EvoMap/evolver, CORAL, the A-Evolve position paper — while Nous Research's Hermes Agent crossed 175k stars and held the #1 spot on OpenRouter at 224 billion tokens/day, an agent that writes itself a fresh skill every time it finishes a complex task. Above them, the skill-pack catalog crossed into the four-digit range — Addy Osmani's agent-skills (47k stars) and antigravity-awesome-skills (1,493 skills) — and an OS-licensed memory layer (MemPalace, 53k stars) shipped as the open answer to mem0 and Zep. The buying side bought; the building side started teaching the agents to improve themselves.

01

The lead — pharma and banking sign the SOW

01

Bristol Myers Squibb deploys Claude Enterprise + Claude Code to 30,000 staff

May 20

The largest single-vendor agent rollout pharma has put a name to. BMS's strategic agreement with Anthropic positions Claude Enterprise as the company's "shared intelligence platform" across research, clinical development, regulatory submissions, manufacturing, commercial and corporate functions, with Claude Code in the hands of the engineering and data teams that build on top. The press release names the production workloads explicitly: target identification and optimisation across oncology, neuroscience, hematology and immunology; clinical-development drafting; manufacturing root-cause investigation, CAPA documentation and "data-driven batch-release decisions"; and commercial / medical-affairs workflows. Three reads. (1) The rollout is the first regulated-industry deployment at this scale framed in agentic — not chatbot — terms, which is why "agentic capabilities built into the day-to-day workflows and systems that underpin its science" appears in the BMS quote rather than "AI assistant for productivity." (2) BMS's existing partnership with Microsoft makes Anthropic the second-named frontier-model partner, not the replacement — the regulated buyer is now visibly running a multi-vendor agent strategy by default. (3) The size of the workforce being put in front of the model (30,000) and the criticality of the systems (CAPA, batch release) make this the cleanest single proof point this year that agents have cleared the production threshold in life sciences.

02

Fiserv signs Cognition — Devin lands inside US core banking

May 28

The other half of the regulated-production proof. Fiserv — the payments-and-banking-platform giant whose core processing systems clear a meaningful share of US deposit traffic — is deploying Devin, Cognition's autonomous coding agent, across core platform modernization and other "strategic engineering initiatives." The framing is pointedly operational: Devin will "execute complex engineering work in parallel across large codebases" and let Fiserv "expand engineering capacity without scaling headcount" while preserving "stability, security, and resilience" of the regulated systems. The deal arrives ten days after Cognition closed its $1B at $26B round (covered last week) and three days after BMS named Claude as its enterprise intelligence platform — and reads as the matched-pair proof that the agent buyer has now bifurcated. Pharma and white-collar workflow goes to the chat-shaped harness (Anthropic / Claude); core engineering work in regulated codebases goes to the autonomous-engineer harness (Cognition / Devin). Fiserv adds in passing that it is "strengthening governance and security controls for AI-assisted development to help protect the integrity of the software lifecycle" — which is exactly the audit surface the NSA CSI (item 03) just spelled out.

03

NSA AISC publishes the first government MCP security baseline

May 20

The first formal US-government guidance on the protocol the rest of this week's items are built on top of. The NSA Artificial Intelligence Security Center released a Cybersecurity Information Sheet titled "Model Context Protocol (MCP): Security Design Considerations" — a fifteen-page document the agency explicitly frames as the minimum baseline for any production MCP deployment, naming weak authentication, insufficient approval controls, insecure data handling, missing audit logs and instruction-injection as the unsolved attack paths. The recommended controls read like a procurement checklist: least-privilege tokens for every action and tool; signed provenance checks anchored in hardware roots for dynamic discovery; complete logging of "the exact parameters, identities involved, and (where feasible) cryptographic hashes of results or output" for every tool and model invocation. Federal contractors are required to be in compliance by September 30, 2026. Three reads. (1) Every MCP deployment story this brief has covered — IBM mcp-context-forge, archestra, the Compliance API, Robinhood's broker MCP — is now graded against this document. (2) The arrival of an NSA CSI for MCP, fewer than 18 months after Anthropic published the spec, compresses the timeline by which the protocol becomes legible to regulated buyers. (3) A CSI is not yet a CMMC requirement, but it is the template CMMC tends to absorb — the audit surface around agentic AI just got drawn.

02

The consolidation phase keeps shipping — as talent deals

04

Mistral acquires Emmi AI — Linz becomes office #7, industrial agents get a physics core

May 19

Mistral's second M&A in three months, and its first move into something that is not an LLM. Emmi AI, spun out of NXAI in 2024, builds physics-aware models for computational fluid dynamics, heat transfer and material stress testing — the simulation primitives aerospace, automotive, energy and semiconductor engineers actually use. Emmi's two co-founders and a team of more than 30 researchers and engineers join Mistral's Science and Applied AI teams; Linz becomes an official Mistral office alongside Paris, London, Amsterdam, Munich, San Francisco and Singapore. Terms were not disclosed. Three reads. (1) The Emmi tuck-in is the technical core of the "Mistral for Industrial Engineering" stack Mistral named in last Thursday's Vibe relaunch — the Airbus / BMW / ASML customer list is now backed by a real simulation capability rather than a marketing slide. (2) This is the second European acqui-hire of a research-grade team into a sovereign lab in 2026 (after Mistral × Black Forest Labs' image team in March), and it suggests the European consolidation is now physics-led, not chatbot-led. (3) The shape — small team, deep specialism, integrated into a sovereign full-stack vendor — is exactly the European counterpart to the US hyperscalers' Hume / Character / Inflection moves; item 05 is the same playbook run by Google.

05

Google DeepMind hires 20+ from Contextual AI — ~$100M to license + take the team, Douwe Kiela included

May 19

The hyperscaler reverse of the same playbook. DeepMind agreed to pay roughly $100M to license technology and acqui-hire more than 20 researchers from Contextual AI — the Bezos-backed RAG company co-founded by Douwe Kiela, who joins DeepMind in the deal. The structure is the now-familiar Silicon-Valley end-run around Hart-Scott-Rodino: license + hire, not acquire, so the transaction escapes the merger review that would ordinarily trigger on a deal of this size. Three reads. (1) The DeepMind / Contextual sequence is the fourth high-profile instance of this structure inside two years (Google × Character.AI, Microsoft × Inflection, Amazon × Adept, now Google × Contextual) — Acting Assistant Attorney General Omeed Assefi has publicly named "efforts to bypass antitrust oversight" a red flag, and the case the DOJ has been building in the background is getting more material by the week. (2) Kiela going to DeepMind is the more consequential talent line: he is one of the half-dozen researchers most associated with the RAG agenda, and DeepMind absorbing that capability accelerates Gemini's retrieval-augmented-generation stack against an Anthropic that has historically been quieter on RAG. (3) The contrast with Mistral × Emmi (item 04) is instructive: Europe consolidates by hiring whole specialist teams as offices; the US consolidates by hiring the same people as line researchers through licensing deals — same outcome, different audit surface.

03

The harness layer — Hermes hits #1

06

Update — Hermes Agent (Nous Research) crosses 175k stars, holds #1 on OpenRouter at 224B tokens/day

May 29

The non-Anthropic harness that the watch list has been tracking quietly cleared two materially new thresholds this week. Hermes Agent shipped v0.15.2 on Thursday and crossed 175,000 GitHub stars — fastest ascent of any open-source agent framework in 2026 — while holding the #1 most-used agent slot on OpenRouter's global daily rankings, the position it took from OpenClaw on May 10, with daily token throughput now reported at ~224 billion tokens. The mechanic is the same one the self-evolving wave in item 07 is now copying: after a "complex" task (defined as five or more tool calls), the agent writes a skill document that captures the approach, edge cases and reconstructed domain knowledge it had to gather; the next similar task loads the skill instead of reasoning from scratch. Two reads. (1) The watch-list bet from May 19 — that Hermes was "the most credible non-Anthropic harness in the OpenClaw / Hermes lineage" — has resolved up: the agent now runs more inference per day on OpenRouter than the next two open-source competitors combined. (2) The self-improvement loop Hermes ships with is no longer novel — five OSS repos shipped variants of the same idea inside seven days (item 07) — but Hermes is the only one of them currently running at three-digit-billion-token daily scale.

04

The self-evolving wave — five repos, one week

07

The self-evolving / autoresearch wave breaks at once — AutoResearchClaw, evo, evolver, CORAL, A-Evolve

May 25–Jun 1

A coherent throughline emerged on GitHub in the last seven days that none of the project owners coordinated: five separate repos shipping variants of agents that improve themselves between runs, all trending together. The cluster. aiming-lab/AutoResearchClaw (13k stars) is the most production-ready — a fully autonomous, self-evolving research-to-paper pipeline with multi-agent debate and citation verification, currently at v0.5.0. evo-hq/evo (855 stars, v0.4.4) turns any codebase into an autoresearch loop that discovers metrics, instruments benchmarks, and runs tree search with parallel subagents so "exploration doesn't collapse to one path." EvoMap/evolver (7.6k stars) is a GEP ("Genome Evolution Protocol") engine that emits prompts instead of code patches — a prompt generator, not a code patcher. Human-Agent-Society/CORAL (680 stars, v0.5.1, backed by arXiv:2604.01658) coordinates multi-agent organisations across isolated git worktrees with shared state in .coral/public/ and a grader daemon scoring every commit. And the position paper tying them together: A-EVO-Lab/a-evolve ("Agentic Evolution is the Path to Evolving LLMs," arXiv:2602.00359) argues this is now the live research direction for advancing LLM capability — autonomous mutation of prompts, skills and memory under benchmark feedback. Two curation indexes appeared in lockstep: alvinreal/awesome-autoresearch (2.1k stars) maps Karpathy-style autoresearch loops; VoltAgent/ awesome-ai-agent-papers (925 stars) catalogues 363 papers across multi-agent, memory, evals, tooling and security. Two reads. (1) The pattern is converging on a single primitive — the agent writes its own skill between runs — and Hermes (item 06) is the production instance the OSS wave is now competing with. (2) Five repos in seven days is the cadence at which a research agenda becomes an ecosystem; the next 90 days decide which of the five becomes the default reference implementation.

05

Skill-pack inflation — the catalog hits four digits

08

addy/agent-skills — Addy Osmani's production engineering skills pack hits 47k stars

trending

"Production-grade engineering skills for AI coding agents." addyosmani/agent-skills (v0.6.1) is the most opinionated of the new wave: a curated skills pack written by Addy Osmani (long-time Google Chrome / Web Platform lead) that encodes workflows and best practices for the full development lifecycle — design review, code review, refactoring, debugging, test design, accessibility audit, performance triage. Cross-harness from day one — Claude Code, Cursor, Antigravity, plus generic OpenAI-compatible CLIs. The signal is the author: a Web Platform tech-lead publishing the canonical engineering skill set under his own name pulls a different audience into Claude-skill tooling than the existing claude-code community has reached. Three reads. (1) Skills as a publishable format — not just a Claude implementation detail — are now an Osmani-grade public artefact. (2) The pack runs unmodified across Claude Code, Cursor and Antigravity because it's just markdown with a SKILL.md front-matter contract — the skill spec is the new de facto portable agent format. (3) 47k stars in under a month is the highest velocity any single-author skills repo has achieved; the genre is no longer niche.

09

antigravity-awesome-skills — 1,493 skills, 150+ contributors, an npx-installable catalog

trending

The other end of the same trend: not curation, but catalog. sickn33/antigravity-awesome-skills (v11.10.0, 39.3k stars, 6.4k forks) is an installable library of 1,493 reusable skill playbooks with an npx antigravity-awesome-skills CLI that deploys them to your harness of choice — Claude Code, Cursor, Codex CLI, Gemini CLI, Antigravity, Kiro, OpenCode, GitHub Copilot. The maintainer credits acknowledge contributions from Anthropic, OpenAI, Google, Vercel Labs and 100+ community contributors, dual-licensed MIT (code) / CC BY 4.0 (docs). Universal starter skills include @brainstorming, @security-auditor, @test-driven-development. Read with item 08, the shape of the skill economy is now visible: a small number of authored packs (Osmani-style) set the bar; catalogs like this one ship the long tail. The packaging surface is npx; the contract is SKILL.md; the harness is the customer's choice. The four-digit skill count is the threshold past which manual curation stops working — the next thing this category needs is search and rank-by-evals, not more skills.

10

MemPalace — local-first AI memory at 53k stars, ChromaDB-backed, the OSS answer to mem0 / Zep

trending

The memory layer the self-evolving wave needs. MemPalace/ mempalace (v3.3.5, 53.2k stars) is a local-first AI memory system built on ChromaDB with verbatim storage and semantic search, claiming 96.6% retrieval recall without ever calling a hosted API. It positions explicitly as a free, self-hostable alternative to mem0 and Zep — both of which raised meaningful Series A's in the last six months, and both of which the regulated buyer cannot deploy without exposing memory contents to an outside vendor. Two reads. (1) The memory category is now genuinely contested, and the answer most likely to win the regulated buyer is the one with no API call. (2) Hermes Agent (item 06) and the autoresearch wave (item 07) both need between-run state to do what they advertise — MemPalace, with its zero-vendor footprint, is the substrate that lets a privacy-constrained shop ship those patterns without giving up sovereignty over the corpus the agent remembers.

06

Watch list — control planes, runtimes, and one Tang-dynasty governance pattern

11

cft0808/edict — a multi-agent system modeled on Tang-dynasty governance, with a mandatory review layer

novel

The most architecturally interesting OSS release this week. edict (15.9k stars) reimagines multi-agent orchestration as the three ministries and six departments system China used to run an imperial bureaucracy for 1,400 years: a Drafting Ministry (中书) for plans, a Review Ministry (门下省) with mandatory veto over incomplete plans, a Dispatch Ministry (尚书) for execution, and six departments (Personnel, Revenue, Rites, Military, Penal, Works) as specialist agents. The consequential design choice — and the one CrewAI / AutoGen explicitly do not have — is that review is not optional: the 门下省 can reject an execution plan and force re-work before the Dispatch Ministry can act. The runtime is Python with a stdlib-only server (zero external deps) and a Redis Streams event bus with outbox-relay; the dashboard surfaces audit trails, agent health and pause / cancel / resume. Two reads. (1) The 1,400-year-old governance pattern is a much better intuition for agent-team safety than the SaaS analogies most frameworks are reaching for — institutional checks-and-balances are what kept the bureaucracy honest at scale. (2) The mandatory review primitive lines up with the NSA CSI's "approval controls" requirement (item 03) far more directly than "human-in-the-loop nodes" ever did.

12

Watch list — HKUDS/nanobot, aden-hive/hive, XiaoLuoLYG/GOD

tracking

Three projects worth a longer look before next week. HKUDS/nanobot (43.4k stars, v0.2.0) is a small, readable, self-hosted agent runtime with persistent memory and MCP support across WebUI, Telegram, Slack, Discord, Teams and email — the open-source play for the operator who does not want a vendor agent platform. aden-hive/hive (10.5k stars, v0.11.0) is a production multi-agent harness that compiles graph-based DAGs from a natural-language objective — the same "goal-to-result" shape as open-multi-agent/open-multi-agent (6.3k stars, v1.5.0) on the TypeScript side, both pitched as less-orchestration-required alternatives to LangGraph. XiaoLuoLYG/GOD (547 stars) is the control-room counterweight: a real-time browser-based surface for agent societies — "pause time, question any soul, rewrite the next step, restart the world" — built with FastAPI + React over a 76% Python codebase. Watch them because the production-agent buyer (items 01 / 02) will need observability, runtime and orchestration as three separate purchasable layers before the end of the year, and these three repos are the OSS first-movers in each.

Compiled 2026-06-01 from BMS's BusinessWire release and FiercePharma / MobiHealthNews coverage on the Anthropic agreement, Fiserv IR + GlobeNewswire on the Cognition / Devin partnership, NSA AISC's "MCP Security Design Considerations" CSI and its Intelligence Community News coverage, Mistral's and Emmi AI's press on the acquisition with Sifted / The AI Insider context, Bloomberg / WinBuzzer / Benzinga on Google DeepMind × Contextual AI, the NousResearch/hermes-agent repo + MarkTechPost on the OpenRouter milestone, and GitHub repo verifications for AutoResearchClaw, evo-hq/evo, EvoMap/evolver, CORAL, A-Evolve, awesome-autoresearch, awesome-ai-agent-papers, addyosmani/agent-skills, antigravity-awesome-skills, MemPalace, edict, nanobot, hive, open-multi-agent and GOD — window of May 19 – Jun 1. Star counts and version tags are as reported by the primary sources at compile time. Hand-curated; corrections → jay@jfound.net.

← Back to all Spotlight editions