AI Agent Digest: Week 37, 2026 - OpenAI Gives Away the Harness, Visa and Mastercard Agree to ID Your Agent, and 77% Claim an Inventory They Do Not Have

Share
AI Agent Digest: Week 37, 2026 - OpenAI Gives Away the Harness, Visa and Mastercard Agree to ID Your Agent, and 77% Claim an Inventory They Do Not Have

Eight stories this week, and a theme that snuck up on everyone: the industry spent the week arguing about identity, inventory and recourse, and barely mentioned capability at all. Nobody shipped a smarter agent. Several people shipped ways to find out whose agent it is.

1. OpenAI put the Codex harness behind one API call

The Agents API went into public beta on Wednesday. It hands developers the same managed harness that runs Codex: long-running sessions, orchestration, context compaction, recovery, sandboxed code execution, file editing, MCP connections, artifact generation and multi-agent delegation. There is no additional API fee beyond model tokens and paid tools. Partner sandboxes include Cloudflare, DigitalOcean, Modal, Oracle, Vercel, E2B, Daytona, Runloop and Blaxel, and you can bring your own.

Hot take: OpenAI just made the harness layer free, and a lot of companies whose entire product was "we manage sessions and context so you don't have to" woke up as a feature. That was always going to happen, because session management is plumbing and plumbing gets absorbed. The interesting question is what is left when the harness costs nothing, and the answer is what the agent knows about your business and what it is allowed to do. The moat was never the loop.

2. Visa and Mastercard agreed to check each other's agents

Ant International, Mastercard and Visa announced a Know-Your-Agent interoperability framework, aligning three protocols that were until now straightforwardly competitive: Visa's Trusted Agent Protocol, Mastercard's Verifiable Intent and Ant's Agentic Mobile Protocol. Each network keeps its own verification and decision-making, but agrees to recognise common trust signals so an agent vouched for in one ecosystem is legible in another. The number in the press release: agents are projected to orchestrate three to five trillion dollars of consumer commerce by 2030.

Hot take: Identity standards are settling well before capability standards do, and that keeps happening. A2A folded in with MCP last month, Okta gave agents an ID in Week 35, and now the card networks. The pattern is consistent and worth internalising: the industry can not agree on what an agent should be able to do, but it is converging fast on who has to answer for it. If you are picking infrastructure, weight identity and audit over feature lists.

3. 72% shop with an AI assistant, 23% trust it with the card

Visa's Trust Index, run by the Harris Poll across 2,065 US consumers, found 72% have used an AI assistant to shop and 23% trust AI to manage payments. When a payments brand is attached, the number moves: 61% said they would trust Visa to handle an agent-initiated transaction.

Hot take: That spread is the whole thesis of agentic commerce in one chart. People are not refusing to let agents buy things, they are refusing to do it without someone to call when it goes wrong. Trust here is not a feeling about accuracy, it is the existence of a chargeback. Build reversibility and you will get permission; build accuracy alone and you will not.

4. Google and Accenture are embedding 1,000 engineers in customers

The Accenture Gemini Enterprise Business Group launched Tuesday, built around a 1,000-person forward deployed engineer workforce sitting directly with clients, on top of Accenture's roughly 50,000 Google Cloud-skilled staff.

Hot take: Three weeks ago OpenAI could not sell an agent deployment without sending engineers. Now the largest cloud vendor and the largest systems integrator on earth are doing it at a thousand-person scale. Two instances make a pattern, and the pattern is an admission: the software does not survive contact with a real company without a human in the room for months. That is fine for the Fortune 500. It is a structural problem for everyone else, and it means the product you can actually adopt without a consultant is a different product, not a cheaper version of the same one.

5. An attacker stood up a credential-theft operation in under six hours

Google Threat Intelligence published findings that adversaries have moved off single-prompt tricks and onto agentic chains that plan, execute and iterate. One financially motivated actor assembled a multi-agent framework out of an AI coding assistant, a prompt and a set of agent instructions, and had a large-scale credential-testing operation running in under six hours, with agents handling scanning, credential testing, troubleshooting and IP rotation between them. A China-linked espionage group was observed building an agent that watches a target, reasons about options, then runs port scans and service analysis.

Hot take: Read Google's own caveat before you panic: they have not seen end-to-end autonomous zero-day discovery and exploitation against real targets. What they have seen is tooling assembly getting cheap, which is a different and more boring problem. The uncomfortable part is that the attacker used the same stack you use, because there is no separate evil version of any of this. Capability is symmetric. The only asymmetry you get to build is the boundary: scoped tools, real permission walls, and a stop that actually stops.

6. Agent security had a funding-and-product week

CrowdStrike shipped Falcon Guardian, which discovers known and shadow agents across Windows and macOS, traces a prompt through tool calls to downstream system actions, and blocks agents nobody approved. It landed alongside research counting 17,800 public AI add-ons across 6.7 million installations pulling instructions from unverified external sources, some impersonating major AI vendors. AIR Security launched the same week with $50 million to build an inline firewall that screens agent instructions.

Hot take: 17,800 add-ons taking instructions from somewhere you cannot see is not an edge case, it is the default install of the modern developer laptop. The reason this category is getting funded this fast is that the discovery problem came first and nobody solved it: you cannot police agents you have not enumerated. Expect every endpoint vendor to ship an agent inventory by Christmas, because it is the cheapest thing to build and the easiest thing to sell.

7. Zscaler gave its SOC agents actual job descriptions

Agentic SOC went global on Thursday, with distinct agents assigned to triage, root-cause investigation, verdicting and automated containment, wired into inline zero-trust telemetry.

Hot take: Skip the product and look at the org chart. Somebody sat down and decided which parts of a security analyst's job an agent owns outright, which parts it prepares for a human, and where containment needs a signature. That decomposition is the actual work, and it is the exercise every business owner is going to have to run on their own roles whether they buy anything or not. Security teams just got there first because the cost of getting the boundary wrong is legible to them.

8. 77% claim a complete agent inventory. 44% run discovery.

The number of the week, from a Harness survey of large organisations: 77% say they have a complete inventory of the agents running in their business, but only 44% operate active discovery tooling. Meanwhile 74% trust their testing to catch agent failures, and 19% have automated gates that actually block a bad release.

Hot take: Both gaps measure the same thing, which is confidence unsupported by machinery. Seventy-seven percent is not an inventory, it is a feeling about an inventory. And the testing number is worse: trusting your tests while having nothing that stops a failing one from shipping means the test is a document, not a control. If you take one action from this week's digest, go and check which of your own certainties has a system behind it. Most organisations are running governance by vibes and will keep getting away with it right up until the week they don't.

Briefly noted

GitHub expanded Copilot Workspace so separate agents handle implementation, testing and documentation on the same codebase, and OpenHands hit 1.0 with production sandboxing at roughly 68% on SWE-bench. Adobe acquired Rilo, pushing agents further into marketing automation. Gnani.ai launched a sovereign agentic stack for Indian enterprises built on a 30-billion-parameter model trained natively in 11 Indian languages. McKinsey's State of AI 2026 found 32% of organisations declined to buy a software product because agentic coding tools let them build it instead. And OpenAI's incident report to the European Commission over the wiki episode went in this week, which we covered in full on Tuesday.

What we're watching next week

Whether anyone builds something serious on the Agents API within days of it going free, because that will tell us how much of the agent startup landscape was harness and how much was product. Whether the KYA framework publishes an actual specification or stays a press release. And whether a second large survey replicates the Harness inventory gap, because if 77% against 44% holds up across samples, it is the governance story of the quarter.

Bottom line

Nothing this week made agents smarter. The Agents API made them cheaper to run, the card networks made them identifiable, the security vendors made them findable, and one survey quietly established that most companies do not know how many they have. That is what a market looks like when it stops asking whether the technology works and starts asking who is holding it.

Want to test the most advanced AI employees? Try it here: https://Geta.Team

Read more