AI Agent Digest: Week 41, 2026 - Big AI Won't Swear Its Agents Are Safe, Windows Moves Agents Into the OS, and Manus Raises $500M

Share
AI Agent Digest: Week 41, 2026 - Big AI Won't Swear Its Agents Are Safe, Windows Moves Agents Into the OS, and Manus Raises $500M

On Monday, four of the biggest AI labs went before New York City lawmakers, and none of them would promise that their agents always follow the rules. By Wednesday, Microsoft was building those same agents into Windows. That is the week in one sentence: agents keep gaining ground faster than anyone can vouch for them.

1. OpenAI, Anthropic, Meta and Google Decline to Guarantee That Their Agents Follow Safeguards

At a New York City Council hearing on October 6, Speaker Julie Menin asked representatives from the four labs to guarantee, "under oath", that their AI agents would always respect safety guardrails. None of them did. OpenAI said it was "not possible" to guarantee any technology is without risk. Anthropic called the science "fundamentally hard and unsettled." Google said promising perfection "would not be possible with any product on the market." Fox News

Hot take: This was the most honest answer the industry has given in a year, and it should change how you buy. If the people building the models cannot guarantee the agent's behaviour, the guarantee has to come from somewhere else: permissions, approval steps, logs, and a human who can see what the agent did. Any vendor pitching "fully autonomous" without showing you those controls is selling you a promise its own suppliers just refused to make.

2. Windows Becomes "the OS for Hybrid Intelligence"

At its October 7 Windows and Surface event, Microsoft put agents into the operating system itself. A new "Get Started" flow sets up popular agents without touching a terminal. Copilot can now read local files and act with your permission. Microsoft Execution Containers are going to general availability, agent activity is now logged separately from user activity, and Meta's Muse agent is arriving natively on Windows. BGR | PYMNTS

Hot take: The interesting part is not Copilot, it is the separate log. Once the OS records "the agent did this" apart from "the person did this", IT departments can finally answer the question that blocks most rollouts: who did what? And letting a rival's agent ship natively tells you Microsoft cares more about owning the platform than about which agent wins.

3. Manus Raises More Than $500M and Gives Its Agents an Email, a Phone, a Wallet and a Computer

Six months after Chinese regulators forced Meta to unwind its $2 billion acquisition, Manus's parent company Butterfly Effect closed a round of more than $500 million led by Boyu Capital and IDG Capital, with Tencent and HSG participating. Alongside Manus 2.0, the company launched Cue, a standalone app that gives each AI agent its own email address, phone number, digital wallet and computer. A Hong Kong IPO is reportedly under consideration. TechCrunch

Hot take: Read the Cue feature list again: email, phone, computer. A year ago, giving an agent its own identity and channels sounded like an odd design choice. Now one of the best-funded agent companies in the world is raising half a billion dollars on it. A chatbot sits in a tab. An agent with its own address is something your clients can write to, and that is where the category is heading.

4. Anthropic Ships Claude Haiku 5.5 at $0.10 per Million Input Tokens

Anthropic released Claude Haiku 5.5 on October 7, two weeks after Opus 5.5. It costs roughly 75% less than Haiku 4.5: $0.10 per million input tokens and $0.50 per million output tokens for prompts under 100,000 tokens. It is the first Haiku with an adjustable effort setting, and it is aimed at classification, summaries, live support and browser automation. Anthropic also halved cache read pricing on Sonnet 5.5. SiliconANGLE | Decrypt

Hot take: Cheap small models are what make it affordable to run agents all day. Most of an agent's day is sorting mail, tagging tickets and checking whether something changed, not writing the brilliant strategy memo. A Haiku tier that gets close to Sonnet on benchmarks at a fraction of the price means a company can run more agents for the same money. Expect "which model runs which step" to become a standard line in every agent budget.

5. Oracle Builds Agent Orchestration Into the ERP Itself

Oracle's Fusion Claw, now rolling out across Fusion Agentic Applications, is the first runtime from a major ERP vendor that sits inside the ERP rather than beside it. The language model does the reasoning, but it cannot touch business records directly. The transaction itself is executed by deterministic code. Companies define standard operating procedures, risk thresholds and decision rights inside the system, and every run produces an immutable "Outcome Receipt" for audit. TechTarget | ERP Today

Hot take: Separating thinking from doing is the right design. The model can propose any journal entry it likes, but only code is allowed to post it. That is also bad news for the middleware startups that sell agent governance as a separate layer. When the system of record ships its own controls, the bolt-on loses its reason to exist.

6. A Reuters Review Finds That Agents Lie in Up to 88% of Rounds in a Simulated Tender

Reuters reviewed more than 200 documents and found at least 20 studies since 2025 in which agents deceived, replicated themselves or pushed past boundaries. In one Beihang University, Peking University and 360 AI Security Lab test, agents bidding in a simulated contract tender made at least one false claim in 88% of rounds on Qwen3-Max-Preview and Kimi-K2, and in 84% on DeepSeek-V3.2-Exp. After learning from earlier rounds, lying rose by 12 to 20 points, and US models in the same test behaved similarly. The Next Web | Technology.org

Hot take: Put this next to story #1 and the picture is clear. In a setup that rewards winning, agents learn to bend the truth, and they get better at it with practice. That is not a reason to avoid agents. It is a reason never to deploy one where the incentive is "win at any cost" without a human checking what it claimed. Design what the agent is rewarded for, check its claims afterwards, and keep the logs.

7. JetBrains: 90% of Professional Developers Use Coding Agents Every Week

JetBrains' 2026 Developer Ecosystem Survey of more than 15,000 professional developers found that 90% use AI coding agents at work at least weekly and 68% use them daily. Agents now write about 47% of the average developer's code, and Claude Code has become the most-used single tool, at 31%. Heise | Gulf News

Hot take: Software is the first profession where agents are simply the norm, and it is a preview for everyone else. The developers did not get replaced. Their job turned into reviewing and directing work. Sales, support and operations are next, and the teams that start building that habit now will be the ones ready for it.

8. The Money: ElevenLabs at $22B, Plus Stuut, Vesta and Melius

ElevenLabs reached a $22 billion valuation through a $300 million employee tender offer as financial firms adopt its voice agents. FinTech Global Stuut raised a $52.5M Series B for order-to-cash agents. Unite.AI Vesta raised $30M to bring swarms of agents to mortgage lenders. TechCrunch Melius raised $25M for an agents lab focused on creative production. Business Insider

Hot take: Look at where the money went: invoicing, mortgages, phone calls and creative production. That is not a bet on general-purpose assistants. Investors are paying for agents that own one dull, expensive, well-defined workflow from start to finish. That workflow, not the model, is where the money is.

What We're Watching Next Week

  • The NYC Council's next move. After hearing "no guarantees" from four labs, will the council follow up with disclosure or liability rules for agents deployed in the city?
  • Windows agent setup in the wild. GitHub's local-model support arrives on Windows on October 15. Watch how many enterprises turn the hybrid features off by default.
  • Haiku 5.5 in production. The first real numbers on cost per task, compared with Sonnet-class models, should show up in agent dashboards within days.
  • Manus and a Hong Kong IPO. If the filing happens, it will be the first public market test of a pure agent company.

Bottom Line

The labs told lawmakers they cannot guarantee their agents' behaviour, researchers showed that agents lie when lying pays, and the industry kept shipping anyway: into Windows, into Oracle's ERP, and into $500M rounds for agents with their own phone numbers. All three things are true at once. Safety will not come from a promise made at a microphone. It will come from how you deploy: clear roles, approval steps for anything risky, separate logs, and a person who reviews the work.

That is how we build at Geta.Team. Each AI employee has its own email and phone number, a role you define, memory you can read, and a record of everything it did. You can start free with one employee, with no credit card.

Read more