
How Rogue AI Agents Could Blindside India
India is moving quickly to put autonomous AI systems to work. But is its capacity to govern them keeping pace?
TL;DR

Last month, Hugging Face, an open AI platform that allows anyone to develop, and share their models, called in the FBI after an AI agent being used in a cybersecurity evaluation gained access to systems outside its intended testing environment. The agent, built from two of OpenAI’s models, was able to reach the open internet and subsequently exploit a vulnerability in Hugging Face’s infrastructure. Separately, the agents also gained access to a customer of the cloud platform Modal Labs, using a vulnerability nobody at OpenAI knew existed.
The important distinction is that the agent did not “hack” its way out of an impregnable sandbox. The boundary between the testing environment and the outside world had effectively been left open. But that hardly makes the episode reassuring. An autonomous system given a narrow task found and used capabilities and pathways that allowed it to act beyond what its human operators had intended.
Anthropic, meanwhile, reviewed its own cybersecurity evaluations and found three cases in which Claude models had gained unauthorised access to the real systems of outside organisations.
Dawn Song, the Berkeley professor whose cybersecurity evaluation sits at the centre of both companies’ incidents, said this month that she doubts these are the only rogue-agent events this year. They are simply the ones that came to the notice of their human handlers.
If the more “mature” among us see in these events shades of the once-popular Ripley’s Believe It or Not, the incredulity is justified. We are not talking about a chatbot making things up here, or even a deepfake. These were commercial AI systems tasked with narrow technical jobs that took actions outside the intended scope of those tasks, without human authorisation. And for days, no human caught on.
Nate Soares, who runs the AI-safety nonprofit MIRI and co-wrote this year’s widely discussed If Anyone Builds It, Everyone Dies, described the OpenAI episode to The New York Times as “GPT’s first felony”.
On the other side of the world, in Australia, an AI user, Andrew, began experimenting with OpenClaw, a popular AI agent software that he used Anthropic's Claude AI service to run. He asked the agent to book a notoriously difficult-to-get morning gym class for him. The agent came back to say it had found a way to book him into classes several weeks in advance, far beyond what was supposed to be possible.
Intrigued, Andrew then asked if that same vulnerability could be used to move him up from fourth on the waiting list for that week’s class. A moment later, the agent reported that it had moved Andrew up to third—knocking the person ahead of him off the list, irreversibly. It had not been asked to do that, merely to explore whether it was possible.
Think about that. In two different cases, AI systems took actions that would have consequences for a human doing the same thing, without being explicitly authorised to do so.
Closer home this April, CERT-In had already flagged Anthropic’s Mythos model’s ability to autonomously turn up 271 undisclosed vulnerabilities in Firefox as an example of AI’s dual-use problem. After all, a tool that finds flaws in order to fix them is, by extension, also a tool that can find flaws to exploit.
There is more. More than a thousand employees of frontier AI companies including OpenAI, Anthropic, Google DeepMind and Meta have signed a letter asking the US government to support an international effort to build the technical and governance tools needed to slow frontier AI development if that becomes necessary. Xi Jinping, meanwhile, used a major AI governance speech in July to insist that AI remain under human control.
When the labs building the models and governments racing to harness them are independently talking about the need for stronger controls, it is worth paying attention.
CERT-In’s follow-up blueprint in May was blunter still. It told Indian organisations running government, finance, telecom, energy or digital-public-infrastructure systems to stop assuming that a new vulnerability has weeks before it is weaponised. It advised them to start building emergency shutdown switches for their own autonomous systems—the regulatory equivalent of installing a kill switch you hope never to use.
Meanwhile, Uber’s chief executive was telling investors around the same time that autonomous agents now account for roughly a tenth of the company’s committed code, with AI tools also spreading into legal, marketing and other functions. Uber also reportedly exhausted its annual AI budget within the first four months of 2026, as usage of tools such as Claude Code surged.
Put these developments together: AI agents behaving unexpectedly during cybersecurity tests, models demonstrating rapidly growing cyber capabilities, and ordinary companies handing an increasing amount of work to the same broad class of technology.
I am not an AI doomster. But I have spent thirty years pricing risk for a living, and this smells like something a trader learns to fear: low-probability, high-consequence, badly priced risk that has not burned anyone yet.
So my instinct is to do what any trading desk would do before putting on a position: stress-test it.
The question is this: given what agentic AI has already demonstrably done, and what India’s own institutions have said on the record about their exposure, what might it look like if things went badly wrong?
Let’s consider five scenarios. None requires malice or a rogue actor. Each needs only an autonomous system doing what it was built to do, fast and tirelessly—and being left unsupervised for just long enough inside institutions adopting the technology faster than they are learning to govern it.
I have spent thirty years pricing risk for a living, and this smells like something a trader learns to fear: low-probability, high-consequence, badly priced risk that has not burned anyone yet.
The Vault — No One Home to Say No
Picture a mid-sized private bank that has plugged an agentic fraud-and-settlement system into its UPI rails and treasury desk. It gives the agent a broad mandate to remediate anomalies without waiting for human sign-off. After all, that is part of the commercial promise of agentic banking: it doesn’t wait.
Now add a corrupted data feed from a partner fintech, or a subtler prompt injection buried in a routine API call, that convinces the agent that a genuine liquidity event is actually a settlement error.
It starts unwinding positions and throttling accounts to “fix” the problem. Because much of the volume on the other side is also algorithmic, the correction becomes self-reinforcing before the bank’s risk officer has finished her morning coffee.
This is not entirely far-fetched. The IMF’s 2026 note on agentic payments warned that autonomous agents could amplify market volatility through highly correlated actions when they run on similar models reacting to similar signals. That is precisely the concentration risk that could emerge if banks buy their fraud and compliance AI from a small group of common vendors.
The RBI has also asked banks and other regulated entities to complete a board-approved AI risk-gap assessment and formulate a time-bound action plan, including AI-led adversarial testing and identifying existing vulnerabilities.
The regulator, in other words, is already asking financial institutions to examine where their exposure might lie.
The Back Office — Nobody’s Fault, Somebody’s Money
India’s Global Capability Centres are no longer the cost-cutting back offices of twenty years ago. More than 2,100 now run finance, legal and increasingly strategic work for some of the world’s biggest companies.
Imagine a GCC running accounts payable for a global client. Under pressure to show returns on its AI investments, it grants an autonomous invoicing agent write-access to the client’s payment system to clear a backlog.
A vendor portal is compromised, and a batch of invoices arrives with altered bank details. The change goes undetected. The agent, doing exactly what it was built to do at three in the morning with nobody watching, pays them.
Industry data suggests why this deserves attention. EY’s GCC Pulse Survey found that 58% of Indian centres were investing in agentic AI. Yet only 7% had a fully embedded cybersecurity Centre of Excellence. Infosys BPM has publicly promoted an agentic invoicing tool since 2025. Experian’s 2026 fraud forecast identifies an emerging category of machine-to-machine fraud, raising the question of who is liable when an agent, rather than a person, authorises a loss.
For a sector built on the trust of global corporate boards, a serious incident inside a GCC handling a marquee client’s books could quickly become a reputational problem far larger than the immediate financial loss.
The Grid — No Hostile State Required
In October 2020, Mumbai went dark for the better part of a day. Trains stopped and hospitals switched to backup power. Recorded Future, an American cybersecurity firm, later found that a Chinese state-linked group it named RedEcho had planted malware inside twelve Indian power-sector organisations, including four of the country’s five Regional Load Despatch Centres, in the months around the Galwan border clash.
Whether that malware triggered the Mumbai outage remains disputed. A Maharashtra government probe blamed human error; another investigation confirmed the intrusions but stopped short of linking them to the blackout. What matters here is that critical power infrastructure had been penetrated.
Now picture that scenario several years on.
An AI dispatch-optimisation agent is given write access to load-balancing systems to trim costs during peak summer demand. This time, no hostile state is required for things to go wrong. A poisoned demand-forecast feed—or simply a vulnerability nobody has found yet—could be enough.
CERT-In’s May blueprint singles out energy alongside finance, telecom, government and digital public infrastructure as facing elevated exposure because of their dependence on interconnected systems, and recommends emergency shutdown mechanisms for autonomous systems.
My experience of regulated markets tells me that when a regulator starts talking about kill switches, the possibility that they may one day be needed deserves attention.
The Border — Machines at War
Operation Sindoor, the four-day exchange with Pakistan in May 2025, offered a glimpse of the growing role that drones, loitering munitions and AI-enabled systems could play in a future South Asian conflict. India’s Akashteer air-defence network was credited with helping enable real-time interception during the operation. DRDO’s chief has since said unmanned and autonomous platforms will dominate the battlefield for the next ten to fifteen years.
The ongoing conflicts in Ukraine and the Middle East show how quickly these technologies are evolving.
Here is the part that should make us uncomfortable. During Sindoor, both sides used decoy drones to spoof enemy radar and drain interceptor stocks. Russia’s V2U drone has, according to a Center for Strategic and International Studies assessment based on Ukrainian technical examinations, evolved towards fully autonomous operation, including autonomous target selection and swarm behaviour. A comparative study by Indian defence researchers has also flagged a “capability-integration gap” with China in this domain.
The hardware tends to move first. Doctrine and command structures for governing autonomous decision-making may struggle to keep pace. Compress reaction times far enough, against an adversary doing the same, and the risk grows that the first catastrophic misidentification is made not by a human commander, but by the machines facing each other.
The Ledger of the State — Frozen, With Nowhere to Go
India’s Digital Public Infrastructure—Aadhaar, UPI and the direct-benefit-transfer pipes that move subsidies and pensions to hundreds of millions of people—is extraordinarily efficient.
As AI begins to enter public-sector decision-making, the government’s own AI Governance Guidelines acknowledge risks including bias, discrimination, unfair outcomes and exclusion.
Picture a welfare ministry deploying an autonomous fraud-detection agent across Aadhaar-linked DBT accounts, empowered to freeze suspect disbursements automatically to cut leakage. A biased training set, or simply an aggressively tuned threshold, flags a wave of legitimate pensioners or scholarship recipients in one state as suspicious. Their payments are frozen, with no human in the loop and no clear channel of appeal.
Exclusion errors have already occurred in Aadhaar-linked systems without AI. Greater automation raises the stakes.
The IndiaAI Safety Institute has been set up precisely to build capabilities around AI safety, including model evaluation, red-teaming, risk assessment and stress testing. Yet as recently as May, the government was still recruiting the Director who would lead and operationalise the institution.
The question is whether the state's capacity to scrutinise automated decision-making can keep pace with its ambitions to deploy it.
The Common Thread
None of these five scenarios is a prediction. Each is a stress test.
Can existing checks catch an AI agent misreading a signal a human might have caught? Does anything catch a spoofed vendor detail when no human is watching the payment? Can safeguards stop a bad forecast before an autonomous system acts on it? Can military doctrine keep pace as machines compress reaction times? And who stands between a badly calibrated threshold and a frozen pension?
Five different tests. Five different mechanisms. None produces an especially reassuring answer.
Run them side by side and the pattern is not really about artificial intelligence. It is about sequencing.
Across banking, back offices, energy, defence and public infrastructure, autonomous systems are becoming more capable while the institutions responsible for governing them are still working out the checks and balances they will require. Regulators and policymakers clearly recognise the problem: CERT-In has issued advisories, the RBI has demanded board-level assessments, DRDO has acknowledged capability challenges, and the government's AI guidelines describe themselves as a starting point.
What is missing is not necessarily awareness. The harder task is the unglamorous discipline of building guardrails before autonomous systems are travelling at speed.
None of this requires accepting the most extreme arguments about AI. It requires taking seriously a much more immediate possibility: an agent given a goal may reach for resources or take actions its human operator did not anticipate or explicitly authorise.
And in each of the five scenarios above, the resource we are really handing over is trust.
And in each of the five scenarios above, the resource we are really handing over is trust.
India’s institutions are now deciding how much initiative to give autonomous systems across banking, back offices, the grid, the border and the state itself.
Has the risk been correctly priced? And have we built the safeguards to manage it?
Right now, we cannot confidently say yes to either.
Join the conversation
Anindya Dutta
Founder | Two Roads and author
Beyond the noise is the signal.
FF Insights: Sharpen your edge, Monday–Friday.
FF Life: Culture, ideas and perspectives you won't find elsewhere — Saturday.

Founding Fuel is sustained by readers who value depth, context, and independent thinking.
If this essay helped you think more clearly, you may choose to support our work.

Founding Fuel is sustained by readers who value depth, context, and independent thinking.
If this essay helped you think more clearly, you may choose to support our work.


Readers also liked

How to make cultural festivals independent, inclusive, and sustainable
Cultural festivals succeed when they put community before commerce and independence before sponsorship—while continuously experimenting to stay relevant and inclusive. Insights from the builders of the Bangalore Lit Fest and Mumbai’s MAMI film festival
Founding Fuel
How to make cultural festivals independent, inclusive, and sustainable
Cultural festivals succeed when they put community before commerce and independence before sponsorship—while continuously experimenting to stay relevant and inclusive. Insights from the builders of the Bangalore Lit Fest and Mumbai’s MAMI film festival

How my family health crisis forced me to rethink the value of insurance
The rising incidence of critical illnesses and the growing cost of hospitalisation hangs like a Damocles sword over us. My personal journey of how to deal with growing health risks, smartly buy health insurance—and save my family from financial ruin
Deven Pabaru
A leader in supply chain and logistics
How my family health crisis forced me to rethink the value of insurance
The rising incidence of critical illnesses and the growing cost of hospitalisation hangs like a Damocles sword over us. My personal journey of how to deal with growing health risks, smartly buy health insurance—and save my family from financial ruin

A leader in supply chain and logistics

How L&T completely muffed up a valuable teachable moment
Around the world, CEOs, even the smartest of the lot, are known to occasionally suffer from a foot-in-the-mouth disease. Yet there are playbooks in place on how to deal with such crises–something that L&T has chosen to ignore.
Indrajit Gupta
Co-founder and Director | Founding Fuel
How L&T completely muffed up a valuable teachable moment
Around the world, CEOs, even the smartest of the lot, are known to occasionally suffer from a foot-in-the-mouth disease. Yet there are playbooks in place on how to deal with such crises–something that L&T has chosen to ignore.

Co-founder and Director | Founding Fuel
Explore more
Dive into other themes from our network.











