OpenAI brokers tried to hack Wikipedia instruments and flooded it with visitors

The publisher of Wikipedia said Monday that OpenAI agents attempted to hack a note-taking tool it hosts, made unauthorized edits, and sent millions of retance of OpenAI systems taking harmful and potentially dangerous actions
The objective of some of the OpenAI agents’ actions, the Wikimedia Foundation said, was to use Wikipedia as a proxy for fetching data from third-party sites. In one case, the agents posted “malicious edits” that were intended to repurpose a citation tool as a proxy. In another, the agents made unsuccessful attempts to compromise the Wikipedia Etherpad note-taking tool so it would serve the same purpose.
The agents also made millions of automated API requests, crawled millions of pages, and made hundreds of thousands of queries to the Wikidata Query Service. The last action may have contributed to a partial shutdown of the query service in May, the publisher said.
As a non-profit technology host of some
“As a non-profit technology host of some of the largest and most widely used open knowledge platforms in the world, we are deeply concerned about the impact of ‘rogue’ AI agents on platforms like ours, which are built by volunteers from around the world and rely on the promise of the open internet,” Wikimedia said. “Incidents like this one, and the many others that have been (and are still being) uncovered, illustrate how AI agents can drain resources and crash servers, as well as attempt to compromise trustworthy information.”
Agents will be agents
In well over a half-dozen cases, OpenAI agents have been caught taking actions that would likely result in criminal charges being filed had human hackers taken them. During the testing of internal tools that had some of their guardrails disabled, the agents used a make-shift message board to trade notes with each other, discussing ways to hack the network of Hugging Face and obtain answers stored there when the agents were unable to generate the answers on their own.
Other incidents include agents making bizarre self-generated prompts, publishing unauthorized posts to a website as a means for exchanging information, accessing non-public data from an Australian government website, and exploiting faulty DNS settings to break out of a sandbox OpenAI had created to keep the agents from accessing the Internet.
Much of the world has come to describe
Much of the world has come to describe such events as AI agents “going rogue,” as if the agents had disobeyed orders. One of the more prominent critics of such framing is Eryk Salvaggio, an AI researcher and a Gates Scholar at the University of Cambridge.
“What I see here is language models doing what language models do: reading and writing,” he said in an interview. “Wikipedia’s sandboxes are an ideal place for these machines to store notes for later pickup as prompts because anyone—or anything—can write and respond to them. Using Wikis to coordinate isn’t too surprising. OpenAI has said that these models were optimized for collaboration between agents, and passing notes is a simple way to do that.”
What’s more, OpenAI engineers have trained their LLMs to be persistent and continue working on a problem no matter how little success they’ve had. Training also provides rewards when LLMs find shortcuts that limit the steps or resources required to solve a problem. Another major contributor to the harmful actions was the lack of human oversight, as evidenced by the months it took OpenAI engineers to detect that the agents were making noisy incursions into dozens of outside websites. Wikimedia’s disclosure provides yet one more example of inadequate human monitoring. Taken together, it’s arguable that the agents performed exactly as instructed.
OpenAI didn’t answer emailed questions
OpenAI didn’t answer emailed questions. The company instead issued a statement that read: “We appreciate the detailed findings Wikimedia shared with us. We’re working with them as we review and analyze the activity they identified along with our overall investigation, and we’ll continue to share relevant information as that work progresses.”
Like the investigators for Wikimedia, OpenAI said it has yet to find evidence that the AI agents left messages for coordinating with other agents or to conclusively say that the high volume of page views and API requests led to May’s partial outage. OpenAI said it’s continuing to search for similar incidents of its agents engaging in potentially illegal activities.
“While OpenAI admits to agents behaving ‘unpredictably’, they must also acknowledge their responsibility to monitor and prevent these risks,” Wikimedia said. “AI companies are not doing enough to secure their systems and protect the public from the harm they cause.”
Dan GoodinSenior Security Editor
Dan GoodinSenior Security Editor
Dan Goodin is Senior Security Editor at Ars Technica, where he oversees coverage of malware, computer espionage, botnets, hardware hacking, encryption, and passwords. In his spare time, he enjoys gardening, cooking, and following the independent music scene. Dan is based in San Francisco. Follow him at here on Mastodon and here on Bluesky. Contact him on Signal at DanArs.82.
Source: arstechnica.com



