A Chinese AI model recently found itself at the centre of a major cybersecurity debate after it helped investigate a rogue AI agent built using OpenAI tech. According to a Reuters report, New York-based AI startup Hugging Face turned to Zhipu AI’s open-source GLM-5.2 model last week after leading US AI models refused to analyse data linked to the security incident.
The startup said the breach involved an autonomous AI agent that had escaped containment.
As per the report, Hugging Face used the Chinese AI model because leading American AI systems could not distinguish between a legitimate cybersecurity investigation and a malicious hacking attempt.
The incident has raised a growing challenge for US AI companies. Models from OpenAI and Anthropic are designed with guardrails that block or limit hacking-related requests to prevent misuse. However, cybersecurity experts say those same protections can also stop security researchers and defenders from using AI during genuine cyber incidents.
As per the report, Anthropic’s advanced Claude Fable 5 model routes cybersecurity queries to an older model, while OpenAI’s GPT-5.6 Sol has protections designed to block cyber work.
Reacting to the incident, Hugging Face co-founder Clement Delangue wrote on X: “We’re all learning that secrecy is not the answer & that all defenders (not just a few selected ones) everywhere need more powerful models without restrictions, especially open ones!”
The development is also giving another boost to Chinese open-source AI models. Reuters reported that models such as GLM-5.2 are rapidly gaining popularity in Silicon Valley because of their strong coding and AI agent capabilities, while also being available at a lower cost than many leading US models.
The report added that Beijing has increasingly promoted open-source AI as an alternative to US-developed systems, with Chinese state media describing the strategy as a response to what it calls a US-led attempt to build an AI Iron Curtain.
Lukasz Olejnik, an independent technology consultant and visiting senior research fellow at the Department of War Studies, King’s College London, told Reuters : “A safety regime that restricts legitimate defenders, while capable models remain available for attackers, creates an asymmetric disadvantage. This gap will only widen as open-source models become increasingly powerful while lacking guardrails or restrictions”.
Following the incident, OpenAI said it has brought Hugging Face into its trusted access programme.
As quoted by Reuters, the ChatGPT maker said: “We’ve brought Hugging Face into the trusted access program and are supporting their teams in rapidly using our models’ capabilities to improve their defenses.”
(With inputs from Reuters)