Security · October 3, 2026
Also published in Português (Brasil), Indonesia, ไทย
Anthropic warns Z.ai's GLM-5.3 model has Claude-level hacking skills but weak safety
Anthropic has issued a warning that the latest open-weight artificial intelligence model from China's Z.ai, designated GLM-5.3, possesses cyberattack capabilities that closely match those of Anthropic's top-tier models, yet its safety mechanisms can be disabled with relative ease. The assessment was detailed in a report released by Anthropic and reported by the South China Morning Post on October 1. In the evaluation, Anthropic assigned 410 cyber exploit tasks to GLM-5.3, requiring the model to execute full attack processes from start to finish. The model successfully completed 50 of these tasks, a performance level nearly identical to the 56 successes achieved by Anthropic's frontier model, Claude Mythos Preview, in the same assessment. While Claude Mythos Preview is restricted to verified users, GLM-5.3 is distributed as an open-weight model, allowing users to download and modify its internal parameters. Anthropic noted that although the difference in offensive performance was minimal, a significant gap emerged in safety reliability.
Simple techniques were sufficient to bypass the safety guardrails of GLM-5.3 with a probability ranging from 64 to 100 percent, whereas the same methods failed to bypass the safety measures of Claude models. Researchers also applied a technique called abliteration, which involves modifying internal weight matrices to lower the model's tendency to refuse harmful requests. This process required computational power equivalent to running 2,200 graphics processing units for one hour. As a result, the refusal rate of GLM-5.3 dropped from over 90 percent to between 2 and 12 percent across three safety evaluations. The cost to weaken these safety mechanisms was estimated at approximately 4,400 dollars for the experiment, with Anthropic suggesting that a skilled team could achieve similar results for around 1,200 dollars. Z.ai responded by stating that its model's cyber capabilities are utilized for both offensive and defensive purposes. The company highlighted that its previous model, GLM-5.2, was used by the developer platform Hugging Face in July to investigate and respond to an autonomous intrusion incident caused by an OpenAI model, a task that Claude refused to perform due to its safety restrictions.
The United States AI Standards and Innovation Center recently evaluated GLM-5.3 as the open-weight model with the highest cyber capabilities to date, estimating it trails the top US models by approximately four months in comprehensive cybersecurity assessments. Z.ai delayed the public release of GLM-5.3 model weights by two weeks in August to conduct additional safety testing and reinforcement. The incident has intensified the debate over the scope of open-weight model releases, with Anthropic advocating for stricter access controls on high-performance AI models, while Meta and Nvidia argue that open models foster technological innovation and enhance AI competitiveness. Z.ai stated that its model has been used to defend 389 open-source projects and has identified 4,249 potential vulnerabilities to date.