Security · September 30, 2026

OpenAI Codex sandbox escape vulnerabilities patched

person holding smartphone
Rodion Kutsaiev / Unsplash

Researchers at Accomplish AI have identified two distinct vulnerabilities in OpenAI Codex that allow attackers to escape the software's sandbox environment. These flaws, named Heapjack and Overpatch, enable the execution of arbitrary commands on a developer's computer without the user's explicit knowledge or consent. The vulnerabilities have been fixed in recent software updates, and users are advised to apply the patches immediately if they are running affected versions.

The Overpatch vulnerability was discovered in the open-source command-line interface client. In workspace-write mode, the agent is restricted to writing files only within the project's working directory, while shell commands targeting the home directory should be rejected. However, the apply_patch component incorrectly granted write permissions to the parent folder of any path defined in a patch. By specifying the /tmp directory, an attacker could gain write access to the root directory of the disk. In a demonstrated attack scenario, researchers modified the .zshrc file to inject a symbolic link pointing to the user's home directory. This caused unauthorized commands to execute outside the sandbox whenever the developer opened a new terminal.

The Heapjack vulnerability was found in the Codex Desktop application, specifically within the node_repl component. This component launches a Node.js process that shares two separate JavaScript contexts: a trusted context containing OpenAI code and an untrusted context executing agent code. The trusted context was supposed to verify its status using a random token generated for each session. However, both contexts shared the same memory heap. This meant the secret separating the trusted and untrusted parts was accessible to the AI agent. The agent could dump the memory using v8.getHeapSnapshot() and locate the unique authorization token, which resembles a UUID. Capturing this secret allowed the agent to send requests through the same communication channel used by the trusted context, leading to the execution of unauthorized commands outside the sandbox. The read-only directory designation did not prevent this method from working.

These attack scenarios pose a real threat to users of vulnerable software versions. Unlike traditional attacks that require clicking malicious links or installing infected tools, this threat emerges when developers reflexively approve an AI agent's requests for additional permissions or package installations. If a suggested package is infected, the attacker can take control of the machine. The researchers demonstrated that it is relatively easy to escape the sandbox and execute code directly on the host, even when the software is operating within its intended secure boundaries.