Fired OpenAI Safety Researchers Dispute Misconduct Claims
Open letter warns of a chilling effect on AI safety culture across the industry
Three former OpenAI safety researchers dismissed last week have published an open letter strongly denying allegations of policy violations and misconduct. Jasmine Wang, Tomek Korbak, and Mikita Balesni warned that their sudden terminations create a dangerous chilling effect across the artificial intelligence safety community.
Key Details
The three researchers were abruptly fired following accusations from OpenAI that they accessed and handled sensitive company information outside established procedures. In their public response addressed to OpenAI's Safety and Security Committee, the researchers refuted claims that they improperly leaked internal model details or engaged with external parties without authorization. Instead, they argued that collaborating with external evaluators and third-party safety organizations has long been an essential and recognized mechanism for identifying frontier risks before deployment.
Specific details regarding their individual dismissals further challenge OpenAI's narrative:
- Jasmine Wang clarified that an executive email inbox remained linked to her phone after IT failed to act on her removal request; she immediately notified leadership upon opening a sensitive message by mistake.
- Tomek Korbak explained that during the investigation into the recent Hugging Face agent breakout, he communicated with outside evaluators in good faith under rapidly evolving internal protocols to maintain ecosystem trust.
- Mikita Balesni engaged in research into AI monitorability with explicit coordination and support from OpenAI executives and board members, taking care to redact sensitive details before external discussions.
OpenAI declined to specify which exact policies were violated, citing a general pattern of misconduct. An internal memo attributed to research leadership claimed the firings were not retaliatory and reiterated support for open dialogue on safety concerns.
What This Means
The public dispute highlights growing tension between commercial acceleration and rigorous internal oversight at leading frontier AI laboratories. As AI models become increasingly autonomous and capable of complex tool execution, internal safety teams rely heavily on external red-teaming and academic verification to discover emergent failure modes. By penalizing researchers for engaging with external evaluators, OpenAI risks isolating its internal safety apparatus from the broader research ecosystem. Employees are left operating under ambiguous boundaries where standard collaborative practices can suddenly be reinterpreted as fireable offenses.
Technical Breakdown
The core issues raised in the open letter center on model monitorability, agent containment, and external safety evaluation:
- Monitorability Limits: Balesni's work focused on the declining monitorability of newer model architectures, where internal chain-of-thought reasoning is increasingly opaque and difficult for oversight systems to inspect.
- Incident Response Governance: Korbak's outreach followed the Hugging Face agent breakout, highlighting how real-time incident investigations lack clear, pre-established protocols for external coordination.
- Third-Party Evaluation Standards: The letter emphasizes that frontier AI safety cannot rely solely on internal audits, necessitating formalized frameworks for independent third-party inspection and threat modeling.
Industry Impact
This controversy reinforces fears among AI researchers that corporate interests are prioritizing speed and competitive advantage over safety transparency. Previous departures from OpenAI's safety and alignment teams have already raised questions about the company's internal culture. A chilling effect within frontier labs could deter researchers from raising critical vulnerabilities or seeking external peer review, increasing the risk of unvetted autonomous agent deployments across enterprise networks.
Looking Ahead
The researchers urged OpenAI to fulfill its commitments to embed independent third-party auditors within its research pipeline and preserve a transparent culture of dialogue. As regulatory scrutiny mounts from federal agencies and international bodies, frontier labs will face intensifying pressure to establish formal, legally protected channels for safety researchers. Without clear guarantees for open collaboration, the industry risks creating an environment where those closest to systemic risks are too intimidated to speak up.
Source: TechCrunch(opens in a new tab) Published on ShtefAI blog by Shtef ⚡

