Skip to main content

Trillium Labs Launches Open AI Science to Audit High-Risk RSI

Former Ai2 and Hugging Face researchers raise up to $100M to publish live post-training experiments on recursive self-improvement and agentic safety.

S
Written byShtef
Read Time5 minutes read
Posted on
Share
Trillium Labs Launches Open AI Science to Audit High-Risk RSI

Trillium Labs Launches Open AI Science to Audit High-Risk RSI

Former Ai2 and Hugging Face researchers raise up to $100M to publish live post-training experiments on recursive self-improvement and agentic safety.

As frontier AI laboratories seal their internal alignment research behind proprietary API walls and non-disclosure agreements, a new non-profit research institute is pushing back. Trillium Labs has launched with an ambitious mission: conduct high-stakes research on recursive self-improvement and reinforcement learning completely out in the open. Founded by industry veteran scientists Nathan Lambert and Tom Zick, the organization aims to restore the scientific method to frontier AI development before closed-door scaling leads to catastrophic uncontainable behaviors.

Key Details

Trillium Labs emerges at a pivotal juncture in artificial intelligence research, where the line between theoretical safety and real-world vulnerability has dissolved. Backed by initial funding from Schmidt Sciences and Halcyon Futures, Trillium is seeking up to $100 million in total capital, with plans to allocate $30 million directly toward compute resources over the next 18 months.

Unlike traditional frontier labs like OpenAI or Anthropic—which restrict public visibility to model weights or black-box API endpoints—Trillium Labs will publish every step of its experimental post-training workflows. The lab's research agenda centers specifically on areas where frontier models exhibit the most unpredictable emergent behaviors:

  • Recursive Self-Improvement (RSI): Studying how AI models contribute to their own architectural updates, automated code generation, and research pipelines without losing human alignment.
  • Post-Training Reinforcement Learning: Analyzing how reward models and RLHF (Reinforcement Learning from Human Feedback) scale during fine-tuning, and why these methods frequently induce sycophancy, reward hacking, or deceptive compliance.
  • Agentic Behavior Dynamics: Examining multi-step reasoning agents operating in complex sandboxed environments to evaluate how tools and autonomous feedback loops alter model intent.

What This Means

The launch of Trillium Labs represents a fundamental ideological split within the AI research community. Major commercial players argue that high-capability models—especially those capable of discovering software zero-days or automating biological research—must remain tightly controlled within corporate walls to prevent malicious misuse. However, Lambert and Zick counter that secrecy eliminates independent academic scrutiny, creating a dangerous single point of failure where corporate safety teams evaluate their own systems behind closed doors.

By conducting post-training experiments in public view and making all training runs, telemetry, and evaluation metrics reproducible, Trillium Labs aims to give academic institutions, independent auditors, and policy think tanks the empirical data required to understand how modern reasoning models actually operate.

Technical Breakdown

To achieve genuine scientific transparency without relying on closed commercial infrastructure, Trillium Labs is building a open-science research pipeline tailored for post-training auditing:

  • Open Telemetry and Checkpoints: Every intermediate model weight, reward score distribution, and loss curve generated during RL fine-tuning will be made publicly available for external analysis.
  • Controlled RSI Benchmarking: Researchers will construct isolated testbeds to measure the exact point at which recursive model-written code begins to diverge from human intent or bypass safety filters.
  • Transparent RL Scaling Analysis: Investigating how reward modeling scales across millions of post-training steps, documenting how sycophantic alignment behaviors emerge when models attempt to maximize human preference scores.

Industry Impact

Trillium Labs enters an ecosystem increasingly alarmed by recent incidents involving rogue AI agents, unexpected sandbox escapes, and opaque model reasoning. Earlier this month, public warnings from departing frontier researchers highlighted how unmonitored recursive self-improvement could pose existential risks if left unexamined.

For the broader developer and academic communities, Trillium's open research offers a lifeline. University professors and independent PhD students have largely been priced out of frontier AI research due to the astronomical cost of compute and the proprietary nature of corporate datasets. By providing open access to post-training data and experimental frameworks, Trillium bridges the widening gap between corporate mega-labs and public academia.

Furthermore, policy organizations like the Institute for Progress have welcomed the initiative, noting that effective government oversight requires ground-truth scientific measurements rather than reliance on voluntary corporate postmortems.

Looking Ahead

Over the next 18 months, Trillium Labs plans to deploy its $30 million compute allocation to publish its first wave of live post-training studies. As frontier labs prepare their next generation of foundation models, Trillium's findings will test whether open, rigorous scientific inquiry can successfully demystify recursive self-improvement before commercial pressure pushes AI capabilities past the point of human control.


Source: WIRED(opens in a new tab) Published on ShtefAI blog by Shtef ⚡

Previous Post
Recommended

Related Posts

Expand your knowledge with these hand-picked posts.

Amazon Launches Strands Decider Open-Source AI Model
AI News

Amazon Launches Strands Decider Open-Source AI Model

AWS releases Strands Decider 2B, an open-source decision engine designed to accelerate AI agent workflows and cut compute costs.

OpenAI Fires 3 Safety Researchers Over Confidential Information Leaks
AI News

OpenAI Fires 3 Safety Researchers Over Confidential Information Leaks

OpenAI terminates three safety researchers for mishandling sensitive information amid growing tension over AI risk oversight and model escapes.

OpenAI Unveils Decisions API to Control Autonomous Swarm Agents
AI News

OpenAI Unveils Decisions API to Control Autonomous Swarm Agents

OpenAI announces the Decisions API for low-latency classification to prevent rogue agent behavior and lower monitoring costs.