Dario Amodei, CEO of artificial intelligence firm Anthropic, is championing a novel approach to ensure the safety and ethical development of advanced AI systems. He proposes the integration of independent external evaluators directly within AI companies, a mechanism he believes is crucial for maintaining accountability.
This initiative comes as the rapidly evolving AI landscape sparks increasing concerns about its societal impact, with Amodei suggesting that such oversight is vital to build public trust and manage the inherent risks of frontier technologies. The discussion also highlights the intricate web of connections between leading AI safety advocates, their companies, and the philosophical movement of “effective altruism.”
Amodei’s push for embedded oversight and METR’s role
Amodei has outlined a vision where accountability for AI development is secured through “embedded evaluators,” independent teams working closely with AI companies. These evaluators would have continuous, employee-like access to development processes, allowing them to verify adherence to safety protocols, report incidents, and assess the alignment of AI models, even during their training phases.
He frequently cites METR, an AI safety testing laboratory, as a prime example of an organization capable of fulfilling this critical role. METR explicitly states its mission is to evaluate cutting-edge AI models, aiding companies and society in comprehending AI capabilities and the potential dangers they present. This model aims to transcend traditional audits, offering a dynamic and transparent layer of scrutiny to a field often shrouded in proprietary secrecy.
Deep roots in effective altruism and AI safety
The proposed oversight model and key organizations like METR are deeply intertwined with the “effective altruism” (EA) movement. This philosophy emphasizes using evidence and careful reasoning to maximize positive impact through one’s time and resources. While METR’s public materials do not explicitly mention EA, its founders have openly framed their work through this lens.
Beth Barnes, METR’s founder and CEO, previously worked alongside Amodei at OpenAI during the early development of ChatGPT. She has articulated a vision for AI safety where dedicated professionals would assess models for existential risks, anticipate potential dangers, and identify early warning signs. Similarly, Paul Christiano, another METR founder who led research on model alignment at OpenAI, has also aligned his approach to AI safety with effective altruism principles.
- Beth Barnes, METR’s founder, envisions a role for evaluators to “decide if they’re going to kill us.”
- Paul Christiano, a co-founder, prioritized a “fundamentally good product” for the effective altruism movement.
These figures provide informal yet significant links to Anthropic, a company that itself emerged with substantial backing from prominent supporters of the effective altruism movement, showcasing a shared ideological foundation within this segment of the AI industry.
Anthropic’s funding and its altruistic links
Anthropic’s inception and growth were significantly bolstered by investments from individuals deeply committed to effective altruism. These financial ties highlight how the movement has influenced the trajectory of major AI research initiatives.
Most notably, Sam Bankman-Fried, the founder of the now-collapsed cryptocurrency exchange FTX and a convicted fraudster, led Anthropic’s 2022 Series B financing round. Before FTX’s downfall, Bankman-Fried was a high-profile advocate of effective altruism, publicly stating that the philosophy shaped his financial strategies and philanthropic endeavors. Another key investor was Skype co-founder Jaan Tallinn, who spearheaded Anthropic’s 2021 Series A financing. Tallinn is a prominent supporter of effective altruism, having spoken at its global conferences, co-founded influential organizations like the Centre for the Study of Existential Risk, and contributed significantly to AI safety research institutes.
Amodei’s journey from OpenAI to Anthropic’s safety mission
Dario Amodei’s journey to advocating for stringent AI oversight is rooted in years of experience at the forefront of AI development. Holding a Ph.D. in biophysics from Princeton, he transitioned from academic research to pioneering roles in technology. His career included significant stints at Baidu, where he focused on speech recognition, and Google Brain, where he began integrating safety considerations into neural-net research.
In 2016, Amodei joined OpenAI, eventually rising to vice president of research. During his tenure, he played a crucial role in developing early iterations of large language models, including ChatGPT 2 and ChatGPT 3. However, as OpenAI began to scale its models, Amodei grew increasingly concerned about the lack of robust safeguards. He ultimately departed in 2020, driven by a conviction that the company was not prioritizing safety sufficiently and that its financial interests might compromise its commitment to responsible development. This departure became a foundational moment for Anthropic, which was established with a strong emphasis on AI safety and ethical development.

