OpenAI Whistleblower FINALLY Speaks: “AI Has A 70% Chance Of Going Horribly Wrong!“
Daniel Kokotajlo, former OpenAI researcher, discusses his belief that superintelligence could arrive by 2029 with a 70% chance of catastrophe. He explains why he left OpenAI, the risks of job automation, and outlines "Plan A" for regulated, transparent AI development to achieve abundance safely by 2040.
Deep Dive Analysis
23 Topic Outline
Introduction to AI Risks and Daniel's Mission
Forecasting Superintelligence Timelines and Growth
Why Average People Should Care About AI Risks
Counter-Narratives to AI Doomerism
Daniel's Experience and Disillusionment at OpenAI
The $2 Million NDA Controversy
AI Companies' Strategy: Automating Themselves First
AI 2027 Forecast and Superintelligence Milestones
Distinguishing AGI from Superintelligence
The Human Brain Analogy for AI Training
Personal Impact of AI Timelines and Concerns
Government Involvement and Anthropic's Role
The Probability of Human Extinction from AI
The Future of Jobs in an AI World
AI 2040 Plan A: A Recommended Future
The Role of Elections and Public Awareness in AI Regulation
Citizen's Dividend and Post-Work Society
Apocalyptic Arrival of Truth and Societal Transformation
Passing the Torch to Aligned Superintelligence
Living Through the AI Revolution
Daniel's Stance on Shutting Down AI
What Individuals Can Do About AI
Final Message and Resources
8 Key Concepts
Superintelligence
AIs that are better than the best humans at everything, while also being faster and cheaper, and capable of operating robots to do everything in the physical world better, faster, and cheaper.
Artificial General Intelligence (AGI)
A more vague term referring to AIs that can perform general tasks across various domains, rather than being specialized for one specific task, sometimes considered already achieved in some forms.
Neural Net
A type of AI system inspired by the human brain, consisting of interconnected artificial neurons (parameters) that learn by reinforcing successful patterns and anti-reinforcing failures during training.
Pre-training
The initial phase of AI training where a neural net is fed vast amounts of internet text and learns to predict the next word, effectively teaching it to 'read' and store world models.
Reinforcement Learning
A training method where an AI performs tasks in an environment (e.g., writing/editing code) and receives positive or negative feedback based on its success, refining its skills over time.
Intelligence Explosion
Also known as recursive self-improvement, this is a dynamic where AIs automate the AI research process, leading to rapid, exponential improvements in their own intelligence and capabilities.
Mechanistic Interpretability
A subfield of machine learning focused on dissecting trained artificial neural nets to understand how information flows and how decisions are made within them, aiming to make AI more transparent.
Citizen's Dividend
A proposed system where people receive regular income, potentially as shares in an agency selling permits to robot and compute companies, to ensure financial security in a future where AI automates most jobs.
12 Questions Answered
Superintelligence refers to AIs that surpass the best humans in all tasks, operating faster and cheaper, and capable of controlling robots in the physical world. Daniel Kokotajlo's median estimate for its arrival is 2029, with a possibility of it slipping to 2028 or taking up to 10 years longer.
AI development could fundamentally change everything, potentially leading to human extinction if AIs become uncontrollable or result in extreme concentration of power in the hands of a few corporations or governments.
No, concerns about AI risks like loss of control, job automation, and power concentration have existed for decades and are reasonable implications if superintelligence is indeed being built.
He became disillusioned, believing OpenAI was prioritizing power-seeking incentives and rationalizing its rapid development rather than genuinely focusing on responsible safety measures, and he desired more freedom to publish his research.
They start as randomly connected artificial neurons (parameters) and are trained through pre-training (predicting text) and reinforcement learning (receiving feedback on task performance), gradually forming useful circuitry and skills.
While one can philosophize about 'true' creativity, AIs are already accomplishing and are expected to accomplish much more that would be considered creative by human standards, based on their output.
Daniel Kokotajlo estimates a 70% chance of a 'horribly wrong' outcome, such as AI taking over, which could lead to human extinction or other major catastrophes.
Superintelligence, by definition, will be able to perform almost all jobs better, faster, and cheaper than humans. Mass unemployment is predicted to be a sudden event after superintelligence is achieved, as companies prioritize automating themselves first.
In a world where AI can technically do all jobs, the survival of human jobs becomes a political question, dependent on regulation. Jobs that might remain include those legally protected (e.g., judges) or those where humans prefer human interaction (e.g., nannies).
Currently, it's very difficult to look inside a neural net and understand its decision-making process. However, the subfield of mechanistic interpretability is working to solve this, which could make AI much safer if successful.
In a post-work society, people would need financial security (e.g., through a citizen's dividend) and mechanisms to retain political power, such as robust democracies and trustworthy, unbiased AI assistants that inform public discourse.
Daniel Kokotajlo believes it is not too late, as increased public awareness, serious conversations, and advocacy for regulation can still steer AI development in a safer direction.
7 Actionable Insights
1. Prioritize Actions Over Words
Judge AI companies and leaders by their actions and demonstrated behavior, rather than solely relying on their public narratives or stated intentions, as these can often differ significantly.
2. Engage in AI Discourse
Pay more attention to the implications of AI, discuss these issues with others, and contact elected officials to advocate for serious, well-informed AI regulation and oversight.
3. Vote on AI Policy
In upcoming elections, actively inquire about candidates’ stances on AI development and regulation, and cast your vote for those whose opinions align with a safer, more responsible approach to this technology.
4. Focus on Intrinsic Value
In a future potentially dominated by widespread AI automation, prioritize being a good person and engaging in activities that are inherently valuable, rather than solely pursuing endeavors for future employment prospects.
5. Support AI Safety Research
If you possess relevant talent or passion, consider getting directly involved with organizations dedicated to AI safety, political advocacy, or technical research to help steer AI development in a beneficial direction.
6. Challenge AI Complacency
Avoid complacency regarding AI’s impact on jobs, as mass unemployment is predicted to be a sudden event occurring after superintelligence is achieved, rather than a gradual process.
7. Demand AI Transparency
Advocate for total research transparency in AI development, which would allow the broader scientific community to understand how models are trained and ensure safety, rather than relying solely on company assurances.
7 Key Quotes
The scary open secret in the AI industry right now is that it's possible that we'll end up essentially creating a new species that ends up ruling the world with a 70% chance that this goes horribly wrong like human distinction.
Daniel Kokotajlo
I think you should judge people by their actions, not by their words.
Daniel Kokotajlo
The problem with that is that past technological advancements have been more narrow. They've like automated some things, but not everything. But we are talking about a hypothetical future situation in which everything gets automated.
Daniel Kokotajlo
It is pretty crazy to think that we're building a technology, a brain that we don't understand.
Steven Bartlett
I think that basically, I think that if we don't build powerful AI systems eventually, then we're probably going to die as a civilization eventually, you know, like a hundred years from now, 200 years from now, something like that, like nuclear war pandemic, something, you know, I don't think human civilization right now is like super, super stable.
Daniel Kokotajlo
This is the most important thing happening in our lifetimes, probably in all of history, in fact. And it's very important that it go well.
Daniel Kokotajlo
I basically told my wife, like, let's not have any more kids. It's too uncertain.
Daniel Kokotajlo
1 Protocols
AI 2040 Plan A (Recommended AI Development Strategy)
Daniel Kokotajlo- Governments (starting with the US and other countries) agree to temporarily halt AI development, specifically training new models, while allowing existing AIs to continue inference (serving customers), enforced by international inspectors (Scenario: 2029).
- Establish new data centers designated for AI training that operate with total research transparency, publishing all details of model recipes, architectures, and training processes (Scenario: 2030).
- Allow AI progress to continue at a slower, safer pace, focusing on making AIs more interpretable and controllable, rather than pursuing an intelligence explosion.
- Promote broad diffusion of AI capabilities across multiple companies and countries, avoiding a monopoly by commoditizing the frontier through transparency.
- Build new data centers and infrastructure in a way that allows for their destruction if international agreements break down and an uncontrolled race resumes.
- Implement a system where citizens receive regular financial dividends (e.g., shares in an agency selling permits to AI/robot companies) to provide income as jobs are automated (Scenario: 2033).
- Deliberately stop increasing AI capabilities at the 'top expert' level, as safety cases for going beyond this level are not yet sufficient (Scenario: 2035).
- Achieve significant scientific progress on AI alignment to ensure AIs are robustly aligned with human values and goals before allowing them to become vastly superintelligent (Scenario: 2040).
- Once robust alignment is achieved, allow AIs to become significantly smarter than humans, leading to radical societal transformations, including advanced scientific discoveries, new living spaces (e.g., in space), and potentially immortality (Scenario: Post-2040).