Anthropic’s first technical PM on token maxing, the jagged edge, and living in the future | Dianne Penn
Dianne Penn, Head of Product for Anthropic’s AI Research and Labs teams, discusses Anthropic's rapid growth, the evolution of Claude's capabilities, and how the product management role is changing in the AI era. She shares insights on eval-driven development, fostering innovation in AI, and finding joy in working with rapidly advancing models.
Deep Dive Analysis
25 Topic Outline
Anthropic's Early Days and Culture
Hidden Inflection Points and Identity Formation
Big Milestones and Frontier Model Development
The Importance of Coding Capabilities for Claude
Operating Inside the AI Exponential Curve
Uncovering Emerging AI Capabilities
Token Spending and Experimentation
Anthropic Labs: Incubation and Discontinuous Bets
The Role of Research and User Feedback
Qualities of Successful AI Researchers
Forward Compatibility and Ambitious Thinking
Frontier Model Safeguards and Scrutiny
Evolving Product Management in the AI Era
Evals as the New PRDs
The Value of Hands-On Leadership in AI
Finding Joy and Collaboration in AI Work
Dianne's Personal Use of Claude
Avoiding Overreliance and Brain Atrophy with AI
Claude's Constitution and Its Pushback Capability
AI Writing and Verifiability
Enduring Value of Human Brains
Navigating AI with Kids
Avoiding Burnout in a Fast-Paced AI Environment
Final Thoughts on PM Role and Culture
Lessons from High-Yield Bond Trading
5 Key Concepts
Emerging Capabilities
These are discontinuous jumps in AI model abilities that occur as more data and compute are added during training. Models can go from being unable to perform a task to reliably performing it, making exact prediction of new capabilities difficult without robust evaluation systems.
Evals as New PRDs
This concept means that evaluations (evals) of AI model performance, particularly in response to specific user pain points, are now the primary way to define and prioritize product work, replacing or supplementing traditional Product Requirements Documents (PRDs) in AI development.
Token Maxing
This refers to the practice of spending a significant amount of money on AI tokens to extensively use and experiment with advanced AI models. The idea is that this allows users to 'live in the future' by experiencing capabilities that will become cheap and widespread later.
Product Overhang
This describes the gap between what current AI models are capable of and what users are actually exploring or solving with them. There's often more potential value in existing models that has yet to be discovered and integrated into user-facing products.
Claude's Constitution
This is a set of principles and guidelines built into Claude's alignment research and safety systems that dictates how the AI should think and operate. Counterintuitively, this framework, including the ability for Claude to push back, makes the model more useful and intelligent.
9 Questions Answered
Anthropic's early culture was very much like a startup, with a strong emphasis on mission, values, and a bottoms-up approach, where engineers and designers would often donate time to work on experimental projects like Golden Gate Claude.
Key inflection points included the development and training of Opus 3, which built internal trust and a clear goal to create a frontier model, and the realization that focusing on coding capabilities could differentiate Claude in the market.
People should prepare by being adaptable, thinking with first principles, and applying that thinking to invest in new products and explain differences to users, as the pace of AI improvement means plans can change rapidly.
Anthropic Labs is an internal organization focused on identifying and pursuing discontinuous, large bets that might not be on the core roadmap, with a culture of experimentation to explore 10x, 100x, or 1,000x ideas like Claude Code and Skills.
Researchers at Anthropic work on a broad vision of the future of AI, while also making iterative improvements to Claude by forming hypotheses, tweaking algorithms, adjusting training, and testing to improve model capabilities, especially in areas with user impact.
As frontier models become more capable, Anthropic evolves its safeguards, red-teaming, and pre-release processes, including building fallback UXs and systems to ensure asymmetrical benefit and minimize severe risks, while aiming for broad accessibility.
Product managers need strong first principles thinking, the ability to translate user feedback into actionable 'evals' for researchers, and a hands-on approach to building and shipping with AI technology, even for those in leadership roles.
Claude's ability to push back, stemming from its alignment and safety research, makes it a better thinking partner. It helps users reach better conclusions by challenging ideas and adding to the conversation, rather than just agreeing or completing delegated tasks.
Human brains will continue to be most valuable in areas requiring nuanced judgment, persistence, proactivity, and deep subject matter expertise (e.g., biology, life sciences), as these involve accumulated experience and traits AI systems have yet to fully develop.
12 Actionable Insights
1. Embrace Adaptability in AI Development
Recognize that new AI models bring unpredictable ’emerging capabilities.’ Be adaptable in your plans and decision-making when faced with new information, rather than sticking to old strategies, to leverage the rapid advancements effectively.
2. Sweat Tokens as Much as Pixels
Treat token spend as an input for experimentation, not just a cost. Be ambitious and use AI models extensively to generate and refine ideas, as there’s no substitute for hands-on interaction with the technology to understand its full potential.
3. Foster Communal AI Discovery
Don’t make AI experimentation an individual sport. Work in public, share ideas, and try variations of what others are doing to accelerate communal discovery of new use cases and unlock magical possibilities with AI tools.
4. Lead with Hands-On AI Experience
If you’re a manager or leader in AI, be hands-on with the technology yourself. Spend time tinkering, building, and shipping with AI to maintain a strong understanding of model capabilities and guide your team effectively.
5. Find Joy in AI Through Collaboration
If you’re struggling to find joy in working with AI, pair with an enthusiastic colleague. Collaborating on a use case you care about can make experimentation feel less like work and more like a shared discovery, fostering a positive feedback loop.
6. Use AI for Better Human Conversations
Leverage AI as a personalized coach for crucial conversations. Use it to brainstorm, refine your approach, and consider different reactions, augmenting your emotional intelligence and improving communication with colleagues.
7. Cultivate Independent Thinking with AI
To avoid over-reliance or ‘brain rot’ from AI, form your own point of view first before engaging with the model. Use AI as a sparring partner to augment and challenge your thinking, rather than letting it dictate your thoughts entirely.
8. Prioritize AI Training for Writing Quality
If AI writing quality is a concern, focus on training improvements to enhance tone and character. Recognize that as other capabilities advance, writing may become a ‘rough edge’ that requires dedicated investment to improve.
9. Develop Strong Human Judgment
Focus on cultivating nuanced judgment, persistence, and proactivity, as these human traits will remain critical even as AI advances. These qualities are essential for making strategic decisions and creating superior experiences.
10. Nurture Curiosity and Inner Voice in Kids
Encourage children to develop curiosity for learning, persistence, and their own inner voice. Foster their ability to form opinions and take a stance, preparing them for a future where independent thought is highly valued.
11. Build a Supportive, Low-Ego Team
Combat burnout and accelerate progress by fostering a culture of radical ownership and team collaboration. Hire individuals who are low-ego, team-oriented, and willing to help each other, especially during critical decision-making and launches.
12. Provide Specific User Feedback to AI Labs
Actively provide detailed feedback on AI models (e.g., thumbs up/down, specific examples of failures). This granular input is crucial for researchers to understand pain points and make targeted improvements to the models.
6 Key Quotes
Evals are the new PRDs.
Dianne Penn
If you're willing to spend $100,000 a year right now in tokens, you are living the way somebody in 2028 is going to live.
Gary Tan (via Dianne Penn)
You have to sweat the tokens as much as you sweat the pixels.
Dianne Penn
One thing we say a lot on the team is you need frontier products in order to have frontier models and for people to feel the magic of frontier models.
Dianne Penn
Proactivity is not necessarily always doing a thing that you are scheduled to do. It is knowing when to come up with a new idea.
Dianne Penn
No matter how far you go, there's always another level.
Dianne Penn's Grandfather
1 Protocols
Eval-Driven Development Loop
Dianne Penn- Understand the user pain point by deeply reading consented user feedback and transcripts, sweating the tokens as much as the pixels.
- Identify the specific trajectory of failure (e.g., hallucination, overconfidence, tool use failure, search synthesis failure).
- Reproduce the issue consistently and determine if it's a significant problem.
- Generate a set of examples (e.g., 30-40) demonstrating the failure, creating an 'eval set' with prompts and expected 'golden answers'.
- Add this eval set to repositories to consistently test future versions of the AI model.
- Measure the quality of new model versions against the eval to track improvements in the identified area.