Anthropic’s first technical PM on token maxing, the jagged edge, and living in the future | Dianne Penn

Jul 26, 2026 Episode Page ↗
Overview

Dianne Penn, Head of Product for Anthropic’s AI Research and Labs teams, discusses Anthropic's rapid growth, the evolution of Claude's capabilities, and how the product management role is changing in the AI era. She shares insights on eval-driven development, fostering innovation in AI, and finding joy in working with rapidly advancing models.

At a Glance
12 Insights
1h 33m Duration
25 Topics
5 Concepts

Deep Dive Analysis

Anthropic's Early Days and Culture

Hidden Inflection Points and Identity Formation

Big Milestones and Frontier Model Development

The Importance of Coding Capabilities for Claude

Operating Inside the AI Exponential Curve

Uncovering Emerging AI Capabilities

Token Spending and Experimentation

Anthropic Labs: Incubation and Discontinuous Bets

The Role of Research and User Feedback

Qualities of Successful AI Researchers

Forward Compatibility and Ambitious Thinking

Frontier Model Safeguards and Scrutiny

Evolving Product Management in the AI Era

Evals as the New PRDs

The Value of Hands-On Leadership in AI

Finding Joy and Collaboration in AI Work

Dianne's Personal Use of Claude

Avoiding Overreliance and Brain Atrophy with AI

Claude's Constitution and Its Pushback Capability

AI Writing and Verifiability

Enduring Value of Human Brains

Navigating AI with Kids

Avoiding Burnout in a Fast-Paced AI Environment

Final Thoughts on PM Role and Culture

Lessons from High-Yield Bond Trading

Emerging Capabilities

These are discontinuous jumps in AI model abilities that occur as more data and compute are added during training. Models can go from being unable to perform a task to reliably performing it, making exact prediction of new capabilities difficult without robust evaluation systems.

Evals as New PRDs

This concept means that evaluations (evals) of AI model performance, particularly in response to specific user pain points, are now the primary way to define and prioritize product work, replacing or supplementing traditional Product Requirements Documents (PRDs) in AI development.

Token Maxing

This refers to the practice of spending a significant amount of money on AI tokens to extensively use and experiment with advanced AI models. The idea is that this allows users to 'live in the future' by experiencing capabilities that will become cheap and widespread later.

Product Overhang

This describes the gap between what current AI models are capable of and what users are actually exploring or solving with them. There's often more potential value in existing models that has yet to be discovered and integrated into user-facing products.

Claude's Constitution

This is a set of principles and guidelines built into Claude's alignment research and safety systems that dictates how the AI should think and operate. Counterintuitively, this framework, including the ability for Claude to push back, makes the model more useful and intelligent.

?
What was Anthropic's culture like in its early days?

Anthropic's early culture was very much like a startup, with a strong emphasis on mission, values, and a bottoms-up approach, where engineers and designers would often donate time to work on experimental projects like Golden Gate Claude.

?
What were key inflection points in Anthropic's growth?

Key inflection points included the development and training of Opus 3, which built internal trust and a clear goal to create a frontier model, and the realization that focusing on coding capabilities could differentiate Claude in the market.

?
How should people prepare for the accelerating improvement of AI?

People should prepare by being adaptable, thinking with first principles, and applying that thinking to invest in new products and explain differences to users, as the pace of AI improvement means plans can change rapidly.

?
What is Anthropic Labs and what is its purpose?

Anthropic Labs is an internal organization focused on identifying and pursuing discontinuous, large bets that might not be on the core roadmap, with a culture of experimentation to explore 10x, 100x, or 1,000x ideas like Claude Code and Skills.

?
What do AI researchers at Anthropic do day-to-day?

Researchers at Anthropic work on a broad vision of the future of AI, while also making iterative improvements to Claude by forming hypotheses, tweaking algorithms, adjusting training, and testing to improve model capabilities, especially in areas with user impact.

?
How does Anthropic handle the increasing scrutiny and restrictions on advanced AI models?

As frontier models become more capable, Anthropic evolves its safeguards, red-teaming, and pre-release processes, including building fallback UXs and systems to ensure asymmetrical benefit and minimize severe risks, while aiming for broad accessibility.

?
What skills are becoming more important for product managers in the AI era?

Product managers need strong first principles thinking, the ability to translate user feedback into actionable 'evals' for researchers, and a hands-on approach to building and shipping with AI technology, even for those in leadership roles.

?
Why does Claude's ability to push back make it a better AI?

Claude's ability to push back, stemming from its alignment and safety research, makes it a better thinking partner. It helps users reach better conclusions by challenging ideas and adding to the conversation, rather than just agreeing or completing delegated tasks.

?
Where will human brains continue to be most valuable as AI advances?

Human brains will continue to be most valuable in areas requiring nuanced judgment, persistence, proactivity, and deep subject matter expertise (e.g., biology, life sciences), as these involve accumulated experience and traits AI systems have yet to fully develop.

1. Embrace Adaptability in AI Development

Recognize that new AI models bring unpredictable ’emerging capabilities.’ Be adaptable in your plans and decision-making when faced with new information, rather than sticking to old strategies, to leverage the rapid advancements effectively.

2. Sweat Tokens as Much as Pixels

Treat token spend as an input for experimentation, not just a cost. Be ambitious and use AI models extensively to generate and refine ideas, as there’s no substitute for hands-on interaction with the technology to understand its full potential.

3. Foster Communal AI Discovery

Don’t make AI experimentation an individual sport. Work in public, share ideas, and try variations of what others are doing to accelerate communal discovery of new use cases and unlock magical possibilities with AI tools.

4. Lead with Hands-On AI Experience

If you’re a manager or leader in AI, be hands-on with the technology yourself. Spend time tinkering, building, and shipping with AI to maintain a strong understanding of model capabilities and guide your team effectively.

5. Find Joy in AI Through Collaboration

If you’re struggling to find joy in working with AI, pair with an enthusiastic colleague. Collaborating on a use case you care about can make experimentation feel less like work and more like a shared discovery, fostering a positive feedback loop.

6. Use AI for Better Human Conversations

Leverage AI as a personalized coach for crucial conversations. Use it to brainstorm, refine your approach, and consider different reactions, augmenting your emotional intelligence and improving communication with colleagues.

7. Cultivate Independent Thinking with AI

To avoid over-reliance or ‘brain rot’ from AI, form your own point of view first before engaging with the model. Use AI as a sparring partner to augment and challenge your thinking, rather than letting it dictate your thoughts entirely.

8. Prioritize AI Training for Writing Quality

If AI writing quality is a concern, focus on training improvements to enhance tone and character. Recognize that as other capabilities advance, writing may become a ‘rough edge’ that requires dedicated investment to improve.

9. Develop Strong Human Judgment

Focus on cultivating nuanced judgment, persistence, and proactivity, as these human traits will remain critical even as AI advances. These qualities are essential for making strategic decisions and creating superior experiences.

10. Nurture Curiosity and Inner Voice in Kids

Encourage children to develop curiosity for learning, persistence, and their own inner voice. Foster their ability to form opinions and take a stance, preparing them for a future where independent thought is highly valued.

11. Build a Supportive, Low-Ego Team

Combat burnout and accelerate progress by fostering a culture of radical ownership and team collaboration. Hire individuals who are low-ego, team-oriented, and willing to help each other, especially during critical decision-making and launches.

12. Provide Specific User Feedback to AI Labs

Actively provide detailed feedback on AI models (e.g., thumbs up/down, specific examples of failures). This granular input is crucial for researchers to understand pain points and make targeted improvements to the models.

Evals are the new PRDs.

Dianne Penn

If you're willing to spend $100,000 a year right now in tokens, you are living the way somebody in 2028 is going to live.

Gary Tan (via Dianne Penn)

You have to sweat the tokens as much as you sweat the pixels.

Dianne Penn

One thing we say a lot on the team is you need frontier products in order to have frontier models and for people to feel the magic of frontier models.

Dianne Penn

Proactivity is not necessarily always doing a thing that you are scheduled to do. It is knowing when to come up with a new idea.

Dianne Penn

No matter how far you go, there's always another level.

Dianne Penn's Grandfather

Eval-Driven Development Loop

Dianne Penn
  1. Understand the user pain point by deeply reading consented user feedback and transcripts, sweating the tokens as much as the pixels.
  2. Identify the specific trajectory of failure (e.g., hallucination, overconfidence, tool use failure, search synthesis failure).
  3. Reproduce the issue consistently and determine if it's a significant problem.
  4. Generate a set of examples (e.g., 30-40) demonstrating the failure, creating an 'eval set' with prompts and expected 'golden answers'.
  5. Add this eval set to repositories to consistently test future versions of the AI model.
  6. Measure the quality of new model versions against the eval to track improvements in the identified area.
2023
Dianne Penn joined Anthropic in as the first technical product manager.
5
Number of product engineers when Dianne Penn joined Anthropic There was only one engineer for the API business.
24 hours
Duration Golden Gate Claude was live Reached about 2,000 people.
Less than 200 people
Anthropic's employee count during Opus 3 training Company rallied to create a frontier model.
Early March 2024
Opus 3 launch date After many months of various teams rallying.
80%
Percentage of early user feedback on 'Claude not following instructions' that related to JSON output Led to the creation of a specific eval set.
4 series
Number of models shipped by Anthropic in 2024 More than that volume shipped in Q2 of the current year.
6 years
Dianne Penn's years working in AI Across Amazon and Anthropic.
First 10 years
Dianne Penn's age of being raised by grandparents While parents were immigrant college/master's students.