My current work is organized around three connected questions:
Understanding and interpreting models
How can we better evaluate and understand the representations and behaviors that emerge in language models?
Improving alignment
How can we design learning objectives and feedback mechanisms that produce robust behavior when human values are diverse, conflicting, or underspecified?
Anticipating emerging risks
How can models and multi-agent systems help us discover failure modes that static evaluations and human red teams have not yet anticipated?