My current work is organized around three connected questions:
Understanding model failures
How do we build evaluations to uncover when and why models exhibit harmful, biased, sycophantic, or otherwise undesirable behavior in real human–AI interactions?
Improving alignment mechanisms
How can feedback, reward design, and learning objectives better represent diverse or conflicting human values and produce robust behavior?
Anticipating emerging risks
How can adversarial, multi-agent, and automated evaluation systems discover failure modes that static evaluations or human red teams may not anticipate?