How can a person, institution, or AI system determine the right thing to do?
I work on normative ethics, medical and research ethics, and AI alignment. I study how these fields move from general principles to concrete decisions, and what those mechanisms can teach us about giving AI agents guidance for particular situations.
Current work
My current projects focus on the specification problem in AI alignment, normative world models, and human-in-the-loop systems.
THE COMPETENCE PROBLEM OF AI ALIGNMENT
Capable AI systems can create a distinctive alignment problem when they treat the reasons behind an institutional rule as permission to reopen a question the institution has legitimately settled.
HUMAN-IN-THE-LOOP SYSTEMS
My forthcoming paper examines the competencies human reviewers need when they use AI systems in research-ethics oversight.
VERY GOOD ROBOT
A newsletter about how we can give AI agents guidance about what they should do.
Very Good Robot
How to help AI agents do the right thing. Very Good Robot asks how we can give AI agents guidance about what they should do in particular situations. It draws on mechanisms developed in medical and research ethics for moving from general principles to concrete decisions in regulated, practice-based institutions shaped by divided responsibility, uncertainty, and correction.