Kshitij Sachan
I'm a member of technical staff at Anthropic, working to ensure the world safely makes the transition through transformative AI.1 I'm currently doing pretraining science. Before that I was on the Frontier Red Team,2 where I co-started our cyber evaluations effort and helped with bio and cyber evaluations for Claude 3.3
Earlier, I worked on AI control4 and interpretability5 at Redwood Research, did reinforcement learning research with George Konidaris while getting a BS/MS at Brown, and interned at Jane Street.
Outside of work, I like ultimate frisbee, biking, and playing the flute.
I'm hiring.
Scholar / LinkedIn / kshitijsachan39@gmail.com
- Anthropic's mission, from Claude's constitution.
- Anthropic's Frontier Red Team.
- Claude 3 model card, section 6.
- AI Control: Improving Safety Despite Intentional Subversion
- Polysemanticity and Capacity in Neural Networks