Anthropic's Fellows Program supports four months of mentored, empirical AI safety research, enabling fellows to work on impactful public projects in areas like interpretability and adversarial robustness.
Funder: Anthropic
Due Dates: July 26, 2026 (Application deadline for July 2026 cohort)
Funding Amounts: ~$3,850/week (US) or equivalent for 4 months; ~$15,000/month compute funding; additional research expenses covered
Summary: 4-month fellowship for researchers and engineers to conduct mentored, empirical AI safety research with public outputs.
Key Information: Full-time work authorization and physical presence in US, UK, or Canada required; no visa sponsorship available.
The Anthropic Fellows Program offers a unique four-month, full-time opportunity for engineers and researchers to collaborate with Anthropic mentors on high-priority AI safety research. Fellows select and shape their projects in areas such as scalable oversight, adversarial robustness, interpretability, AI security, model organisms, and model welfare, aiming to produce impactful public outputs (e.g., papers, open-source tools). The program is designed to foster innovation and support the next generation of technical AI safety leaders. Past fellows have contributed to significant advancements in empirical AI safety, including methods for rapid response to LLM jailbreaks, tracing internal model reasoning, and studying misalignment in simulated environments.