Research Engineer, Code RL (Reinforcement Learning)

Anthropic — New York, NY

Posted: 2026-09-04

Job Description

• Research and engineer reinforcement-learning systems that teach AI models to write, edit, test, debug, and ship real software safely and efficiently.
• Design RL environments, coding tasks, reward signals, and verifiers; run training experiments on frontier models and interpret results.
• Build scalable RL infrastructure and improve the speed, reliability, and performance of ML pipelines.
• Requires strong software engineering and deep Python expertise, including async/concurrent programming; relevant RL, RLHF, LLM fine-tuning, PyTorch, distributed training, GPU/TPU, or sandboxing experience is valuable.
• New York, NY hybrid role requiring office attendance at least 25% of the time; bachelor's degree or equivalent required, visa sponsorship may be available, and compensation is $500,000–$850,000 USD annually.

View job and apply