Siddarth Venkatraman
RL at Mistral · PhD at Mila, Quebec AI Institute
I’m an RL intern at Mistral AI and a final-year PhD student at Mila, co-supervised by Glen Berseth and Esmeralda Whitammer. I work closely with Yoshua Bengio and was an academic collaborator at LawZero, helping develop safe and controllable AI systems.
I also work with LLNL on scaling off-policy RL for large reasoning models. Recently, I finished an internship at Valence Labs, where I trained flow bridges for molecular systems.
My research focuses on the science of RL for LLMs and inference scaling. Right now, I’m most excited (and concerned) about self-improvement and long-horizon RL. I’m also fascinated by the Machine Minds we’re spawning. I want to understand them better.
ls research/
LE CRITIQUE: Privileged Value Functions for LLM Reinforcement Learning
Preprint, 2026
Recursive Self-Aggregation Unlocks Deep Thinking in Large Language Models
Preprint
A Comedy of Estimators: On KL Regularization in RL Training of LLMs
COLM 2026
Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training
NeurIPS 2025