Research Scientist — Reinforcement Learning (Foundation Models)
The role
We’re hiring a Research Scientist (Reinforcement Learning) to help bring RL into the core of how foundation models are adapted and improved for industrial use.
In industry, models don’t live in isolation: domain experts validate, correct, and act on model outputs. We want RL to leverage that expert feedback not only as a post‑training patch, but increasingly earlier in the pipeline, shaping objectives, training signals, and adaptation strategies.
What you’ll work on
Human‑in‑the‑loop reinforcement learning
Turn expert validation/correction into a reliable learning signal.
Design feedback interfaces/signals that are practical in real operational settings.
RL for industrial foundation models
Develop RL methods that sit on top of (or integrate with) foundation models used in production.
Explore ways for RL to intervene earlier in the chain (not just after deployment).
From research to deployment
Build evaluation protocols aligned with real constraints: robustness, uncertainty reduction, safety, auditability, and cost of error.
Work closely with scientists/engineers to ship demonstrators that connect benchmarks to field outcomes.
What we’re looking for
PhD (preferred) or equivalent research experience in Reinforcement Learning / Machine Learning.Strong foundations in RL (e.g., policy optimization, off‑policy learning, offline RL, exploration, credit assignment).Ability to design rigorous experiments, debug failure modes, and iterate fast with scientific discipline.Strong programming skills (Python; deep learning stack such as PyTorch).Nice to haveExperience with real‑world RL constraints (noisy/limited feedback, safety requirements, deployment considerations).Comfort with complex data modalities (time series, scientific/industrial signals, multimodal setups).Publications or open research artifacts in RL / sequential decision‑making.
Why join
Work on RL problems that matter in the real world: expert feedback loops, uncertainty reduction, and mission‑critical constraints.A research culture that values clarity, rigor, and humility and that connects fundamental ideas to deployable systems.High ownership in a small team: you’ll shape direction, not just execute tasks.
#J-18808-Ljbffr