Research Engineer, Forge
PARIS, 75
il y a 1 jour
- Turn real customer requirements into reliable training and deployment workflows
- Work end-to-end across model adaptation and post-training, evaluation, data, and infrastructure
- Build and improve post-training and evaluation workflows for CPT, SFT, RL, and distillation
- Develop tools and pipelines for synthetic data generation, data curation, training, evaluation, and deployment
- Debug and harden large-scale ML systems, including distributed training, scheduling/execution, checkpointing, observability, and reproducibility
- Improve the Forge codebase through clear APIs, tests, documentation, and maintainable abstractions
- Advance the RL training stack, including high-throughput asynchronous rollout and scalable post-training systems
- Ensure Forge deployment is seamless and adaptable across diverse clients, hardware, software stacks, cloud, and on-premises environments
- Partner with researchers and infrastructure engineers to translate bottlenecks into concrete system improvements
- Collaborate with scientists, engineers, product, and customer-facing teams to ship maintainable and trusted Forge projects
Requirements
- Strong Python engineering skills and experience working in large codebases, including testing, code review, CI, and operational ownership
- Hands-on experience with PyTorch, JAX, or similar
- Strong systems and infrastructure fundamentals
- Experience with LLM training or post-training, including fine-tuning, RL, distillation, evaluation, and/or data pipelines
- Excellent debugging skills in distributed jobs, data issues, quality regressions, and infrastructure failures
- Clear communication with technical and non-technical stakeholders
- High agency, low ego, and comfort in fast-moving, under-specified environments
- Nice to have: distributed training experience with FSDP, DeepSpeed, Megatron, or similar
- Nice to have: cluster/orchestration experience with SLURM, Ray, Kubernetes, Kueue, Karpenter, Skypilot, or similar
- Nice to have: experience building reliable ML infrastructure, evaluation systems, or large-scale data processing pipelines
- Nice to have: research experience in LLMs, agents, multimodal models, reasoning, code, or domain adaptation
- Nice to have: open-source contributions, publications, or widely used internal tooling
- Nice to have: experience training multi-billion-parameter models and on petabyte- and exabyte-scale datasets
- Nice to have: ability to identify bottlenecks across the stack and drive improvements from first principles
Core Competencies
Demonstrates strong Python engineering skills and hands-on experience with PyTorch or JAX, focusing on building and improving ML workflows, including training, evaluation, and deployment. Capable of collaborating with cross-functional teams to enhance system performance and reliability in fast-paced environments.
Highest-signal resume keywords
- Python Engineering
- ML Workflow Development
- Distributed Training
- Post-Training Evaluation
- Clear Communication
ATS Optimization Keywords
Hard Skills
- Python
- PyTorch
- JAX
- LLM Training
- Debugging
- Data Pipelines
- Testing
- Code Review
- Operational Ownership
- Infrastructure Fundamentals
Soft Skills
- Clear Communication
- High Agency
- Low Ego
Industry Keywords
- Model Adaptation
- Synthetic Data Generation
- Large-Scale ML Systems
- Observability
- Reproducibility
- Maintainable Abstractions
- Scalable Systems
- Bottleneck Identification
Tools & Technologies
- FSDP
- DeepSpeed
- Megatron
- SLURM
- Ray
- Kubernetes
- Kueue
- Karpenter
- Skypilot
Entreprise
Jobtailor
Plateforme de publication
WHATJOBS
Offres pouvant vous intéresser
PARIS, 75
il y a 2 jours
PARIS, 75
il y a 5 jours
PARIS, 75
il y a 5 jours
PARIS, 75
il y a 5 jours