Chargement en cours

HPC Engineer

PARIS, 75
il y a 2 jours

Design and operate Arlequin's cloud-native, sovereign HPC platform powering our AI products and research

5+ years, Senior

Full-time

Paris (hybrid / full-remote / on-site flexible)

About Arlequin

Arlequin AI is both a research lab in topological deep learning and an AI platform. Our first product, HuDex, converts massive volumes of raw, unstructured, multilingual data into strategic decisions in minutes instead of days, providing an operational advantage for governments and businesses.~30 people, post-seed, Series A in progress. On-site in Paris.

The role

We are building a cloud-native, sovereign, and end-to-end automated high-performance computing platform to support the scientific computing workloads of our products and research. Our products are also deployed on-premise, in air-gapped environments, for clients with the strongest security requirements.

The Platform team owns the technical foundations Arlequin AI runs on: DevOps, MLOps, FinOps, FieldOps, Security/Compliance, compute and IT. Our role is not to build infrastructure for its own sake — we build self-service tools so engineers can ship without waiting on us, Forward Deployed Engineers can deploy clients without our help, C-levels understand the real cost of what they sell, and researchers don’t have to deal with engineering questions to run their experiments. Today we are a team of 3, and we aim to grow the team to 8 people by December.

As an HPC Engineer, you join the Platform team to design and operate our compute platform covering all of Arlequin’s compute workloads, ensuring a performant, scalable, and reliable platform — 100% cloud-native and orchestrated by Kubernetes. Concretely, you operate GPU and CPU clusters on Kubernetes and optimize end-to-end performance (GPU, high-throughput network interconnect, high-performance storage), while building the tools and abstractions that make teams autonomous in using, launching, and monitoring their workloads.

Your responsibilities

Infrastructure: design, deploy, and maintain infrastructure on Scaleway; industrialize IaC with OpenTofu and Terragrunt

Data storage: design and operate high-performance storage (parallel/distributed file systems, data caches, object storage) and optimize end-to-end I/O for training and inference

Compute platform: design, deploy, and operate a Kubernetes compute cluster sized for heterogeneous workloads; set up scheduling (queues, priorities, gang scheduling, fair-sharing); optimize GPU sharing and utilization; deploy and maintain device plugins

Workloads: identify and eliminate bottlenecks across compute, network, and I/O; design and operate model serving (real-time, streaming, batch, progressive rollouts); operate large-scale batch pipelines and training pipelines with the research and MLOps teams

Instrumentation & optimization: monitor compute and GPUs; instrument inference services (latency, throughput, error rate, SLOs); optimize compute costs with unit metrics; configure alerting; improve utilization and right-sizing

Platform engineering: collaborate with research/product on the platform roadmap; support capacity planning; build self-service abstractions; document the platform and runbooks

On-call: no rotation today, incidents are handled during business hours. When a rotation becomes necessary, it will never exceed one week on-call out of five, and will be compensated

Stack

GPU: NVIDIA GPU Operator, CUDA, DCGM

Serving & batch: SGLang, Kueue

Distributed computing: NCCL, MPI, RDMA / RoCE

Storage: Blob Storage

Languages: Python, Bash

What we’re looking for

5+ years of experience in infrastructure / HPC / compute platforms, including production experience

Advanced proficiency with Kubernetes in a production environment

Hands-on experience with GPU workloads and batch schedulers: scheduling, resource management, and sharing

Experience operating a production service under latency and availability constraints: autoscaling, load management, SLOs, incident management

Strong IaC experience (Terraform / OpenTofu, ideally Terragrunt)

Good understanding of distributed computing (multi-node / multi-GPU) and large-scale batch processing

Autonomy and a strong sense of ownership over the production scope, excellent communication, a feedback culture, technical curiosity, and pragmatism

Bonus

Experience with Scaleway

In-depth knowledge of the NVIDIA / ROCm / TPU ecosystems

Hands-on experience with inference optimization: quantization, continuous batching, KV-cache / prefix caching, speculative decoding, prefill/decode disaggregation

RDMA / InfiniBand / RoCE and low-latency networking

FinOps awareness, high-performance networking and storage knowledge, background in scientific computing or research, experience optimizing compute code

Process

First-fit interview (30 min)

Technical interview with the Platform team (coding + design)

Meeting with the hiring manager and the founders

Offer

Package

Competitive compensation including equity stake

Flexible remote policy: hybrid / full-remote / full on-site

Training, conference, and certification budget, with time dedicated to CNCF/LF open source contributions

#J-18808-Ljbffr
Entreprise
Arlequin AI
Plateforme de publication
WHATJOBS
Offres pouvant vous intéresser
PARIS, 75
il y a 1 jour
PARIS, 75
il y a 7 jours
PARIS, 75
il y a 7 jours
Soyez le premier à postuler aux nouvelles offres
Soyez le premier à postuler aux nouvelles offres
Créez gratuitement et simplement une alerte pour être averti de l’ajout de nouvelles offres correspondant à vos attentes.
* Champs obligatoires
Ex: boulanger, comptable ou infirmière
Alerte crée avec succès