Platform Reliability Engineer
Sonepar is an independent family-owned group and a world leader in the distribution of electrical equipment, solutions and related services to the trade . In 2025, Sonepar generated a turnover of €33.6 billion. With a presence in 40 countries and an extensive network of retail outlets, the Group is undertaking an ambitious transformation to make life easier for its customers by offering them an omnichannel experience and a range of sustainable solutions for the industrial, construction and energy sectors. Our 46,000 employees are committed to the electrification of the world and united by a shared Purpose: "Power Progress for Future Generations.".
How will you shape our tomorrow
The Platform Reliability Engineer is responsible for ensuring the reliability, availability, and performance of the organization’s digital platform. Working under the guidance of the Platform SRE Lead, this role applies Site Reliability Engineering principles to day-to-day operations, incident management, and continuous improvement initiatives.
This position combines strong hands-on operational expertise with an engineering mindset, focused on automation, resilience, and scalable operations across the platform.
Key responsibilities
Platform Operations & Reliability
- Monitor, operate, and maintain platform health, availability, and performance.
- Apply SRE principles to improve system reliability, scalability, and fault tolerance.
- Contribute to the definition, tracking, and reporting of Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets.
- Continuously improve operational processes to enhance platform stability and customer experience.
Incident & Problem Management
- Participate in the resolution of incidents and support cross-functional incident response.
- Perform root cause analysis and contribute to preventive and corrective actions.
- Use monitoring, observability tools, and data analysis to detect and mitigate issues before they impact service availability.
- Contribute to post-incident reviews and operational documentation.
Automation & Tooling
- Develop and maintain automation scripts to support deployment, monitoring, and recovery processes.
- Reduce manual operational work through Infrastructure as Code (IaC) and CI/CD pipelines.
- Improve operational efficiency by automating repetitive tasks and workflows.
Observability
- Implement and enhance monitoring, logging, and alerting solutions for real-time system visibility.
- Contribute to distributed tracing and telemetry to support root cause analysis and troubleshooting.
- Ensure observability data supports SLO tracking and operational decision-making.
Resilience Engineering
- Participate in resilience testing and chaos engineering experiments to validate system robustness.
- Contribute to improving failover mechanisms and disaster recovery procedures.
- Identify single points of failure and propose reliability improvements.
Performance & Scalability
- Assist in optimizing platform performance, resource utilization, and cost efficiency.
- Collaborate with development teams to embed reliability and operability into application design.
- Support scalability testing and performance validation activities.
- Work closely with the Platform SRE Lead, platform engineers,