Expert, Senior
Take a step forward and let Edenred surprise you.
Every day, we deliver innovative solutions to improve the life of millions of people, connecting employees, companies, and merchants all around the world.
We know there are hundred ways for you to grow. With us, you will expand your skills in a multicultural, challenging, and dynamic environment.
Dare to join Edenred and get ready to thrive in a global company that will offer you endless opportunities.
Edenred is all about meritocracy. You come as you are, and you contribute. Indeed, the Edenred Group recognizes, recruits and develops all talents and singularities.
We are committed to preventing all forms of discrimination and to providing all our candidates with equal opportunities regardless of their gender and gender expression, disability, origin, religious belief and sexual orientation or any other criteria.
In close collaboration with platform, assets, integration, and product teams, the Data & AI Platform OPS & SRE Lead acts as the guardian of production excellence for Edenred’s Data & AI ecosystem.
The role is responsible for defining, implementing, and coaching Site Reliability Engineering, operational, and FinOps practices across Data, AI, and Agentic workloads—while preserving a strict “you build it, you run it” mindset. As a player-coach, the role is both operationally hands-on and structurally influential, contributing directly to system design, CI/CD, and architectural decisions.
Define and operate SRE standards for Data, AI, and Agentic systems (SLOs, error budgets, reliability targets)
Own observability frameworks (metrics, logs, traces, data freshness, AI signals)
Own FinOps practices for Data & AI workloads (cost attribution, optimisation, guardrails)
Define and run incident, problem, and service management practices
Co-own CI/CD pipelines and guardrails (policy-as-code, quality gates, ..)
Key member of the Design Authority, validating operability by design
Player-coach: contribute hands-on to critical systems and coach teams on production readiness
Share on-call responsibilities with builders, enforce blameless postmortems
Ensure discovery & fast-track workloads are observable, cost-controlled
Continuously improve platform reliability, scalability, operational maturity
Strong SRE / production engineering background (Databricks multi-workspace/metastore) ecosystem
Experience operating data platforms, AI pipelines, and agentic systems in production
Deep understanding of observability, incident management, and reliability engineering
Solid FinOps knowledge for cloud-native and AI workloads
Comfortable with CI/CD, automation, and infrastructure-as-code
Hands-on contributor able to dive into complex incidents and designs
Strong coaching mindset; raises the operational bar across teams
Ability to embed SRE thinking without centralising ownership
Pragmatic and automation-first
Ability to influence design decisions without hierarchical authority
Clear communicator in high-pressure production contexts
Comfortable arbitrating trade-offs between reliability, speed, and cost
Trusted partner of engineers, architects, and product leaders
Apply now and Vibe with Us!
Sign up to apply and find out right away if you're a fit.
Your agent will tell you — in seconds.
Sign up and I'll tell you right away how well Edenred matches you — what you already have, and what's missing. Then I stay on it: I search for you and only write when I find something worth your time.