Skip to main content
Posted August 21, 2026

Lead Site Reliability Engineer

Mastercard
O Fallon, Missouri 63368, United States Full-Time
155000.00 - 205000.00
Reference: 3157050269

Mastercard is seeking a Lead Site Reliability Engineer to drive reliability, scalability, and security for mission-critical financial services platforms. You will design and optimize cloud-native, highly available systems, implement SRE best practices, and lead incident response and postmortems. Partnering with IT and Cybersecurity teams, you'll automate deployments, observability, and resilience testing while mentoring engineers. Ideal candidates bring deep experience with cloud, CI/CD, infrastructure-as-code, and securing large-scale, distributed systems in a regulated environment.

Responsibilities

  • Lead design and operation of highly available, secure, and scalable financial services platforms.
  • Define and implement SRE best practices, including SLOs, SLIs, and error budgets.
  • Architect and maintain cloud-native infrastructure using infrastructure-as-code and automation.
  • Own incident response, root cause analysis, and postmortems for critical production issues.
  • Drive observability across systems with robust monitoring, logging, and alerting solutions.
  • Collaborate closely with IT and Cybersecurity teams to embed security and compliance into the stack.
  • Optimize performance, capacity planning, and cost management for large-scale distributed systems.
  • Mentor and guide engineers on SRE principles, tooling, and operational excellence.
  • Continuously improve CI/CD pipelines and deployment strategies for safer, faster releases.
  • Champion a culture of reliability, innovation, and continuous improvement within the team.

Required Skills

  • Site Reliability Engineering (SRE)
  • Public cloud platforms (AWS, GCP, or Azure)
  • Kubernetes and container orchestration
  • Linux systems engineering and administration
  • Infrastructure as Code (Terraform, Cloud
  • Formation, or similar)
  • CI/CD pipelines (Jenkins, Git
  • Lab CI, Git
  • Hub Actions, or similar)
  • Monitoring and observability (Prometheus, Grafana, Datadog, Splunk, etc.)
  • Scripting/programming (Python, Go, or similar)
  • Security and compliance for financial/regulated environments
  • Incident management and on-call operations

Sign up for Job Alerts