Posted August 21, 2026
Senior Site Reliability Engineer
Mastercard
O Fallon, Missouri 63368, United States
Full-Time
135000.00 - 180000.00
Reference: 3157050273
Mastercard is seeking a Senior Site Reliability Engineer to enhance reliability, scalability, and performance of our critical IT & Data Management platforms. You will design resilient architectures, automate deployments, and champion observability to ensure always-on services. Collaborating with cross-functional teams, you'll identify and resolve production issues, implement robust incident management, and drive SRE best practices. This role offers the opportunity to work with cutting-edge cloud and container technologies in a culture that values innovation, collaboration, and continuous growth.
Responsibilities
- Design and maintain highly available, scalable, and secure cloud infrastructure for Mastercard's IT & Data Management platforms.
- Build and improve automation for deployments, configuration management, and infrastructure provisioning.
- Implement and refine monitoring, logging, and alerting to ensure service reliability and rapid incident detection.
- Lead and participate in incident response, root cause analysis, and post-incident reviews to drive continuous improvement.
- Partner with development and data teams to embed SRE best practices, including SLIs/SLOs and capacity planning.
- Optimize system performance and cost efficiency across distributed, cloud-native environments.
- Develop tools and scripts to reduce toil and improve operational excellence.
- Contribute to security, compliance, and governance standards within production environments.
Required Skills
- Site Reliability Engineering (SRE)
- Cloud platforms (AWS, Azure, or GCP)
- Kubernetes and containerization (Docker)
- Infrastructure as Code (Terraform/Cloud
- Formation)
- CI/CD pipelines (Jenkins, Git
- Hub Actions, Git
- Lab CI)
- Linux systems administration
- Monitoring and observability (Prometheus, Grafana, Datadog, New Relic)
- Scripting/programming (Python, Go, Bash)
- Distributed systems and microservices
- Incident management and on-call operations
