BizOps Engineer II
Mastercard is seeking a BizOps Engineer II to join our Information Technology & Data Management team in Financial Services. In this role, you will optimize and support mission-critical platforms that power secure, high-volume transactions worldwide. You'll design and implement robust monitoring, automation, and incident management solutions to ensure high availability, performance, and scalability. Collaborating closely with software engineers, data engineers, and product teams, you will troubleshoot complex issues, analyze system metrics, and drive continuous improvement across infrastructure and applications. This position offers the opportunity to work with cutting-edge cloud and data technologies while influencing operational best practices and reliability standards. You'll contribute to incident response, root-cause analysis, and post-incident reviews, helping to build more resilient systems. Mastercard's culture emphasizes innovation, collaboration, and continuous learning, giving you room to experiment with new tools and approaches. If you are passionate about system reliability, automation, and bridging the gap between development and operations in a dynamic, global environment, this role provides a chance to make a tangible impact on secure digital payments worldwide.
Responsibilities
- Design, implement, and maintain monitoring, alerting, and observability for mission-critical applications and infrastructure.
- Automate operational tasks, deployments, and remediation workflows to improve reliability and reduce manual intervention.
- Collaborate with software and data engineering teams to optimize system performance, scalability, and resilience.
- Lead and participate in incident response, troubleshooting complex production issues, and driving timely resolution.
- Conduct root-cause analysis and implement long-term fixes to prevent recurrence of incidents.
- Support CI/CD pipelines and release processes to enable safe, rapid, and reliable deployments.
- Analyze system metrics and logs to identify bottlenecks, trends, and optimization opportunities.
- Contribute to reliability best practices, runbooks, and operational documentation.
- Partner with security and compliance teams to ensure systems meet regulatory and security requirements.
- Mentor junior team members and help foster a culture of continuous improvement and learning.
Required Skills
- Site Reliability Engineering (SRE) practices
- Cloud platforms (AWS, GCP, or Azure; AWS preferred)
- Infrastructure as Code (Terraform, Cloud
- Formation, or similar)
- Linux systems administration
- Containerization and orchestration (Docker, Kubernetes)
- CI/CD pipelines (Jenkins, Git
- Lab CI, or similar)
- Monitoring and observability (Splunk, Prometheus, Grafana, Cloud
- Watch)
- Scripting/programming (Python, Bash, or similar)
- SQL and basic data querying/analysis
- Incident management and root-cause analysis
