Site Reliability Engineer
Contract Details
- Notice - 1-2 weeks
- 1099 contract / contractor / B2B
- Term - 3-6 months with possible extens
Our client is a fast-growing global fintech company building a modern digital payments platform that enables fast, secure, and compliant international money transfers. Serving millions of customers across multiple markets, the company operates a highly available cloud-native infrastructure where reliability, scalability, and operational excellence are core business priorities.
We are looking for an experienced Site Reliability Engineer to help strengthen platform reliability, improve production operations, and drive automation across the engineering organization.
About the Role
As a Site Reliability Engineer, you will own and improve the reliability, availability, and operational health of a large-scale cloud platform. You'll collaborate closely with Software Engineers, Infrastructure, Customer Operations, and Product teams while helping evolve production support processes and operational standards.
The role combines traditional SRE responsibilities with modern AI-assisted engineering practices, leveraging AI tools to improve incident response, documentation, operational workflows, and engineering productivity.
This position is ideal for someone who enjoys solving complex production challenges, improving observability, automating repetitive operational work, and building scalable reliability practices.
Responsibilities:
-
Act as the primary technical escalation point for critical production incidents, providing hands-on support during high-severity outages.
-
Improve platform reliability by reviewing new product launches, infrastructure changes, and production readiness before release.
-
Design, implement, and optimize monitoring, alerting, and observability solutions across cloud infrastructure and applications.
-
Analyze operational metrics, recurring alerts, and incident trends to reduce alert fatigue and improve overall system health.
-
Lead incident investigations and post-mortems, ensuring root causes are identified and preventative actions are implemented.
-
Collaborate with Engineering, Infrastructure, Customer Operations, and external support teams to coordinate incident response and customer communications.
-
Participate in capacity planning, peak traffic readiness, disaster recovery exercises, and system performance reviews.
-
Develop and maintain operational runbooks, documentation, and incident response procedures.
-
Improve internal reliability tooling and automate operational workflows using modern AI-assisted development tools.
-
Drive continuous improvements in operational excellence through automation, standardization, and proactive reliability initiatives.
Requirements
-
4+ years of experience as a Site Reliability Engineer, DevOps Engineer, Production Engineer, or a similar infrastructure-focused role.
-
Strong experience supporting production systems running on AWS.
-
Hands-on experience with monitoring and observability platforms such as Datadog, AWS CloudWatch, New Relic, or similar.
-
Experience with incident management platforms such as PagerDuty.
-
Strong understanding of production incident management, root cause analysis, and post-incident review processes.
-
Experience working with ticketing and documentation platforms such as Jira and Confluence.
-
Familiarity with operational dashboards and reporting tools (Looker or similar BI platforms).
-
Experience building operational documentation, runbooks, and support processes.
-
Comfortable working outside regular business hours when critical production incidents require senior engineering support.
-
Experience using AI-assisted engineering tools (Claude, GitHub Copilot, Cursor, or similar) to improve engineering workflows, automate documentation, incident triage, reporting, or operational tasks.
-
Strong scripting or automation skills (Python, Bash, or similar) are considered a plus.
Nice to Have
-
Experience working in fintech, payments, financial services, or other high-availability environments.
-
Experience with Infrastructure as Code (Terraform, CloudFormation, or similar).
-
Familiarity with Kubernetes and containerized environments.
