Site Reliability Engineer SRE DevOps Engineer Fintech


Savanna HR


DevOps


a day ago

3-6 Yrs


Gurgaon


Description

Job Title: Site Reliability Engineer (SRE) / DevOps Engineer

Location: Delhi NCR

Experience: 3–6 Years

Employment Type: Full-time

Work Mode: On-site

Industry: Fintech

About the Role

We are seeking a highly skilled and motivated Site Reliability Engineer (SRE) / DevOps Engineer to join our dynamic team. You will be instrumental in building, automating, and maintaining a highly scalable, secure, and reliable infrastructure for our cutting-edge fintech platform. This critical role involves close collaboration with Engineering, Product, Security, and Data teams to ensure our systems consistently deliver high availability, exceptional performance, robust scalability, and unwavering reliability. Your contributions will directly impact our customers and the success of our business.

Key Responsibilities

- Design, deploy, and manage highly available and scalable cloud infrastructure.
- Build and maintain robust, automated CI/CD pipelines for seamless application deployment.
- Manage containerized applications efficiently using Docker and Kubernetes.
- Automate infrastructure provisioning and management through Infrastructure as Code (IaC) tools like Terraform.
- Implement comprehensive monitoring, logging, alerting, and observability solutions across all applications and infrastructure components.
- Continuously monitor system performance, availability, capacity, and overall reliability.
- Troubleshoot and resolve production issues, actively participating in incident management and providing on-call support.
- Conduct thorough Root Cause Analysis (RCA) for critical incidents and implement preventive measures to avoid recurrence.
- Drive improvements in system reliability, scalability, fault tolerance, and recovery mechanisms.
- Collaborate closely with development teams to optimize application deployment and operational processes.
- Automate repetitive operational tasks to reduce manual intervention and increase efficiency.
- Implement and maintain effective backup, disaster recovery, and business continuity processes.
- Perform rigorous capacity planning and infrastructure optimization to ensure optimal resource utilization.
- Optimize cloud infrastructure and operational costs without compromising performance or reliability.
- Ensure all infrastructure adheres to stringent security and compliance best practices.

Required Skills & Experience

- 3–6 years of professional experience in Site Reliability Engineering (SRE), DevOps, Cloud Engineering, or Platform Engineering.
- Strong hands-on experience with major cloud platforms such as AWS, Azure, or GCP.
- Deep understanding of Linux operating systems and fundamental networking concepts.
- Practical experience managing containerized environments with Docker and Kubernetes.
- Proficiency with CI/CD tools like Jenkins, GitHub Actions, or GitLab CI/CD.
- Hands-on experience with Infrastructure-as-Code (IaC) tools such as Terraform.
- Experience with monitoring and observability tools including Prometheus, Grafana, ELK Stack, Datadog, or New Relic.
- Strong scripting skills in Python, Shell/Bash, or similar languages.
- Solid understanding of microservices architecture, APIs, databases, and distributed systems.
- Proven experience troubleshooting complex production environments.
- Comprehensive understanding of high availability, scalability, fault tolerance, and disaster recovery principles.

Preferred Experience

- Candidates with prior experience in the Fintech, Payments, Banking, Digital Lending, or Insurtech industries, or other high-transaction environments will be highly regarded.
- Exposure to the following technologies and concepts will be considered an advantage: – Payment or transaction processing systems – Banking / financial APIs – Kafka or other distributed messaging systems – PCI-DSS or similar security/compliance frameworks – High-volume transaction platforms – Cloud security / DevSecOps practices

Good to Have

- Experience with AWS services like EKS, EC2, RDS, S3, CloudWatch, and IAM.
- Familiarity with Helm and ArgoCD.
- Knowledge of GitOps principles.
- Kubernetes administration experience.
- Exposure to service mesh technologies.
- Experience with cloud cost optimization strategies.
- Understanding of MLOps / AI infrastructure.

What We’re Looking For

We are looking for a proactive, hands-on individual with a strong passion for automation and a dedication to ensuring system reliability. You should be adept at solving complex production challenges and thrive in a fast-paced environment where uptime, security, scalability, and performance are paramount to business success. As a well-funded, high-growth company, we offer an exciting opportunity to make a significant impact.

Key Success Metrics

- Maintain high system availability and reliability targets.
- Ensure faster and safer application deployments.
- Minimize the frequency of production incidents.
- Achieve rapid incident resolution times.
- Drive continuous improvements in infrastructure automation.
- Enhance system performance and scalability.
- Optimize cloud infrastructure costs effectively.

Notice Period: Immediate to 30 Days preferred.