
DevOps Observability (SRE) Engineer
Skills & requirements
About the role
DevOps Observability (SRE) Engineer
This Senior SRE Engineer role focuses on ensuring high availability, fault tolerance and system reliability through comprehensive observability. The successful candidate will design and maintain high-availability architectures while managing incident response and post-mortems.
Key Responsibilities
Design and implement high-availability and fault-tolerant system solutions
Develop and maintain observability strategy using tools such as OpenTelemetry, Grafana, ELK, Loki, Tempo, Mimir, and VictoriaMetrics
Configure monitoring systems, dashboards and alerting mechanisms
Establish incident management processes and develop response procedures and playbooks
Respond to critical incidents via PagerDuty and participate in on-call rotations
Conduct incident post-mortems and root cause analyses
Collaborate with development teams on system reliability and train them on observability tools
Requirements
3+ years of experience as an SRE Engineer
Proficiency with observability and monitoring tools including OpenTelemetry, Grafana stack, ELK, and related technologies
Hands-on experience with Kubernetes and Docker containerization
Infrastructure as code experience using Terraform and Ansible
CI/CD systems experience with GitLab CI/CD and ArgoCD
Experience supporting .NET backend teams
Demonstrated ability handling critical incidents and conducting post-mortems
Strong collaboration skills and proactive mindset for system improvements