
Staff Software Engineer
Skills & requirements
About the role
Role Overview
Join Wrike's Backend Reliability (BRE) team as a Staff Software Engineer, architecting core reliability solutions that power a distributed work management platform serving millions of users. This role combines individual contributor expertise with leadership responsibility for shaping how Wrike scales, performs, and recovers from failure.
Key Responsibilities
Design, build, and maintain critical reliability infrastructure including HTTP rate limiters, database schema migration tools, circuit breakers, and distributed Redis-based caching systems
Troubleshoot complex production issues, optimize PostgreSQL performance, and ensure distributed systems remain stable under high load
Lead preliminary investigations during severe production incidents to identify root causes, assess impact, and propose mitigation strategies
Create scalable, reusable tools and frameworks enabling other engineering teams to build more resilient services
Influence reliability best practices across engineering through design reviews, knowledge sharing, and setting high technical standards
Leverage AI-powered development tools and coding agents to accelerate development and automate error-prone tasks
Required Qualifications
Strong expertise building scalable, high-performance backend systems using Java/JVM
Deep understanding of distributed systems including high availability, CAP theorem, and fault tolerance
Extensive experience with PostgreSQL, Redis, Docker, and Kubernetes in production environments
Practical experience with message brokers such as RabbitMQ or Kafka
Proven ability to work independently with minimal supervision, applying critical thinking to validate decisions
Excellent English communication skills for international collaboration
Standout Qualities
Background in infrastructure engineering or SRE practices, experience leading technical initiatives and mentoring engineers, familiarity with observability tools (Graylog, Zabbix, Grafana) or BigQuery, and demonstrated understanding of complex system failure modes and graceful recovery design.