Role description
Role Description
Responsibilities
- Design and maintain distributed database systems providing low-latency, strongly consistent data access
- Implement and optimize replication, consensus, and caching mechanisms to meet availability and performance goals
- Operate production systems, including participating in the on-call rotation, ensuring high availability and data durability
- Collaborate with infrastructure and product teams to assess current and future use cases and requirements, supporting the development of a mid- to long-term roadmap that reflects these needs
- Contribute to system design reviews, postmortems, and reliability improvements
- Write high-quality, efficient code in Go and Rust for performance-critical systems
On-call work may be necessary occasionally to help address bugs, outages, or other operational issues, with the goal of maintaining a stable and high-quality experience for our customers.
Requirements
- BS, MS, or PhD in Computer Science or related technical field involving coding (e.g., physics or mathematics), or equivalent technical experience
- 5+ years of professional software development experience
- Experience designing and implementing software using distributed systems fundamentals: replication, consistency, partitioning, and fault tolerance
- Experience building databases, storage systems, or large scale data infrastructure
- Proficient in programming and debugging across a range of languages such as Python, Go, Rust, C/C++, or Java
- Familiarity with consensus and coordination systems (e.g. Raft, Paxos, ZooKeeper, etc)
- Strong debugging and performance analysis skills
Preferred Qualifications
- Experience building distributed databases or storage systems
- Practical experience with and deep understanding of data structures used in storage systems (e.g. LSM trees, B-trees, Hash Indexes)