Senior Site Reliability Engineer

4 days ago

Apply Now
Logo of Cribl

Cribl

501 - 1000 employees

Founded 2017

☁️ SaaS

Description

• Engage with teams and improve service delivery and reliability across their entire lifecycle • Measure and monitor all production systems with an eye towards availability, latency and overall system health • Seek out the cause of errors and instability in our production cloud services and drive teams towards better operational excellence • Engage with product and platform teams to improve and evolve systems by lobbying for changes that improve reliability, resilience, and observability • Help Identify and drive down toil with creative innovation and automation • On-call responsibilities

Requirements

• Extensive experience with enterprise scale continuous delivery environments • Development with JavaScript/Node.js/TypeScript in a Linux/Mac environment • Experience with sustainable incident response in a blameless environment • Experience with Configuration Management Tools like Terraform (preferred) or Puppet, Chef, Ansible • Knowledge of cloud platforms (prefer AWS) and container + orchestration technologies • Experience with APM and Observability and related tools such as, New Relic, Splunk, CloudWatch, Prometheus, Grafana/Kibana, Sentry etc. • Background in Linux Systems Engineering • Experience with Incident response related tools for instance, PagerDuty, FireHydrant, Blameless etc. • Comfortable with a high level of autonomy and working with a distributed team

Apply Now

Similar Jobs

December 20, 2024

Join SOFTSWISS as a Senior DevOps/System Engineer, enhancing their sports betting platform through advanced infrastructure management and automation.

December 20, 2024

Work with Hippo Insurance to revolutionize home insurance through innovative DevOps practices. Collaborate closely with teams to enhance productivity and scalability.

Built by Lior Neu-ner. I'd love to hear your feedback — Get in touch via DM or lior@remoterocketship.com