Senior Site Reliability Engineer (Observability & Analytics) – Platform Infra

ElasticCanadaPlatform - SRE

From another job board

NEL has not scanned this employer's domain, and NEL payment protection does not apply. You apply on the employer's own site.

Listed on Elastic’s public greenhouse board and shown here for discovery. NEL is not involved in this hiring process.

About this role

Elastic, the Search AI Company, enables everyone to find the answers they need in real time, using all their data, at scale — unleashing the potential of businesses and people. The Elastic Search AI Platform, used by more than 50% of the Fortune 500, brings together the precision of search and the intelligence of AI to enable everyone to accelerate the results that matter. By taking advantage of all structured and unstructured data — securing and protecting private information more effectively — Elastic’s complete, cloud-based solutions for search, security, and observability help organizations deliver on the promise of AI. What is The Role Platform Observability & Analytics runs the infrastructure that tells Elastic the truth about its own platform. The observability clusters show Cloud engineers how production is behaving right now, and the analytics pipelines show the business how the platform and the products get used over time. This role sits on the observability side. We run 200+ hosted deployments across every supported cloud region, ingesting logs, metrics and traces for all of Elastic Cloud, plus the SLA and SLO monitoring for ESS and Serverless. When Cloud engineering needs to know what production is doing, they're looking at something we run. What You Will Be Doing • Owning end-to-end delivery of moderate-to-high complexity projects on the team’s roadmap , with minimal day-to-day direction. • Operating and hardening shared Elastic Cloud infrastructure (ECH, ECE, and ECK) as Infrastructure as Code – writing and reviewing the Terraform, Python, and Go that other engineers depend on. • Carrying a 24/7 on-call rotation: responding to incidents, driving them to resolution, and writing clear RCAs/postmortems that lead to lasting fixes rather than repeat pages. • Reviewing others’ code and designs, and being a trusted second set of eyes on production changes to critical infrastructure. • Mentoring less experienced engineers, and proactively raising risks, ideas, and improvements in team discussions. • Improving runbooks, documentation, and operational processes so the on-call load gets lighter over time. What You Bring • 5+ years of SRE, platform engineering, or infrastructure engineering experience • Proficiency with Terraform; comfortable owning large, multi-workspace configurations in a team setting • Strong software engineering fundamentals in Python; comfort with Go is a plus. • Deep Linux systems knowledge and experience operating containerized workloads in production. • Experience carrying a 24/7 on-call rotation, resolving incidents under pressure , and writing RCAs that hold up under review. • A track record of consistently delivering end-to-end projects of moderate-to-high complexity with minimal oversight, and being a valuable code/design reviewer for your team. • Comfort thinking about the security implications of the infrastructure you build, not just its reliability – you don’t need to be a security specialist, but you default to a security-conscious mindset. • A pattern of mentoring less experienced engineers and speaking up with ideas and concerns in team discussions. • Clear written and verbal communication – you document what you build and can explain it to both engineers and non-engineers. • Comfort working across time zones, in both real-time and asynchronous contexts. Bonus Points • Experience with the Elastic Stack (Elasticsearch, Logstash, Beats, Kibana) in production. • Experience with GitOps-style deployment tooling (ArgoCD, Helm) or policy-as-code frameworks (e.g., Kyverno) on Kubernetes. • Experience with secrets management (Vault) or access-control/bastion tooling (Teleport). • Experience with configuration management tools (e.g., Puppet, Ansible) at fleet scale. • Exposure to FedRAMP, GovCloud, or other regulated/compliance-driven infrastructure. Compensation for this role is in the form of base salary. This role does not have a variable compensation component. The…

Sign in to open the employer’s application

This role is hosted on Elastic’s own site, and NEL is not part of that hiring process. An account is free, takes a minute, and lets you keep the roles you are following in one place.

You can sign up as a professional looking for work or as a client hiring for one. Either opens this application.

Looking for escrow-protected work?

Roles posted directly on NEL are paid through escrow, with funds held and released on agreed milestones.