Senior Site Reliability Administrator

 

Description:

As a Senior Site Reliability Administrator, you will play a critical hands-on role within a globally distributed SRE team focused on ensuring the reliability, performance, and stability of the data services that power our client-facing SaaS platforms. You’ll work across key components of our distributed systems ecosystem, including Kafka, Elasticsearch, Cassandra, Solr, Redis, and OpenSearch supporting both on-premises infrastructure and public cloud environments (AWS, Azure, GCP).

This role is perfect for engineers who thrive on solving complex operational challenges, driving automation, and improving system reliability at scale. You’ll collaborate closely with cross-functional teams, strengthen your expertise in distributed systems, and contribute to building resilient, high-quality services that power critical applications for client around the world.

What The Role Offers
 

  • Operate, maintain, and scale distributed data platforms including Kafka, Elasticsearch, Cassandra, Solr, Redis, and OpenSearch across on-premises and public cloud environments (AWS, Azure, GCP).
  • Build, enhance, and support infrastructure using Infrastructure-as-Code (IaC) tools such as Terraform and Ansible.
  • Perform patching, upgrades, and routine maintenance to ensure systems remain secure, stable, and compliant with internal standards.
  • Collaborate with SRE and engineering teams to design, deploy, and monitor data platforms and supporting infrastructure.
  • Participate in incident response, troubleshooting, and root cause analysis, while contributing to a 24x7 on-call rotation for critical services.
  • Support capacity planning, performance tuning, and system health assessments to ensure optimal reliability and scalability.
  • Develop and maintain technical documentation, including operational procedures, change plans, and incident reports.
  • Contribute to automation and reliability improvements, support service requests, help meet SLA/OLA commitments, and actively participate in team knowledge sharing and training initiatives.
     

What You Need To Succeed
 

  • Bachelor’s degree in Computer Science, Engineering, Information Systems, or a related field - or equivalent practical experience.
  • 4+ years of experience in Information Technology supporting large‑scale enterprise systems.

Organization Open Text
Industry IT / Telecom / Software Jobs
Occupational Category Senior Site Reliability Administrator
Job Location Ontario,Canada
Shift Type Morning
Job Type Full Time
Gender No Preference
Career Level Experienced Professional
Experience 4 Years
Posted at 2026-07-24 3:41 pm
Expires on 2026-09-07