Skip to Main Content

Opportunities

presentation

Maryland

Senior Site Reliability Engineer

We are Systematix and we are currently looking for a Senior Site Reliability Engineer – AI & Data Platform to establish the reliability, observability, and operational engineering capabilities supporting an emerging enterprise AI and Data ecosystem for one of our key clients.

ABOUT THE PROJECT
Our client is a global leader in science and technology, supporting a diverse portfolio of businesses across healthcare, life sciences, diagnostics, manufacturing, and industrial innovation. As investment in AI, machine learning, and advanced data capabilities continues to accelerate, the organization is building the engineering foundation required to operate these environments reliably and at enterprise scale.
Current monitoring, alerting, observability, and automated operational capabilities are still evolving. The successful candidate will have an opportunity to establish modern SRE practices rather than simply inherit an existing mature environment. Working closely with Platform Engineering, DevOps, and MLOps teams, this individual will help build the standards, tooling, automation, and operating practices required to deliver highly available, resilient, and observable AI and Data platforms.

ABOUT THE RESPONSIBILITIES

  • Design and implement monitoring, observability, and alerting capabilities across Microsoft Azure cloud infrastructure and AI/ML environments.
  • Establish Site Reliability Engineering practices, standards, operating models, and engineering principles.
  • Define and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and appropriate platform reliability metrics.
  • Build dashboards, telemetry, and actionable alerting that provide meaningful visibility into platform health and performance.
  • Implement observability solutions using technologies such as Prometheus, Grafana, OpenTelemetry, or comparable platforms.
  • Automate operational processes, remediation, recovery, and repetitive support activities wherever possible.
  • Improve platform availability, scalability, resilience, performance, and recoverability.
  • Develop proactive infrastructure health, capacity, and performance monitoring capabilities.
  • Support the reliability of production AI/ML applications, services, and GPU-intensive workloads.
  • Partner with Platform Engineering teams to incorporate reliability and resilience into cloud infrastructure architecture.
  • Partner with DevOps engineers to improve deployment reliability and automate operational processes.
  • Partner with MLOps engineers to establish monitoring and observability for model, inference, and supporting infrastructure.
  • Develop incident response processes, runbooks, operational tooling, and automated recovery capabilities.
  • Support capacity planning and performance management across cloud and AI infrastructure.
  • Lead and contribute to root-cause analysis and post-incident reviews.
  • Identify recurring operational issues and develop engineering solutions that eliminate or significantly reduce manual intervention.
  • Drive continuous improvement of platform reliability, observability, and operational efficiency.

ABOUT THE REQUIREMENTS

  • Extensive senior-level experience in Site Reliability Engineering, Production Engineering, Systems Engineering, or a closely related engineering discipline.
  • Strong hands-on experience with Microsoft Azure and enterprise cloud infrastructure.
  • Deep expertise in monitoring, observability, telemetry, and alerting.
  • Hands-on experience with Prometheus, Grafana, OpenTelemetry, or comparable observability technologies.
  • Demonstrated experience defining and implementing SLIs, SLOs, reliability metrics, and actionable alerting.
  • Strong Infrastructure as Code experience.
  • Advanced automation and scripting capabilities.
  • Strong experience with Docker, Kubernetes, and containerized production environments.
  • Experience with CI/CD pipelines and modern software delivery practices.
  • Demonstrated experience operating and improving complex production systems at enterprise scale.
  • Experience automating operational processes, remediation, and recovery activities.
  • Strong understanding of distributed systems, cloud architecture, scalability, resilience, and performance.
  • Exceptional troubleshooting, root-cause analysis, and systems-thinking capabilities.
  • Strong software engineering mindset with a demonstrated preference for solving operational problems through engineering and automation rather than manual processes.
  • Strong communication and collaboration skills with the ability to work across Platform Engineering, DevOps, MLOps, cloud, security, data, and application engineering teams.

PREFERRED QUALIFICATIONS

  • Experience supporting AI and machine learning infrastructure in production environments.
  • Experience supporting GPU, accelerated computing, or high-performance computing environments.
  • Strong Python development or scripting experience.
  • Experience with Azure Machine Learning and related Azure AI/data services.
  • Experience with MLOps platforms and machine learning lifecycle infrastructure.
  • Experience implementing observability for machine learning models, inference services, and supporting infrastructure.
  • Experience establishing an SRE capability, standards, or operating model within an immature or evolving engineering environment.
  • Experience working within large, global, or federated enterprise environments.

ABOUT THE ROLE
This is a remote contract opportunity supporting a strategic enterprise AI, Data, and Cloud engineering initiative.
The successful candidate must be based within, or able to consistently work, Eastern Time Zone business hours.
This is a senior, highly hands-on engineering position. The successful candidate will help establish the organization's SRE capability while personally designing and implementing the observability, automation, reliability, and operational tooling required to support the platform.
We are looking for a software and engineering-oriented SRE rather than a traditional operations or support professional. The ideal candidate approaches recurring operational problems as opportunities to engineer and automate them away.

AI DISCLOSURE
As part of our recruitment process, Systematix may use artificial intelligence (AI) tools to assist with resume screening, candidate matching, and recruitment administration. All hiring decisions are ultimately made by our recruitment and hiring teams.

APPLY NOW
If you are interested in finding out more, please contact us or submit your resume to jobs@systematix.com.
Know someone who would be a great fit? We welcome referrals of qualified candidates and are always interested in connecting with talented technology professionals.

ABOUT SYSTEMATIX
Systematix is a Canadian-owned Global Consulting and Resourcing firm with nearly 50 years of experience delivering technology solutions to clients across North America and the United Kingdom. We provide the highest-caliber consulting solutions to a diverse client base across all levels of government and private industry. Systematix is committed to creating a diverse, inclusive environment and is proud to be an equal opportunity employer. At Systematix, we value diverse perspectives, experiences, and backgrounds.
Systematix. Solutions Focused. People Driven.

 

BH 22418

Apply

Senior Site Reliability Engineer

Contact Us!