Observability Engineer

Job Summary
As part of the Discovery Central Services - Technology Services, the Observability Engineer role at Discovery Limited is responsible for the implementation, delivery, support, maintenance, and continuous enhancement of observability, AIOps, and IT Operations Management (ITOM) capabilities that enable the performance, availability, scalability, reliability, and operational visibility of technology services. Working closely with senior engineers, specialists, and operational teams, the role ensures effective monitoring and management of infrastructure, applications, networks, and services through the configuration, optimisation, and integration of observability platforms and ITOM solutions.
The role is accountable for designing, maintaining, and improving telemetry pipelines, supporting cloud-native and containerised environments, and delivering AIOps capabilities that leverage analytics, automation, and intelligent event correlation to improve operational efficiency and reduce service disruption. This includes enhancing monitoring coverage, alert effectiveness, automated remediation, service mapping, discovery, and integration of observability capabilities using various software technologies.
The Observability Engineer collaborates across development, infrastructure, platform, and operations teams to support incident management, problem management, and continuous service improvement initiatives. The position requires strong technical expertise, analytical thinking, and a proactive approach to identifying risks, optimising system performance, and improving service reliability through observability data and operational insights.
The role plays a key part in advancing observability, AIOps, and ITOM maturity across the organisation, enabling data-driven decision-making, improving operational resilience, and fostering a culture of reliability, automation, and continuous improvement throughout the technology landscape.
Key Responsibilities
- Lead the delivery, implementation, support, configuration, and continuous improvement of Observability, AIOps, and ITOM platforms and capabilities to enhance service performance, reliability, and operational visibility.
- Design, implement, and maintain scalable telemetry pipelines that collect, process, and analyse metrics, logs, traces, and events across infrastructure, applications, networks, and services.
- Configure, optimise, and manage observability and monitoring solutions, including technologies such as Dynatrace, Prometheus, Grafana, ELK Stack, Datadog, OpenTelemetry, and related platforms.
- Analyse telemetry and operational data to proactively identify trends, anomalies, performance bottlenecks, and reliability risks, leveraging insights to support AIOps-driven operations, service optimisation, and continuous improvement initiatives.
- Define, implement, and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to measure service health, guide engineering priorities, and support reliability objectives.
- Support and enhance ITOM capabilities, including Discovery, Service Mapping, Event Management, and automation initiatives to improve operational efficiency and reduce service disruption.
- Collaborate with development, infrastructure, platform, and operations teams to improve monitoring coverage, alert quality, system resilience, and end-to-end operational visibility across cloud, hybrid, and containerised environments.
- Participate in incident, problem, and change management activities by providing observability insights, performing root cause analysis, supporting post-incident reviews, and driving preventative improvements.
- Develop and implement automation, scripts, and operational workflows that streamline observability processes, reduce manual effort, and improve service reliability and operational effectiveness.
- Maintain observability standards, dashboards, documentation, and best practices while providing observability, AIOps, and ITOM expertise by promoting knowledge sharing, operational excellence, and continuous service improvement.
Stakeholder Engagement:
- Internal Stakeholders: Development, Infrastructure & Operation teams (across the Discovery Group)
- External Stakeholders: Vendor support teams for observability tools, Third-party service providers for cloud and monitoring solutions
Skills
- Agent Deployment & configuration
- Dynatrace Alert handling
- Data Integrity
- Data Collection
- Monitoring
- Dashboard Development / Dashboards
- Infrastructure Automation
- Incident and Problem Tracking
- CI/CD
- Infrastructure as Code
- Amazon Web Services / AWS
- Microsoft Azure
- Kubernetes
- Docker
- Big Data
Qualifications
Certifications
Observability Platform Certification
Work Experience
5+ years' experience in IT Operations, Infrastructure, Monitoring, Observability, SRE, or a related technology field.
3+ years' hands-on experience with observability platforms such as Dynatrace, Prometheus, Grafana, ELK, OpenTelemetry, Splunk, or similar.
Experience implementing and supporting Observability, AIOps, and ITOM capabilities, including Event Management, Discovery, and Service Mapping.
Experience monitoring enterprise infrastructure, applications, cloud platforms, and containerised environments.
Strong experience in incident management, problem management, root cause analysis, and service reliability improvement.
Practical experience with automation and scripting using Python, PowerShell, Bash, DQL, SQL or APIs.
Practical experience with networking technologies
Experience integrating observability solutions within ITSM platforms such as ServiceNow.
Proven ability to collaborate across development, infrastructure, and operations teams to improve service performance, availability, and operational efficiency.
Preferred
- Experience with Dynatrace and ServiceNow ITOM.
- Knowledge of Kubernetes, cloud-native technologies (Azure, AWS, GCP), and SRE practices.
EMPLOYMENT EQUITY
The Company’s approved Employment Equity Plan and Targets will be considered as part of the recruitment process. As an Equal Opportunities employer, we actively encourage and welcome people with various disabilities to apply.