Observability Engineer Specialist

Discovery – Group Information Systems - Digital Channels
Observability Engineer
About Discovery
Discovery’s core purpose is to make people healthier and to enhance and protect their lives. We seek out and invest in exceptional individuals who understand and support our core purpose, and whose own values align with those of Discovery. Our fast-paced and dynamic environment enables smart, self-driven people to be their best. As global thought leaders, Discovery is passionate about innovating in order to not only achieve financial success, but to ignite positive and meaningful change within our society.
About Digital Channels
Working in a high performance organization that prides itself in attracting the finest talent, we challenge ourselves to find solutions that make a difference in the world. Our environment is always buzzing with energy and smart, motivated people working on finding the best way to move forward.
The Digital Channels team works on dynamic new projects and product enhancements within the web and mobile platforms in order to improve business inefficiencies, gain competitive advantage on our products and ultimately to provide better service to our clients. Using knowledge of the organization’s technology infrastructure and specific software applications, Application Platform Services helps the business to address changes through technologies.
Key Purpose
The Observability Specialist at Discovery Limited plays a critical role in implementing and maintaining observability solutions that provide end-to-end visibility into system performance and reliability. The role involves deploying and optimising observability platforms, designing telemetry pipelines, and ensuring integration with cloud and container environments. It requires strong technical expertise, collaboration across engineering teams, and the ability to drive continuous improvement in monitoring, alerting, and automation practices. This position supports incident response, root cause analysis, and the definition of reliability metrics, contributing to a culture of operational excellence and resilience.
Areas of responsibility may include but not limited to
- Deploy, configure, and maintain observability platforms such as Prometheus, Grafana, Dynatrace, ELK Stack, and OpenTelemetry to ensure reliable and scalable monitoring solutions. Decisions in platform setup directly impact system visibility and operational performance.
- Design and optimize observability systems to provide comprehensive visibility across infrastructure, applications, and services. These design choices influence monitoring accuracy and system resilience.
- Monitor system health and analyze telemetry data, including metrics, logs, and traces, to detect anomalies and performance issues. Insights from this analysis guide proactive remediation and reliability improvements.
- Define and track SLIs, SLOs, and error budgets to align observability practices with reliability objectives. These metrics inform prioritization of engineering tasks and risk management strategies.
- Participate in incident response activities, contributing to root cause analysis and post-incident reviews to improve system resilience. Decisions during incidents affect recovery time and future prevention measures.
- Implement observability solutions in cloud and containerized environments, ensuring coverage across distributed systems. These implementations impact scalability and monitoring consistency.
- Automate observability workflows and integrate observability into CI/CD pipelines using tools such as Jenkins, GitHub Actions, and GitLab CI. Automation decisions reduce operational overhead and improve deployment efficiency.
- Write scripts in Python, Bash, or PowerShell to automate observability tasks and support operational workflows. These scripts enhance efficiency and reduce manual intervention.
- Maintain observability documentation, including architecture diagrams, configuration standards, and operational playbooks, to ensure consistency and knowledge sharing across teams.
- Collaborate with cross-functional teams to enhance monitoring coverage, improve alerting accuracy, and promote observability best practices. Engagement includes working with senior engineers, technical leads, and operations managers to align observability strategies with business objectives.
Stakeholder Engagement: - Internal Stakeholders: Development teams (junior to senior engineers), Infrastructure and operations teams (technical leads and managers), Site Reliability Engineering (SRE) teams, CI/CD pipeline owners and automation specialists
- External Stakeholders: Vendor support teams for observability tools, Third-party service providers for cloud and monitoring solutions
Personal Attributes and Skills
Behavioral Skills
- Excellent written and oral communication skills (English)
- Ability to work in a self-driven, complex environment with multiple and changing priorities
- Ability to focus on deadlines and deliverables
- Ability to think abstractly
- Ability and desire to quickly learn new technologies
- Clean code thinking
Technical Skills
- Expertise in Prometheus, Grafana, Dynatrace, ELK Stack, and OpenTelemetry for monitoring, logging, and tracing.
- Ability to analyse metrics, logs, and traces to identify anomalies, trends, and root causes.
- Experience automating observability workflows and embedding observability into CI/CD pipelines using Infrastructure as Code tools.
- Strong skills in Python, Bash, or PowerShell for automation and operational support
- Ability to interpret complex telemetry data and recommend effective solutions.
- Maintains effectiveness under pressure and adapts to evolving tools, processes, and organisational needs.
- Actively seeks opportunities to enhance observability practices and adopt emerging technologies
- Works effectively with development, infrastructure, and operations teams to achieve shared goals.
- Supports team enablement, knowledge sharing, and onboarding of new engineers.
Education and Experience
Minimum
- Bachelor’s degree in computer science or information technology, or related field
- AWS Certified Solutions Architect or equivalent
- ITIL Foundation Certification
- Advanced scripting skills in Python or Bash
- 5+ years of experience in observability, monitoring, or site reliability engineering
- Experience with observability platforms and telemetry pipeline design
- Familiarity with cloud platforms and container orchestration
Advantageous
- Certified Kubernetes Administrator (CKA)
- Experience in Agile or DevOps environments
- Working experience within IT infrastructure environment
EMPLOYMENT EQUITY
The Company’s approved Employment Equity Plan and Targets will be considered as part of the recruitment process. As an Equal Opportunities employer, we actively encourage and welcome people with various disabilities to apply.