Observability Technical Lead

Apply

Sign up to receive career updates before completing the application

Note: You will complete the application on the next page


Skip & Continue

Job Number: 11487

External Description:

Key Responsibilities

Enterprise Observability Platform Engineering

  • Serve as the senior SME for enterprise observability, defining architectures, standards, integration patterns, and reusable solutions
  • Design and optimize observability across cloud, on-prem infrastructure, networks, Kubernetes, containers, applications, APIs, databases, middleware, and enterprise platforms
  • Standardize metrics, logs, traces, events, topology, dashboards, alerting, instrumentation, and service health
  • Provide technical leadership for LogicMonitor, Splunk, OpenTelemetry, Grafana, Prometheus, Tempo, VictoriaMetrics, Loki, and related technologies
  • Lead OpenTelemetry adoption, including instrumentation, collectors, telemetry pipelines, distributed tracing, context propagation, and vendor-neutral standards
  • Design telemetry pipelines that route metrics, logs, and traces across multiple platforms
  • Develop API-driven and Observability-as-Code capabilities for onboarding, configuration, validation, and lifecycle management
  • Establish governance for RBAC, telemetry standards, alerting, retention, integrations, configuration management, and data lifecycle
  • Improve observability coverage, telemetry quality, scalability, reliability, performance, and cost efficiency

Observability Automation, AIOps & Reliability

  • Drive automation and self-service through APIs, Infrastructure-as-Code, CI/CD, GitOps, and reusable observability pattern
  • Integrate Edwin AI and AIOps capabilities for anomaly detection, event correlation, root-cause analysis, investigation, and operational intelligence
  • Connect metrics, logs, traces, events, topology, CMDB, service ownership, and operational context to deliver end-to-end visibility
  • Establish reliability standards including SLIs, SLOs, error budgets, monitoring coverage, alert quality, MTTD, and MTTR
  • Reduce alert fatigue through intelligent correlation, dynamic thresholds, suppression, automation, and event-management practices
  • Evaluate eBPF, continuous profiling, Kubernetes observability, dependency mapping, and auto-instrumentation
  • Advance operations from reactive monitoring toward proactive and predictive operations

Global Technical Leadership

  • Collaborate with engineering and operations teams across North America, Europe, and Asia
  • Lead architecture reviews, platform evaluations, workshops, and technical working sessions
  • Influence observability strategy and standards across teams without direct authority
  • Mentor engineers and promote observability, SRE, and reliability best practices
  • Evaluate emerging technologies and recommend adoption based on interoperability, scalability, business value, and cost
  • Translate strategy into standards, reference architectures, and reusable implementation patterns

Reasonable schedule flexibility is required to support global collaboration.

Required Qualifications

  • Bachelor’s Degree
  • 7+ years of experience in observability, monitoring, APM, SRE, DevOps, platform engineering, or related professional experience
  • 7+ years of experience designing, implementing, operating, and maintaining enterprise-scale observability platforms using observability-as-code, monitoring-as-code, infrastructure-as-code, GitOps, and CI/CD practice

Preferred Qualifications

  • Experience with AWS, Azure, GCP, and large-scale Kubernetes environments
  • Experience designing OpenTelemetry Collector architectures and telemetry pipelines
  • Expertise with Grafana, Tempo, Loki, and Prometheus-compatible platforms
  • Experience with VictoriaMetrics or other large-scale time-series databases.
  • Knowledge of eBPF, continuous profiling, auto-instrumentation, and cloud-native telemetry
  • Experience integrating observability platforms with ServiceNow, ITSM, CMDB, incident management, and automation platforms
  • Understanding of SRE practices including SLIs, SLOs, error budgets, and incident management
  • Experience enabling developer self-service and internal developer platform integrations
  • Knowledge of RBAC, secrets management, governance, compliance, and telemetry data protection

Job Number: 30215676

Community / Marketing Title: Observability Technical Lead

Location_formattedLocationLong: Florida, US