Applied Observability for AIOps

Applied Observability for AIOps

Introduction

The LFN Applied Observability for AIOps Working Group is an open-source collaboration under the Linux Foundation Networking (LFN) umbrella. The WG brings together subject matter experts from telecommunications operators, technology vendors, and the LFN CTO office to address a critical gap in the industry: there is no standardized, production-proven body of best practices for end-to-end observability and AI Driven Operations across modern telco network stacks.

The challenge is structural. A modern operator runs a Telco Data Estate that spans the radio access network, core network, transport, OSS, BSS, customer experience systems, security, IoT, and an increasingly distributed cloud and edge footprint. Each layer emits its own telemetry, often through its own protocol and tooling, with different owners and retention policies. The primary problem is not the lack of data, but identifying which pieces are relevant for which AIOps use cases, and then putting them to work safely.

The industry context is direct. McKinsey reports a $5.5T global skills deficit driving automation. Operators applying AI to network operations are seeing roughly 30 percent cost savings. Private 5G is projected to grow 4X by 2028. The longer-term direction is moving from connectivity providers to intelligence service providers, where AI becomes an intrinsic property of the network rather than an external service bolted onto 5G, ahead of 6G.

This Working Group is positioned to produce the open, validated reference operators needed to move from static silos to a living, self-healing network. The output is a Best Practice Guide validated through a simulated proof-of-concept environment.

Mission Statement

Develop and publish an open-source best practice guide for end-to-end observability and AI-driven operations across telco network stacks (RAN, transport, 5G/6G core, OSS, BSS, edge/MEC, NTN, enterprise IT), validated through a simulated proof-of-concept environment, enabling operators and vendors to adopt production-ready patterns on the path to true Autonomous Networks.

Core Method: The Four-Step AIOps Strategy

The WG aligns on a single method that applies across every domain in scope.

  1. Find the Data: Map silos and instrument the environment with OpenTelemetry and eBPF for kernel-level harvesting with zero code changes. Identify which signals are relevant for each AIOps use case.

  2. Apply Domain Expertise: Use a CRISP-DM style methodology to select the right features for high-value use cases such as fault prediction, energy optimization, and customer experience.

  3. Place Models on the Grid: Right-size models (4B vs 7B vs 70B vs 405B) and deploy them at the appropriate tier (Far Edge, Edge/RAN, Regional, Core/Cloud) based on latency budget and use-case fit.

  4. Govern the Agents: Use protocols like MCP (Model Context Protocol) and A2A (Agent-to-Agent) to manage drift, control latency, and ensure agents collaborate securely across the infrastructure.

Key Objectives

  1. Produce an observability landscape assessment grounded in operator reality (D2).

  2. Design an end-to-end reference architecture covering metrics, logs, traces, and events across every domain in scope (D3).

  3. Catalog and validate AIOps use cases with operator-relevant ROI signals (D4).

  4. Build a simulated proof-of-concept environment using synthetic telemetry seeded by operator profiles (D5).

  5. Publish a comprehensive best practice guide suitable for adoption by LFN member organizations and the broader telco community (D6).

North-Star KPIs for the WG output

  • TM Forum Autonomous Networks level demonstrated in PoC: target L3 (Conditional Autonomy) on at least one end-to-end use case.

  • Reference ROI on at least one validated use case from the catalog: 5G Fault Prediction (40 percent fewer outages), Customer Churn Reduction (15-25 percent), Revenue Assurance / RAFM (2-5 percent recovery), Energy Optimization (20-30 percent savings), Service Assurance Latency Prediction.

Scope

In Scope

Telco domains:

  • Radio Access Network (RAN): O-RAN components, CU/DU/RU observability, RIC telemetry (Near-RT and Non-RT), fronthaul and midhaul monitoring, AI-RAN.

  • Transport: xHaul, IP/MPLS/Segment Routing, optical and DWDM, microwave, time-sensitive networking, SDN controllers.

  • 5G/6G Core: control-plane and user-plane function observability, service-based architecture (SBA) tracing, network function lifecycle, IMS, packet core.

  • OSS: service orchestration, inventory and CMDB, service assurance, observability platforms, AIOps.

  • BSS: CRM, order management, billing and charging, product catalog, customer experience.

  • Edge and MEC: telco-grade Kubernetes, edge UPF distribution, hyperscaler edge offerings (Wavelength, Edge Zones, Distributed Cloud Edge).

  • Non-Terrestrial Networks (NTN): satellite and HAPS integration, latency-aware telemetry, 3GPP NTN Rel 17/18.

  • Cloud Infrastructure: NFVI (VM and K8s), hybrid cloud and sovereign deployment patterns, GPU and AI accelerator infrastructure.

Cross-cutting technologies:

  • Observability: OpenTelemetry MELT (Metrics, Events, Logs, Traces), eBPF, gNMI, streaming telemetry, distributed tracing for SBA.

  • AI and ML platform: vLLM, llm-d distributed inference, KServe, MCP, A2A, LlamaStack, model catalog and registry.

  • Infrastructure: Kubernetes-native platforms (OpenShift, Rancher, Wind River Studio, etc.), multi-cluster observability, hybrid cloud.

  • Security: zero-trust telemetry access, PII masking at source, identity and policy (RBAC, OPA/Gatekeeper), security observability spanning UE, O-RAN, 5G Core, and Cloud (per ETSI ISG MEC threat model).

  • Models: Classic ML (XGBoost, Random Forest, BERT, Regression), Generative AI (open-weight LLMs and SLMs), Hybrid AI patterns (Classic detects, GenAI explains).

Out of Scope

  • Vendor-specific proprietary tool implementations or product recommendations.

  • Production deployment of observability infrastructure at operator sites.

  • Standardization or specification work. This is a best-practice guide, not a formal standard.

  • IMS, VoLTE and VoNR voice services may be addressed in a future phase.

Governance and Operations

 

Stakeholders and Roles

Name

Organization

Role

Responsibilities

Ravi Sharma

 

Red Hat

Project Technical Lead

Overall technical direction, cloud-native observability, PoC lead, best practice guide principal author.

Fatih E. NAR

Red Hat

WG Vice Chair

Stakeholder alignment,scope governance, escalation path between WG and LFN CTO, cross-organizational coordination.

Tony Hansen

AT&T

Security and Observability TAC Lead

Security observability, ONAP-derived patterns, OpenSSF best practices, AT&T sponsor.

Murat Parlakisik

Amazon (AWS)

AI/ML Architecture Lead

AI/ML architecture, NTN domain, reference architecture lead, 6G sandbox host for D5 PoC.

Ranny Haiby

LFN

LFN CTO / Governance Sponsor

LFN alignment, governance oversight, final presentation sponsor, and community outreach.

(TBD)

Turkcell

Operator Representative

Operator validation, use-case input, data seeds for synthetic telemetry.

(TBD)

Verizon

Operator Representative

Operator validation, use-case input, data seeds for synthetic telemetry.

Deliverables and Scope

The WG will produce seven deliverables. Each carries a defined scope, lead owner, supporting stakeholders, and target completion date. Sections 6.1 through 6.7 below detail each deliverable with method and concrete artifact list.

D1 - WG Charter and Governance

Target: May 2026    Lead: Vice Chair

Scope: This document plus governance artifacts. Mission statement, roles and responsibilities, meeting cadence, decision-making process, IP policy, communication channels, repository setup.

Artifacts: Charter v1.x (Google Doc), Charter Deck v1.x (Google Slides), GitHub repo bootstrap, mailing list, wiki home, scheduled meeting series.

D2 - Observability Landscape Assessment

Target: June 2026    Lead: Security and Observability TAC Lead

Scope: Survey current observability tools, protocols, and gaps across RAN, Transport, Core, OSS, BSS, Edge, NTN, and Cloud Infra. Triangulated view from AT&T, Verizon, Turkcell.

Method: Per-operator fill-in templates capturing vendor and tech stack, data silos and telemetry, AIOps data features, pain points, and TM Forum AN roadmap. Outputs consolidated by WG editors into an anonymized landscape view.

Artifacts: Three per-operator workbooks (one each for AT&T, Verizon, Turkcell) returned to WG. Consolidated landscape report. Pain-points heatmap. Gap analysis feeding into D3.

D3 - Reference Architecture

Target: July 2026    Lead: AI/ML Architecture Lead (AWS)

Scope: End-to-end observability and AIOps reference covering the full stack.

Five-layer Cloud-Native AI Infrastructure Stack:

  • L5 Runtime: diagnostic, planning, validation agents.

  • L4 Protocols: MCP for model-to-tool integration, A2A for peer-to-peer agent collaboration.

  • L3 Inference Engine: vLLM + KServe for high-throughput open-weight model serving; llm-d for distributed inference across tiers.

  • L2 Platform: Kubernetes-native AI/ML platform, hybrid cloud.

  • L1 Hardware: NVIDIA, AMD, Intel, with open silicon choice.

Distributed AI Grid (4 tiers):

  • Tier C - Core DC: 400-500ms, 70B-405B params, FP16 / BF16. Role: train, fine-tune, governance, audit.

  • Tier R - Regional DC: 20-50ms, 7B-70B FP8 quantized. Role: distill, digital twin, sovereign RAG.

  • Tier E - Edge / RAN: 5-30ms, 1B-8B Domain SLMs, INT8 / INT4. Role: quantize, E2 xApps, beamforming, energy optimization.

  • Tier F - Far Edge / Device: 1-15ms, less than or equal to 4B params, INT4. Role: on-device inference, API handoff, IoT.

Intelligence lifecycle: Train at Core to Distill to Region to Quantize to Edge to Handoff to Device.

Governance Guardrails:

  • Drift Monitoring: auto-retrain when accuracy drops.

  • Latency Control: meet 3GPP URLLC where required (under 1ms user plane).

  • Hallucination Control: domain-grounded RAG plus human-in-the-loop QA.

  • Right-Sizing principle: do not use a 70B when a 7B gets the job done. Memory equals parameters times precision times 1.1x overhead.

Closed-loop reference flow: sub-200ms full closure. Device hint (~40ms), RAN xApp localizes (~60ms), regional digital twin (~120ms), core governance check (~150ms), edge policy push (~180ms).

Artifacts: Architecture document (http://draw.io and Markdown), reference deployment topology, protocol and API map.

D4 - AIOps Use Case Catalog

Target: July to August 2026    Lead: AI/ML Architecture Lead with operator co-leads

Scope: Operator-validated use cases with technique-to-problem mapping. The WG explicitly opposes hype: classic ML where it wins, GenAI where unstructured data and NLP are involved, Hybrid (Classic detects, GenAI explains) for production telco.

Five anchor use cases with target ROI signal:

  • 5G Fault Prediction: 40 percent fewer outages (XGBoost + BERT mixture of experts).

  • Customer Churn Reduction: 15-25 percent reduction (regression on CRM + NPS).

  • Revenue Assurance (RAFM): 2-5 percent recovery (Random Forest + Neural Transformers, Hybrid pattern).

  • Energy Optimization: 20-30 percent savings (AI-RAN + smart grid, Classic ML).

  • Service Assurance Latency Prediction (Neural Nets, Classic ML).

Method: Each use case is documented with data sources, feature engineering, model choice, deployment tier, governance considerations, and KPI of success.

Artifacts: Use Case Catalog (Markdown plus workbook), per-use-case data feature spec, Telco-AIX experiment links.

D5 - Simulated PoC Environment

Target: August to September 2026    Lead: AI/ML Architecture Lead with Vice Chair

Scope: Working sandbox that validates a subset of the architecture and use cases end-to-end on synthetic telemetry seeded by operator data profiles.

Foundation: AWS / LFN AI TAC contributor open-source 6G sandbox repository. Telco-AIX repository 26+ pre-built AI experiments serve as starting baselines.

Validation goals:

  • Prove sub-200ms closed loop on at least one use case.

  • Demonstrate at least one ROI-positive run from the catalog (e.g. 5G Fault Prediction with measurable detection lead time).

  • Show hybrid-cloud deployment across at least two of the four AI Grid tiers.

  • Show closed-loop remediation triggered without human action on a known pattern, with novel patterns escalated to human-in-the-loop.

Artifacts: Sandbox repo (forks / branches under Open-Experiments), synthetic telemetry generators with statistical operator profiles, validation harness, demo recordings.

D6 - Best Practice Guide

Target: September to October 2026    Lead: Vice Chair (principal author) with TAC leads

Scope: Consolidated publication suitable for LFN community adoption.

Content map:

  • Background and operator context.

  • The 4-step AIOps strategy.

  • 4-stage data pipeline pattern (Observe / Instrument / Govern / Unify) with technology choices and trade-offs.

  • 5-layer AI infrastructure stack.

  • Distributed AI Grid (4 tiers) with model-to-tier mapping rules.

  • Governance Guardrails for production AI.

  • Sub-200ms closed-loop reference.

  • Five validated use cases with implementation playbooks.

  • Security observability and zero-trust pattern (UE, O-RAN, 5G Core, Cloud).

  • Sovereign AI considerations for regulated regions.

  • Sustainability and cost discipline (GPU utilization, FinOps for AI).

  • Lessons learned from D5 PoC.

Artifacts: PDF and HTML guide under Apache 2.0, executive summary deck, reference diagrams (http://draw.io sources).

D7 - Final Review, Demo and LFN Presentation

Target: November / December 2026    Lead: WG Vice Chair plus LFN CTO

Scope: PoC demo with simulated data, guide walkthrough, stakeholder sign-off, LFN board presentation, publication.

Artifacts: Recorded demo (~10 min), final stakeholder sign-off, LFN board deck, blog announcement, project page on LFN site.

Timeline and Milestones

The project spans Q2 through Q4 of 2026 across five phases.

Month

Phase

Key Activities

May 2026

Foundation

Charter ratification and kickoff. Set up shared repos and meeting series.

June 2026

Foundation

D2 landscape assessment from operator returns. Begin reference architecture.

July 2026

Architecture and Design

Deliver D3 reference architecture and start D4 use case catalog. Mid-project checkpoint.

August 2026

Validation

D5 PoC sandbox operational. Begin validation runs for selected use cases.

September 2026

Validation to Publication

Finalize D5 results. Draft D6 guide incorporating PoC learnings.

October 2026

Publication

Peer review on D6, incorporate feedback, finalize.

November / December 2026

Delivery

D7 final demo, sign-off, LFN board presentation, community publication.

 

The AI Observability Working Group. This meeting is open to the public; however, voting and the agenda will be set by the members of the working group.

TBD — 60 minute meeting

Meeting Registration / Join

Zulip

Request a Topic at an upcoming TSC Meeting:

You will need to Sign In and click the Edit button near the top of the screen to add agenda items.

Type your topic here, using "@" to assign to a user and "//" to select a due date

LF Anti-Trust Policy

Meeting Recordings (look in Past Meetings tab)

All TSC Minutes (by year)

Register to participate

Create from Template Pro (Inline)