Benjamin Ryzman
2 October 2026
AI infrastructure for telcos
Why moving from pilots to production changes the platform
Artificial intelligence has moved beyond the demonstration phase. Telco operators are already exploring it for fault diagnosis, energy management, capacity planning, customer care, and radio access network (RAN) optimization. Enterprise customers, meanwhile, are asking for private inference, local data processing, and AI services that meet stricter requirements for latency, sovereignty, and control. For telco organizations, the opportunity that AI provides is real.
The difficult part is transforming this promising model into a dependable service. In an isolated, experimental environment, a notebook can detect an anomaly with ease. However, a production system must collect trustworthy data, run in the right location, survive software and model updates, fit established security controls, and tell an operator what went wrong at two in the morning. Moreover, telcos offering AI services are building them into estates which already carry critical workloads. There’s little room for error.
So how do telcos turn their assets – networks, sites, data, and operational discipline – into reliable services, without risking pre-existing systems?
Telcos already own part of the answer to the AI challenge
Telco operators begin with assets that many AI providers would struggle to reproduce: distributed sites near customers, high-capacity transport, access to live network signals, and the long-term experience of operating infrastructure under pressure. Those strengths matter when an industrial customer needs video analytics inside a factory, a public body needs data to remain within a jurisdiction, or a network team needs a prediction while there is still time to prevent an outage.
However, proximity alone is not a product. The advantage appears when locality, connectivity, compute, and operations are packaged as a repeatable service. That calls for a common foundation across sites, with supported hardware profiles, consistent deployment paths, and clear rules for data, identity, and service ownership.
One platform can support two routes to value
The first route runs inward. AI can help operators spot degrading equipment earlier, find patterns across alarms and tickets, forecast capacity, reduce energy waste, and shorten the path from a fault to a safe response. These gains come from closing the loop between data, inference, decision, and action. The model is only one step in that loop.
The second route faces the market. Telcos can offer private AI environments, distributed inference, and industry-specific services where a generic public cloud feels too far away, too opaque, or too detached from the network. Manufacturing, logistics, healthcare, and the public sector all have workloads shaped by local data, predictable latency, and firm control over where processing happens. Those needs fit the operator footprint rather well.
Both routes can share the same platform principles. Physical infrastructure should be provisioned through automation. The operating system must keep pace with CPUs, GPUs, data processing units, and their drivers. Virtual machines, containers, and physical servers should each be available where they make sense. Data and machine learning operations (MLOps) tools then turn the infrastructure into something that application teams can use safely and repeatedly.
The second deployment is the real test
When a team builds and deploys their very first AI product, it rarely runs smoothly on code and automation alone. Instead, it relies heavily on tribal knowledge: undocumented context and manual workarounds held in the heads of key team members. One person knows which dataset is safe. Another knows the driver version that works. Someone else watches the model after each update. That human scaffolding is useful while the team learns, but it becomes operational debt when the next service copies only part of the pattern.
Repeatability is therefore a better measure of production-readiness than a polished demonstration. Can a second team deploy the service without assembling the original experts? Can it be patched, audited, and rolled back through supported paths? Can the on-call engineer distinguish a network fault from stale data or a change in model behaviour? Production begins when those answers stop depending on institutional memory.
Build for change as well as scale
AI technology will keep moving. Accelerator generations, model-serving patterns, and regulatory expectations will change, often faster than the underlying facilities. A useful telco platform must absorb that movement, without forcing every site or service to start again. Open source helps by giving operators visibility into the stack, freedom across hardware, and more control over how software is integrated, maintained, and audited.
Canonical brings together many of the layers needed for that foundation: Ubuntu for the host, MAAS for automated server provisioning, Kubernetes and OpenStack for different workload models, MicroCloud for smaller sites, and operated data and machine learning tooling above them. The individual products matter, but consistency is the larger point. Teams need fewer one-off paths from the data centre to the edge, and from hardware installation to model operation.
Telcos already know how to run distributed, regulated, and highly available systems. AI does not make that experience obsolete, it makes it more valuable. The next challenge is to extend the same discipline across data, accelerators, models, and the services built around them, creating a platform that operators can trust internally and that customers can buy with confidence.
Build an AI foundation that can carry the load
Read the whitepaper for a practical view of AI-ready telco architecture, the use cases most likely to create value, and the operating model needed to take them beyond the pilot stage.