Three musts before AI models scale

May 22, 2026
5
mins read

In this week’s newsletter, Rakuten Mobile Chief Consultant for AI and Data Petrit Nahi reflects on decades of research and more than six years of work scaling AI deployments, revealing key requirements for data platform, organizational structure and operational system readiness.

AI models are ready. They have been for quite some time.

I know because decades ago, my research work focused on distributed multi-agent systems for dynamic network coverage that are not substantially different from what’s finally being implemented today. While architectures, tooling, accessibility and compute power have evolved dramatically, the foundational principles behind these systems remain fundamentally unchanged.

The data platforms and operational systems gap is now finally closing, in part because networks have evolved in terms of data availability and accessibility. Yet not every organization is building the data platforms, operational systems and organizational structures required for AI to be successful.

I recently discussed this journey with Fierce Network at the FutureNet event in London, where I shared Rakuten Mobile’s experience overcoming the foundational hurdles that are required to enable true AI progress.

Even as a so-called greenfield operator working with a clean slate, we were not immune to the challenges of getting the foundations right on the first try.

What we learned is that autonomy is not primarily an AI problem, but rather a data, organizational and platform problem. These support structures and operational pipelines need to be ready as well as the models.

Own the data or lose the loop

One of the most consequential decisions we made was to centralize access to data as well as the data itself. This fundamental building block formed the support structure for everything else we wanted to do related to planned increased levels of automation.

We established a data science department and gave it ownership of all critical telemetry, traces, probes and performance data across the RAN and core.

Note, I didn’t say the RAN team or the core team. There is one data science team.

Historically the data collection tools would be deployed and owned by separate departments operating in isolation. The problem with this approach is that if you don’t have the data, you don’t have the information you need to make the right decisions.

These departmental separations exist today in tier-one telco organizations. We frequently see fragmentation across OSS, BSS and domain-specific teams, each running their own domains of responsibility. They often have the additional challenge of being vendor islands that speak different languages than other islands. We proactively avoided domain fragmentation, but we also depend on the same vendors, for all operators there is a need to create a data “Esperanto” so there is one common data language independent of vendor.

All these consequences span beyond just inefficiency. Fragmented data ownership institutes a firm ceiling on how far autonomy can go. No model, however sophisticated, can make good decisions if it is flying blind from data it cannot reach.

The data you’re collecting isn’t always the data you need

Centralizing data is only half of the battle. The other half is understanding what specific data matters for autonomous decision-making.

For a long time in telecom, there was reliance on sharpening the network element view of how the network was performing. The past decade has seen a move towards a subscriber-centric view—understanding what is that overall experience and considering the answer as critical data input in the decision-making process.

The shift from network-element-centric to subscriber-centric customer experience data is an exponential jump in volume, cost and platform requirements.

The result is a many-fold increase in data volume that must be processed. Data that looks much different from traditionally aggregated network element data. When you evolve to the subscriber-centric view you start to deal with experiences at the individual subscriber level. The cost to process, analyze and learn from this data, and prep it to inform decisioning is not insignificant.

In Rakuten Mobile, we have hundreds of thousands of cells from multiple vendors that are constantly feeding data into predictive and reactive AI models. While our virtual, cloud-native environment has made scaling more manageable than it might be in a traditional network, the numbers are still daunting by traditional standards. This is invaluable insight as the subscriber-centric view is not just about more data but the right data.

It is what gives autonomous systems the context required to act with confidence versus aggregated proxies that smooth over the signals you need.

Closing maintenance windows and pivoting to continuous action

In the old world we made changes during scheduled maintenance windows via controlled, rules-based actions. Yes, there were SON solutions, but they were narrow and specific.

AI-powered environments demand intent-driven applications operate continuously. Intent such as “I want to save energy” is defined, degrees of freedom are specified and the application takes a range of actions to configure the network to conserve energy when there is no traffic.

Parallel applications run simultaneously, each tuned to fulfill the intent, energy saving, coverage compensation and so on. They are coordinated to avoid clashes between rules and intents.

Human engineers can’t execute at this scale and speed.

But this is also where complexity multiplies. Each application is pursuing its own intent simultaneously. An energy saving application cannot be putting cells into power saving mode while a coverage compensation application is trying to recruit those same cells to cover an outage. The coordination layer between these applications is what prevents autonomous action from devolving into chaos.

Importantly, humans must be able to see and trust the process and outcomes.

At Rakuten Mobile, our confidence building journey started with small cells, then maintenance windows only, then carefully selected non-sensitive regions before expanding more broadly.

From an autonomy perspective, confidence is established from trusting the inventory and topology information in your operational environment, making sure they truly represent the state of the network.

A human engineer acting on a single network element makes one decision at a time. These applications are evaluating and acting across the entire network continuously, resulting in exponential speed growth.

In practice, we are operating at roughly a thousand times the pace a human-driven operation could sustain.

What breaks when autonomy scales

We didn’t move overnight from maintenance windows to autonomous operations at scale. We had to put guardrails in. We needed to know the right data existed and was accessible, understand and push the limits of the northbound and southbound network system capabilities and capacities and so on.

This meant constantly slowing things down and speeding back up. If the speed at which we were getting information wasn’t sufficient to make decisions in a timely fashion, we’d have to go back and rescale the OSS and other downstream systems.

Another stressor was the continued process of setting, changing and taking actions in the network. A cloud-native network afforded greater flexibility to scale in response to these demands, but that flexibility was not uniform. Further downstream, where systems were never designed for this speed and volume, challenges still arose.

What helped was our use of open RAN and a perspective that it is not just a technology but actually a mindset. When there are clean, open interfaces across the network, you can get the data you need at the speed you need it, without fighting the architecture every step of the way.

The first use case is the hardest

Given where we are in the journey, today we have several solutions running autonomously in the network at Rakuten Mobile.

At Mobile World Congress, we announced Rakuten Mobile received world first Level 4 autonomy validation from TM Forum for RAN energy efficiency optimization in a live, open RAN network. We’ve got other autonomous workflows live and in progress.

The sequencing of use cases matters as much as the foundations that support them. Energy efficiency was the right place for us to start. We determined that the blast radius of a wrong decision would be small and could be recovered from quickly.

This made it the ideal environment to build system confidence and organizational trust before moving into higher-stakes territories.

The first use case is always the hardest because you are building the supporting system in tandem. But once it’s in place? Each new use case is not starting from scratch because you are now transferring learning to a framework that has already been validated.

Three foundational musts, no matter where you start

Regardless of whether you are starting from scratch or transforming a legacy network, the path to autonomous network operations comprises the same three requirements:

  1. Consolidate your data platform first. This is the foundation everything else depends on because if the data is fragmented, siloed or inaccessible, no model will compensate for it. At FutureNet, operators noted they were still working on this critical step.
  2. Restructure the organization around the data. A consolidated data platform is only as strong as the organizational structure around it. Ownership and accountability for data instrumentation have to live in one place, with the structure following the data, not the other way around.
  3. Ensure your systems can handle AI-era demands. The volume and speed of autonomous action will stress-test every system upstream and downstream, necessitating the platforms underneath AI be built for that load before it arrives.

Operators that treat autonomy as an AI problem will keep stalling at the same points. The ones that treat it as a data, organizational and platforms opportunity will find that once the foundations are right, the pace of progress changes substantially.

AI
Data
Related Newsletter
Tracking AI usage metrics that matter
At DTW Ignite 2026 in Copenhagen, Abe Nejad of The Network Media Group hosted Rakuten Symphony SVP Faiq Khan and Telenor SVP, Global Business Security Officer and Head of Network and Cloud Technology Strategy Terje Jensen to discuss KPIs telcos should watch closely in the AI era. This article explores why activity-based AI metrics like token consumption and transaction volume are the wrong measure of success, what enterprises are already demanding as a result and how outcome-based measurement is starting to take shape.
August 7, 2026
5
MINUTES
AI and autonomy building blocks begin with cloud-native
On a recent episode of Zero-Touch Live, Boost Mobile's Dawood Shahdad, SVP of Hybrid MNO Network, and Sruthi Nair, Director of Voice Core Engineering, spoke with Rakuten Symphony's Anshul Bhatt about the journey of building one of the few operator networks running fully cloud-native on public cloud. In this week’s Zero-Touch newsletter, we dive deeper into the role cloud-native plays in powering autonomy and the essential building blocks that must accompany it.
June 25, 2026
4
MINUTES
AI and autonomy building blocks begin with cloud-native
On a recent episode of Zero-Touch Live, Boost Mobile's Dawood Shahdad, SVP of Hybrid MNO Network, and Sruthi Nair, Director of Voice Core Engineering, spoke with Rakuten Symphony's Anshul Bhatt about the journey of building one of the few operator networks running fully cloud-native on public cloud. In this week’s Zero-Touch newsletter, we dive deeper into the role cloud-native plays in powering autonomy and the essential building blocks that must accompany it.
June 25, 2026
6
MINUTES
Winning automation execution strategies for LATAM
In a region increasingly defined by its diversity and cost pressures, autonomous networks are on every operator’s roadmap. In this week’s newsletter, Mauricio Gonzalez Nappa, telco lead at systems integrator Isbel, and Sachin Mahajan of Rakuten Symphony examine what shapes the diverse LATAM market, why the path to Level 3 and Level 4 automation requires laser-focused execution and where automation is already paying off.
June 11, 2026
4
MINUTES