NWDAF Reinforcement Learning Support for 5G Operation Recommendations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The existing 5G telecommunication systems lack support for reinforcement learning (RL) to provide operation recommendation functions, hindering network automation and optimization.

Innovation Solution

Implementing a network data analytics function (NWDAF) that supports RL by training RL models based on subscription requests, managing reward information, and providing operation recommendations through an MTLF and AnLF, with features like RL capability registration and model training services.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If reinforcement learning techniques are adopted to provide operation recommendation functions, then system automation capability is improved, but the complexity of the network system increases due to the need for new functionality to manage reward information and train RL models

Engineering Contradiction:
Improvesystem automation capabilityVSAvoidnetwork system complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The NWDAF is enhanced to perform multiple functions: traditional network data analytics, RL model training, and reward information management. This allows a single existing network function to support both conventional analytics and new RL-based automation, reducing the need for separate dedicated components and thereby managing complexity while improving automation capability

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The NWDAF acts as an intermediary between the consumer NF and the RL model training process. It collects data from various sources, trains the RL model, manages reward information, and provides recommendations to the consumer NF. This intermediary role centralizes the complexity management within a single function rather than distributing it across multiple components

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If RL model training is implemented in the network, then operation recommendation accuracy is improved, but the training time and computational resources required increase

Engineering Contradiction:
Improveoperation recommendation accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs RL model training in advance during off-peak periods or when resources are available, so that trained models are ready for deployment when needed. The NWDAF can accumulate training data over time and perform batch training, reducing the time required when actual recommendations are needed while maintaining high accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The RL model training is performed periodically rather than continuously, allowing the system to balance between maintaining up-to-date accurate models and conserving computational resources. The training can be triggered by periodic schedules or by accumulation of sufficient training data, ensuring accuracy while managing time and resource constraints

Inventive Principle:
Principle #19Periodic action

3Reliability

If comprehensive data collection is performed for RL model training, then model training quality is improved, but the data management complexity and storage requirements increase

Engineering Contradiction:
Improvemodel training qualityVSAvoiddata management complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The NWDAF collects data with different quality requirements for different purposes: some data is collected for RL model training with high quality requirements, while other data is collected for basic analytics with lower quality requirements. This selective data collection approach ensures model training quality while reducing unnecessary data management complexity and storage requirements

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12615192B2Method of supporting reinforcement learning in mobile communication system and devices for performing the same
Publication Date: 2026.04.28 ELECTRONICS & TELECOMM RES INST
  • US12615192B2 patent drawing
  • US12615192B2 patent drawing
  • US12615192B2 patent drawing

AI summary

A method of supporting reinforcement learning (RL) in a mobile communication system and devices for performing the same are disclosed. The method of supporting RL in a mobile communication system includes receiving a subscription request to train an RL model from a consumer network function (NF), training the RL model by collecting data in response to the subscription request, and when training of the RL model is completed based on a training completion condition, transmitting a notification including RL model information to the consumer NF.