Deep Reinforcement Learning for Real-Time RAN Parameter Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current software-defined networking and self-optimizing network technologies face challenges in dynamically optimizing radio access network parameters due to complex and dynamic user traffic patterns, limited observability of network status, and the need for rapid algorithm tuning to improve user experience and reduce operational costs.

Innovation Solution

A reinforcement learning framework with multiple sub-agents, each comprising a neural network, is used to determine optimal settings for radio access network parameters in real-time, leveraging live data streaming and vendor APIs for closed-loop control, allowing for simultaneous optimization of multiple parameters and resolving conflicts between different policies.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional SDN and SON technologies are used to optimize RAN parameters, then network control capability is maintained, but optimization speed and adaptability to dynamic traffic patterns deteriorate

Engineering Contradiction:
Improveadaptability to dynamic traffic patternsVSAvoidoptimization speed
Core Design Contradiction:
Adaptability or versatilityVSSpeed

Solution Approach 1:

The system employs reinforcement learning agents that autonomously learn and optimize RAN parameters without human intervention. The agents continuously interact with the network environment, receiving feedback through reward functions and self-adjusting their policies to maximize network performance metrics, thereby enabling self-directed optimization that adapts rapidly to changing traffic conditions

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent implements dynamic optimization by training reinforcement learning models that can adapt their behavior in real-time based on current network states. The agents use experience replay buffers and continuous learning mechanisms to dynamically adjust parameter settings, allowing the system to respond flexibly to evolving traffic patterns and network conditions rather than relying on static configurations

Inventive Principle:
Principle #15Dynamics

2Reliability

If complex algorithms are developed to handle dynamic network conditions, then optimization performance improves, but development time and costs increase

Engineering Contradiction:
Improveoptimization performanceVSAvoiddevelopment time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system implements a closed-loop feedback mechanism where reinforcement learning agents receive continuous performance feedback through reward functions based on network KPIs. This feedback drives iterative improvement of the agents' policies through trial-and-error learning, allowing the system to automatically refine optimization performance without requiring extensive manual algorithm development and tuning

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent uses simulation environments to create virtual copies of the radio access network for training reinforcement learning agents. By developing and testing algorithms in simulated environments first, the system can validate optimization performance before deploying to production networks, significantly reducing development time and risks associated with direct implementation on live networks

Inventive Principle:
Principle #26Copying

3Measurement precision

If manual tuning of RAN parameters is performed, then control precision is maintained, but response time to changing conditions increases

Engineering Contradiction:
Improveparameter control precisionVSAvoidresponse time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical tuning processes with automated reinforcement learning systems. The agents use intelligent algorithms to precisely select and adjust RAN parameters based on real-time network state assessments, eliminating the delays inherent in manual intervention while maintaining or improving control precision through systematic exploration of the parameter space

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Measurement precision

If more observability of network status is achieved, then decision accuracy improves, but system complexity increases

Engineering Contradiction:
Improvedecision accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system extracts only the most relevant network state features and performance indicators needed for reinforcement learning decision-making. By selectively observing and processing key network parameters rather than attempting to monitor all possible network variables, the system achieves sufficient decision accuracy while avoiding the complexity overhead of comprehensive network observability

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11494649B2Radio access network control with deep reinforcement learning
Publication Date: 2022.11.08 AT&T INTELLECTUAL PROPERTY I L P
  • US11494649B2 patent drawing
  • US11494649B2 patent drawing
  • US11494649B2 patent drawing

AI summary

A processing system including at least one processor may obtain operational data from a radio access network (RAN), format the operational data into state information and reward information for a reinforcement learning agent (RLA), processing the state information and the reward information via the RLA, where the RLA comprises a plurality of sub-agents, each comprising a respective neural network, each of the neural networks encoding a respective policy for selecting at least one setting of at least one parameter of the RAN to increase a respective predicted reward in accordance with the state information, and where each neural network is updated in accordance with the reward information. The processing system may further determine settings for parameters of the RAN via the RLA, where the RLA determines the settings in accordance with selections for the settings via the plurality of sub-agents, and apply the plurality of settings to the RAN.