Distributed Reinforcement Learning for Radio Access Network Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current radio access network optimization methods, such as rule-based and reinforcement learning approaches, face challenges in scalability and effectiveness due to reliance on centralized architectures or independent agents, leading to suboptimal configurations and interference issues in dense networks.

Innovation Solution

A distributed and coordinated reinforcement learning algorithm is employed to dynamically optimize radio access network configurations, utilizing a coordination graph to find globally optimal joint antenna configurations while learning locally, allowing for decentralized adaptation and efficient data management.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If centralized architectures are used for radio access network optimization, then coordination between nodes is improved, but scalability and system complexity deteriorate

Engineering Contradiction:
Improvecoordination between nodesVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent divides the centralized optimization function into distributed components at each network node. Each node runs independent reinforcement learning agents that make local decisions, eliminating the need for a centralized controller while maintaining coordination through peer-to-peer communication and shared learning mechanisms.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from a vertical centralized control architecture to a horizontal distributed architecture where nodes operate at the same level. This dimensional shift allows nodes to maintain autonomy while achieving coordination through lateral communication and collective learning, reducing hierarchical complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Device complexity

If independent agents are used for network optimization, then device complexity is reduced, but interference issues and suboptimal configurations increase

Engineering Contradiction:
Improvedevice complexityVSAvoidnetwork performance
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent combines independent local agents with a collaborative learning mechanism. While each node maintains its own reinforcement learning agent for local decision-making, nodes share experiences and coordinate through communication interfaces, merging individual capabilities into a collectively optimized system that avoids interference issues.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements feedback loops where nodes exchange performance metrics and configuration information with neighboring nodes. This feedback mechanism allows independent agents to adjust their configurations based on network-wide performance, preventing suboptimal decisions and interference while maintaining local autonomy.

Inventive Principle:
Principle #23Feedback

3Productivity

If dense network configurations are deployed, then network capacity is increased, but interference between nodes worsens

Engineering Contradiction:
Improvenetwork capacityVSAvoidinterference between nodes
Core Design Contradiction:
ProductivityVSObject-generated harmful factors

Solution Approach 1:

The patent employs dynamic configuration adjustment where network parameters such as antenna tilt, azimuth, and power settings are continuously optimized by reinforcement learning agents based on real-time performance feedback. This dynamic adaptation allows dense networks to automatically adjust configurations to minimize interference while maintaining high capacity.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes key operational parameters (antenna tilt angles, transmission power levels, frequency allocations) through learned policies that optimize network performance. By dynamically adjusting these parameters based on network conditions and interference levels, the system maintains high capacity in dense deployments while mitigating harmful interference effects.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240022950A1Decentralized coordinated reinforcement learning for optimizing radio access networks
Publication Date: 2024.01.18 TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
  • US20240022950A1 patent drawing
  • US20240022950A1 patent drawing
  • US20240022950A1 patent drawing

AI summary

A method of a node in a radio access network that optimizes radio access network operations. The method of the node includes determining a topology of the radio access network, exchanging optimization information and network metric information with neighbor nodes in the topology, determining an updated local configuration for the node based on a negotiated optimization with the neighbor nodes, and updating an optimization function based on collected updated network metric information of the node executing the updated local configuration.