RIS-Assisted RAN Control Using Hierarchical Reinforcement Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing wireless communication systems face challenges in achieving energy efficiency in cellular networks, particularly with the deployment of reconfigurable intelligent surfaces (RIS), where conventional reinforcement learning methods suffer from long convergence issues and inefficient exploration.

Innovation Solution

Implementing a hierarchical reinforcement learning (HRL) algorithm with a meta-controller for small base station sleep control and sub-controllers for transmission power control, combined with RIS to optimize signal propagation and reduce energy consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Use of energy by moving object

If conventional reinforcement learning methods are used for sleep control in RIS-assisted RAN, then the system can achieve energy efficiency optimization, but the convergence time becomes excessively long and exploration efficiency is poor

Engineering Contradiction:
Improveenergy efficiencyVSAvoidconvergence time
Core Design Contradiction:
Use of energy by moving objectVSLoss of time

Solution Approach 1:

The patent divides the reinforcement learning control into two hierarchical levels: a meta-controller that handles high-level sleep control decisions for base stations, and sub-controllers that manage low-level transmission power adjustment. This segmentation allows each controller to focus on specific aspects of optimization, reducing the overall convergence time while maintaining energy efficiency improvements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a hierarchical dimension to the reinforcement learning structure, transitioning from a single-level flat architecture to a two-level hierarchical architecture. The meta-controller operates at the strategic level while sub-controllers execute at the operational level, adding a temporal and functional dimension that accelerates convergence without sacrificing optimization performance.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of energy

If conventional reinforcement learning methods are used for sleep control in RIS-assisted RAN, then the system can optimize energy consumption, but the exploration of the environment becomes inefficient

Engineering Contradiction:
Improveenergy consumptionVSAvoidexploration efficiency
Core Design Contradiction:
Loss of energyVSProductivity

Solution Approach 1:

The patent segments the exploration function between meta-controller and sub-controllers. The meta-controller performs strategic exploration of sleep control policies, while sub-controllers handle tactical exploration of power levels. This division of exploration responsibilities improves overall exploration efficiency by preventing redundant searches and focusing computational resources on promising regions of the solution space.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The hierarchical structure acts as an intermediary between high-level energy optimization goals and low-level control actions. The meta-controller translates energy efficiency objectives into sleep control decisions, which then guide sub-controller power adjustments. This intermediary layer mediates the exploration process, making it more directed and efficient rather than random and exhaustive.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If RIS is deployed to improve signal propagation and increase channel capacity, then the system achieves better signal quality and connectivity, but the system complexity increases

Engineering Contradiction:
Improvesignal qualityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements self-service control where the hierarchical reinforcement learning system automatically adapts the RIS configuration and base station sleep states based on real-time channel conditions. The meta-controller and sub-controllers continuously monitor signal quality metrics and autonomously adjust system parameters, eliminating the need for manual configuration and reducing operational complexity despite the advanced RIS capabilities.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent manages RIS complexity by dynamically changing operational parameters rather than fixed configurations. The hierarchical RL system continuously adjusts phase shifters, reflection coefficients, and sleep states based on channel conditions. This parameter-based adaptability allows the system to handle RIS complexity through software control rather than hardware complexity, maintaining signal quality while managing system burden.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250350538A1System and method for reconfigurable intelligent surface (RIS)-assisted energy-efficient (EE) radio access network (RAN) using hierarchical reinforcement learning
Publication Date: 2025.11.13 TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
  • US20250350538A1 patent drawing
  • US20250350538A1 patent drawing
  • US20250350538A1 patent drawing

AI summary

A method, system and apparatus for reconfigurable intelligent surface-assisted energy-efficient radio access networks using hierarchical reinforcement learning are disclosed. A method in a network node operating as a meta-controller and configured to communicate with a wireless device and a plurality of sub-controllers is provided. The method includes determining a state of the meta-controller including traffic load ratios of the plurality of sub-controllers. The method also includes receiving an extrinsic reward that is based on an energy efficiency of a cell including the plurality of sub-controllers. The method further includes selecting a goal according to a policy, the selected goal being an on/off state of each of the plurality of sub-controllers, the policy being selected to increase the extrinsic reward. The method includes configuring the plurality of controllers with the selected goal and an indication of the policy for selecting the goal.