Reinforcement Learning Routing for IAB Network Energy Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current techniques for managing routing in Integrated Access and Backhaul (IAB) networks do not adequately consider dynamicity and energy-related factors, leading to suboptimal routing decisions that do not account for the heterogeneity of mobile traffic and energy consumption variations across different nodes.

Innovation Solution

A reinforcement learning system is trained to optimize routing by acquiring observations of energy and traffic performance, predicting future network states, and reconfiguring routing tables to minimize energy consumption and maximize throughput, using a machine learning component triggered by new node onboarding or periodic link status reports.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Use of energy by moving object

If traditional routing techniques are used in IAB networks, then routing decisions are simple to implement, but energy consumption is not optimized and routing performance is suboptimal

Engineering Contradiction:
Improveenergy consumptionVSAvoidrouting system complexity
Core Design Contradiction:
Use of energy by moving objectVSDevice complexity

Solution Approach 1:

The reinforcement learning system enables the routing infrastructure to automatically optimize its own operations by learning from observed network states and energy consumption patterns, making routing decisions that minimize energy usage without requiring manual configuration or external optimization systems

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements continuous feedback loops where the RL agent observes network state (traffic load, energy consumption, link quality), executes routing actions, receives rewards based on performance metrics, and updates its policy accordingly, creating a closed-loop control system that adapts to changing network conditions

Inventive Principle:
Principle #23Feedback

2Productivity

If reinforcement learning is implemented for routing optimization, then energy consumption is reduced and routing performance is improved, but system complexity increases

Engineering Contradiction:
Improverouting efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The routing system transitions from static, pre-configured routing tables to dynamic, adaptive routing decisions where the RL agent continuously learns and adjusts routing policies based on real-time network conditions, traffic patterns, and energy consumption observations

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the parameter space by incorporating energy consumption metrics and traffic load characteristics as observable state variables, allowing the RL agent to learn policies that optimize routing decisions based on these additional dimensions beyond traditional network state

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If routing decisions consider dynamic traffic and energy factors, then routing optimisation is improved, but measurement and detection difficulty increases

Engineering Contradiction:
Improverouting decision accuracyVSAvoidstate observation complexity
Core Design Contradiction:
Measurement precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The observation mechanism is designed to collect multiple types of data (traffic load, energy consumption, link quality, network state) through a unified interface and processing pipeline, allowing the same system components to serve multiple measurement functions simultaneously

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240406835A1Energy-aware routing based on reinforcement learning
Publication Date: 2024.12.05 TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
  • US20240406835A1 patent drawing
  • US20240406835A1 patent drawing
  • US20240406835A1 patent drawing

AI summary

There is provided a method for training a reinforcement learning system for optimising routing for a network including a plurality of Integrated Access and Backhaul (IAB) nodes connected to an IAB donor. The method includes acquiring observations characterising a current state of the plurality of IAB nodes, determining an action to be performed based on latest acquired observations, executing the action by initiating update of the routing information based on the determined action, acquiring observations characterising an updated state of the plurality of IAB nodes, determining a reward for the determined action, based on the updated state of the plurality of IAB nodes, storing an experience set, and training the reinforcement learning system to maximise reward with respect to an optimisation objective, using the one or more stored experience sets in the buffer.