Reinforcement Learning Routing for IAB Network Energy Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current techniques for managing routing in Integrated Access and Backhaul (IAB) networks do not adequately consider dynamicity and energy-related factors, leading to suboptimal routing decisions that do not account for the heterogeneity of mobile traffic and energy consumption variations across different nodes.
Innovation Solution
A reinforcement learning system is trained to optimize routing by acquiring observations of energy and traffic performance, predicting future network states, and reconfiguring routing tables to minimize energy consumption and maximize throughput, using a machine learning component triggered by new node onboarding or periodic link status reports.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Use of energy by moving object
If traditional routing techniques are used in IAB networks, then routing decisions are simple to implement, but energy consumption is not optimized and routing performance is suboptimal
Solution Approach 1:
The reinforcement learning system enables the routing infrastructure to automatically optimize its own operations by learning from observed network states and energy consumption patterns, making routing decisions that minimize energy usage without requiring manual configuration or external optimization systems
Solution Approach 2:
The system implements continuous feedback loops where the RL agent observes network state (traffic load, energy consumption, link quality), executes routing actions, receives rewards based on performance metrics, and updates its policy accordingly, creating a closed-loop control system that adapts to changing network conditions
2Productivity
If reinforcement learning is implemented for routing optimization, then energy consumption is reduced and routing performance is improved, but system complexity increases
Solution Approach 1:
The routing system transitions from static, pre-configured routing tables to dynamic, adaptive routing decisions where the RL agent continuously learns and adjusts routing policies based on real-time network conditions, traffic patterns, and energy consumption observations
Solution Approach 2:
The system changes the parameter space by incorporating energy consumption metrics and traffic load characteristics as observable state variables, allowing the RL agent to learn policies that optimize routing decisions based on these additional dimensions beyond traditional network state
3Measurement precision
If routing decisions consider dynamic traffic and energy factors, then routing optimisation is improved, but measurement and detection difficulty increases
Solution Approach 1:
The observation mechanism is designed to collect multiple types of data (traffic load, energy consumption, link quality, network state) through a unified interface and processing pipeline, allowing the same system components to serve multiple measurement functions simultaneously
Data Source
AI summary
There is provided a method for training a reinforcement learning system for optimising routing for a network including a plurality of Integrated Access and Backhaul (IAB) nodes connected to an IAB donor. The method includes acquiring observations characterising a current state of the plurality of IAB nodes, determining an action to be performed based on latest acquired observations, executing the action by initiating update of the routing information based on the determined action, acquiring observations characterising an updated state of the plurality of IAB nodes, determining a reward for the determined action, based on the updated state of the plurality of IAB nodes, storing an experience set, and training the reinforcement learning system to maximise reward with respect to an optimisation objective, using the one or more stored experience sets in the buffer.


