Wireless Link Change Decisions Using Reinforcement Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current handover decision-making in wireless communication systems relies on instantaneous parameters, which may not accurately predict long-term suitability of links, leading to suboptimal device performance and system efficiency.

Innovation Solution

Implementing reinforcement learning to make link change decisions based on rewards and outcomes of past decisions, allowing for device-specific optimization and consideration of long-term performance metrics such as quality of service and time spent on target links.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If handover decision-making is based on instantaneous parameters, then handover decision speed is improved, but long-term link suitability prediction accuracy deteriorates

Engineering Contradiction:
Improvehandover decision speedVSAvoidlong-term link suitability prediction accuracy
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The patent implements feedback mechanisms where the outcomes of previous handover decisions are tracked and fed back into the decision-making process. The network equipment monitors whether devices remain connected to target links, experience dropped calls, or perform well, and uses this feedback information to continuously improve future handover decisions. This closes the loop between instantaneous decisions and long-term performance.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary actions by tracking and recording outcomes of handover decisions over time before making new decisions. Historical data about which target links proved successful or unsuccessful in the long-term is accumulated in advance, allowing the system to learn from past experiences and make more informed predictions about future handover suitability.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If device-specific reinforcement learning is implemented, then device performance optimization is improved, but system complexity increases

Engineering Contradiction:
Improvedevice performance optimizationVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the handover decision-making process into device-specific instances, where each wireless device maintains its own reinforcement learning model and tracked outcomes. This segmentation allows personalized optimization for each device based on its unique characteristics, movement patterns, and service requirements, while the overall system complexity is managed through modular architecture where each device's learning process is independent.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11963047B2Link change decision-making using reinforcement learning based on tracked rewards and outcomes in a wireless communication system
Publication Date: 2024.04.16 TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
  • US11963047B2 patent drawing
  • US11963047B2 patent drawing
  • US11963047B2 patent drawing

AI summary

Decision-making equipment (22) is configured for link change decision-making using reinforcement learning. The decision-making equipment (22) is configured to track rewards (30-1, . . . 30-M) earned for, and outcomes (28-1, . . . 28-M) of, respective link change decisions (26-1, . . . 26-M). In some embodiments, possible outcomes of a link change decision to change a serving link of a wireless device to a target link include at least: a change of the serving link of the wireless device from the target link to another link; and a network-initiated disconnect of the wireless device from the target link. Regardless, the decision-making equipment (22) is also configured to make a link change decision (28-(M+1)) based on the tracked rewards (30-1, . . . 30-M) and outcomes (28-1, . . . 28-M).