Reinforcement Learning for Wireless Network Resource Allocation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing network management techniques rely on static rules that do not account for context information, leading to inefficient resource allocation and suboptimal user experience in wireless communication networks, especially with the increasing diversity of devices and applications.

Innovation Solution

The implementation of reinforcement learning techniques to dynamically manage network resources by observing network states and adjusting parameters based on observed conditions, such as throughput, latency, and user mobility, to optimize resource allocation and improve quality of service.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If static rules are used for network management, then implementation is simple, but resource allocation efficiency deteriorates

Engineering Contradiction:
Improveimplementation simplicityVSAvoidresource allocation efficiency
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The patent transforms static network management rules into dynamic reinforcement learning-based management. The system continuously learns optimal resource allocation strategies by observing network states and receiving feedback on allocation outcomes, enabling adaptive decision-making that improves resource efficiency while maintaining implementation feasibility through modular RL integration.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent implements feedback mechanisms where the reinforcement learning system observes network performance metrics, evaluates the effectiveness of resource allocation decisions, and uses this feedback to continuously refine management policies. This closed-loop approach enables the system to learn from past decisions and improve resource allocation efficiency over time.

Inventive Principle:
Principle #23Feedback

2Device complexity

If context information is not considered, then management complexity is low, but user experience deteriorates

Engineering Contradiction:
Improvemanagement complexityVSAvoiduser experience quality
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent segments context information into distinct observable states (e.g., network load, user mobility patterns, application requirements) that the reinforcement learning system can process independently. This segmentation allows the system to consider multiple context factors without proportionally increasing management complexity, as each segment can be evaluated and weighted based on its relevance to user experience.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent dynamically adjusts management parameters based on observed context information. The reinforcement learning system modifies resource allocation parameters, handoff thresholds, and quality of service settings in response to changing network conditions and user behaviors, thereby improving user experience while keeping complexity manageable through parameter-based adaptation rather than structural complexity.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If dynamic reinforcement learning is implemented, then resource allocation efficiency improves, but computational complexity increases

Engineering Contradiction:
Improveresource allocation efficiencyVSAvoidcomputational complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies reinforcement learning selectively to critical network management decisions rather than all decisions uniformly. The system focuses computational resources on high-impact areas such as resource allocation and handoff management, while using simpler rules for less critical functions. This partial application of complex RL techniques improves resource allocation efficiency without requiring full-system computational overhead.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The reinforcement learning system serves itself by automatically learning optimal policies from network operations data without requiring extensive manual configuration or external intervention. The system self-adjusts its parameters and strategies based on observed outcomes, reducing the need for complex external control mechanisms and minimizing the effective computational complexity that must be managed by external systems.

Inventive Principle:
Principle #25Self-service

4Stability of the object's composition

If static management rules are used, then system stability is high, but adaptability to changing conditions deteriorates

Engineering Contradiction:
Improvesystem stabilityVSAvoidadaptability to changing conditions
Core Design Contradiction:
Stability of the object's compositionVSAdaptability or versatility

Solution Approach 1:

The patent implements continuous learning and adaptation through reinforcement learning, where the system continuously observes network states, evaluates outcomes, and refines management policies. This continuous action loop maintains system stability by building upon learned patterns while simultaneously improving adaptability to changing conditions through ongoing optimization, rather than relying on fixed static rules.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The reinforcement learning system performs preliminary learning during periods of stable network conditions, building up knowledge and optimized policies in advance. When changing conditions occur, the system can quickly adapt by applying previously learned patterns and making incremental adjustments, rather than reacting from scratch. This preliminary action approach maintains stability while preparing the system for future adaptability.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10225772B2Mobility management for wireless communication networks
Publication Date: 2019.03.05 AT&T INTELLECTUAL PROPERTY I L P
  • US10225772B2 patent drawing
  • US10225772B2 patent drawing
  • US10225772B2 patent drawing

AI summary

Mobility management for wireless communication networks is provided. A method can include determining, by a device comprising a processor, expected utilities respectively for a group of network actions on behalf of a user equipment in a communication network based on an observed state of network devices of the communication network, selecting, by the device, a network action from the group of network actions having at least a threshold expected utility, resulting in a selected network action, determining, by the device, an adjustment parameter for the selected network action based on a conformity of the selected network action to a network policy applicable to the communication network, and based on the adjustment parameter, adjusting, by the device, an expected utility of the expected utilities corresponding to the selected network action to account for the observed state of the network devices of the communication network.