QoE Control Policy Adaptation via Reinforcement Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing PCC mechanisms in wireless communication networks are unable to efficiently control user data traffic to achieve a desired Quality of Experience (QoE) due to static QoS settings that do not directly map to actual user QoE, and lack of access to actual user QoE, making it difficult to adapt to dynamic changes in network and user data traffic characteristics.

Innovation Solution

A method where a node in a wireless communication network receives data indicating a desired quality of experience level for user data traffic, determines a control policy rule based on this data, obtains data on the estimated quality of experience level, and adapts the control policy using machine learning processes, specifically reinforcement learning algorithms, to optimize QoS settings and achieve the desired QoE.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If static QoS settings are used to control user data traffic, then the network control is simple and stable, but the Quality of Experience (QoE) cannot be optimized for dynamic network conditions and user characteristics

Engineering Contradiction:
ImproveQoE consistencyVSAvoidAdaptability to dynamic changes
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent applies dynamics by transitioning from static QoS settings to dynamic QoS parameter adjustment. The system continuously adapts QoS parameters based on real-time network conditions and user characteristics, allowing the control mechanism to respond to changing environments while maintaining reliable QoE delivery.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent implements feedback mechanisms where the system monitors actual QoE metrics and network conditions, then uses this information to adjust QoS parameters. The feedback loop enables continuous optimization of QoS settings based on observed performance, resolving the contradiction between stability and adaptability.

Inventive Principle:
Principle #23Feedback

2Reliability

If machine learning algorithms are introduced to adapt control policies, then the QoE optimization capability is improved, but the system complexity increases

Engineering Contradiction:
ImproveQoE achievementVSAvoidControl system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies self-service by enabling the system to automatically adapt QoS parameters using machine learning algorithms without requiring manual intervention. The control system autonomously learns from data and adjusts parameters, improving QoE while managing complexity through automated decision-making processes.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If operator access to actual user QoE is enabled, then the QoS parameter optimization is improved, but the network privacy and security requirements become more stringent

Engineering Contradiction:
ImproveQoE measurement accuracyVSAvoidPrivacy and security risks
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent introduces an intermediary layer that enables the operator to access QoE measurement data without directly exposing sensitive user information. This mediator mechanism allows accurate QoE measurement for optimization purposes while maintaining privacy and security through controlled data access and processing.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12256263B2Machine learning based adaptation of QoE control policy
Publication Date: 2025.03.18 TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
  • US12256263B2 patent drawing
  • US12256263B2 patent drawing
  • US12256263B2 patent drawing

AI summary

A node of a wireless communication network receives first data indicating a desired quality of experience level for user data traffic of a user of the wireless communication network. Based on a control policy and the desired quality of experience level, the node determines a rule for controlling the user data traffic. Further, the node obtains second data indicating an estimated quality of experience level for the user data traffic subject to control according to the rule. Based on the first data and the second data, the node adapts the control policy, e.g., using a reinforcement learning, RL, mechanism.