Reinforcement Learning Agent for Real-Time Communication Quality

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Bandwidth estimation, congestion control, and video quality optimization for real-time communication remain challenging due to changing network conditions and application requirements, leading to a degraded end-user experience from slow updates.

Innovation Solution

Implementing reinforcement learning in real-time communications, where an agent interfaces with sending and receiving computing devices to automatically adjust audio and video transmission parameters based on network conditions and user-perceived quality, using a reinforcement learning model with a control policy and state-action value function to maximize expected user-perceived quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional bandwidth estimation and congestion control methods are used, then the system can maintain basic communication functionality, but the end-user quality of experience degrades due to slow updates and inability to respond to changing network conditions

Engineering Contradiction:
Improveend-user quality of experienceVSAvoidupdate speed
Core Design Contradiction:
ReliabilityVSSpeed

Solution Approach 1:

The patent implements a feedback mechanism where the receiving computing device sends feedback information about actual user-perceived quality to the agent. This feedback loop enables continuous learning and adaptation, allowing the system to respond to changing network conditions in real-time rather than relying on slow traditional updates. The agent uses this feedback to adjust transmission parameters dynamically, resolving the contradiction between reliability and update speed.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The reinforcement learning agent operates autonomously to optimize transmission parameters without requiring manual intervention or slow centralized control updates. The agent independently learns from network conditions and user feedback, making self-driven adjustments to bandwidth estimation and congestion control parameters, thereby achieving fast response times while maintaining high user quality of experience.

Inventive Principle:
Principle #25Self-service

2Productivity

If reinforcement learning continuously adjusts transmission parameters, then user-perceived quality is optimized in real-time, but the system complexity increases due to the need for monitoring, learning, and adaptation mechanisms

Engineering Contradiction:
Improvequality optimization efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The reinforcement learning agent performs self-service by autonomously learning from network conditions and user feedback without requiring complex external management systems. The agent independently manages bandwidth estimation, congestion control, and parameter adjustments, simplifying the overall system architecture while maintaining high optimization efficiency. This self-driven approach resolves the contradiction between productivity and complexity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The feedback mechanism provides the agent with direct information about user-perceived quality, enabling efficient learning and adaptation. This feedback loop simplifies the system by replacing complex manual monitoring and control mechanisms with an automated agent that continuously optimizes parameters based on real-time feedback, thereby improving productivity without proportionally increasing complexity.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11699084B2Reinforcement learning in real-time communications
Publication Date: 2023.07.11 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11699084B2 patent drawing
  • US11699084B2 patent drawing
  • US11699084B2 patent drawing

AI summary

An agent interfaces with a sending computing device and a receiving computing device to automatically adjust one-way or two-way real-time audio and real-time video transmission parameters responsive to changing network conditions and/or application requirements. The agent incorporates a reinforcement learning model that adjusts transmission parameters to maximize an expected value of a sum of future rewards; the expected value of the sum of future rewards is based on a current state of the sending computing, a current action (e.g. a current set of transmission parameters) at the sending computing device and a reward provided by the receiving computing device. The reward is representative of a user-perceived quality of experience at the receiving computing device.