Reward Autoencoder for Reinforcement Learning Network Policies

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Defining rewards for reinforcement learning models in network performance management is challenging due to high-dimensional data and incorrect reward definitions, leading to wastage of computing and networking resources.

Innovation Solution

A reward autoencoder platform generates rewards by embedding network performance data into a lower-dimensional space, calculating reconstruction errors, and determining convex hulls to define optimal network policies and metrics, thereby improving network performance and conserving resources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If reinforcement learning is used for network performance management, then automation and adaptability are improved, but difficulty in defining rewards and modeling actions increases

Engineering Contradiction:
Improveautomation of network performance managementVSAvoidcomplexity of defining rewards and modeling actions
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The system enables the reinforcement learning model to automatically define its own reward functions and action models based on observed network states and outcomes, eliminating the need for manual specification of rewards and actions by operators

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

An intermediary layer is introduced that automatically translates network performance data into meaningful reward signals and action representations, bridging the gap between raw data and reinforcement learning requirements

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If high-dimensional network performance data is used, then measurement precision is improved, but loss of outlier detection and computational efficiency decreases

Engineering Contradiction:
Improveprecision of network performance measurementVSAvoidloss of outlier detection capability
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system extracts and separates outlier data points from the main high-dimensional dataset, applying specialized processing to identify and flag anomalies while maintaining the integrity of the overall performance measurement

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Outlier detection and handling is performed as a preliminary step before main reinforcement learning processing, ensuring that anomalous data points are identified and managed before they can interfere with normal operation

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If manual reward definition is used, then control and precision are improved, but productivity and resource efficiency decrease

Engineering Contradiction:
Improveprecision of reward definitionVSAvoidproductivity of network performance management
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The reinforcement learning system automatically generates and refines its own reward definitions based on observed network performance patterns, eliminating manual intervention while maintaining adaptive precision

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Reward definitions are made dynamic and adaptive, automatically adjusting to changing network conditions and performance requirements without manual reconfiguration

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11063841B2Systems and methods for managing network performance based on defining rewards for a reinforcement learning model
Publication Date: 2021.07.13 VERIZON PATENT & LICENSING INC
  • US11063841B2 patent drawing
  • US11063841B2 patent drawing
  • US11063841B2 patent drawing

AI summary

A device may receive network policies of a network, and network performance data identifying KPIs of the network, and may generate an embedded space of reconstructed data that is embedded in an original space that includes the KPIs. The device may calculate reconstruction errors based on differences between the reconstructed data and the network performance data, and may calculate a convex hull of the original space. The device may calculate a convex hull of the embedded space, and may determine reward metrics based on the reconstruction errors, the convex hull of the original space, and the convex hull of the embedded space. The device may define performance baselines associated with portions, and may generate a new reward for a portion based on a particular reconstruction error, a particular convex hull of the embedded space, and a particular performance baseline. The device may perform actions based on the new reward.