Adversarial RL Application Manager Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The management and administration of complex distributed computing systems are hindered by scaling issues, including communication overheads and component failures, making traditional automated management systems impractical and inefficient, prompting the need for alternative methodologies such as machine-learning-based approaches.

Innovation Solution

An automated reinforcement-learning-based application manager is trained using adversarial training, which selects disadvantageous actions less frequently to explore a larger subset of the system-state space, resulting in a more robust and complete optimal control policy compared to traditional training methods.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional automated management systems are used to manage distributed computing systems, then the systems can perform basic management functions, but the systems become impractical and inefficient due to scaling issues and communication overheads

Engineering Contradiction:
Improvemanagement efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical automated management systems with a machine-learning-based approach. The application manager uses reinforcement learning to learn optimal management policies through interaction with the computing environment, substituting rigid rule-based control with adaptive intelligent control that scales better with system complexity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameter of control from fixed rule-based decisions to dynamic learned policies. The application manager learns to adjust management parameters adaptively based on system state, transforming the management approach from static to dynamic and enabling better handling of scaling issues

Inventive Principle:
Principle #35Parameter changes

2Reliability

If the automated reinforcement-learning-based application manager is trained using traditional methods with simulators or recorded actions, then the training process is simpler, but the learned control policy is less robust and incomplete

Engineering Contradiction:
Improvecontrol policy robustnessVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by proactively selecting potentially disadvantageous actions during training to explore the system-state space before final policy convergence. This forward-looking exploration strategy ensures more comprehensive coverage of possible states and transitions, leading to more robust learned policies

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses feedback mechanisms where the application manager receives rewards or penalties based on the outcomes of selected actions. This feedback drives the reinforcement learning process, allowing the system to learn from both successful and unsuccessful actions, thereby improving policy robustness through iterative refinement

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If the application manager selects disadvantageous actions at higher frequency during training, then the exploration of system-state space increases, but the training efficiency decreases and convergence slows down

Engineering Contradiction:
Improvesystem-state space coverageVSAvoidtraining efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent applies partial action by selecting disadvantageous actions at a controlled, moderate frequency rather than exclusively. This balanced approach allows sufficient exploration of the system-state space while maintaining reasonable training efficiency, avoiding the extremes of either pure exploitation or excessive exploration

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10977579B2Adversarial automated reinforcement-learning-based application-manager training
Publication Date: 2021.04.13 VMWARE INC
  • US10977579B2 patent drawing
  • US10977579B2 patent drawing
  • US10977579B2 patent drawing

AI summary

The current document is directed to automated reinforcement-learning-based application managers that that are trained using adversarial training. During adversarial training, potentially disadvantageous next actions are selected for issuance by an automated reinforcement-learning-based application manager at a lower frequency than selection of next actions, according to a policy that is learned to provide optimal or near-optimal control over a computing environment that includes one or more applications controlled by the automated reinforcement-learning-based application manager. By selecting disadvantageous actions, the automated reinforcement-learning-based application manager is forced to explore a much larger subset of the system-state space during training, so that, upon completion of training, the automated reinforcement-learning-based application manager has learned a more robust and complete optimal or near-optimal control policy than had the automated reinforcement-learning-based application manager been trained by simulators or using management actions and computing-environment responses recorded during previous controlled operation of a computing-environment.