Automated Reward Function Learning for Distributed Computing Management

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing complexity of distributed computing systems has led to management and administration challenges, including significant inefficiencies and computational overheads, making traditional automated management systems impractical, prompting the need for alternative methodologies such as machine-learning-based approaches.

Innovation Solution

An automated reinforcement-learning-based application manager that learns and improves a reward function to optimize policies in distributed computing systems, initially relying on human input and subsequent self-improvement using accumulated state/action trajectories.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional automated management systems are used for distributed computing systems, then initial policy implementation is achieved, but the systems suffer from significant inefficiencies and high computational overheads due to increasing complexity

Engineering Contradiction:
Improvemanagement efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system implements self-service through automated reinforcement learning where the application manager autonomously learns optimal management policies by interacting with the distributed computing environment. The system accumulates state-action trajectories and automatically improves its reward function without requiring manual reconfiguration, enabling it to adapt to increasing system complexity while maintaining management efficiency.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system employs feedback mechanisms by continuously monitoring the distributed computing environment, accumulating state-action trajectories, and using this feedback to iteratively improve the reward function. This closed-loop feedback enables the system to learn from past experiences and adapt to changing conditions, resolving the contradiction between managing complex systems and maintaining efficiency.

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If machine-learning-based approaches are adopted to manage complex distributed systems, then adaptability and optimization improve, but development costs and implementation complexity increase

Engineering Contradiction:
Improvepolicy adaptabilityVSAvoiddevelopment cost
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The system applies preliminary action by pre-defining a template-based reward function structure that incorporates common distributed computing management objectives. This preliminary framework reduces development costs by providing a head-start configuration, while the reinforcement learning component enables subsequent adaptability through automated learning from accumulated trajectories without requiring extensive custom development.

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If the reward function is manually configured, then initial management policies can be implemented, but the system cannot adapt to changes and improve over time

Engineering Contradiction:
Improveinitial setup easeVSAvoidpolicy improvement capability
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The system implements dynamics by transitioning from a static manually-configured reward function to a dynamic reward function that automatically evolves through reinforcement learning. The reward function adapts its parameters based on accumulated state-action trajectories, enabling the system to maintain ease of initial setup while gaining the capability to improve policies over time in response to changing conditions.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10963313B2Automated reinforcement-learning-based application manager that learns and improves a reward function
Publication Date: 2021.03.30 VMWARE INC
  • US10963313B2 patent drawing
  • US10963313B2 patent drawing
  • US10963313B2 patent drawing

AI summary

The current document is directed to automated reinforcement-learning-based application managers that learn and improve the reward function that steers reinforcement-learning-based systems towards optimal or near-optimal policies. Initially, when the automated reinforcement-learning-based application manager is first installed and launched, the automated reinforcement-learning-based application manager may rely on human-application-manager action inputs and resulting state/action trajectories to accumulate sufficient information to generate an initial reward function. During subsequent operation, when it is determined that the automated reinforcement-learning-based application manager is no longer following a policy consistent with the type of management desired by human application managers, the automated reinforcement-learning-based application manager may use accumulated trajectories to improve the reward function.