Multi-Agent RL Workload Placement for Cloud SLA Compliance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Cloud computing providers face challenges in efficiently allocating resources to meet service level agreements (SLAs) due to dynamic workload demands and varying resource requirements, leading to inefficiencies and potential SLA violations.

Innovation Solution

Implementing a multi-agent reinforcement learning-based system that dynamically places and migrates workloads across virtual machines to optimize resource usage while ensuring SLA compliance, using a placement engine that evaluates resource states and rewards to make placement recommendations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a provider dedicates a static amount of resources to each user, then service level agreements (SLAs) can be ensured, but resource utilization efficiency deteriorates due to idle resources and inability to adapt to dynamic workload demands

Engineering Contradiction:
ImproveSLA complianceVSAvoidresource utilization efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements dynamic resource allocation by transitioning from static dedicated resources to a system where workload placement is continuously adjusted based on real-time resource states and SLA requirements. The dynamic placement engine monitors resource utilization and migrates workloads between virtual machines to optimize both SLA compliance and resource efficiency simultaneously

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the parameter of resource allocation from fixed/static to variable/dynamic. By using reinforcement learning agents that adapt their placement strategies based on changing resource states and workload characteristics, the system optimizes the balance between reliability (SLA compliance) and productivity (resource utilization) under varying conditions

Inventive Principle:
Principle #35Parameter changes

2Reliability

If excessive resources are allocated to a single workload, then SLA compliance is improved, but the number of workloads that can be served in parallel deteriorates due to reduced spare resources

Engineering Contradiction:
ImproveSLA complianceVSAvoidnumber of workloads served
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent merges multiple workload placements into a unified optimization framework. Instead of allocating resources to each workload independently, the system considers the global state of all workloads and resources simultaneously, using multi-agent reinforcement learning to find placement configurations that maximize both SLA compliance and the number of workloads served

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The dynamic placement engine serves multiple functions simultaneously: it ensures SLA compliance for individual workloads while also optimizing overall resource utilization to maximize the number of workloads that can be served. The system adapts its behavior based on the current state of the entire system rather than optimizing for a single objective

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Device complexity

If static resource allocation is used, then implementation simplicity is improved, but adaptability to dynamic execution environments deteriorates when new workloads compete for resources

Engineering Contradiction:
Improveallocation system complexityVSAvoidadaptability to dynamic demands
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent implements feedback mechanisms where the placement engine continuously monitors resource states, workload performance, and SLA compliance. This feedback is used by reinforcement learning agents to dynamically adjust workload placements, enabling the system to adapt to changing conditions while maintaining a relatively simple architectural framework

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20240012685A1Swarm multi-agent reinforcement learning-based pipeline for workload placement
Publication Date: 2024.01.11 DELL PROD LP
  • US20240012685A1 patent drawing
  • US20240012685A1 patent drawing
  • US20240012685A1 patent drawing

AI summary

Multi-agent reinforcement learning-based workload placement is disclosed. A placement engine is configured to use the state of a system and actual rewards to generate expected rewards that correspond to actions. Agents can take actions for corresponding workloads based on the expected rewards output by the placement engine. This allows workloads to be placed in a manner that conservers power relative to load placement policies while helping avoid service level agreement violations.