DAG-Based Workload Remediation for SLA Failure Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional workload provisioning systems face significant challenges in remediating failures to satisfy Service Level Agreements (SLAs) during workload performance, often requiring time-consuming server device re-initialization or resource reconfiguration, which exacerbates the failure to meet quick recovery requirements.
Innovation Solution
An Information Handling System (IHS) equipped with a resource management engine that generates a Directed Acyclic Graph (DAG) to identify and configure resource devices to satisfy SLAs, allowing for rapid remediation by reallocating resources to ensure SLA compliance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional workload provisioning systems perform re-initialization or full rebuild of server devices to remediate SLA failures, then the system can restore SLA compliance, but the recovery time increases significantly
Solution Approach 1:
The patent segments the server device into multiple independent resource components (storage resources, networking resources, processing resources) that can be individually monitored and remediated. When an SLA failure occurs, the system identifies which specific resource component caused the failure and remediates only that component rather than performing a full re-initialization of the entire server device, thereby reducing recovery time while maintaining SLA compliance.
Solution Approach 2:
The patent implements continuous monitoring of resource devices against SLA requirements in advance, allowing the system to detect SLA failures immediately when they occur. This preliminary detection and immediate remediation action prevents prolonged non-compliance and reduces the overall recovery time by acting at the moment of failure rather than after degradation has occurred.
2Reliability
If the system performs full rebuild of resources to satisfy SLA requirements, then SLA compliance is achieved, but the complexity and time consumption increase
Solution Approach 1:
The patent divides the server device into distinct resource components (storage, networking, processing) with independent monitoring and remediation capabilities. When an SLA failure is detected, the system identifies the specific failing component and performs targeted remediation only on that component, avoiding the complexity of reconfiguring the entire server device while still achieving SLA compliance.
Solution Approach 2:
The patent extracts the failing resource component from the overall server device configuration and performs remediation independently on that extracted component. This allows the system to isolate the problem area and apply fixes only where needed, reducing the overall complexity of the remediation process compared to a full rebuild approach.
Data Source
AI summary
A workload resource device Service Level Agreement (SLA) failure remediation system includes a resource management system coupled to resource devices. The resource management system receives a workload intent for performing a workload associated with SLA(s), and generates a Directed Acyclic Graph (DAG) that identifies a first resource device and second resource device(s) for performing the workload. Based on the DAG, the resource management system configures the first resource device and the second resource device(s) to perform the workload, and stores the DAG in at least one database. If the resource management system determines that the first resource device is not satisfying the SLA(s) during the performance of the workload, it uses s portion of the DAG associated with the first resource device to configure at least one of the resource devices to operate with the second resource device(s) to subsequently perform the workload such that the SLA(s) are satisfied.


