Data Center Self-Healing Engine for Autonomous Fault Mitigation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Modern data centers face complex fault diagnosis and mitigation challenges due to their intricate infrastructure, requiring highly trained personnel to trace issues across multiple layers of network abstraction and encapsulation, necessitating expert human intervention.
Innovation Solution
An AI-based self-healing engine that employs a hybrid reasoning model combining human-defined skills and large language models (LLMs) for automated troubleshooting and self-healing, utilizing a decentralized system to detect, diagnose, and mitigate faults in data centers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If automated AI-based troubleshooting is implemented, then expert intervention is reduced, but system complexity increases due to the need for hybrid reasoning models combining LLMs with human-defined skills
Solution Approach 1:
The patent introduces a self-healing engine as an intermediary component that sits between monitoring systems and data center operations. This engine absorbs the complexity of hybrid reasoning models (combining LLMs with human-defined skills) while presenting a simplified interface for automated troubleshooting, thereby reducing the need for expert intervention without exposing users to underlying system complexity
Solution Approach 2:
The self-healing engine enables the system to diagnose and resolve faults autonomously using a hybrid reasoning model that combines large language models with human-defined troubleshooting skills. The system serves itself by automatically executing troubleshooting steps, modifying configurations, and recovering from faults without requiring continuous expert intervention
2Speed
If decentralized fault detection and diagnosis is implemented, then response time is improved, but coordination complexity increases across multiple autonomous agents
Solution Approach 1:
The patent segments the data center monitoring and management system into multiple decentralized autonomous agents, each responsible for specific devices or functions. Each agent independently detects and diagnoses faults in its domain, enabling parallel processing and faster response times while the self-healing engine coordinates actions across agents through standardized interfaces
Data Source
AI summary
A computer system for use with a data center includes a monitoring system configured to receive telemetry data from the data center and to generate alert data in response to a data center fault indicated by the telemetry data; a data center orchestrator coupled to the monitoring system that is configured to manage operation of the data center; and a self-healing engine that operates via an application programming interface (API) configured to receive the alert data, topology data corresponding to a topology of the data center, and user intent data and to select and execute one or more skills in conjunction with the data center orchestrator to correct the data center fault.


