Data Center Self-Healing Engine for Autonomous Fault Mitigation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern data centers face complex fault diagnosis and mitigation challenges due to their intricate infrastructure, requiring highly trained personnel to trace issues across multiple layers of network abstraction and encapsulation, necessitating expert human intervention.

Innovation Solution

An AI-based self-healing engine that employs a hybrid reasoning model combining human-defined skills and large language models (LLMs) for automated troubleshooting and self-healing, utilizing a decentralized system to detect, diagnose, and mitigate faults in data centers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If automated AI-based troubleshooting is implemented, then expert intervention is reduced, but system complexity increases due to the need for hybrid reasoning models combining LLMs with human-defined skills

Engineering Contradiction:
Improveautomated troubleshootingVSAvoidsystem complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent introduces a self-healing engine as an intermediary component that sits between monitoring systems and data center operations. This engine absorbs the complexity of hybrid reasoning models (combining LLMs with human-defined skills) while presenting a simplified interface for automated troubleshooting, thereby reducing the need for expert intervention without exposing users to underlying system complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The self-healing engine enables the system to diagnose and resolve faults autonomously using a hybrid reasoning model that combines large language models with human-defined troubleshooting skills. The system serves itself by automatically executing troubleshooting steps, modifying configurations, and recovering from faults without requiring continuous expert intervention

Inventive Principle:
Principle #25Self-service

2Speed

If decentralized fault detection and diagnosis is implemented, then response time is improved, but coordination complexity increases across multiple autonomous agents

Engineering Contradiction:
Improveresponse timeVSAvoidcoordination complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent segments the data center monitoring and management system into multiple decentralized autonomous agents, each responsible for specific devices or functions. Each agent independently detects and diagnoses faults in its domain, enabling parallel processing and faster response times while the self-healing engine coordinates actions across agents through standardized interfaces

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250328412A1System and methods for data center fault mitigation
Publication Date: 2025.10.23 METALSOFT CLOUD INC
  • US20250328412A1 patent drawing
  • US20250328412A1 patent drawing
  • US20250328412A1 patent drawing

AI summary

A computer system for use with a data center includes a monitoring system configured to receive telemetry data from the data center and to generate alert data in response to a data center fault indicated by the telemetry data; a data center orchestrator coupled to the monitoring system that is configured to manage operation of the data center; and a self-healing engine that operates via an application programming interface (API) configured to receive the alert data, topology data corresponding to a topology of the data center, and user intent data and to select and execute one or more skills in conjunction with the data center orchestrator to correct the data center fault.