Data Center Power Supply Failure Prediction via Management Controller
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data centers face challenges in predicting and preventing power supply failures, which can lead to downtime and reduced operational efficiency due to variations in environmental conditions and the lack of effective predictive mechanisms for redundant power supplies.
Innovation Solution
A system and method that utilize a management controller to detect power supply failures, receive operational data from both primary and redundant power supplies, and calculate a probability of failure based on environmental and operational factors, allowing for proactive switching and maintenance to prevent failures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional power supply monitoring is used, then failure detection is achieved, but predictive capability and prevention are lacking
Solution Approach 1:
The system performs preliminary actions by collecting and analyzing operational data, environmental conditions, and failure patterns before actual power supply failures occur. The management controller continuously monitors multiple parameters and calculates failure probabilities in advance, enabling proactive maintenance scheduling and redundant power supply switching before failures impact system operation, thus reducing downtime while improving reliability
Solution Approach 2:
The system implements feedback mechanisms by continuously collecting operational data from power supplies, comparing it against historical failure patterns and environmental conditions, and using this feedback to update failure probability calculations. The management controller receives status information from power supplies and adjusts maintenance schedules or triggers redundant supply activation based on real-time feedback, creating a closed-loop system that improves reliability while minimizing downtime
2Reliability
If redundant power supplies are implemented, then system availability is improved, but failure prediction capability remains insufficient
Solution Approach 1:
The management controller serves as an intermediary that simplifies the management of redundant power supplies. It automatically collects operational data from multiple power supplies, compares it against historical failure patterns and environmental conditions, calculates failure probabilities, and makes intelligent decisions about which redundant supply to activate. This intermediary layer abstracts the complexity from the overall system, maintaining high availability while reducing management burden
Solution Approach 2:
The system implements self-service capabilities by enabling power supplies to automatically report their operational status, environmental conditions, and performance metrics to the management controller. The controller autonomously analyzes this self-reported data, compares it against known failure patterns, and automatically manages redundant power supply activation without requiring manual intervention, thereby reducing management complexity while maintaining system availability
3Measurement precision
If environmental monitoring is added, then failure prediction accuracy is improved, but system complexity increases
Solution Approach 1:
The management controller is designed with multi-functionality, serving both as the primary controller for power supply management and as the environmental monitoring system. It collects operational data from power supplies and environmental data (temperature, humidity, airflow) using the same controller infrastructure, eliminating the need for separate dedicated monitoring hardware. This universal approach improves failure prediction accuracy through enhanced data collection while minimizing increases in system complexity
Data Source
AI summary
An information handling system may include a first power supply for a first system, a second power supply for a second system, and a management controller. The management controller may detect that the first power supply has failed, receive first information from the first system related to the operation of the first power supply prior to the failure of the first power supply, receive second information from the second system associated with the second power supply, and determine a probability of failure of the second power supply based upon a comparison of the first information with the second information.


