Automated Device Failure Management via Historical Data Querying
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In data centers, device failures are complex and require significant human intervention and technical knowledge for analysis and fixing, leading to inefficient and repetitive processes, especially with the increasing number of devices.
Innovation Solution
A method and apparatus for automatically detecting device failures, generating failure reports, querying a device object repository for historical failure information and fix solutions, and applying the most relevant fix solutions to resolve issues quickly, reducing manual labor and improving efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If separate manual analysis is conducted for each device failure, then professional technical knowledge can be applied, but the process becomes repetitive and consumes large amounts of human operations and time
Solution Approach 1:
The system performs preliminary action by automatically collecting and storing device failure information and fix solutions in a database before failures occur. When a failure happens, the system queries the database for historical failure patterns and recommended fix solutions, eliminating the need for repetitive manual analysis and reducing failure analysis time while maintaining fixing accuracy.
Solution Approach 2:
The system implements feedback by continuously collecting device failure information, analyzing patterns from historical data, and providing feedback in the form of recommended fix solutions. This closed-loop feedback mechanism allows the system to learn from past failures and improve future failure resolution efficiency without requiring repeated manual intervention.
2Productivity
If more devices are installed in the data center, then service capacity increases, but device management becomes more complicated and resource consumption increases
Solution Approach 1:
The system applies universality by creating a unified device management platform that handles multiple device types (servers, storage devices, network devices) through a single automated failure management system. This universal approach allows the data center to scale up device installation while keeping management complexity controlled through standardized automated processes rather than device-specific manual procedures.
Solution Approach 2:
The system enables self-service by automatically detecting device failures, querying historical failure data, and providing fix solutions without requiring human intervention. This self-service capability allows the data center to manage increasing numbers of devices autonomously, reducing the complexity burden on human operators as device counts grow.
3Reliability
If human operations are used for failure analysis, then professional technical knowledge is applied, but large amounts of human resources and operational effort are required
Solution Approach 1:
The system introduces an intermediary - an automated failure management system with database querying capabilities - that bridges the gap between device failures and fix solutions. This intermediary automatically collects failure information, queries historical data, and provides recommended solutions, maintaining high failure analysis quality through systematic data analysis while eliminating the need for extensive human operational effort.
Data Source
AI summary
Embodiments of the present disclosure relate to a method and apparatus for managing a failure of a device. The method comprises detecting whether a failure occurs in a device, and generating a failure report for the failure in response to the failure occurring in the device. The method further comprises querying a device object repository with the failure report, and the object device repository stores historical failure information associated with the device and a fix solution corresponding to the historical failure information. The method further comprises obtaining the fix solution from the device object repository based on a comparison between the failure report and the historical failure information. Embodiments of the present disclosure can manage the failure of the device more effectively.


