Fault analysis method and system for computer room equipment
By integrating and analyzing the operating data of multiple types of equipment in the computer room, using machine learning and multi-source data fusion technology, the problem that traditional fault diagnosis methods are difficult to integrate multi-source data is solved, efficient and accurate fault analysis and processing is achieved, and the intelligent level of computer room operation and maintenance is improved.
Patent Information
- Application Number
- CN202510144984.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
There are many types of equipment in the computer room and the monitoring data formats are different, which makes it difficult for traditional fault diagnosis methods to effectively integrate multi-source heterogeneous data, cannot accurately identify abnormal patterns, and it is difficult to accurately locate the source of the fault, affecting the timeliness of fault recovery and system stability.
By obtaining the operating data of computer room equipment, including device logs, sensor data, network traffic information and power consumption data, cleaning, format conversion and time synchronization processing are carried out to form a unified analysis data set. Use machine learning models to perform real-time analysis, identify abnormal data patterns, and generate preliminary fault warnings. Combining multi-source data fusion technology and historical fault databases, traceability analysis is used to use association analysis algorithms to determine the faulty equipment, and evaluate its impact on the entire computer room system.
It realizes efficient and accurate fault analysis, can accurately identify abnormal patterns, accurately locate the source of faults, improve the timeliness of fault recovery and system stability, and improves the automation and intelligence level of computer room operation and maintenance.
Smart Images

Figure CN120067756A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer room equipment, and particularly relates to a method and system for fault analysis of computer room equipment. Background Art
[0002] With the rapid development of computer technology, computer rooms have become the core infrastructure of various enterprises, government agencies, and data centers. Their stability and reliability directly affect the continuity of business and data security. A computer room usually consists of multiple subsystems such as servers, storage devices, network devices, and environmental control systems (such as power supply, refrigeration, and fire protection). These devices are in a high-load operation state for a long time and are easily affected by factors such as hardware aging, software failures, environmental factors (such as temperature and humidity changes), and human operation errors. Traditional fault diagnosis methods mostly rely on manual inspections and log analysis, with low efficiency and difficulty in detecting potential faults in a timely manner, resulting in equipment downtime, data loss, and even serious business interruptions. Therefore, there is an urgent need for an efficient and intelligent fault analysis method to improve the operation and maintenance management level of computer room equipment.
[0003] The existing technologies have the following deficiencies:
[0004] During the fault diagnosis process of a computer room, there are a wide variety of devices in the computer room, and the monitoring data formats generated by different devices are diverse, including log information, sensor data, network traffic, power consumption records, etc. The data sources are scattered and the dimensions are complex. In practical applications, traditional data fusion methods often have difficulty effectively integrating these multi-source heterogeneous data, resulting in the system being unable to accurately identify abnormal patterns and even possibly having false alarms or missed alarms. In addition, in the face of sudden faults (such as multiple devices having anomalies simultaneously in a short period of time), the existing methods are difficult to accurately locate the fault source, thus affecting the timeliness of fault recovery and system stability. Therefore, how to achieve efficient and accurate fault analysis by combining artificial intelligence, big data analysis, and machine learning technologies in the complex environment of a computer room is a technical problem that the current industry urgently needs to solve. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for fault analysis of computer room equipment to solve the deficiencies in the background art.
[0006] To achieve the above purpose, the present invention provides the following technical solution: A method for fault analysis of computer room equipment, comprising the following steps:
[0007] S1: Obtain the operation data of the computer room equipment, where the operation data includes equipment logs, sensor data, network traffic information, and power consumption data;
[0008] S2: Clean, convert the format and synchronize the time of the collected operation data to form a unified analysis data set, perform real-time analysis on the data set based on a machine learning model, identify abnormal data patterns, and generate preliminary fault warnings;
[0009] S3: Use multi-source data fusion technology, combine the historical fault database with the association analysis algorithm, conduct a traceability analysis on the abnormal data, and determine the faulty equipment;
[0010] S4: Based on the analysis of the topological relationship and business relevance of computer room equipment, evaluate the impact scope of the faulty equipment on the entire computer room system;
[0011] S5: Combine the fault case library to automatically generate an optimized processing plan for the faulty equipment, and after the faulty equipment is processed, continuously monitor the status of the faulty equipment to verify whether the fault is eliminated.
[0012] Preferably, in S3, align the timestamps of log data, sensor data, network traffic, and power consumption data, adopt a device topology model, associate the data with the corresponding device IDs, query the historical fault database, find similar abnormal data patterns, calculate the similarity between the current abnormal data and the historical data, calculate the fault propagation probability between devices, establish the association relationship between devices, and calculate the fault propagation path; use the Apriori algorithm to mine frequent fault patterns, check whether the current fault conforms to the known patterns through rule matching, and assist in judging the root cause device; combine the multi-source data characteristics, historical fault similarity, and association rules to calculate the fault score of each device. If the device log error rate > 80%: +0.4, device temperature > 90°C: +0.3, device overload historical similarity > 85%: +0.2, device current anomaly > 10%: +0.1, devices with a final score > 0.7 are determined to be faulty devices.
[0013] Preferably, in S2, after measuring the influence degree of a certain device fault between different levels through a graph traversal algorithm, generate a cross-level influence diffusion index. The method for obtaining the cross-level influence diffusion index is as follows:
[0014] Use all the devices in the computer room as nodes to construct a directed weighted graph G(V, E), where V is the set of devices, E is the connection relationship between devices, and the weight of the edge represents the influence degree on device v i For device v j , layer the topological graph by level, set the faulty source device v s and perform a depth-first search or breadth-first search to traverse all possible affected devices, and record the cumulative value of the path weight W(v s , v t ); calculate the level span L(v s , vt ),i.e., device v s affects v t After passing through how many different levels, calculate the impact score I(v t ) of a single affected device, and the expression is: Where: F is the set of fault source devices, α is the level attenuation coefficient, indicating the degree of impact attenuation of cross-level propagation, is the cross-level impact factor; the cross-level impact diffusion index CIEI of the entire computer room, and the expression is: Where: |V| is the total number of affected devices.
[0015] Preferably, after measuring the impact degree of the faulty device on data integrity and business continuity, generate a business data loss risk index. The acquisition method of the business data loss risk index is:
[0016] Define the business data dependency: The device set V = {v 1 , v 2 ,..., v n}, where v n represents the key device in the computer room, and the business data set D = {d 1 , d 2 ,..., d m}, where d m represents the business data, calculate the failure probability P(v i ) of device v i , and the expression is: Where: λ i represents the failure rate of device v i , T is the observation time window, calculate the failure impact factor F(v i ): F(v i ) = P(v i ) × W i ; where: W i is the importance weight of the device; calculate the data loss risk probability P(d j ), and the expression is: Where: M- 1 (d j ) represents all the devices that affect data d j ;
[0017] Calculate the data integrity impact score I(d j ): I(d j ) = P(d j ) × (1 - R j ); where: R j is the data backup and recovery probability, and for data without backup, R j= 0, full backup data R j = 1; Calculate the business continuity impact score C(d j ), and the calculation expression is: C(d j ) = I(d j ) × B j ; where: B j is the business importance weight; The overall business data loss risk index BDLRI of the computer room, the expression is: I D| is the total number of affected business data.
[0018] Preferably, normalize the cross - layer impact diffusion index and the business data loss risk index so that they are both in the range of [0, 1]. Calculate the impact range and impact level Impact Score of the faulty device on the entire computer room system through weighted calculation. The expression is: Impact Score = 0.6 × CIEI + 0.4 × BDLRI.
[0019] Preferably, if Impact Score > 0.75, trigger an emergency warning and notify all operation and maintenance teams; if 0.5 < ImpactScore ≤ 0.75, trigger a high - level warning and notify database and network engineers to intervene; if Impact Score ≤ 0.5, only trigger a regular alarm and arrange for regular inspections.
[0020] Preferably, in S5, after the computer room equipment fault is processed, it is necessary to continuously monitor the status of the faulty device to verify whether the fault is completely eliminated. Specifically, the fault recovery index FRI calculation formula is: where: FRI is the fault recovery index, S diff is the key status parameter change rate, T var is the time - series stability change amount, P anomaly is the abnormal pattern matching probability, and ImpactScore is the impact level, indicating the impact degree of the device on the entire computer room system.
[0021] Preferably, the key status parameter change rate S diff , and the calculation expression is: X post is the key operating parameter of the device after fault repair, X normal is the average operating parameter of the device in the normal state; The time - series stability change amount T var , the expression is: σ post is the standard deviation of the device status parameter after repair, σ normal is the historical standard deviation in the normal state; The abnormal pattern matching probability Panomaly The calculation expression is: P anomaly = AnomalyDetector(X post ); Use a machine learning anomaly detection model to calculate the similarity between the current device status and historical fault data, and the output range is [0, 1].
[0022] The present invention also provides a fault analysis system for computer room equipment, including a data acquisition module, a data processing module, a fault tracing module, a fault impact level determination module, and an intelligent fault handling module;
[0023] Data acquisition module: Obtain the operation data of computer room equipment, and the operation data includes equipment logs, sensor data, network traffic information, and power consumption data;
[0024] Data processing module: Clean, format convert, and time synchronize the acquired operation data to form a unified analysis data set, perform real-time analysis on the data set based on a machine learning model, identify abnormal data patterns, and generate a preliminary fault warning;
[0025] Fault tracing module: Use multi-source data fusion technology, combine the historical fault database with the association analysis algorithm, perform traceability analysis on abnormal data, and determine the faulty equipment;
[0026] Fault impact level determination module: Based on the topological relationship of computer room equipment and business relevance analysis, evaluate the impact range of the faulty equipment on the entire computer room system;
[0027] Intelligent fault handling module: Combine the fault case library, automatically generate an optimized fault handling plan for the faulty equipment, and after the faulty equipment is handled, continuously monitor the status of the faulty equipment to verify whether the fault is eliminated.
[0028] In the above technical solution, the technical effects and advantages provided by the present invention are:
[0029] 1. Through multi-source data fusion, machine learning analysis, and intelligent fault diagnosis, the present invention solves the problems of inconsistent data formats, inaccurate fault identification, and low traceability analysis efficiency in traditional computer room fault analysis. By constructing a unified data analysis framework, standardize the processing of multi-dimensional information such as equipment logs, sensor data, network traffic, and power consumption, and use a machine learning model to detect abnormal data patterns in real time and quickly generate a preliminary fault warning. At the same time, through topological relationship modeling and association analysis algorithm, combined with the historical fault database, accurately locate the faulty equipment, calculate its cross-level impact diffusion index and business data loss risk index on the entire computer room system, so as to quantitatively evaluate the impact range of the fault, and based on the weighted calculated impact level, perform intelligent hierarchical warning to ensure that the fault is efficiently responded to and processed.
[0030] 2. The present invention introduces an intelligent optimized fault handling solution and a continuous status monitoring mechanism, automatically recommends the optimal repair strategy using a fault case library, and ensures that the handling solution is highly targeted and the repair efficiency is high. After the fault is repaired, the system continuously monitors the device status, calculates the Fault Recovery Index (FRI), and dynamically evaluates whether the fault is completely eliminated by combining the change rate of key device parameters, the change amount of time series stability, the probability of abnormal pattern matching, and the impact level, so as to avoid the residue of latent faults. Through the present invention, the operation and maintenance of the computer room can achieve efficient fault detection, accurate traceability, intelligent early warning, optimal repair, and continuous monitoring closed-loop management, greatly improving the automation and intelligence level of the computer room operation and maintenance, and ensuring the stability of the business system and the integrity of data. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0032] Figure 1 It is a flowchart of the method of the present invention.
[0033] Figure 2 It is a system module diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0035] Embodiment 1. Please refer to Figure 1 As shown, a fault analysis method for computer room equipment in this embodiment includes the following steps:
[0036] S1: Obtain the operation data of the computer room equipment, where the operation data includes equipment logs, sensor data, network traffic information, and power consumption data;
[0037] S2: Clean, format convert, and time synchronize the collected operation data to form a unified analysis data set, perform real-time analysis on the data set based on a machine learning model, identify abnormal data patterns, and generate a preliminary fault warning;
[0038] S3: Using multi-source data fusion technology, combined with the historical fault database and association analysis algorithm, conduct traceability analysis on the abnormal data to determine the faulty equipment;
[0039] S4: Based on the analysis of the topological relationship and business relevance of computer room equipment, evaluate the impact scope of the faulty equipment on the entire computer room system;
[0040] S5: Combining with the fault case library, automatically generate an optimized processing plan for the faulty equipment, and after the faulty equipment is processed, continuously monitor the status of the faulty equipment to verify whether the fault is eliminated.
[0041] In S1, in the computer room, the equipment operation data is the basis for fault monitoring and analysis. Obtaining this data requires real-time collection and recording through a variety of hardware and software means to ensure the accuracy and integrity of fault analysis. Specifically, it includes the following types of data:
[0042] Device logs are record files automatically generated by computer room equipment during operation, including event records, status information, error logs, etc. Acquisition methods: Log services built into the system (such as syslog for Linux, event logs for Windows, etc.); Logs generated by application programs or middleware (such as database logs, server management software logs); Log interfaces or SNMP protocols of network devices (such as routers, switches). Example data content: Startup and shutdown records, user login and access records, system warning, error, crash and other abnormal information, configuration change or service status change records. Device log data is usually used to record system behavior, detect sudden errors and abnormal states, and is one of the core bases for fault location.
[0043] Sensor data is data generated by physical sensors in the computer room environment and equipment status monitoring, including parameters such as temperature, humidity, voltage, current, wind speed, and rotation speed. Acquisition methods: Obtained through environmental monitoring sensors installed in the computer room (such as temperature and humidity monitors, smoke alarms, etc.). Server, UPS power supply, air conditioner and other equipment are equipped with built-in sensor modules to obtain data through communication protocols (such as Modbus, BACnet, SNMP). Example data content: Computer room temperature and humidity, air outlet temperature of air conditioning equipment, voltage and current values of power supply equipment, server fan rotation speed, hard disk temperature. Sensor data is used to monitor the computer room environment and equipment operation status, and can effectively identify abnormalities caused by environmental factors, such as equipment overheating faults caused by high temperature.
[0044] Network traffic information refers to the data transmission situation between network devices (such as switches, routers, firewalls) and servers in the computer room, including traffic volume, network latency, packet loss rate, etc. Acquisition methods: Real-time collection through network monitoring devices (such as network probes, traffic analyzers); network management protocols (such as SNMP, NetFlow, sFlow, etc.); interfaces provided by network device logs and monitoring systems. Example data content: Network interface input / output traffic (bps), packets transmitted per second (pps), packet loss rate, latency time, network port status and error statistics. Network traffic information can be used to detect security events such as network bottlenecks and DDoS attacks, as well as identify faults caused by network congestion or communication anomalies.
[0045] Power consumption data refers to the power consumption of devices or systems in the computer room, including voltage, current, power, and energy consumption records, etc. Intelligent power monitoring systems (such as PDU power distribution units), UPS power management systems, power monitoring sensors and power management interfaces (Modbus, BACnet, SNMP, etc.). Example data content: Real-time voltage, current, and power values, UPS battery status and discharge time, total energy consumption of the computer room and energy consumption distribution of each device, power interruption and switching records. Power consumption data is used to monitor the stability of the computer room power supply system and can help identify faults caused by abnormal power supply (such as voltage fluctuations, overload). In addition, energy consumption optimization is also one of the important goals in operation and maintenance management.
[0046] S2: Clean, format-convert, and time-synchronize the collected operation data to form a unified analysis dataset. Based on a machine learning model, perform real-time analysis on the dataset, identify abnormal data patterns, and generate preliminary fault warnings.
[0047] Remove redundant, incorrect, and missing data to improve data quality. Remove duplicate data: Check whether there are duplicate entries in logs, sensor data, and network traffic data and perform deduplication. Fill in missing values: For a small amount of missing data, mean filling, interpolation method, or KNN interpolation can be used for completion; if the data is severely missing, discard the data point. Outlier detection: Use box plot analysis (IQR), Z-score method, or based on historical statistical rules to remove obviously abnormal error data points (such as extreme values caused by sensor failures).
[0048] Convert data from multiple sources into a unified structured format for easy analysis. Unified data format: Convert data from different sources (such as JSON, CSV, Syslog, etc.) into a unified database format, such as SQL tables, Parquet files, or storage based on a time series database (TSDB). Data standardization: For numerical data, use Min-Max normalization or Z-score normalization to make it fall within a unified range (such as 0-1 or -1-1) to eliminate the influence of dimensionality.
[0049] Unify the timestamps of different data sources to ensure the temporal consistency of the data. Based on NTP (Network Time Protocol), synchronize the timestamps of all devices to ensure the consistency of time data. Divide the data into fixed time windows (such as 1 second, 5 seconds, or 1 minute), and align the multi-source data within the same time window to form a time series dataset.
[0050] Extract key features to provide effective input for machine learning models, including: Time series features: Calculate time series features such as moving average, standard deviation, skewness, and kurtosis. Statistical features: Extract historical distribution statistical information of device logs, sensor data, and network traffic. Behavioral pattern features: Combine the computer room topology to construct the dependency relationship between devices (such as whether a sudden increase in network traffic is accompanied by an increase in CPU load). Anomaly scoring: Introduce indicators such as MAE (Mean Absolute Error) and SSE (Sum of Squared Errors) to evaluate the degree of deviation of the current data from the normal mode.
[0051] Train a model based on historical data to detect abnormal data in real time. Unsupervised learning methods (suitable for unknown faults): Isolation Forest: Detect abnormal points based on a tree model, such as a sudden increase in device power consumption or a sharp increase in the log error rate. Principal Component Analysis (PCA): Used for dimensionality reduction and discovery of abnormal points, such as temperature and humidity deviating from the normal mode. Autoencoder: Use a neural network to learn the normal mode and identify anomalies when the error is large. Supervised learning methods (suitable for known faults): Random Forest: Train a classification model based on historical fault data to identify specific types of device anomalies. LSTM (Long Short-Term Memory Network): Suitable for time series data prediction, such as predicting abnormal loads on servers. Logistic Regression: Used for binary classification problems, such as determining whether a device is about to fail. Based on the anomaly detection results, determine the alarm level and trigger a warning.
[0052] Set the alarm threshold based on the Anomaly Score, e.g., Normal range (0 - 0.5): No anomaly; Slight anomaly (0.5 - 0.7): Prompt alarm (suggested to check); Severe anomaly (0.7 - 1.0): Emergency alarm (handle immediately). Multi-factor alarm strategy: Combine multiple data sources for joint judgment to avoid misjudgment of a single indicator. For example: Server CPU load is higher than 90% and fan speed is abnormal → It may be a heat dissipation failure. Network traffic surges and the packet loss rate of multiple ports increases → It may be a DDoS attack.
[0053] The early warning report includes: Alarm time: YYYY - MM - DD HH:MM:SS; Fault device: Server ID, Switch ID, etc.; Anomaly type: High temperature, abnormal power consumption, traffic surge, etc.; Affected scope: List of affected services and devices; Suggested measures: Check the fan operation, adjust the load balance, restart the device, etc.
[0054] S3: Use the multi-source data fusion technology, combine the historical fault database with the association analysis algorithm, and conduct traceability analysis on the abnormal data to determine the faulty device.
[0055] Align the timestamps of various data such as log data, sensor data, network traffic, and power consumption to ensure the time consistency of different data sources. Adopt the device topology model to associate the data with the corresponding device ID for subsequent analysis (e.g., Abnormal log of Server A → Abnormal temperature of Server A → Abnormal CPU load of Server A). Normalize different types of data (such as CPU utilization 0 - 100%, current 0 - 10A, log error count 0 - 10000) to the same numerical range (such as 0 - 1) for subsequent algorithm processing. Use PCA (Principal Component Analysis) or Autoencoder for dimensionality reduction to remove redundant features and extract key fault patterns. Set the analysis time window (such as the past 5 minutes, the past 1 hour, the past 1 day) to check the data change trend within this time period to judge the evolution process of the fault.
[0056] Query the historical fault database to find similar abnormal data patterns, such as the temperature, power consumption, and log features of a certain device having a 90% similarity with a previous fault. Use methods such as KNN (K-Nearest Neighbor), cosine similarity, and DTW (Dynamic Time Warping) to calculate the similarity between the current abnormal data and the historical data. Statistically calculate the fault occurrence probability of the same type of devices (e.g., This brand of UPS has had 3 battery failures in the past 6 months). Calculate the fault propagation probability between devices (e.g., A certain switch fails → The packet loss rate of multiple servers increases).
[0057] Adopt methods such as Bayesian networks, Markov chains, and graph neural networks (GNNs) to establish the correlation relationships between devices and calculate the fault propagation paths. Example: Sudden increase in switch traffic → Network packet loss in downstream servers → Increased CPU occupancy of multiple servers → Error reported in the log of a certain server. Traceability analysis result: It may be a network attack or a switch port failure.
[0058] Use the Apriori algorithm to mine frequent fault patterns, such as: 「UPS overload」→「Increased room temperature」→「Server failure」(support 80%, confidence 90%); Through rule matching, check whether the current fault conforms to the known patterns to assist in judging the root cause device.
[0059] Through time series data analysis, judge whether an anomaly is a causal relationship with another anomaly rather than a random correlation. For example: Increase in server CPU utilization rate → Temperature increase after 2 minutes (possibly due to CPU load causing temperature increase). Temperature increase first → CPU utilization rate increase after 5 minutes (possibly due to abnormal operation of the computer room air conditioner causing CPU frequency reduction).
[0060] Combine multi-source data characteristics, historical fault similarity, and association rules to calculate the fault score (Anomaly Score) for each device.
[0061] Example scoring criteria: Device log error rate > 80%: +0.4
[0062] Device temperature > 90°C: +0.3
[0063] Device overload historical similarity > 85%: +0.2
[0064] Device current anomaly > 10%: +0.1
[0065] Devices with a final score > 0.7 are determined to be high-probability faulty devices.
[0066] Generate a fault traceability report: Determine the source device of the fault (such as Server A, Switch B). Predict possible fault types (such as heat dissipation failure, hardware failure, network attack). Provide the fault propagation path (such as UPS voltage fluctuation → Server power failure → Database downtime).
[0067] S4: Based on the topological relationship and business relevance analysis of computer room equipment, evaluate the impact scope of faulty devices on the entire computer room system and provide a fault level assessment report.
[0068] Obtain the topological structure of the computer room equipment, and collect the connection relationships of devices such as servers, storage devices, switches, UPSs, power systems, and refrigeration systems. Use protocols such as SNMP, Modbus, and IPMI to obtain the upstream and downstream dependencies of the devices. Construct a topological relationship diagram and define the hierarchy between devices (such as core switch → business server → database storage).
[0069] Use a graph model (Graph Database, such as Neo4j) to represent the connection weights of devices: network dependencies (switch → server), power dependencies (UPS → server), and cooling dependencies (air conditioner → server temperature).
[0070] Determine the mapping relationship between the business system and physical devices (such as a Web application depending on the backend database, and the database depending on the storage system). Through log analysis and API call tracing, determine the mutual dependencies of different business modules. Use distributed tracing tools (such as Jaeger, Zipkin) to analyze the business traffic path. Business Continuity Level (BCL): Assessment of the importance of the business, such as payment system > internal mail system. Business Critical Path (BCP): Identify whether the faulty device is at the core node of the critical business flow.
[0071] After measuring the impact degree of a device failure between different levels (network layer, computing layer, storage layer, power supply layer) through a graph traversal algorithm, generate a cross-level impact diffusion index. The method for obtaining the cross-level impact diffusion index is as follows:
[0072] Use all the devices in the computer room as nodes to construct a directed weighted graph G(V, E), where V is the set of devices, such as servers, switches, storage devices, UPSs, etc. E is the connection relationship between devices, such as network links, power supply links, storage access paths, etc. The weight Wij of the edge represents the impact degree on device v i For device v j (based on historical data, fault propagation probability, etc.). Layer the topological graph according to levels (network layer, computing layer, storage layer, power supply layer), and each level may have internal connections and cross-level connections.
[0073] Set the faulty source device v s and perform a depth-first search (DFS) or breadth-first search (BFS) to traverse all possible affected devices. Record the cumulative value of the path weight W(v s , v t (the propagation impact value from the faulty source device v s to the target device v t ).
[0074] Calculate the level span L(v s , vt ), i.e., device v s affects v t How many different levels (such as network layer → computing layer → storage layer) have passed through. Calculate the impact score I(v t ), and the expression is: Where: F is the set of fault source devices, α is the level attenuation coefficient (empirical value, such as 0.5 - 1.0, which can be optimized according to historical data), representing the degree of impact attenuation during cross-level propagation, is the cross-level impact factor, and as the level span increases, the impact value decays exponentially.
[0075] The cross-level impact diffusion index CIEI of the entire computer room, and the expression is: Where: |V| is the total number of affected devices. The higher the CIEI value, the greater the fault impact range and the more serious the cross-level impact.
[0076] After measuring the degree of impact of the faulty device on data integrity and business continuity, generate the business data loss risk index. The method for obtaining the business data loss risk index is:
[0077] Define the business data dependency relationship: The device set V = {v 1 , v 2 ,..., v n}, where v n represents the key devices in the computer room, such as storage servers, database servers, application servers, etc. The business data set D = {d 1 , d 2 ,..., d m}, where d m represents business data, such as database tables, file storage, log data, etc. The mapping relationship M between devices and business data: V → D, that is, the device v n affects the data d m .
[0078] Calculate the failure probability P(v i ) of the device v i ), and the expression is: Where: λ i represents the failure rate of the device v i (which can be calculated based on historical data or the manufacturer's MTBF). T is the observation time window (such as the past 24 hours). Calculate the fault impact factor: F(v i ) = P(v i ) × W i ; where: W i is the importance weight of the device (such as a database server is more important than an ordinary web server).
[0079] Calculate the probability of data loss risk P(d j ), the loss risk of business data d j is affected by all devices that support it, and the expression is: Where: M -1 (d j ) represents all devices that affect data d j .
[0080] Calculate the impact score of data integrity I(d j ). If data d j has a backup, the impact score is reduced: I(d j ) = P(d j ) × (1 - R j ); Where: R j is the data backup recovery probability (such as the recovery rate of snapshots and redundant storage). For data without a backup, R j = 0, and for fully backed up data, R j = 1.
[0081] Calculate the impact score of business continuity C(d j ). The loss of business data may cause business interruption, and the calculation expression is: C(d j ) = I(d j ) × B j ; Where: B j is the business importance weight (such as the core database is more important than the log file).
[0082] The business data loss risk index BDLRI of the entire computer room is expressed as: |D| is the total number of affected business data. The value range of BDLRI is 0 to 1. The closer it is to 1, the higher the data loss risk.
[0083] Normalize the cross - level impact diffusion index and the business data loss risk index so that they are both in the range of [0, 1]. Obtain the impact range and impact level Impact Score of the faulty device on the entire computer room system through weighted calculation of the normalized cross - level impact diffusion index and the business data loss risk index. The expression is: ImpactScore = 0.6×CIEI + 0.4×BDLRI; If Impact Score > 0.75, trigger an emergency warning and notify all operation and maintenance teams; If 0.5 < Impact Score ≤ 0.75, trigger a high - level warning and notify database and network engineers to intervene; If ImpactScore ≤ 0.5, only trigger a regular alarm and arrange for regular inspections.
[0084] S5: Combine with the fault case library to automatically generate an optimized processing solution for the faulty device. After the faulty device is processed, continuously monitor the status of the faulty device to verify whether the fault is eliminated.
[0085] During the process of handling computer room equipment failures, the traditional manual troubleshooting method is inefficient and has a high misjudgment rate, making it difficult to resume services in a timely manner. To improve the accuracy and speed of fault handling, the fault case library can be combined, and technologies such as historical data matching and machine learning analysis can be used to automatically generate the optimal processing solution for the faulty device. By constructing an intelligent case matching model, the system can identify the current fault characteristics, retrieve similar fault histories from the case library, extract the best repair strategies, and form a multi-level fault handling solution for the operation and maintenance personnel to refer to or execute automatically.
[0086] When optimizing the processing solution, the system will dynamically adjust the repair steps according to factors such as the severity of the fault and the scope of business impact, in combination with historical successful cases. For example, for low-risk faults, it is recommended to adjust system parameters or load balancing methods online; for medium-risk faults, it is recommended to switch the business traffic to standby devices or implement temporary remedial measures; if it is a high-risk fault, operations such as emergency shutdown maintenance, hardware replacement, or data recovery need to be performed. At the same time, the system will be optimized in combination with the real-time computer room environment status (such as temperature, power load, network traffic) to ensure the feasibility of the solution and minimize the business interruption time.
[0087] Finally, the automatically generated optimized solution can not only improve the fault handling efficiency, but also continuously update the case library through continuous learning and optimization mechanisms, improving the accuracy of future fault diagnosis and repair. After the operation and maintenance personnel execute the processing solution, the system will automatically record the fault handling process, recovery effect, and optimization suggestions, providing data support for subsequent optimization. In the long run, this intelligent solution can greatly reduce human misjudgment, improve the intelligent level of computer room operation and maintenance, and ensure the stability of the business system and data security.
[0088] After the computer room equipment failure is handled, it is necessary to continuously monitor the status of the faulty device to verify whether the fault is completely eliminated and evaluate its impact on the overall computer room system. By comparing key monitoring indicators, historical data, and impact assessment, calculate the fault recovery index to ensure the stable operation of the system.
[0089] The calculation formula for the Fault Recovery Index FRI is: ImpactScore; where: FRI is the Fault Recovery Index, used to measure whether the faulty device has returned to normal, with a value range of [0, 1]. The closer the value is to 1, it indicates that the fault has been completely eliminated and the device status has returned to normal. S diff is the change rate of key state parameters, calculating the deviation degree of the key operating indicators (such as CPU temperature, power consumption, I / O throughput) of the repaired device relative to the normal state. Tvar is the change amount of time - series stability, which is used to measure the fluctuation of the device operation state. If the volatility is still high after repair, it indicates that there may be potential hazards. P anomaly is the probability of abnormal pattern matching. The similarity between the current device operation state and the historical abnormal state is calculated through a machine - learning anomaly detection model. If the similarity is high, it means that the fault may not be completely repaired. Impact Score is the impact level, which represents the degree of influence of the device on the entire computer room system. The value range is [0, 1]. The larger the influence range, the higher this value is.
[0090] Among them, the change rate S of the key state parameter diff , and the calculation expression is: X post is the key operation parameter of the device after fault repair (such as CPU temperature, current, disk I / O, etc.). X normal is the average operation parameter of the device in the normal state (which can be calculated from historical data). The closer this value is to 0, the more normal the device is restored; the larger the value, the more abnormal the device still is.
[0091] The change amount of time - series stability T var , and the expression is: σ post is the standard deviation of the device state parameter after repair, indicating its fluctuation. σ normal is the historical standard deviation in the normal state. If T var > 1, it means that the device state fluctuates greatly after repair and there may still be a risk of failure.
[0092] The probability of abnormal pattern matching P anomaly The calculation expression of is: P anomaly = AnomalyDetector(X post ); Use a machine - learning anomaly detection model (such as IsolationForest, LSTM time - series prediction) to calculate the similarity between the current device state and historical fault data, and the output range is [0, 1]. If P anomaly is close to 1, it means that there is still an abnormal pattern after the device is repaired and it may not be completely restored.
[0093] In this embodiment, by obtaining the operation data of computer room equipment, including equipment logs, sensor data, network traffic information, and power consumption data, the collected data is cleaned, format-converted, and time-synchronized to form a unified analysis data set, and the data is analyzed in real time based on a machine learning model to identify abnormal data patterns and generate preliminary fault warnings. Subsequently, using multi-source data fusion technology, combined with the historical fault database and correlation analysis algorithm, the abnormal data is traced and analyzed to accurately locate the faulty equipment, and based on the topological relationship and business relevance analysis of computer room equipment, the impact scope of the faulty equipment on the entire computer room system is evaluated. Finally, an optimized processing solution for the faulty equipment is automatically generated in combination with the fault case library, and after the fault is processed, continuous status monitoring is performed on the faulty equipment to verify whether the fault is completely eliminated and ensure the stable operation of the system.
[0094] Example 2, please refer to Figure 2 As shown, a fault analysis system for computer room equipment in this embodiment includes a data acquisition module, a data processing module, a fault tracing module, a fault impact level determination module, and an intelligent fault processing module;
[0095] Data acquisition module: Obtain the operation data of computer room equipment, where the operation data includes equipment logs, sensor data, network traffic information, and power consumption data;
[0096] Data processing module: Clean, format-convert, and time-synchronize the collected operation data to form a unified analysis data set, and perform real-time analysis on the data set based on a machine learning model to identify abnormal data patterns and generate preliminary fault warnings;
[0097] Fault tracing module: Use multi-source data fusion technology, combined with the historical fault database and correlation analysis algorithm, to trace and analyze the abnormal data to determine the faulty equipment;
[0098] Fault impact level determination module: Based on the topological relationship and business relevance analysis of computer room equipment, evaluate the impact scope of the faulty equipment on the entire computer room system;
[0099] Intelligent fault processing module: Combine the fault case library to automatically generate an optimized processing solution for the faulty equipment, and after the faulty equipment is processed, perform continuous status monitoring on the faulty equipment to verify whether the fault is eliminated.
[0100] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to obtain a formula that is closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0101] It should be understood that the term "and / or" in this text is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this text generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship. The specific meaning can be understood by referring to the context before and after.
[0102] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this text can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0103] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application.
Claims
1. A method for analyzing computer room equipment failures, characterized in that: It includes the following steps: S1: Obtain the operation data of computer room equipment, where the operation data includes equipment logs, sensor data, network traffic information, and power consumption data; S2: Clean, format-convert, and time-synchronize the collected operation data to form a unified analysis data set. Based on a machine learning model, perform real-time analysis on the data set, identify abnormal data patterns, and generate preliminary fault warnings; S3: Use multi-source data fusion technology, combine the historical fault database with the association analysis algorithm, perform traceability analysis on the abnormal data, and determine the faulty equipment; S4: Based on the analysis of the topological relationship and business relevance of computer room equipment, evaluate the impact scope of the faulty equipment on the entire computer room system; S5: Combine the fault case library to automatically generate an optimized processing plan for the faulty equipment. After the faulty equipment is processed, continuously monitor the status of the faulty equipment to verify whether the fault is eliminated.
2. A method for analyzing computer room equipment failure according to claim 1, characterized in that: In S3, align the timestamps of the log data, sensor data, network traffic, and power consumption data. Adopt the equipment topology model to associate the data with the corresponding equipment ID, query the historical fault database, find similar abnormal data patterns, calculate the similarity between the current abnormal data and the historical data, calculate the fault propagation probability between equipment, establish the association relationship between equipment, and calculate the fault propagation path; Use the Apriori algorithm to mine frequent fault patterns. Through rule matching, check whether the current fault conforms to the known pattern to assist in judging the root cause equipment. Combine the multi-source data characteristics, historical fault similarity, and association rules to calculate the fault score of each equipment. If the equipment log error rate > 80%: +0.4, equipment temperature > 90°C: +0.3, equipment overload historical similarity > 85%: +0.2, equipment current anomaly > 10%: +0.1, the equipment with a final score > 0.7 is determined to be faulty equipment.
3. A method for analyzing computer room equipment failure according to claim 2, characterized in that: In S2, after measuring the influence degree of a certain equipment fault between different levels through the graph traversal algorithm, generate a cross-level influence diffusion index. The method for obtaining the cross-level influence diffusion index is: Take all the devices in the computer room as nodes and construct a directed weighted graph G(V,E), where V is the device set, E is the connection relationship between devices, and the weight of the edge represents the device v. i For device v j The impact degree is determined by dividing the topology into layers and setting the fault source device v s And perform a depth-first search or breadth-first search to traverse all potentially affected devices and record the cumulative path weight value W ( v s ,v t) ; Calculate the level span L that affects the spread ( v s ,v t) , that is, device v s Affects v t How many different levels have been passed to calculate the impact score I of a single affected device ( v t) , the expression is: Where: F is the set of fault source devices, α is the level attenuation coefficient, which indicates the attenuation degree of the impact of cross-level propagation. is the cross-level impact factor; the cross-level impact diffusion index CIEI of the computer room as a whole is expressed as: in: | V | is the total number of affected devices.
4. A method for analyzing computer room equipment failure according to claim 3, characterized in that: After measuring the influence degree of the faulty equipment on data integrity and business continuity, generate a business data loss risk index. The method for obtaining the business data loss risk index is: Define business data dependencies: Device set V = {v1, v2, ..., v n }, where v n Represents the key equipment in the computer room, and the business data set D = {d1, d2, ..., d m }, where d m Represents business data, computing equipment v i The failure probability P ( v i) , the expression is: Where: i Indicates device v i The failure rate, T is the observation time window, and the failure impact factor F is calculated. ( v i) : F ( v i) =P ( v i) ×W i ;W i is the importance weight of the device; calculate the data loss risk probability P(d j ), the expression is: Where: M -1( d j ) Indicates the impact data d j All equipment; Calculate the data integrity impact score I(d j ):I ( d j ) =P ( d j ) × ( 1-R j ) ; Among them: R j is the probability of data backup recovery, and R is the probability of data without backup. j =0, complete backup data R j =1; Calculate business continuity impact score C ( d j ) , the calculation expression is: C ( d j ) =I ( d j ) ×B j ; Among them: B j is the business importance weight; the overall business data loss risk index BDLRI of the computer room is expressed as: D | The total number of affected business data.
5. A method for analyzing computer room equipment failure according to claim 4, characterized in that: Normalize the cross-level influence diffusion index and the business data loss risk index so that they are both within [0, 1]. Obtain the impact range impact level Impact Score of the faulty equipment on the entire computer room system through weighted calculation of the normalized cross-level influence diffusion index and the business data loss risk index. The expression is: Impact Score = 0.6 × CIEI + 0.4 × BDLRI.
6. A method for analyzing computer room equipment failure according to claim 5, characterized in that: If ImpactScore > 0.75, trigger an emergency warning and notify all operation and maintenance teams; if 0.5 < Impact Score ≤ 0.75, trigger a high-level warning and notify the database and network engineers to intervene; if Impact Score ≤ 0.5, only trigger a regular alarm and arrange for regular inspections.
7. A method for analyzing computer room equipment failure according to claim 1, characterized in that: In S5, after the computer room equipment fault is handled, it is necessary to continuously monitor the faulty equipment to verify whether the fault is completely eliminated. Specifically, the fault recovery index FRI calculation formula is: Where: FRI is the fault recovery index, S diff is the rate of change of key state parameters, T var is the stability change of the time series, P anomaly is the abnormal pattern matching probability, and ImpactScore is the impact level, which indicates the impact of the device on the entire computer room system.
8. A method for analyzing computer room equipment failure according to claim 7, characterized in that: Among them, Key state parameter change rate S diff , the calculation expression is: X post is the key operating parameter of the equipment after fault repair, X normal is the average operating parameter of the equipment under normal conditions; the time series stability change T var , the expression is: σ post is the standard deviation of the equipment status parameters after repair, σ normal is the historical standard deviation under normal conditions; the abnormal pattern matching probability P anomaly The calculation expression is: anomaly =AnomalyDetector ( X post ) ; Use the machine learning anomaly detection model to calculate the similarity between the current device status and historical fault data, with an output range of [0,1].
9. A computer room equipment fault analysis system, used to implement a computer room equipment fault analysis method according to any one of claims 1 to 8, characterized in that: It includes data acquisition module, data processing module, fault tracing module, fault impact level determination module and intelligent fault processing module; Data collection module: obtains the operating data of the computer room equipment, including equipment logs, sensor data, network traffic information and power consumption data; Data processing module: cleans, converts the format and synchronizes the time of the collected operation data to form a unified analysis data set, performs real-time analysis on the data set based on the machine learning model, identifies abnormal data patterns, and generates preliminary fault warnings; Fault tracing module: Utilizes multi-source data fusion technology, combines historical fault database and correlation analysis algorithm to trace the abnormal data and identify the faulty equipment; Fault impact level determination module: Based on the topological relationship of computer room equipment and business correlation analysis, it evaluates the impact of faulty equipment on the entire computer room system; Intelligent fault handling module: Combined with the fault case library, it automatically generates optimized fault equipment handling solutions, and after the fault equipment is handled, it continuously monitors the status of the faulty equipment to verify whether the fault has been eliminated.
Citation Information
Cited By
Equipment fault early warning system and method based on Internet of Things
CN120406407A
Campus self-service laundry equipment abnormity identification system based on Internet of Things
CN120408474A
System and method for monitoring dynamic environment and equipment state of high-voltage power distribution station house
CN120489255A
Fault prediction method of server, electronic equipment, medium and product
CN120560900A
Weak current intelligent management system based on multi-source data fusion
CN120659303A