Data center equipment maintenance method and device

By building a device health scoring system and knowledge graph to analyze the causes of failures, the problem of inefficiency in traditional operation and maintenance methods is solved, intelligent equipment management and fault prediction are realized, and the operation and maintenance efficiency and reliability of data center equipment are improved.

CN120560883APending Publication Date: 2025-08-29CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510661500.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The traditional computer room equipment operation and maintenance methods rely on regular inspections and passive fault responses, resulting in low operation and maintenance efficiency, inability to predict potential faults in advance, lack of intelligent analysis and prediction capabilities, resulting in equipment performance degradation and business interruption.

Method used

By collecting multi-dimensional data, a device health scoring system is built, analyzing the causes of failures using knowledge graphs and graph neural networks, dynamically adjusting maintenance plans, and realizing intelligent operation and maintenance of equipment.

Benefits of technology

It improves operation and maintenance efficiency, reduces the passive response time of equipment failures, improves the intelligence and prediction capabilities of equipment management, and reduces the risk of equipment downtime.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120560883A_ABST
    Figure CN120560883A_ABST
Patent Text Reader

Abstract

The invention discloses a data center equipment maintenance method and device. The method comprises the steps that multi-dimensional data of target equipment is collected, and the multi-dimensional data at least comprises operation data of the target equipment; determining a health degree score of the target equipment according to the operation data of the target equipment; determining whether the target equipment is in a fault state or not according to the health degree score of the target equipment, and determining a fault reason according to a pre-constructed knowledge graph when the target equipment is in the fault state; and determining a maintenance scheme according to the fault reason, wherein the maintenance scheme is used for maintaining the target equipment. The technical problem that the operation and maintenance efficiency of a machine room equipment operation and maintenance mode in the related technology is low is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of operation and maintenance management technology, and specifically, to a data center equipment maintenance method and device. Background Art

[0002] With the rapid development of data centers and cloud computing technologies, the reliability and maintenance of computer room equipment have become critical to ensuring the stable operation of information infrastructure. Data center rooms typically house servers, storage devices, network equipment, power management systems (UPS), and air conditioning and cooling systems. These devices, when operating under high load for extended periods, are susceptible to temperature fluctuations, power supply fluctuations, hardware aging, and environmental factors, leading to performance degradation or even failure, resulting in business interruptions or data loss.

[0003] Traditional computer room equipment operation and maintenance mainly relies on regular inspections, manual monitoring and passive fault response. That is, when equipment malfunctions or crashes, the operation and maintenance personnel receive alarm information and conduct troubleshooting and repairs, resulting in low operation and maintenance efficiency. Summary of the Invention

[0004] The embodiments of the present application provide a data center equipment maintenance method and apparatus to at least solve the technical problem of low operation and maintenance efficiency of computer room equipment operation and maintenance methods in related technologies.

[0005] According to one aspect of an embodiment of the present application, a data center equipment maintenance method is provided, comprising: collecting multi-dimensional data of a target device, wherein the multi-dimensional data includes at least: operating data of the target device; determining a health score of the target device based on the operating data of the target device; determining whether the target device is in a fault state based on the health score of the target device, and if the target device is in a fault state, determining a cause of the fault based on a pre-constructed knowledge graph; and determining a maintenance plan based on the cause of the fault, wherein the maintenance plan is used to maintain the target device.

[0006] Optionally, determining the health score of the target device based on the operating data of the target device includes: extracting the current temperature of the target device, the power consumption of the target device, and the load of the target device from the operating data of the target device; obtaining the historical fault data of the target device from the multi-dimensional data; determining the temperature score of the target device, the power consumption score of the target device, the utilization score of the target device, and the fault history score of the target device based on the current temperature of the target device, the power consumption of the target device, the load of the target device, and the historical fault data of the target device, respectively; determining the health score of the target device based on the temperature score of the target device, the power consumption score of the target device, the utilization score of the target device, and the fault history score of the target device.

[0007] Optionally, when the health score of the target device is greater than a first threshold, the target device is determined to be in a healthy state; when the health score of the target device is in a target interval, the target device is determined to be in a first type of fault state; when the health score of the target device is less than a second threshold, the target device is determined to be in a second type of fault state, wherein the first threshold is greater than the maximum value of the target interval and the second threshold is less than the minimum value of the target interval.

[0008] Optionally, the method also includes: obtaining historical data, the historical data including at least: device identification, device failure type and failure cause; determining a device triple based on the historical data, the device triple being used to represent the relationship between the device, the device failure type and the failure cause; and constructing the knowledge graph based on the device triple.

[0009] Optionally, determining the cause of the fault based on a pre-constructed knowledge graph includes: using a graph neural network to analyze the operating data of the target device to obtain the cause of the fault of the target device, wherein the graph neural network is trained based on the knowledge graph.

[0010] Optionally, the method further includes: performing feature extraction on each node in the knowledge graph to obtain a feature vector of each node; and using the feature vector of each node to train an initial graph neural network to obtain the graph neural network.

[0011] Optionally, historical fault data and historical operation data are extracted from the multi-dimensional data of the target device; the historical fault data and the historical operation data are preprocessed to obtain time series data; the fault prediction model is trained using the time series data to obtain a trained fault prediction model, and the trained fault prediction model is used to predict the probability of failure of the target device within the prediction period.

[0012] According to another aspect of an embodiment of the present application, a data center equipment maintenance device is also provided, including: an acquisition module for acquiring multi-dimensional data of a target device, wherein the multi-dimensional data includes at least: operating data of the target device; a first determination module for determining a health score of the target device based on the operating data of the target device; a second determination module for determining whether the target device is in a fault state based on the health score of the target device, and when the target device is in a fault state, determining the cause of the fault based on a pre-constructed knowledge graph; a maintenance module for determining a maintenance plan based on the cause of the fault, and the maintenance plan is used to maintain the target device.

[0013] According to another aspect of the embodiment of the present application, a computer device is provided, including: a memory and a processor, wherein the memory is used to store program instructions; the processor is connected to the memory and is used to execute the above-mentioned data center equipment maintenance method.

[0014] According to another aspect of the embodiments of the present application, a computer program product is provided, including computer instructions, which implement the above-mentioned data center equipment maintenance method when executed by a processor.

[0015] In an embodiment of the present application, multi-dimensional data of a target device is collected, wherein the multi-dimensional data includes at least: operating data of the target device; determining a health score of the target device based on the operating data of the target device; determining whether the target device is in a faulty state based on the health score of the target device, and determining the cause of the fault based on a pre-constructed knowledge graph when the target device is in a faulty state; determining a maintenance plan based on the cause of the fault, the maintenance plan is used to maintain the target device, and by evaluating the health score of the device based on the multi-dimensional data, and when the health score indicates a device fault, quickly identifying the cause of the fault based on the knowledge graph and determining the operation and maintenance plan, the purpose of avoiding manual passive determination of the operation and maintenance plan is achieved, thereby achieving the technical effect of improving operation and maintenance efficiency, and further solving the technical problem of low operation and maintenance efficiency of the operation and maintenance mode of the computer room equipment in the related art. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0017] Figure 1 This is a hardware structure block diagram of a computer terminal for implementing a data center equipment maintenance method according to an embodiment of the present application;

[0018] Figure 2 is a flow chart of a data center equipment maintenance method according to an embodiment of the present application;

[0019] Figure 3 is a flow chart of another data center equipment maintenance method according to an embodiment of the present application;

[0020] Figure 4 is a sequence diagram of another data center equipment maintenance method according to an embodiment of the present application;

[0021] Figure 5 This is a structural diagram of a data center equipment maintenance method according to an embodiment of the present application. DETAILED DESCRIPTION

[0022] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0023] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0024] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are explained as follows:

[0025] Multi-source Data Fusion: This refers to obtaining information from multiple different data sources (such as sensor data, equipment logs, environmental parameters, etc.) and integrating it into a more consistent and information-rich data set through data preprocessing, feature extraction, fusion algorithms, etc., to improve the accuracy and robustness of equipment health status analysis.

[0026] Time-series anomaly detection: This refers to identifying anomalous data points or trends in time-varying data streams through statistical analysis, machine learning, or deep learning methods. For example, LSTM (Long Short-Term Memory) or ARIMA (Autoregressive Integrated Moving Average) models can be used to predict normal operating patterns and calculate error residuals to identify anomalies.

[0027] Equipment Health Score (EHS): A quantitative health scoring system based on factors such as equipment operating status, historical failure rates, and environmental impacts. This scoring system typically uses a multi-metric weighted calculation or machine learning approach to help data center operators intuitively understand equipment health and plan maintenance plans in advance.

[0028] Knowledge Graph: This structured association of information such as equipment, fault types, and historical maintenance records allows for fault pattern identification and root cause analysis. In computer room equipment management, knowledge graphs can help infer potential causes of equipment failures and provide corresponding operational and maintenance recommendations.

[0029] The information collected in the embodiments of the present application is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with the relevant laws, regulations and standards of the relevant regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or reject the automated decision results; if the user chooses to reject, the expert decision-making process will be entered.

[0030] In order to solve the problems existing in the related art, the embodiment of the present application provides a data center equipment maintenance method, which can be run on Figure 1 In the computer terminal shown, the computer terminal is explained below.

[0031] The data center equipment maintenance method embodiment provided in the embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 The hardware structure block diagram of a computer terminal for implementing a data center equipment maintenance method is shown. Figure 1As shown, the computer terminal 10 may include one or more (illustrated by 102a, 102b, ..., 102n in the figure) processors (the processor may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions connected via a wired and / or wireless network. In addition, it may also include: a display, a keyboard, a cursor control device, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, and a BUS bus. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0032] It should be noted that the one or more processors and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10. As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0033] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the data center equipment maintenance method in the embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned data center equipment maintenance method. The memory 104 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0034] The transmission module 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission module 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission module 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0035] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 .

[0036] It should be noted that, in some optional embodiments, the above Figure 1 The computer terminal shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. Figure 1 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the computer terminal described above.

[0037] In the above-mentioned operating environment, an embodiment of the present application provides an embodiment of a data center equipment maintenance method. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0038] Figure 2 is a flow chart of a data center equipment maintenance method according to an embodiment of the present application, such as Figure 2 As shown, the method includes the following steps:

[0039] Step S202: collecting multi-dimensional data of the target device, wherein the multi-dimensional data at least includes: operating data of the target device;

[0040] Step S204: determining a health score of the target device based on the operating data of the target device;

[0041] Step S206: determining whether the target device is in a fault state based on the health score of the target device (device health score), and if the target device is in a fault state, determining the cause of the fault based on the pre-built knowledge graph;

[0042] Step S208: determining a maintenance plan according to the fault cause, wherein the maintenance plan is used to maintain the target device.

[0043] It's important to note that traditional O&M approaches can lead to the following problems: Passive response mode: Most current data center management systems use a monitoring strategy based on fixed thresholds, triggering alarms only when certain key parameters (such as temperature, current, and voltage) exceed preset ranges. However, this approach fails to predict potential failures in advance and often requires only post-problem resolution, which can easily lead to sudden equipment downtime and impact business continuity. Low data utilization: Although data center computer rooms are equipped with a large number of monitoring devices, such as temperature and humidity sensors, current and voltage monitoring modules, and server logging systems, this data is typically stored and managed independently, lacking unified data fusion and analysis capabilities. This results in a large number of potential failure characteristics not being fully explored and utilized. Inability to accurately predict equipment health: Existing computer room monitoring systems lack intelligent analysis and predictive capabilities, making it difficult to quantify equipment health. For example, the operating life of equipment is affected by multiple factors (such as load pressure, ambient temperature, and historical failure rates), but existing methods often rely solely on a single monitoring parameter for analysis, failing to provide a comprehensive and accurate health assessment. Disconnected energy efficiency management and failure prediction: Energy efficiency management measures, such as cooling systems and power scheduling in data centers, often operate independently and fail to integrate with equipment health. For example, when certain devices are overloaded and their temperatures rise, the cooling system fails to be dynamically adjusted in a timely manner, causing the equipment to overheat and accelerate aging or failure.

[0044] Through the above steps S202 to S208, multi-dimensional data of the target device is collected, wherein the multi-dimensional data includes at least: the operating data of the target device; determining the health score of the target device based on the operating data of the target device; determining whether the target device is in a faulty state based on the health score of the target device, and if the target device is in a faulty state, determining the cause of the fault based on a pre-built knowledge graph; determining a maintenance plan based on the fault cause, the maintenance plan is used to maintain the target device. By evaluating the health score of the device based on the multi-dimensional data, and if the health score indicates a device fault, quickly identifying the cause of the fault based on the knowledge graph and determining the operation and maintenance plan, the purpose of avoiding manual passive determination of the operation and maintenance plan is achieved, thereby achieving the technical effect of improving operation and maintenance efficiency, and further solving the technical problem of low operation and maintenance efficiency of the computer room equipment operation and maintenance method in the related art. The following is a detailed description.

[0045] Figure 3 Another data center equipment maintenance method is shown. Figure 3As shown, the process includes: Step 1, multi-source data collection and fusion: Collecting equipment room operating parameters, environmental monitoring data, fault logs, etc. Data preprocessing methods such as standardization, noise reduction, and feature extraction are used to improve data quality and obtain multi-dimensional data. Step 2, equipment failure prediction based on time series analysis: Using machine learning models (such as LSTM and GRU) to model equipment status and predict failures. Calculating the probability of equipment anomalies and providing early warning of possible failure risks. Step 3, equipment health scoring and intelligent operation and maintenance strategy: Constructing a multi-factor health scoring system to evaluate equipment status. Dynamically adjusting inspection and maintenance plans based on health scores to improve maintenance efficiency. Step 4, root cause analysis based on knowledge graphs: Combining historical failure data to construct an equipment failure knowledge graph and automatically identify failure modes. Through reasoning analysis, the root cause of the failure and optimization suggestions are provided. Step 5, adaptive energy efficiency optimization management: Optimizing cooling system power and task allocation based on equipment load, health status, and environmental factors.

[0046] In the technical solution provided in step S204 of the above-mentioned data center equipment maintenance method, the health score of the target device can be determined in the following manner: determining the health score of the target device based on the operating data of the target device, including: extracting the current temperature of the target device, the power consumption of the target device, and the load of the target device from the operating data of the target device; obtaining the historical fault data of the target device from the multi-dimensional data; determining the temperature score of the target device, the power consumption score of the target device, the utilization score of the target device, and the fault history score of the target device based on the current temperature of the target device, the power consumption of the target device, the load of the target device, and the historical fault data of the target device; determining the health score of the target device based on the temperature score of the target device, the power consumption score of the target device, the utilization score of the target device, and the fault history score of the target device. By obtaining the operating data of the target device in real time, the system can determine whether the target device is in a healthy state based on the historical fault data of the target device, thereby avoiding the lag of fault warning and reducing the risk of equipment downtime.

[0047] Among them, when the health score of the target device is greater than a first threshold, the target device is determined to be in a healthy state; when the health score of the target device is in a target interval, the target device is determined to be in a first type of fault state; when the health score of the target device is less than a second threshold, the target device is determined to be in a second type of fault state, wherein the first threshold is greater than the maximum value of the target interval, and the second threshold is less than the minimum value of the target interval.

[0048] It is understandable that multi-dimensional data can be obtained through the following methods: Step 1: Multi-source data acquisition: First, deploy a variety of sensors and monitoring systems in the data center computer room to collect the following key data in real time: Equipment operation data (a direct reflection of equipment health): Computing resource usage: CPU occupancy C(t), memory utilization M(t), disk I / O speed D(t). Power consumption indicators: Equipment power consumption P(t), current I(t), voltage V(t). Heat change: Equipment surface temperature T_d(t). Environmental data (external factors affecting equipment health): Temperature and humidity: ambient temperature T_e(t), relative humidity H(t). Cooling system: cooling power C_p(t), wind speed F_s(t). Historical fault data (used for model training and fault pattern recognition): Fault occurrence time t_f: the specific time point of the equipment failure. Fault type F_k: hardware failure, software failure, power supply abnormality, etc. Maintenance record R(t): equipment maintenance time, maintenance method, and replacement parts information. Step 2: Data preprocessing and cleaning. Since the collected data comes from different devices and sensors, their formats, units, and sampling frequencies may differ. They must be standardized to ensure data comparability and consistency.

[0049] Specifically, data standardization: the normalization method is used to convert all numerical data into the range of 0 and 1; the sliding average filtering method is used to remove noise and improve data smoothness; and the triple standard deviation method is used to eliminate abnormal data points.

[0050] Since different data sources have different sampling frequencies, for example, server CPU load data is updated once per second, while computer room environment data may be collected once per minute, time alignment and feature extraction are required to ensure data consistency.

[0051] Linear interpolation can be used to align all data to the same timestamp. In addition, feature extraction can be performed on the collected data, for example, obtaining the power consumption change rate of each device, the temperature fluctuation trend of each device, and the load balancing indicator of each server (the standard deviation of the CPU load of different servers).

[0052] It can be understood that in the embodiment of the present application, by integrating multi-dimensional information such as equipment operation data, computer room environment data, historical fault data, and combining data preprocessing, feature extraction and time alignment technology, high-quality data input is provided for subsequent equipment fault prediction and health management. This method can improve the integrity, accuracy and time consistency of the data, so that the intelligent analysis model can learn the equipment operation rules more accurately, thereby improving the accuracy of fault prediction and operation and maintenance decisions. The multi-dimensional data obtained in the embodiment of the present application has comprehensive data fusion: comprehensive equipment operation, environmental factors and historical failures are integrated to improve the accuracy of the prediction model. Data quality improvement: The credibility of the data is improved through noise reduction, standardization, anomaly detection and other methods. Time consistency optimization: The interpolation method is used to align the data to ensure time consistency during model training and inference. Accurate feature mining: extract core indicators (such as temperature fluctuation rate, load balancing) to improve the effect of intelligent analysis.

[0053] In the management of equipment in data center computer rooms, traditional operation and maintenance strategies often rely on regular maintenance or post-failure response, and lack a dynamic adjustment mechanism based on the real-time status of equipment. In order to improve the reliability and operation and maintenance efficiency of computer room equipment, Figure 4 A sequence diagram of another data center equipment maintenance method is shown, as shown in FIG. Figure 4 As shown, it includes: Step 1: Calculation of equipment health score: obtained through weighted calculation, each weight coefficient represents the degree of influence of the factor on the health status of the equipment. The calculation method is shown in the following formula:

[0054] HS=w1·S 温度 +w2·S 功耗 +w3·S 使用率 +w4·S 故障历史

[0055] Where: S 温度 : The current temperature score of the device, normalized based on the set ideal temperature range. 功耗 : Equipment power consumption score, reflecting the energy efficiency status of the equipment. 使用率 : CPU, storage or network usage score, indicating the device load. 故障历史 : Scoring based on historical fault records. If the device has experienced a fault in the past, points will be deducted accordingly.

[0056] w1+w2+w3+w4=1 All scores are normalized to [0,100], so that the final health score HS is also within this range.

[0057] For example: The current measurement values ​​of a device are as follows:

[0058] The temperature is 50°C (standard range is 20°C–70°C), corresponding to a score of 85.

[0059] Power consumption is 200W (standard range 150W–300W), corresponding to a score of 80.

[0060] The CPU usage is 75% (the standard range is 0%–100%), corresponding to a score of 75.

[0061] 1 failure occurred in the past 6 months, corresponding to a score of 90.

[0062] Assuming that the weight coefficients are w1 = 0.3, w2 = 0.2, w3 = 0.3, and w4 = 0.2, the health score of the device is calculated as follows:

[0063] HS=(0.3×85)+(0.2×80)+(0.3×75)+(0.2×90)=25.5+16+22.5+18=82

[0064] The HS value of this device is 82, which is in a healthy state and does not require immediate maintenance.

[0065] Step 2: By calculating the device health score (HS), the operating status of the device can be judged, and corresponding operation and maintenance strategies can be formulated based on different score ranges.

[0066] For example, if the health score HS is greater than 80, the device is in good condition, operating stably, and does not require additional maintenance. In this case, the system will not issue a maintenance reminder, and the device can continue to operate at normal load.

[0067] When the health score (HS) is between 60 and 80, the device enters a subhealthy state. While it can continue to operate, some of its performance indicators have deviated from the optimal operating range. In this case, the system recommends preventive maintenance, such as cleaning the cooling system, optimizing load distribution, or making minor adjustments, to prevent further deterioration of the device's condition.

[0068] When the health score (HS) falls below 60, the device may be about to fail, posing a significant operational risk. At this point, the system triggers an emergency maintenance task, advising operations personnel to take immediate action, such as adjusting device operating parameters, inspecting the cooling system, or replacing potentially defective hardware components, to prevent device downtime and overall business stability.

[0069] In addition, the intelligent feature of this strategy lies in that the system can combine historical data and equipment operation trends to predict possible anomalies in the future and conduct operation and maintenance intervention in advance, thereby achieving a shift from "passive repair" to "active optimization" and improving the efficiency and reliability of computer room equipment management.

[0070] For example, if the calculated HS of device A is 85, it is in a healthy state, and the system will not trigger a maintenance plan. If the calculated HS of device B is 72, it is in the sub-healthy range, and the system will prompt the operation and maintenance personnel to perform non-emergency maintenance (such as optimizing the cooling system and adjusting the load). If the calculated HS of device C is 55, which is below the set threshold of 60, the system will trigger an emergency maintenance task and recommend replacing or adjusting the device's operating status.

[0071] The method proposed in this application embodiment dynamically adjusts maintenance plans based on the latest operating status of the equipment, improving operation and maintenance efficiency. It also uses multidimensional data to calculate health scores, reducing false positives and missed positives, and improving equipment management accuracy. This effectively reduces the risk of equipment downtime and improves the overall operational efficiency of the data center.

[0072] In some examples of the present application, a knowledge graph can be determined by: obtaining historical data, wherein the historical data includes at least: a device identifier, a device fault type, and a fault cause; determining a device triple based on the historical data, wherein the device triple is used to represent the relationship between the device, the device fault type, and the fault cause; and constructing the knowledge graph based on the device triple. Determining the fault cause based on the pre-constructed knowledge graph includes: analyzing the operating data of the target device using a graph neural network to obtain the fault cause of the target device, wherein the graph neural network is trained based on the knowledge graph.

[0073] Specifically, a data center knowledge graph is constructed to form a triple relationship of equipment-failure mode-environmental factors; graph embedding calculations are performed using graph neural networks (GNNs) to extract potential correlation patterns of equipment failures; real-time equipment status data is input and combined with the knowledge graph to infer the possible root causes of the failures; intelligent diagnosis results are output and specific operation and maintenance suggestions are provided, such as load optimization, cooling adjustment, or component replacement.

[0074] The specific steps to build a knowledge graph are as follows:

[0075] Set three types of entities: equipment (equipment identification), fault type (alarm category), and fault cause (temperature, humidity, power load, etc.).

[0076] Form a triple relationship (device, fault type, fault cause), for example:

[0077] (Server A, abnormal operation, high temperature failure), (Air conditioner B, insufficient heat dissipation, increased temperature in the computer room), the above content can be stored in RDF (Resource Description Framework) or a graph database (such as Neo4j) to obtain a knowledge graph.

[0078] In some embodiments of the present application, the graph neural network can be trained in the following manner: extracting features from each node in the knowledge graph to obtain a feature vector for each node; and training the initial graph neural network using the feature vector of each node to obtain the graph neural network.

[0079] Get the initial feature vector of each node Indicates basic information about the device, environment, or faulty node, such as:

[0080]

[0081] Calculate the feature update of each node and integrate the information of its neighboring nodes. The formula is as follows:

[0082]

[0083] in: represents the device features after the kth iteration. N(v) is the set of neighbor nodes of node v (i.e., the devices, fault types, or fault causes associated with it). W represents the trainable weight matrix, b represents the bias value, and σ represents the activation function.

[0084] In actual application scenarios, graph neural networks can also determine the risk score of each device:

[0085]

[0086] Where L is the number of layers of GNN iteration and f(·) is the classification function (such as Softmax).

[0087] It should be noted that the risk score of a device is used to indicate the probability of device failure.

[0088] Based on a trained graph neural network and knowledge graph, the system automatically infers the possible root causes of device failures by inputting device operating data. The system obtains the current device status (such as temperature, power consumption, and load), matches the knowledge graph, and searches for similar historical failure patterns. It calculates potential failure paths and assesses the likelihood of failure. It then outputs the possible root causes and recommends remediation measures.

[0089] For example, if the temperature of server A is abnormally high, the GNN calculation results show:

[0090] Device A may experience high temperature due to a fan failure (70% probability) or abnormal heat dissipation in the equipment room (30% probability).

[0091] Combined with the knowledge graph, the system recommends checking the fan components first and suggests reducing the server load.

[0092] Based on the root cause analysis results, the system automatically generates targeted operation and maintenance suggestions to help operation and maintenance personnel quickly resolve the problem: For example:

[0093] High temperature alarm: Possible causes: fan failure, poor heat dissipation.

[0094] Recommended measures: Check the fan operating status and replace the fan or optimize the air conditioning direction if necessary.

[0095] CPU overload:

[0096] Possible causes: abnormal task scheduling and unbalanced load.

[0097] Recommended measures: Optimize the load balancing strategy and migrate computing tasks to low-load servers.

[0098] Server power failure:

[0099] Possible causes: abnormal UPS power supply, battery aging.

[0100] Recommended measures: Check the UPS power supply, replace aging batteries or repair circuit faults.

[0101] Network fluctuations:

[0102] Possible causes: Switch port abnormality, physical connection failure.

[0103] Recommended Action: Redistribute port traffic, check and repair physical connection issues.

[0104] The system combines knowledge graphs with real-time data to provide accurate operation and maintenance recommendations, improve data center operation and maintenance efficiency, and reduce the impact of equipment failures.

[0105] In other embodiments of the present application, the process of troubleshooting a server high temperature fault is as follows:

[0106] Symptom: Server A has a high temperature alarm, and the temperature exceeds 85°C.

[0107] Knowledge graph query: Query the historical fault records associated with device A and find that there have been two fan failures in the past three months.

[0108] Combined with the current temperature data, the probability of fan failure is calculated to be 75%.

[0109] Possible root causes: Fan failure (priority 1) and abnormal heat dissipation in the equipment room (priority 2).

[0110] Operation and maintenance suggestion: Prioritize checking the fan assembly and replace it if any problem is found; if the fan is normal, optimize the air direction of the computer room air conditioner.

[0111] In the embodiments of this application, by constructing a knowledge graph and using GNN for root cause analysis, intelligent diagnosis and prediction of data center equipment failures are achieved. Compared with traditional methods, this approach offers the following advantages: Accurate fault diagnosis: Based on historical data analysis, it reduces false positives and improves fault location accuracy. Real-time performance: By combining device status data, it can quickly infer possible fault causes, reducing troubleshooting time. Intelligent operation and maintenance recommendations: The system provides specific maintenance recommendations, optimizes equipment operation and maintenance efficiency, and reduces manual intervention costs.

[0112] In other embodiments of the present application, an intelligent operation and maintenance recommendation system based on multi-source data analysis and machine learning is also provided, which can combine real-time monitoring data, historical fault records and knowledge graphs to automatically generate targeted fault diagnosis and processing suggestions, improve operation and maintenance efficiency, and reduce the risk of downtime.

[0113] The specific steps include the following:

[0114] 1) Data collection and preprocessing: Collect operating data of key equipment such as servers, network equipment, and cooling systems, and perform data cleaning and standardization.

[0115] 2) Anomaly Detection: Statistical methods and machine learning models are used to detect equipment anomalies and determine whether further analysis is needed.

[0116] 3) Root cause analysis: Use knowledge graphs and causal reasoning models to determine the possible root causes of equipment anomalies.

[0117] 4) Intelligent O&M suggestion generation: Based on the analysis results and combined with the optimization model, targeted troubleshooting suggestions are provided to O&M personnel.

[0118] 5) Feedback optimization: After the operation and maintenance personnel implement the suggestions, the system records the effect data and continuously optimizes the recommended strategy.

[0119] The specific implementation process is as follows:

[0120] The system collects operational data from data center equipment through sensors, log analysis, and SNMP (Simple Network Management Protocol). This data includes temperature, current, voltage, CPU / memory load, and more. Raw data may contain noise and missing values, so it requires data cleaning, deduplication, and standardization for subsequent analysis.

[0121] The system uses statistical methods (such as the mean-standard deviation method) and machine learning models (such as LSTM and IsolationForest) to detect abnormal data points. If the device's operating status exceeds a set threshold or the model's predicted health score falls into the warning range, the anomaly detection mechanism is triggered.

[0122] Equipment health score calculation formula:

[0123] H=w1T+w2C+w3N

[0124] Where: H is the health score, ranging from [0 to 1]. A lower score indicates a worse device status. T is the temperature fluctuation coefficient, indicating whether the temperature is abnormal. C is the CPU / memory load indicator, reflecting computing resource usage. N is the network traffic fluctuation, used to determine whether there is a network anomaly. w1, w2, and w3 are weighting coefficients, whose weights are determined based on historical data.

[0125] When a device anomaly is detected, the system uses knowledge graphs and causal reasoning models to analyze the possible causes of the failure. For example, if a server's CPU load is too high and it is found that the server has previously experienced performance degradation due to poor heat dissipation, there may be a cooling system anomaly.

[0126] The system uses Bayesian networks for causal reasoning to calculate the probabilities of different fault causes:

[0127]

[0128] Where: P(R i |E) is the root cause of the equipment failure R after the abnormal event E occurs. i The probability of P(E|R i ) is the probability of an abnormal event when a specific fault occurs; P(R i ) is the prior probability of the fault occurring; P(E) is the comprehensive probability of all abnormal events.

[0129] The system ranks the possible root causes based on the calculated failure probability and selects the most likely root cause as the preliminary diagnosis result.

[0130] Based on the root cause analysis results, the system combines optimization strategies to generate targeted operation and maintenance suggestions, such as:

[0131] High temperature alarm: It is recommended to check the cooling system of the equipment room and adjust the air conditioning power or wind direction.

[0132] CPU overload: It is recommended to optimize task scheduling and migrate computing tasks to low-load servers.

[0133] Server power outage: It is recommended to check the UPS power supply, replace aging batteries or repair circuit faults.

[0134] Network fluctuations: It is recommended to redistribute port traffic and check and repair physical connection problems.

[0135] The priority of the operation and maintenance recommendation is determined by the scope of the fault and the urgency. The calculation formula is as follows:

[0136] S=α1I+α2U+α3D

[0137] Where: S is the priority score of the operation and maintenance recommendation; I is the fault impact scope, which measures the number of affected devices and services; U is the fault urgency, which reflects the impact of the problem on business continuity; D is the historical occurrence frequency, indicating whether this type of fault occurs frequently; α1, α2, and α3 are weight parameters determined by data training.

[0138] After maintenance personnel implement maintenance recommendations, the system records the actual results and uses them to optimize the recommended strategy. For example, if a recommended approach to a particular fault fails multiple times, the system adjusts the causal inference model and optimizes the root cause analysis process to improve the accuracy of future fault diagnosis.

[0139] The system uses reinforcement learning to optimize operation and maintenance strategies, with the goal of minimizing fault handling time:

[0140]

[0141] Where: T f is the total fault handling time; C t A is the recommended operation and maintenance cost for the tth time; t The execution rate of the suggestions accepted by the operation and maintenance personnel (0 or 1).

[0142] Through continuous optimization, the system can improve the accuracy and practicality of intelligent operation and maintenance recommendations, making data center operation and maintenance more efficient and intelligent.

[0143] In some embodiments of the present application, historical fault data and historical operation data are extracted from the multi-dimensional data of the target device; the historical fault data and the historical operation data are preprocessed to obtain time series data; and a fault prediction model is trained using the time series data to obtain a trained fault prediction model, and the trained fault prediction model is used to predict the probability of failure of the target device within a prediction period.

[0144] The specific training process for the fault prediction model is as follows: Step 1, Data Collection and Preprocessing: Multi-dimensional device data, such as CPU load, temperature, power supply power, and room temperature and humidity, is collected. The data is normalized and denoised to eliminate abnormal data, reduce data volatility, and improve data quality. A sliding window method is used to divide the continuous time series data into multiple subsequences to extract features from each subsequence. The sliding window size is adjusted based on device characteristics and data type.

[0145] The extracted features include:

[0146] Temperature change rate: reflects whether the heat dissipation of the equipment is normal. The calculation formula is:

[0147]

[0148] Where T(t) represents the current temperature of the device, and T(t-1) represents the temperature at the previous moment.

[0149] Power consumption fluctuation rate: Calculate the rate of change of power consumption. The formula is:

[0150]

[0151] Where P(t) is the current power of the device, and P(t-1) is the power at the previous moment.

[0152] Abnormal log frequency: Counts the number of abnormal logs generated by the device within a certain time window, which is used to reflect whether the device has a fault or the risk of impending fault.

[0153] Step 2: Model training.

[0154] After feature extraction is complete, these features are fed into a machine learning model for training. Long Short-Term Memory (LSTM) networks are used to model and predict time series data. LSTMs are particularly well-suited for processing and predicting data with temporal relationships. LSTM networks consist of multiple time steps, with the output of each time step dependent on the current input and the output of the previous step. This temporal dependency is captured through a feedback mechanism. The network input is the extracted feature vector, and the output is the health status or failure probability of the device within a certain timeframe.

[0155] Loss function calculation:

[0156] In order to optimize the model parameters, the mean square error (MSE) is used as the loss function to measure the error between the model prediction and the actual fault state:

[0157]

[0158] Among them, y i is the actual fault status, is the fault state predicted by the model, and n is the number of samples.

[0159] By minimizing the loss function, an LSTM model that can predict the probability of equipment failure is trained.

[0160] Step 3: Prediction results and fault warning.

[0161] Once the model is trained, it can be used to predict the health status of the equipment in real time. Based on the prediction results, the system will assess the risk of equipment failure and provide early warning when necessary.

[0162] Failure probability calculation:

[0163] For each moment, the LSTM model predicts the future health status of the device based on historical data and real-time input data. The output is the failure probability P fail (t), represents the probability that the equipment will fail in the future.

[0164] Fault warning:

[0165] Set a threshold θ, when the predicted failure probability P fail (t) When the threshold is exceeded, the system will trigger a fault warning and notify the operation and maintenance personnel to conduct inspection and maintenance.

[0166] Dynamically adjust thresholds based on historical failure data to improve prediction accuracy.

[0167] The equipment failure prediction method based on time series analysis provided in the embodiment of the present application can identify the failure risk of equipment in advance, significantly reduce the risk of computer room downtime and improve the health management efficiency of equipment through multi-source data collection, sliding window feature extraction, LSTM modeling and fault warning. Compared with traditional methods, it has the following advantages: Multi-dimensional data fusion: By integrating equipment operation data and environmental factors, the accuracy of fault prediction is improved. Time series relationship modeling: LSTM is used to capture the dynamic characteristics of equipment status over time to enhance prediction capabilities. Early warning mechanism: Through fault probability prediction and real-time warning, it ensures that operation and maintenance personnel respond in a timely manner and reduce the occurrence rate of failures.

[0168] In other embodiments of the present application, a computer room fault prediction and simulation model based on digital twins can also be constructed. For example, on the basis of the current method, digital twin technology is introduced to construct a virtual mapping model of the computer room equipment. Through the real-time collection of equipment operation data, combined with historical failure modes, possible failure situations can be simulated in a virtual environment, and potential risks can be predicted in advance. This solution can optimize operation and maintenance strategies and improve fault prevention capabilities through simulation analysis before the equipment actually fails. In order to reduce the data center's dependence on cloud computing resources, edge computing nodes can be deployed on key equipment to achieve local data processing and intelligent alarms. Edge devices can use lightweight machine learning models to complete data analysis locally, quickly detect anomalies and make preliminary responses, while only uploading important events to the central management system, thereby reducing delays and optimizing network bandwidth utilization.

[0169] Building on the data center equipment maintenance method provided in the embodiments of this application, AI-powered O&M technology is introduced to achieve automated O&M decision-making and self-healing capabilities. When the system detects an anomaly, it not only provides O&M recommendations but also, in conjunction with the policy engine, automatically executes some O&M tasks, such as dynamically adjusting resource allocation, restarting abnormal equipment, and switching to redundant equipment. This solution reduces manual intervention, improves O&M efficiency, and minimizes losses caused by human delays or errors.

[0170] Figure 5 A data center equipment maintenance device is shown, the device comprising:

[0171] The acquisition module 50 is configured to acquire multi-dimensional data of the target device, wherein the multi-dimensional data includes at least: operating data of the target device;

[0172] A first determining module 52 is configured to determine a health score of the target device based on the operating data of the target device;

[0173] A second determination module 54 is configured to determine whether the target device is in a fault state based on the health score of the target device, and if the target device is in a fault state, determine the cause of the fault based on a pre-built knowledge graph;

[0174] The maintenance module 56 is configured to determine a maintenance plan based on the fault cause, wherein the maintenance plan is used to maintain the target device.

[0175] The method adopts the method of collecting multi-dimensional data of the target device, wherein the multi-dimensional data includes at least: the operating data of the target device; determining the health score of the target device based on the operating data of the target device; determining whether the target device is in a fault state based on the health score of the target device, and when the target device is in a fault state, determining the cause of the fault based on a pre-built knowledge graph; determining a maintenance plan based on the cause of the fault, and the maintenance plan is used to maintain the target device. By collecting the multi-dimensional data of the target device, the technical effect of detecting whether the device is faulty in real time based on the multi-dimensional data and determining the corresponding maintenance plan in the case of a fault is achieved, thereby solving the technical problems in the related technology that the operation and maintenance mode of the computer room equipment cannot monitor the equipment status changes in real time, resulting in delayed warning, inaccurate prediction, data isolation, lack of health assessment and disconnection of energy efficiency management.

[0176] The maintenance module 56 includes a determination submodule, a test submodule, an identification submodule and a training submodule, wherein the determination submodule is used to determine the health score of the target device based on the operating data of the target device, including: extracting the current temperature of the target device, the power consumption of the target device, and the load of the target device from the operating data of the target device; obtaining the historical fault data of the target device from the multi-dimensional data; determining the temperature score of the target device, the power consumption score of the target device, the utilization score of the target device, and the fault history score of the target device based on the current temperature of the target device, the power consumption of the target device, the load of the target device, and the historical fault data of the target device respectively; and determining the health score of the target device based on the temperature score of the target device, the power consumption score of the target device, the utilization score of the target device, and the fault history score of the target device.

[0177] The determination submodule is further used to determine that the target device is in a healthy state when the health score of the target device is greater than a first threshold; to determine that the target device is in a first type of fault state when the health score of the target device is in a target interval; and to determine that the target device is in a second type of fault state when the health score of the target device is less than a second threshold, wherein the first threshold is greater than the maximum value of the target interval and the second threshold is less than the minimum value of the target interval.

[0178] The determination submodule includes a construction unit for obtaining historical data, wherein the historical data includes at least: device identification, device failure type and failure cause; determining a device triple based on the historical data, wherein the device triple is used to represent the relationship between the device, the device failure type and the failure cause; and constructing the knowledge graph based on the device triple.

[0179] The construction unit includes: a determination subunit and a training subunit. The determination subunit is used to determine the cause of the fault based on a pre-constructed knowledge graph, including: using a graph neural network to analyze the operating data of the target device to obtain the cause of the fault of the target device, wherein the graph neural network is trained based on the knowledge graph.

[0180] The training subunit is used to extract features from each node in the knowledge graph to obtain a feature vector for each node; and to train the initial graph neural network using the feature vector of each node to obtain the graph neural network.

[0181] The training submodule is used to extract historical fault data and historical operation data from the multi-dimensional data of the target device; preprocess the historical fault data and the historical operation data to obtain time series data; use the time series data to train the fault prediction model to obtain a trained fault prediction model, and the trained fault prediction model is used to predict the probability of failure of the target device within the prediction period.

[0182] It should be noted that Figure 5 The data center equipment maintenance device shown is used to perform Figure 2 The data center equipment maintenance method shown in the figure, therefore the relevant explanations in the above data center equipment maintenance method are also applicable to the data center equipment maintenance device, and will not be repeated here.

[0183] An embodiment of the present application also provides a computer device, including: a memory and a processor, wherein the memory is used to store program instructions; the processor is connected to the memory and is used to execute the above-mentioned data center equipment maintenance method.

[0184] An embodiment of the present application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the data center equipment maintenance method in the present application.

[0185] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0186] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0187] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0188] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0189] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0190] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0191] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A data center equipment maintenance method, characterized in that: include: Collecting multi-dimensional data of the target device, wherein the multi-dimensional data at least includes: operating data of the target device; Determining a health score of the target device based on the operating data of the target device; Determining whether the target device is in a fault state based on the health score of the target device, and if the target device is in a fault state, determining the cause of the fault based on a pre-built knowledge graph; A maintenance plan is determined according to the fault cause, and the maintenance plan is used to maintain the target device.

2. The method according to claim 1, characterized in that Determine a health score of the target device based on the operating data of the target device, including: extracting the current temperature of the target device, the power consumption of the target device, and the load of the target device from the operating data of the target device; Acquire historical fault data of the target device from the multi-dimensional data; determining a temperature score of the target device, a power consumption score of the target device, a usage score of the target device, and a failure history score of the target device according to the current temperature of the target device, the power consumption of the target device, the load of the target device, and the historical failure data of the target device, respectively; The health score of the target device is determined according to the temperature score of the target device, the power consumption score of the target device, the usage score of the target device, and the fault history score of the target device.

3. The method according to claim 2, characterized in that The method further comprises: When the health score of the target device is greater than a first threshold, determining that the target device is in a healthy state; When the health score of the target device is within a target range, determining that the target device is in a first type of failure state; When the health score of the target device is less than a second threshold, it is determined that the target device is in a second type of failure state, wherein the first threshold is greater than the maximum value of the target interval and the second threshold is less than the minimum value of the target interval.

4. The method according to claim 1, wherein The method further comprises: Acquiring historical data, wherein the historical data includes at least: device identification, device failure type, and failure cause; Determine a device triplet based on the historical data, wherein the device triplet is used to represent a relationship between a device, a device fault type, and a fault cause; The knowledge graph is constructed based on the device triples.

5. The method according to claim 4, characterized in that Determine the cause of the failure based on the pre-built knowledge graph, including: A graph neural network is used to analyze the operating data of the target device to obtain the cause of the failure of the target device, wherein the graph neural network is trained based on the knowledge graph.

6. The method according to claim 5, characterized in that The method further comprises: Perform feature extraction on each node in the knowledge graph to obtain a feature vector for each node; The initial graph neural network is trained using the feature vector of each node to obtain the graph neural network.

7. The method according to claim 1, characterized in that The method further comprises: Extracting historical fault data and historical operation data from the multi-dimensional data of the target device; Preprocessing the historical fault data and the historical operation data to obtain time series data; The time series data is used to train a fault prediction model to obtain a trained fault prediction model, and the trained fault prediction model is used to predict the probability of failure of the target device within a prediction period.

8. A data center equipment maintenance device, characterized in that: include: A collection module, configured to collect multi-dimensional data of a target device, wherein the multi-dimensional data includes at least: operating data of the target device; A first determining module is configured to determine a health score of the target device based on the operating data of the target device; a second determination module, configured to determine whether the target device is in a fault state according to the health score of the target device, and, if the target device is in a fault state, determine the cause of the fault according to a pre-built knowledge graph; A maintenance module is used to determine a maintenance plan according to the fault cause, and the maintenance plan is used to maintain the target device.

9. A computer device, characterized in that: include: A memory and a processor, wherein the memory is used to store program instructions; The processor is connected to the memory and is used to execute the data center equipment maintenance method according to any one of claims 1 to 7.

10. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the data center equipment maintenance method according to any one of claims 1 to 7 is implemented.