Predictive maintenance method and device for modular data machine room and server
By dynamically adjusting the equipment combination of modular data centers using graph neural networks and multi-objective genetic algorithms, and combining long short-term memory networks and gradient boosting tree models for predictive maintenance, the problems of low prediction accuracy and high operation and maintenance costs in existing technologies are solved, thereby improving the reliability and energy efficiency of data center operation and maintenance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-13
- Publication Date
- 2026-03-31
AI Technical Summary
Existing data center operation and maintenance methods cannot effectively reflect the dynamic combination of equipment and load changes in modular data centers, resulting in low prediction accuracy, high operation and maintenance costs, and poor reliability.
Graph neural networks are used for correlation analysis. Through multi-objective genetic algorithms and constraint screening, the combination relationship between functional modules is dynamically adjusted. Long short-term memory networks and gradient boosting tree models are combined for predictive maintenance to achieve system-level optimization and evaluation.
It significantly improves the reliability and predictive accuracy of data center operation and maintenance, reduces operation and maintenance costs, reduces the risk of unplanned downtime, and optimizes energy consumption.
Smart Images

Figure CN121764510A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of data center operation and maintenance, and in particular to a predictive maintenance method, apparatus and server for modular data centers. Background Technology
[0002] Currently, predictive maintenance in data center operations suffers from low accuracy due to the inability to determine the optimal combination of functional modules within the data center and the lack of relevant predictive maintenance parameters. Existing data center operations primarily rely on manual inspections, threshold alarms, and scheduled maintenance strategies. However, these methods, being scheduled maintenance, are difficult to adjust based on the actual operating status of equipment, easily leading to over-maintenance or delayed maintenance, thus increasing operational costs. Furthermore, existing solutions mainly focus on maintenance of single devices, failing to reflect the coupling relationships between multiple devices and struggling to adapt to the dynamic combination of equipment and frequent load changes in modular data centers. Consequently, false alarms are prone to occur, resulting in low reliability of data center operations. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a predictive maintenance method, device and server for modular data centers, which can significantly improve the reliability of data center operation and maintenance.
[0004] In a first aspect, embodiments of the present invention provide a predictive maintenance method for a modular data center. The method includes: classifying functional modules according to the operation and maintenance requirements of the data center; combining related functional modules into independent closed systems; collecting operational data of each functional module; determining the operational data as prediction parameters; performing correlation analysis on the operational data using a graph neural network to determine key correlation points based on the correlation between each functional module and the closed system; and monitoring the real-time operational status of each functional module in real time. When the real-time operational status meets the triggering conditions for combinatorial optimization, dynamically adjusting the combinatorial relationship between each functional module based on the key correlation points and the real-time operational status using a multi-objective genetic algorithm and constraint screening to obtain a target combinatorial scheme and executing the target combinatorial scheme to perform predictive maintenance on the modular data center.
[0005] In one implementation, the step of performing correlation analysis on the operational data using a graph neural network to determine key correlation points based on the correlation between various functional modules and the closed system includes: performing time-domain feature extraction, frequency-domain feature extraction, and time-frequency-domain feature extraction on the operational data to obtain multi-dimensional feature vectors, and constructing a system correlation graph based on the multi-dimensional feature vectors, functional modules, and the closed system; performing weight identification processing on the system correlation graph, and identifying nodes with edge weights greater than a preset weight threshold as key correlation points between functional modules.
[0006] In one implementation, the step of constructing a system association graph based on multidimensional feature vectors, functional modules, and closed systems includes: performing association relationship analysis on each functional module and closed system through a graph neural network, identifying each functional module and closed system as nodes, calculating the edge weights between nodes based on multidimensional feature vectors, determining the edge weights as connection relationships, and constructing the system association graph.
[0007] In one implementation, the combined optimization triggering condition is a triple activation mechanism, which includes threshold triggering, event triggering, and periodic triggering. After the step of real-time monitoring of the real-time operating status of each functional module, the process includes: when the real-time operating status meets any one of the preset threshold triggering condition, preset event triggering condition, and preset periodic triggering condition, determining that the real-time operating status meets the combined optimization triggering condition, and after recording the triggering type and triggering time, entering the combined optimization process.
[0008] In one implementation, the steps of dynamically adjusting the combination relationship between various functional modules based on key correlation points and real-time operating status through a multi-objective genetic algorithm and constraint screening to obtain a target combination scheme include: performing node aggregation processing on functional modules with strong correlation relationships corresponding to key correlation points to obtain combination units; performing iterative calculation processing on the combination units based on the optimization objective function through a multi-objective genetic algorithm and preset operation and maintenance constraint rules to obtain a set of candidate combination schemes; and screening the set of candidate combination schemes based on preset screening conditions to obtain the target combination scheme, wherein the optimization objective function includes: minimizing energy consumption target, maximizing cooling efficiency target, and balancing load distribution target.
[0009] In one implementation, after the step of executing the target combination scheme to perform predictive maintenance on the modular data center, the method includes: using a predictive model constructed collaboratively by a long short-term memory network model and a gradient boosting tree model to perform multi-index evaluation processing on the operating data within a preset time interval during the predictive maintenance process, obtaining evaluation results, and then evaluating and optimizing the predictive capability based on the evaluation results.
[0010] In one implementation, the step of evaluating and optimizing the predictive capability based on the evaluation results includes: when the evaluation results do not meet the preset evaluation capability standards, performing fault tree diagnosis processing from the data level, model level and hardware level respectively, and generating corresponding optimization suggestions.
[0011] Secondly, embodiments of the present invention also provide a predictive maintenance device for a modular data center. The device includes: a system construction module, which classifies functional modules according to the operation and maintenance requirements of the data center, combines related functional modules into an independent closed system, and collects the operating data of each functional module, determining the operating data as prediction parameters; an association analysis module, which performs association analysis on the operating data through a graph neural network to determine key association points based on the association between each functional module and the closed system, and monitors the real-time operating status of each functional module; and a combination optimization module, which, when the real-time operating status meets the combination optimization triggering conditions, dynamically adjusts the combination relationship between each functional module based on the key association points and the real-time operating status through a multi-objective genetic algorithm and constraint screening to obtain a target combination scheme, and executes the target combination scheme to perform predictive maintenance on the modular data center.
[0012] Thirdly, embodiments of the present invention also provide a server, including a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement any of the methods provided in the first aspect.
[0013] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement any of the methods provided in the first aspect.
[0014] The embodiments of the present invention bring the following beneficial effects: This invention provides a predictive maintenance method, apparatus, and server for modular data centers. The method first categorizes functional modules according to the data center's operational needs, combining related functional modules into independent closed systems. It then collects operational data from each functional module, using this data as prediction parameters. A graph neural network is then used to analyze the correlations between the operational data and the closed systems, identifying key correlation points. The real-time operational status of each functional module is monitored. Finally, when the real-time operational status meets the trigger conditions for combinatorial optimization, a multi-objective genetic algorithm and constraint screening are used to dynamically adjust the combination relationships between functional modules based on the key correlation points and the real-time operational status, obtaining a target combination scheme. This target combination scheme is then executed to perform predictive maintenance on the modular data center. This invention significantly improves the reliability of data center operations and maintenance.
[0015] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.
[0016] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0017] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0018] Figure 1 A flowchart illustrating a predictive maintenance method for a modular data center provided in an embodiment of the present invention; Figure 2 A schematic diagram of a modular system assembly structure provided in an embodiment of the present invention; Figure 3 A schematic diagram illustrating the specific process of a predictive maintenance method for a modular data center provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of an AI-driven correlation analysis method provided in an embodiment of the present invention; Figure 5 A schematic diagram of a predictive maintenance device for a modular data center provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of a server provided in an embodiment of the present invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Currently, with the rapid development of cloud computing, big data, and artificial intelligence, data centers are expanding in scale, IT equipment power density is continuously increasing, and the operational complexity of cooling, power supply and distribution, and environmental protection systems is significantly increasing. To ensure business continuity, data centers have higher requirements for operational reliability, energy efficiency, and maintenance response speed. Existing data center maintenance methods mainly rely on manual inspections, threshold alarms, and periodic maintenance strategies. This approach has the following shortcomings: First, traditional alarm mechanisms are mostly based on setting thresholds for single devices or single parameters, making it difficult to reflect the coupling relationship between cooling systems, environmental systems, and IT loads, and easily leading to false alarms or missed alarms. Second, periodic maintenance methods cannot be adjusted according to the actual operating status of equipment, easily resulting in over-maintenance or maintenance delays, increasing maintenance costs. Third, some existing predictive maintenance methods analyze single devices or single models, lacking modeling of the relationships between modular systems, making it difficult to adapt to the dynamic combination of equipment and frequent load changes in modular data centers. Fourth, at the combination optimization and execution level, most systems rely on manual experience for adjustments, lacking systematic optimization strategies and effect verification mechanisms.
[0021] Based on this, the predictive maintenance method, device and server for modular data centers provided by the present invention can take into account the overall operating characteristics of modular data centers, comprehensively consider the relationship between cooling systems, environmental sensing systems and IT load systems, and have predictive capabilities, dynamic combination optimization capabilities and closed-loop evaluation capabilities, which can improve the operational reliability, energy efficiency level and intelligent operation and maintenance level of data centers.
[0022] See Figure 1 The diagram shows a flowchart of a predictive maintenance method for a modular data center. This method is applied to a modular system. (See attached diagram.) Figure 2 The diagram shows a modular system architecture, which includes an AI decision-making layer, a data interaction layer, and a physical device layer. The method mainly includes the following steps S102 to S106: Step S102: Based on the operation and maintenance requirements of the data center, the functional modules are divided into types, related functional modules are combined into independent closed systems, and the operation data of each functional module is collected and determined as prediction parameters.
[0023] In one implementation, functional modules can be divided according to the data center's operation and maintenance needs. Similar and related functional modules can be combined to form a closed system, and operational data from each module can be collected as predictive parameters. The functional modules include: a cooling module, an environmental sensing module, an environmental control module, and an IT load module. The cooling module includes: an air-cooled module, a liquid-cooled module, and an auxiliary cooling unit. The environmental sensing module includes: temperature and humidity equipment inside and outside the data center, a water immersion sensor, a smoke detector, a differential pressure sensor, a hydrogen quality sensor, a PM2.5 quality sensor, and a carbon dioxide quality sensor, etc. The environmental control module includes: a constant humidity machine. Operational data includes: compressor frequency, refrigerant pump frequency, fan speed, high / low pressure, and coolant flow rate of the cooling module; temperature, relative humidity, water immersion status, smoke status, differential pressure, and air quality of the environmental sensing module; and real-time current, voltage, power consumption, and heat density data of the IT load module.
[0024] Step S104: The operation data is processed by graph neural network to analyze the relationship between the operation data, so as to determine the key correlation points according to the relationship between each functional module and the closed system, and to monitor the real-time operation status of each functional module.
[0025] In one implementation, after constructing a multi-model fusion AI analysis framework, the operational data of individual functional modules and closed systems formed by different modular combinations are input. AI-driven methods are used to mine key correlations between systems, and the modular combination method is dynamically adjusted based on the correlation strength. Specifically, data preprocessing is performed first, using Z-score standardization to eliminate dimensional differences between different modules, using the isolated forest algorithm to remove outliers, and using linear interpolation to fill in missing values. Then, feature engineering is initiated to extract time-domain features (mean, standard deviation, kurtosis), frequency-domain features (fault frequency band energy), and time-frequency-domain features (wavelet entropy) to construct a multi-dimensional feature vector. Finally, a graph neural network (GNN) is used to construct a system correlation graph, where nodes in the system correlation graph represent functional modules or closed systems. The edge weights between two nodes can be calculated using the Pearson correlation coefficient. Correspondingly, by identification, two nodes with edge weights greater than 0.7 are considered strongly correlated key correlation points.
[0026] Step S106: When the real-time operating status meets the combinatorial optimization triggering conditions, the combination relationship between various functional modules is dynamically adjusted based on key correlation points and real-time operating status through multi-objective genetic algorithm and constraint screening to obtain the target combination scheme and execute the target combination scheme to perform predictive maintenance on the modular data center.
[0027] In one implementation, based on the strong correlation between key related points obtained from correlation analysis and the real-time operating status of each node, a multi-objective genetic algorithm combined with operation and maintenance rule constraints can be used to achieve dynamic combination optimization of functional modules and closed systems. Through resource collaborative scheduling, overall operating efficiency can be improved and energy consumption can be reduced. In addition, to ensure the reliable operation of the modular combination system and AI analysis model, a closed-loop process of model building, multi-index evaluation, problem diagnosis and optimization output can be constructed. Through multi-dimensional index quantification and prediction capabilities, the root causes of non-compliance can be accurately located and targeted suggestions can be generated to ensure the supporting value of prediction results for operation and maintenance decisions.
[0028] The predictive maintenance method for modular data centers provided in this invention can predict and warn of equipment operating status before a failure occurs through modular system construction, AI-driven correlation analysis, dynamic combination optimization, and predictive capability assessment. It can also dynamically optimize system combination while ensuring safety, thereby reducing the risk of unplanned downtime, reducing energy consumption, and improving operation and maintenance efficiency.
[0029] See Figure 3 The diagram illustrates a specific process flow for a predictive maintenance method for a modular data center. This invention also provides an implementation method for predictive maintenance of a modular data center. Existing technical solutions have low prediction accuracy in data center scenarios. This is partly due to system fragmentation; traditional methods monitor only a single device (such as an air conditioner), while a data center is a unified whole. Overheating of a server could be caused by a neighboring air conditioner malfunction, excessive load on the server itself, and poor ventilation, making accurate judgment impossible from a single perspective. Another issue is model rigidity; a well-trained model can quickly become ineffective after new equipment is deployed or business load patterns change. Furthermore, because data centers have numerous devices and complex interrelationships, key parameters required for predictive maintenance are easily missing, and manual monitoring makes it difficult to quickly resolve these missing parameters.
[0030] Therefore, this invention first constructs a modular system, and uses modular construction and correlation analysis to enable AI to understand the internal structure and dynamic correlation of the complex system of the computer room. Then, when the system is in a sub-healthy state (such as local overheating or decreased efficiency), the AI will proactively generate a system-level conditioning plan (i.e., a target combination plan). For example, it can direct several air conditioners to work together to precisely cool down the overheated server, thereby proactively adjusting the system state before the failure occurs to achieve predictive maintenance. Finally, the model can be trained to evaluate and optimize its predictive capabilities. If the target is met, the operation will be continuously monitored. If the target is not met, targeted optimization will be performed and optimization suggestions will be generated. For example, suggestions on the types of sensors to be added, suggestions on relearning new correlation relationships, suggestions on updating the model or knowledge base, etc. See (1) to (5) below for details: (1) After constructing the modular system, the operating data of the modular system is processed by time-domain feature extraction, frequency-domain feature extraction and time-frequency-domain feature extraction to obtain multi-dimensional feature vectors. Based on the multi-dimensional feature vectors, functional modules and closed systems, a system association graph is constructed. In one implementation, the association relationship of each functional module and closed system can be analyzed by graph neural network. Each functional module and closed system is identified as a node. The edge weights between nodes are calculated based on the multi-dimensional feature vectors to determine the connection relationship and construct the system association graph. The nodes of the system association graph are divided into two levels. The first-level nodes are independent functional module instances, and the second-level nodes are closed system instances formed by modular combination. All nodes are accompanied by a unique identifier (ID) and core attribute label.
[0031] Specifically, the first-level nodes (functional module examples) include: air-cooled precision air conditioners, liquid-cooled CDU (Coolant Distribution Unit) units, temperature and humidity sensors for the west racks, temperature and humidity monitoring equipment on the north side of the computer room, CPU load monitoring units for the north racks, air quality sensors, auxiliary cooling units, and disk I / O monitoring modules for the east racks. Each node is labeled with its module type (e.g., "cooling-air-cooled", "environmental sensing-temperature and humidity inside the computer room"), installation location (e.g., "rack 3 in column A of the west area", "north side wall outside the computer room"), and core monitoring parameters (e.g., "compressor speed" and "dry bulb temperature").
[0032] Secondary nodes (examples of closed systems): West Zone cooling closed system (related to air-cooled air conditioners, liquid-cooled CDUs, temperature and humidity sensors of all racks in the West Zone and monitoring equipment on the west side outside the computer room), East Zone IT load-cooling collaborative system (related to air-cooled air conditioners, liquid-cooled CDUs, load monitoring units of racks in the East Zone and temperature and humidity equipment on the east side outside the computer room), etc. Each node is labeled with the system coverage area (e.g., "West Zone racks 1-30 area"), core functions (e.g., "cooling supply and load adaptation") and a list of primary nodes included.
[0033] (2) Perform weight identification processing on all edges in the system association graph, and determine the group of nodes connected by the edges whose edge weights are greater than the preset weight threshold as a group of key association points between functional modules with strong correlation. In one implementation, the edge weights are calculated by combining the Pearson correlation coefficient with the time series consistency correction algorithm. The specific steps are as follows: (2-1) Data synchronization preprocessing: Extract the data of the same time window of the running parameters of all nodes (the default window length is 2 hours and the step size is 5 minutes) to ensure that the data timestamps are completely aligned and remove data segments with a time synchronization deviation of more than 5 seconds.
[0034] (2-2) Basic Correlation Calculation: For any two nodes' operating parameter time series (such as "liquid-cooled CDU coolant flow rate" and "North Zone rack real-time power consumption"), calculate the Pearson correlation coefficient r, using the following formula:
[0035] Where n is the number of time series data points, and x and y are the time series of the running parameters of the two nodes, respectively; (2-3) Time series consistency correction: Introduce a time delay factor τ, calculate the maximum correlation coefficient r_max under different delays. If the delay time exceeds 30 seconds, multiply r_max by a decay coefficient of 0.8 (because the correlation with excessive delay has no practical operational significance), and obtain the initial weight w0 of the edge after correction.
[0036] (2-4) Weight normalization: Map the initial weight w0 to the interval [0,1] to obtain the final edge weight w, as shown in the formula:
[0037] in, It is the minimum of all initial weights. It is the maximum value among all initial weights.
[0038] (2-5) Weight constraint: If two nodes belong to modules with no physical connection or logical association (such as "air quality sensor" and "disk I / O monitoring unit"), the edge weight is set to 0 directly and they do not participate in the subsequent identification of key association points.
[0039] In another implementation, the storage and representation of key association points: (2-6) Definition of key association points: Select node pairs with edge weight w greater than 0.7 as key association points. Each key association point contains five-dimensional attributes: "source node ID, target node ID, association weight, core association parameters, and time characteristics".
[0040] (2-7) Storage structure: A hybrid storage method combining key-value pairs and relational database is adopted. Specifically, the memory cache layer stores the five-dimensional attributes of key related points and the association strength sorting number with "source node ID and target node ID" as the key. A Redis cluster is used to realize millisecond-level queries and support high-concurrency access. The persistent storage layer creates a "strong association table" in the PostgreSQL database. The fields include source node ID, target node ID, association weight, core association parameters, average association delay time, weight update timestamp, and whether it is a dynamic association (association flag triggered by adding a new module). It supports multi-condition retrieval by node type, association weight range, and update time.
[0041] (2-8) Representation: In the GNN model, it is represented by directed edges with weight labels. The thickness of the edge is mapped to the weight (weight ≥ 0.8 is a thick edge, and 0.7 ≤ weight < 0.8 is a medium-thick edge). In the operation and maintenance visualization interface, it is displayed in the text format of "node A (parameter X) to [w=0.85] to node B (parameter Y)". It also supports exporting to a CSV format list of relationships, which contains complete five-dimensional attribute information, making it convenient for offline analysis and rule base updates.
[0042] (3) When the real-time running status meets any one of the preset threshold triggering conditions, preset event triggering conditions and preset period triggering conditions, it is determined that the real-time running status meets the combined optimization triggering conditions, and after recording the triggering type and triggering time, it enters the combined optimization process.
[0043] In one implementation, the triggering condition activation mechanism for the combinatorial optimization process employs a triple activation mechanism of threshold triggering, event triggering, and periodic triggering to ensure the timeliness and necessity of dynamic combination. The specific triggering conditions and activation logic are as follows: (3-1) Threshold Trigger: Real-time monitoring of core operating parameters and system-level indicators of each node. When any indicator exceeds the preset threshold range, the process is activated immediately. Core indicators include: cooling efficiency (e.g., liquid-cooled CDU heat exchange efficiency ≤ 85%, air-cooled air conditioner COP value ≤ 2.8); load matching (e.g., IT equipment CPU load ≥ 80% and corresponding area cooling power ≤ 60% of rated value, cabinet power density fluctuation ≥ 30% / 15 minutes); energy consumption (e.g., data center PUE (Power Usage Effectiveness) ≥ 1.4, single closed system energy consumption exceeds the benchmark value by 20%); environment (e.g., local area temperature ≥ 26℃, humidity deviating from the range of [45%, 65%]).
[0044] Threshold triggering uses a "parameter priority weighting" mechanism. For example, the combined trigger of "CPU load exceeding the standard + local temperature exceeding the standard" has a higher priority than single parameter triggering and can skip the system's default 30-second buffer period to activate directly.
[0045] (3-2) Event Triggering: When a preset key event occurs in the system, the combined optimization process is triggered, including: equipment status change events (such as cooling unit failure shutdown, deployment of new IT equipment, and offline sensor replacement); operation and maintenance events (such as manual adjustment of cooling parameters not achieving the expected effect, and performing data center area expansion operations); and sudden external environmental events (such as extreme high temperatures in summer causing the external temperature of the data center to be ≥35℃, and power grid voltage fluctuations causing unstable equipment power). After the event is triggered, the system will automatically collect the historical operating data and current status of the event-related nodes as the basic input for combined optimization.
[0046] (3-3) Periodic Trigger: To avoid missing potential optimization opportunities due to sudden triggers, a periodic trigger mechanism is set up. The default trigger period is 2 hours, which can be adjusted to 1 hour or 4 hours according to operation and maintenance needs. During periodic triggering, the system will comprehensively evaluate the operating efficiency, energy consumption cost, and load balancing of all closed systems. If the evaluation score (out of 100) is lower than 80, the combined optimization process will be initiated. If the score is lower than 75 for three consecutive periods, the "deep optimization mode" will be triggered, and the historical best combined solution library will be introduced for comparative analysis.
[0047] In another implementation, the above activation logic works in synergy, and the triple triggering mechanism adopts the deduplication priority principle. If multiple triggering conditions are met at the same time, the triggering type with the highest priority (event triggering > threshold triggering > periodic triggering) is used to avoid repeated execution of the process. After triggering, a "trigger log" is generated immediately to record the triggering type, triggering indicator / event details, activation time and a list of associated nodes, which is convenient for subsequent operation and maintenance traceability.
[0048] (4) The functional modules with strong correlations corresponding to the key correlation points are processed by node aggregation to obtain the combined unit. Then, the combined unit is iteratively calculated based on the optimization objective function through a multi-objective genetic algorithm and preset operation and maintenance constraint rules to obtain a set of candidate combination schemes. The set of candidate combination schemes is then filtered based on preset screening conditions to obtain the target combination scheme. The optimization objective function includes: minimizing energy consumption target, maximizing cooling efficiency target, and balancing load distribution target.
[0049] In one implementation, see Figure 4The diagram illustrates an AI-driven correlation analysis method. It uses key correlation points obtained through correlation analysis as the core link, combined with multi-objective optimization goals (minimizing energy consumption, maximizing cooling efficiency, and balancing load distribution) and operational constraints, to generate new combined solutions through a four-step process: correlation clustering, solution generation, constraint screening, and effect prediction. (4-1) Clustering of Key Association Points and Division of Combined Units: Based on the five-dimensional attributes of key association points (source node ID - target node ID - association weight - core association parameters - time characteristics), density clustering algorithm (DBSCAN) is used to divide association clusters. During the clustering process, the core condition is "association weight greater than 0.7". Nodes with direct strong association or indirect strong association (such as nodes A and B being strongly associated, and B and C being strongly associated, then A, B, and C are grouped into the same cluster) are aggregated into a combined unit. For example, the liquid-cooled CDU - North Zone cabinet CPU load monitoring unit - North Zone cabinet CPU load monitoring unit - air-cooled air conditioner (the association weight between them is ≥0.75) are aggregated into "North Zone IT load - cooling coordination unit". At the same time, the clustering results must meet the constraints of "physical location proximity" (the installation distance of nodes in the same combined unit is ≤50 meters) and "functional complementarity" (including load monitoring and cooling execution nodes, avoiding the aggregation of single-function nodes).
[0050] (4-2) Generation of multi-objective combination schemes: Based on the combination units, an improved non-dominated sorting genetic algorithm (NSGA-Ⅲ) is used to generate an initial set of combination schemes. The optimization objective function is set as follows: Minimize energy consumption target:
[0051] Where E is energy consumption, and the load factor is dynamically adjusted according to the correlation delay time of the key correlation point (when the delay is ≤10 seconds, the load factor takes the actual value, and when the delay is >10 seconds, it is multiplied by a penalty factor of 1.05).
[0052] The goal is to maximize cooling efficiency.
[0053] Wherein, η is the cooling efficiency, and the baseline value of η ≥ 90% must be guaranteed.
[0054] The goal of load balancing distribution:
[0055] Where N is the number of cabinets in the combined unit, σ is the load distribution, and σ must be ≤ 5kW fluctuation threshold.
[0056] During the algorithm iteration process, the "core correlation parameters" of key correlation points are used as constraints (such as the correlation ratio between "liquid-cooled CDU coolant flow rate" and "IT equipment power consumption" should be maintained at 1.2:1±5%), and 20 candidate combination schemes are initially generated.
[0057] (4-3) Constraint Screening and Solution Optimization: A maintenance rule base is introduced to perform secondary screening of candidate solutions. The rule base includes hard constraints and flexible constraints. Hard constraints (cannot be violated): such as the total capacity of cooling equipment in the combined unit being ≥ 1.1 times the total heat dissipation requirement of IT equipment, the safe distance between high-voltage equipment and sensors being ≥ 1 meter, and new combinations not exceeding the upper limit of the power supply capacity of the computer room.
[0058] Flexible constraints (priority requirements): such as prioritizing the use of node combinations with historical correlation stability scores ≥90, and prioritizing the retention of the original correlation relationships of "core protection nodes" marked by maintenance personnel (such as monitoring units corresponding to critical business servers).
[0059] After constraint screening, 3-5 groups of qualified schemes are retained, and the entropy weight method is used to calculate the comprehensive score of each scheme (energy consumption accounts for 40%, efficiency accounts for 35%, and load balancing accounts for 25%). The scheme with the highest score is selected as the final new combination scheme.
[0060] (4-4) Validation of Scheme Effect Prediction: Based on the LSTM prediction model constructed from historical operating data, the operating effect of the final new combination scheme is predicted, and the energy consumption change curve, cooling efficiency fluctuation range, and load balance prediction value are output for the next hour. If the prediction results meet the optimization objectives of "energy consumption reduction ≥10%", "efficiency improvement ≥5%", and "load fluctuation ≤3kW", the scheme officially takes effect; if not, the process is backtracked to the clustering stage to adjust the correlation cluster partitioning parameters and regenerate the scheme.
[0061] In another implementation, the execution of the new combination scheme adopts a "tiered decision-making" mechanism. Based on the type of equipment involved, the scope of adjustment, and the degree of impact, it is divided into two categories: "automatic system execution" and "execution after confirmation by maintenance personnel." This ensures a balance between operational safety and flexibility. The execution method classification criteria are as follows: three decision thresholds are set, and the execution method is determined based on three dimensions: "number of core devices involved in the scheme adjustment," "power adjustment range," and "number of cabinets covered by the affected area": Level 1 Decision (Automatic Execution): Involves ≤2 core devices, power adjustment range ≤10%, affects ≤5 cabinets, and does not involve high-voltage equipment or node combinations corresponding to critical business operations; Level 2 Decision (Execution after Confirmation): Involves 3-5 core devices, power adjustment range 10%-20%, affects 6-20 cabinets, or involves high-voltage equipment for non-critical business operations; Level 3 Decision (Execution after Collective Review): Involves >5 core devices, power adjustment range >20%, affects >20 cabinets, or involves the combined adjustment of critical business servers and core cooling systems.
[0062] Example illustration: The system automatically executes an instance where the trigger condition is "the temperature and humidity sensor in the west server rack detects a temperature rise to 27℃ (threshold 26℃), and the association weight with the air-cooled air conditioner is 0.82". The system generates a "air-cooled air conditioner - west server rack - temperature and humidity sensor" combined unit through clustering, and uses the NSGA-Ⅲ algorithm to generate a solution: lower the air supply temperature of the air-cooled air conditioner from 18℃ to 17℃, and increase the air supply velocity by 5%. This solution involves one core device (air-cooled air conditioner), an 8% power adjustment, and affects three server racks, and is classified as a Level 1 decision. The system automatically executes the adjustment command and simultaneously provides real-time feedback on the execution status (e.g., "Air-cooled air conditioner parameter adjustment completed, current air supply temperature 17℃, west server rack temperature reduced to 25.2℃") to the operations and maintenance platform.
[0063] After confirmation by operations and maintenance personnel, the following example was executed: The trigger condition was "the CPU load of the North Zone rack increased to 85% (threshold 80%), the heat exchange efficiency of the strongly associated liquid-cooled CDU was 83% (threshold 85%), and the trigger event was 'two new servers added in the North Zone'". The generated combined solution was: increasing the coolant flow rate of the liquid-cooled CDU from 80L / min to 95L / min; adding an auxiliary cooling unit associated with the North Zone rack, and adjusting the airflow direction to face the rack's air inlet. This solution involves two core devices (CDU and auxiliary cooling), a power adjustment range of 18%, and affects four racks, belonging to the secondary decision level. The system pushed the solution to the operations and maintenance personnel's workbench in the form of graphics and data, including solution details, expected effects (e.g., "North Zone IT load heat dissipation efficiency increased by 12%, energy consumption reduced by 8%), and historical execution records of similar solutions. After the operations and maintenance personnel confirmed the solution through the platform (supporting the addition of adjustment opinions, such as increasing the flow rate to 92L / min), the system executed the final solution and recorded the operations and maintenance operation log.
[0064] Regardless of the execution method used, the system will initiate a 30-minute effect verification period after the plan is executed, monitoring changes in key indicators in real time. If the expected results are not achieved, the system will automatically trigger plan backtracking and optimization adjustments to ensure the effectiveness of the combined optimization.
[0065] (5) A prediction model constructed by combining a long short-term memory network model and a gradient boosting tree model is used to evaluate the operational data within a preset time interval during predictive maintenance. The evaluation results are then used to evaluate and optimize the prediction capability. When the evaluation results do not meet the preset evaluation capability standards, fault tree diagnosis is performed at the data level, model level, and hardware level to generate corresponding optimization suggestions. See below for details: (5-1) Prediction Model Construction: The prediction system is constructed using a "dual-model collaboration" architecture. The core of the system is the complementary function of the LSTM time series network and the gradient boosting tree model. The LSTM time series network focuses on the prediction of the remaining life of equipment (RUL). It achieves long-term prediction by extracting the time series features of the equipment's full life cycle operation data. The input layer contains equipment operation parameters (such as the compressor start-stop frequency of air-cooled air conditioners and the temperature difference between the inlet and outlet of the coolant in liquid-cooled CDU), environmental parameters (corresponding regional temperature and humidity, and power grid voltage fluctuations), and historical fault repair records. The parameters are captured by three layers of LSTM units. The hidden layer uses the Dropout mechanism to prevent overfitting. The output layer outputs the predicted RUL value and confidence interval for the next month through the Sigmoid activation function. The gradient boosting tree model focuses on identifying short-term faults and energy efficiency anomalies. It uses core correlation parameters (such as the ratio of IT equipment power consumption to corresponding cooling power) as input features, and builds an integrated decision tree model through multiple iterations. This enables the classification and identification of fault types (such as "compressor jamming" and "sensor drift") and energy efficiency anomalies (such as "abnormal periods with PUE increases exceeding 5%)." During model training, 5-fold cross-validation is used to optimize tree depth and learning rate parameters, ensuring an accuracy rate of ≥92%. The outputs of the two models are weighted and integrated by the data fusion module to form a complete prediction output of "long-term lifespan prediction + short-term anomaly warning."
[0066] (5-2) Multi-indicator evaluation: The evaluation adopts a "tiered evaluation and sliding window verification" mechanism to achieve both accurate positioning and comprehensive evaluation. The specific implementation method is as follows: (5-2-1) Layered Assessment Objects: Based on the "closed system" as the basic assessment unit, a dedicated fault prediction sub-model and energy efficiency prediction sub-model are constructed for each closed system, and assessments are conducted independently. At the same time, a data center-level aggregated assessment model is constructed to weight and summarize the prediction results of each closed system (the weight is positively correlated with the system's energy consumption ratio) to form an overall data center prediction capability assessment conclusion. For example, the fault prediction capability of the "West Zone Cooling Closed System" and the energy efficiency prediction capability of the "East Zone IT Load-Cooling Collaborative System" are assessed separately, and then aggregated to obtain the overall prediction level of the data center.
[0067] (5-2-2) Sliding window execution logic: A sliding time window with a length of 1 day and a step size of 4 hours is adopted. All prediction results and actual data within the period are extracted in each window (such as RUL prediction value and actual equipment retirement time, fault prediction type and actual fault cause). Evaluation indicators are calculated window by window. Finally, the average value of indicators of 3 consecutive windows is used as the final evaluation basis to avoid evaluation deviation caused by data fluctuation in a single window.
[0068] (5-2-3) Index Calculation and Judgment Criteria: When MAPE ≤ 20% and R² ≥ 0.8, it is determined that the prediction ability of the model is good; when 20% < MAPE ≤ 30% or 0.6 ≤ R² < 0.8, it is determined that the prediction ability needs to be optimized; when MAPE > 30% or R² < 0.6, it is determined that the prediction ability does not meet the standard, and the optimization process is immediately started.
[0069] Among them, the Mean Absolute Percentage Error (MAPE): It is used to quantify the relative deviation between the predicted value and the actual value, and the formula is:
[0070] Among them, is the i-th actual value, is the i-th predicted value, and n is the number of data samples within the window.
[0071] The Mean Absolute Error (MAE): It is used to quantify the absolute deviation between the predicted value and the actual value and is used as a supplementary index to MAPE (to avoid the distortion of MAPE when the actual value approaches 0), and the formula is:
[0072] Among them, n is the number of data samples within the window, and i is the serial number of the data sample.
[0073] Coefficient of Determination (R²Score): It is used to measure the ability of the model to explain the data changes, and the formula is:
[0074] Among them, y is the average value of the actual values, n is the number of data samples within the window, and i is the serial number of the data sample.
[0075] The prediction ability evaluation index system is shown in Table 1 below: Table 1: Prediction Ability Evaluation Index System
[0076] (5-3) Generation of Optimization Suggestions: Adopt the "fault tree diagnosis and multi-dimensional verification" logic to locate the root cause of the unqualified prediction from three levels of data, model, and hardware, and generate implementable optimization suggestions. The specific judgment process and suggestions are as follows: (5-3-1) Diagnostic Decision Process: Step 1, initiate data quality verification: check the completeness (missing rate), accuracy (deviation from the baseline), and consistency (timestamp synchronization) of the input data within the sliding window; Step 2, if the data quality verification passes, conduct feature importance analysis: determine whether core predictive features are missing by increasing the feature contribution of the tree model output through gradient boosting; Step 3, if there are no problems with the data and features, conduct model fit analysis: compare the differences in the model's metrics between the training set and the test set to determine whether there is overfitting or underfitting; Step 4, combine the results of the first three steps to locate the root cause and generate corresponding optimization suggestions.
[0077] (5-3-2) Targeted optimization suggestions include: judgments and suggestions regarding insufficient data dimensions, judgments and suggestions regarding data quality issues, and judgments and suggestions regarding model adaptation issues. Specifically: Judgment and suggestions for insufficient data dimensions: The judgment criteria are "in feature importance analysis, the contribution of three or more core features (such as the frosting state of refrigeration equipment and sealing performance parameters) is 0, and there are no collection records for the corresponding parameters", while the data missing rate is less than or equal to 5% (excluding data quality issues). Optimization suggestions include: Hardware level: Add a frosting sensor to the evaporator of the air-cooled air conditioner and a pressure sensor to the liquid-cooled CDU pipeline to supplement the collection of sealing performance-related parameters; Data level: Define the collection range of the new parameters (such as frosting thickness 0-5mm) and the accuracy requirements of the new parameters (±0.1mm), and associate them with the attribute information of the corresponding key correlation points; Model level: After one week of data collection and accumulation of the new parameters, retrain the LSTM model and adjust the input layer dimensions.
[0078] Data quality issue assessment and recommendations: The assessment criteria are "data missing rate > 10%, or outlier percentage > 8% (e.g., sudden data jumps from sensors), or timestamp synchronization deviation > 5 seconds," with complete feature dimensions. Optimization recommendations include: Acquisition optimization: Increase the acquisition frequency of core parameters from 5 seconds / time to 1 second / time, while maintaining non-core parameters at 10 seconds / time to balance resources; Preprocessing optimization: Use a combination algorithm of "Kalman filter + 3σ criterion," first smoothing data fluctuations through Kalman filtering, then removing outliers through the 3σ criterion, and supplementing the missing data with linear interpolation; Hardware troubleshooting recommendations: Push a sensor maintenance checklist, prompting users to check the wiring and power supply status of sensors with high data missing rates, and replace sensors with frequent outliers.
[0079] Model adaptation issues and recommendations: The criteria for judgment are "data quality and feature dimensions meet the standards, but the difference between the model training set and test set metrics is >20% (overfitting), or the model's prediction error on newly deployed devices is 30% higher than historical data (poor adaptability)". Optimization recommendations include: Overfitting optimization: Use a transfer learning framework, using the trained model parameters as pre-training weights, and fine-tune them using public datasets of similar devices (such as the refrigeration equipment dataset from the IEEE PHM challenge), while increasing the Dropout ratio of the LSTM model to 0.3; Adaptability optimization: For newly deployed devices, extract their first 24 hours of operating data as incremental samples, and update the model parameters using online learning to avoid full retraining; Model structure optimization: If the device type is a newly added category (such as magnetic levitation air conditioners), it is recommended to add an attention mechanism layer after the LSTM model to enhance the feature extraction capability of parameters specific to this type of device.
[0080] After optimization suggestions are generated, they are pushed to the operation and maintenance platform in the form of "problem diagnosis report and operation manual". The priority of the suggestions is clearly defined (such as sensor replacement as first priority and model parameter fine-tuning as second priority). The implementation progress of the optimization measures is tracked. The evaluation process is restarted after the measures are implemented to form a closed-loop optimization.
[0081] In practical applications, the present invention provides the following embodiments. Embodiment 1: Predictive maintenance of data center cooling system. In this embodiment, for a medium-sized data center (containing 100 standard server racks, using a hybrid cooling mode of air cooling and cold plate liquid cooling), an AI-driven predictive maintenance method is implemented. The specific steps are as follows: I. Modular System Construction: 1. Functional modules are divided into: cooling module (including 15 air-cooled precision air conditioners and 5 liquid-cooled CDU units), environmental sensing module (100 temperature and humidity sensors inside the rack, 4 temperature and humidity sensors outside the server room, and 8 air quality sensors), and IT load module (CPU and memory monitoring units in each rack).
[0082] 2. Combined closed system: Combine the air-cooled module, liquid-cooled CDU unit, temperature and humidity sensor of the corresponding area cabinet and temperature and humidity sensor on the west side of the computer room to form the west zone cooling closed system. Similarly, construct three closed systems for the east zone, south zone and north zone.
[0083] 3. Data Acquisition: Data from each module is collected through a unified data access gateway at a frequency of 1Hz. The data includes 28 parameters such as compressor speed (rps) and condensing pressure (MPa) for air-cooled air conditioners, coolant flow rate (L / min) and inlet / outlet temperature (°C) for liquid-cooled CDUs, and cabinet temperature (°C) and relative humidity (%RH).
[0084] II. AI-driven correlation analysis and combinatorial optimization: 1. Data preprocessing: The collected data were standardized by Z-score, and 3.2% of the abnormal data caused by sandstorm weather were removed by the isolated forest algorithm. Linear interpolation was used to complete the 0.8% of missing values caused by offline sensors.
[0085] 2. Feature Engineering: Extract the time-domain features (mean, standard deviation), frequency-domain features (fault frequency band energy), and time-frequency-domain features (wavelet entropy) of each parameter to construct a 64-dimensional feature vector.
[0086] 3. Correlation Analysis: A system correlation graph was constructed using GNN to identify key correlation points such as "liquid-cooled CDU flow rate - high-load cabinet temperature" and "outdoor humidity of the computer room - risk of condensation in air-cooled air conditioners" (with weights of 0.82 and 0.76, respectively).
[0087] 4. Combination Optimization: When the CPU load of a certain cabinet in the North Zone exceeds 85%, the system automatically recombines the cabinet with the adjacent No. 2 liquid-cooled CDU unit, and at the same time increases the air supply angle of the air-cooled air conditioner to achieve precise distribution of cooling capacity.
[0088] III. Predictive Capability Assessment and Optimization: 1. Model Training: An LSTM model is built using TensorFlow. The model is trained by inputting 180 days of historical operating data (including 23 failure records) and combining it with an XGBoost model to identify failure types.
[0089] 2. Capability Assessment: After one month of operation, the assessment showed that the West Zone refrigeration system had a MAPE of 15.3%, a MAE of 0.4℃, and an R² of 0.85, indicating good predictive capability. The East Zone system, due to its proximity to the computer room entrance, experienced large temperature and humidity fluctuations, resulting in a MAPE of 32.1%, indicating that its predictive capability did not meet the standards.
[0090] 3. Optimization Implementation: Based on the recommendations, two wind speed sensors were added near the main entrance of the East Zone computer room, and the data acquisition frequency was increased to 2Hz. Transfer learning was used to update the model parameters. One week after optimization, the MAPE of the East Zone system dropped to 18.7%, meeting the operation and maintenance requirements.
[0091] 4. Implementation Results: The early warning time for the data center's cooling system failure was extended from 15 minutes to 2 hours, the air conditioning failure rate was reduced by 85%, cooling energy consumption was reduced by 32%, and monthly maintenance costs were reduced by approximately 12,000 yuan.
[0092] Example 2: Supplementary scenario for predictive capability optimization. When a large data center adds an AI training cluster (rack power density increased to 80kW / rack), the original predictive model exhibits a MAPE of 41.5%, and optimization is performed to address this issue: 1. Assessment and diagnosis: Residual analysis revealed that the model had a significant bias in predicting liquid cooling module failures under high loads. The core reason was the lack of data on coolant pressure pulsation and chip heat flux density.
[0093] 2. Optimization suggestions: Add a pressure sensor to the liquid cooling plate outlet and a heat flux density sensor near the AI server GPU to collect three additional key parameters.
[0094] 3. Model update: Incorporate the new data into the feature vector (expanded to 72 dimensions) and use adversarial training to optimize the LSTM model.
[0095] 4. Effect verification: After optimization, the model's MAPE dropped to 16.8%, and it successfully provided an early warning of a liquid-cooled CDU pump failure 4 hours in advance, avoiding cluster downtime losses.
[0096] In summary, the present invention has the following beneficial effects: 1. Expand the prediction range and improve accuracy: By modularly combining isolated equipment into a closed system, such as a refrigeration system that integrates air cooling, liquid cooling and internal and external environmental equipment, multi-parameter collaborative analysis can be achieved, which can improve the fault prediction accuracy to more than 85%, which is 40% higher than the prediction accuracy of a single device.
[0097] 2. Dynamic adaptation enhances scenario adaptability: AI-driven correlation analysis can identify key correlation points in the system in real time. When new modules are added to the data center or the load changes, the combination method is automatically optimized, and the model adaptation time is shortened from 1-2 weeks in the traditional method to within 24 hours.
[0098] 3. Closed-loop optimization ensures prediction effectiveness: The prediction capability is quantified through a multi-indicator evaluation system, and targeted data collection optimization suggestions are given, so that the prediction model can still maintain an accuracy of MAPE≤20% in complex scenarios and reduce unplanned downtime losses by more than 80%.
[0099] 4. Improve energy efficiency and reduce operation and maintenance costs: Optimized combination based on strong correlation analysis can improve the energy efficiency of the cooling system by 30%. Combined with energy efficiency anomaly prediction, it can realize refined energy consumption management, and save more than 100,000 yuan in electricity costs per year per computer room.
[0100] 5. It can significantly reduce the risk of unplanned downtime, reduce the intensity of manual operation and maintenance, and improve the intelligent operation and maintenance level of data centers, which has good engineering application value and promotion prospects.
[0101] Regarding the predictive maintenance method for modular data centers provided in the foregoing embodiments, this invention provides a predictive maintenance device for modular data centers, see [link to relevant documentation]. Figure 5 The diagram shows a structural schematic of a predictive maintenance device for a modular data center, which includes the following components: System construction module 502 divides the types of functional modules according to the operation and maintenance requirements of the data center, combines related functional modules into independent closed systems, collects the operation data of each functional module, and determines the operation data as prediction parameters. The correlation analysis module 504 uses a graph neural network to perform correlation analysis on the running data, in order to determine key correlation points based on the correlation between various functional modules and the closed system, and to monitor the real-time running status of various functional modules. The combinatorial optimization module 506, when the real-time operating status meets the combinatorial optimization triggering conditions, dynamically adjusts the combinatorial relationships between various functional modules based on key correlation points and real-time operating status through a multi-objective genetic algorithm and constraint screening to obtain a target combinatorial scheme and execute the target combinatorial scheme to perform predictive maintenance on the modular data center.
[0102] The predictive maintenance device for modular data centers provided in this application embodiment can significantly improve the reliability of data center operation and maintenance.
[0103] In one embodiment, when performing the step of analyzing the correlation of operational data using a graph neural network to determine key correlation points based on the correlation between various functional modules and the closed system, the correlation analysis module 504 is further configured to: perform time-domain feature extraction, frequency-domain feature extraction, and time-frequency-domain feature extraction on the operational data to obtain a multi-dimensional feature vector, and construct a system correlation graph based on the multi-dimensional feature vector, functional modules, and the closed system; perform weight identification processing on the system correlation graph, and determine nodes with edge weights greater than a preset weight threshold as key correlation points between functional modules.
[0104] In one embodiment, when performing the step of constructing a system association graph based on multidimensional feature vectors, functional modules, and closed systems, the aforementioned association analysis module 504 is further configured to: perform association relationship analysis on each functional module and closed system through a graph neural network, determine each functional module and closed system as nodes, calculate the edge weights between nodes based on multidimensional feature vectors, determine the edge weights as connection relationships, and construct a system association graph.
[0105] In one implementation, the combined optimization triggering condition is a triple activation mechanism, which includes threshold triggering, event triggering, and periodic triggering. In the step of real-time monitoring of the real-time operating status of each functional module, the above-mentioned correlation analysis module 504 is further used to: determine that the real-time operating status meets the combined optimization triggering condition when the real-time operating status meets any one of the preset threshold triggering condition, preset event triggering condition, and preset periodic triggering condition, and after recording the triggering type and triggering time, enter the combined optimization process.
[0106] In one embodiment, when performing the step of dynamically adjusting the combination relationship between various functional modules based on key correlation points and real-time operating status through multi-objective genetic algorithm and constraint screening to obtain a target combination scheme, the aforementioned combination optimization module 506 is further configured to: perform node aggregation processing on functional modules with strong correlation relationships corresponding to key correlation points to obtain combination units; perform iterative calculation processing on combination units based on optimization objective function through multi-objective genetic algorithm and preset operation and maintenance constraint rules to obtain a set of candidate combination schemes, and perform screening processing on the set of candidate combination schemes based on preset screening conditions to obtain a target combination scheme, wherein the optimization objective function includes: minimizing energy consumption target, maximizing cooling efficiency target, and balancing load distribution target.
[0107] In one embodiment, after executing the target combination scheme to perform predictive maintenance on the modular data center, the combination optimization module 506 is further configured to: perform multi-index evaluation processing on the running data within a preset time interval during the predictive maintenance process using a prediction model jointly constructed by a long short-term memory network model and a gradient boosting tree model, obtain evaluation results, and then evaluate and optimize the prediction capability based on the evaluation results.
[0108] In one embodiment, when performing the step of evaluating and optimizing the predictive capability based on the evaluation results, the combined optimization module 506 is further configured to: when the evaluation results do not meet the preset evaluation capability standards, perform fault tree diagnosis processing from the data level, model level and hardware level respectively, and generate corresponding optimization suggestions.
[0109] The device provided in this embodiment of the invention has the same implementation principle and technical effect as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the device embodiment can be referred to the corresponding content in the aforementioned method embodiment.
[0110] This invention provides a server, specifically, the server includes a processor and a storage device; the storage device stores a computer program, which, when run by the processor, executes the method described in any of the above embodiments.
[0111] Figure 6 This is a schematic diagram of the structure of a server provided in an embodiment of the present invention. The server 100 includes: a processor 60, a memory 61, a bus 62, and a communication interface 63. The processor 60, the communication interface 63, and the memory 61 are connected through the bus 62. The processor 60 is used to execute executable modules, such as computer programs, stored in the memory 61.
[0112] The memory 61 may include high-speed random access memory, or it may also include unstable memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 63 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc.
[0113] Bus 62 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 6 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0114] The memory 61 is used to store programs. After receiving an execution instruction, the processor 60 executes the program. The method executed by the device for defining the flow process disclosed in any of the foregoing embodiments of the present invention can be applied to the processor 60 or implemented by the processor 60.
[0115] Processor 60 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 60 or by instructions in software form. Processor 60 can be a general-purpose processor, including a central processing unit (CPU), network processor, etc.; it can also be a digital signal processor, application-specific integrated circuit (ASIC), off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable memory (EPR), registers, or other mature storage media in the art. This storage medium is located in memory 61. Processor 60 reads information from memory 61 and, in conjunction with its hardware, completes the steps of the above method.
[0116] The computer program product of the readable storage medium provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the foregoing method embodiments. For specific implementation, please refer to the foregoing method embodiments, which will not be repeated here.
[0117] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.
[0118] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method of predictive maintenance of a modular data center, the method comprising: The method comprises: According to the operation and maintenance requirements of the data room, the types of the functional modules are divided and processed, the functional modules of the same type are combined into independent closed systems, the operation data of each functional module is collected, and the operation data is determined as a prediction parameter; Through graph neural network, the operation data is analyzed and processed to determine the key association points according to the association relationship between each functional module and the closed system, and the real-time operation state of each functional module is monitored in real time; When the real-time operation state meets the combination optimization trigger condition, the combination relationship between each functional module is dynamically adjusted according to the key association points and the real-time operation state through multi-objective genetic algorithm and constraint screening, a target combination scheme is obtained, and the target combination scheme is executed to perform predictive maintenance on the modular data room.
2. The method of predictive maintenance of a modular data center of claim 1, wherein, The step of analyzing and processing the operation data through the graph neural network to determine the key association points according to the association relationship between each functional module and the closed system comprises: The operation data is processed to extract time domain features, frequency domain features and time-frequency domain features to obtain a multi-dimensional feature vector, and a system association graph is constructed according to the multi-dimensional feature vector, the functional modules and the closed system; The system association graph is processed to identify the weight, and the nodes with edge weight greater than a preset weight threshold are determined as the key association points between the functional modules.
3. The method of predictive maintenance of a modular data center of claim 2, wherein, The step of constructing the system association graph according to the multi-dimensional feature vector, the functional modules and the closed system comprises: Through the graph neural network, the association relationship between each functional module and the closed system is analyzed and processed, each functional module and the closed system is determined as a node, and the edge weight between nodes is calculated according to the multi-dimensional feature vector to determine the connection relationship and construct the system association graph.
4. The method of predictive maintenance of a modular data center of claim 1, wherein, The combination optimization trigger condition is a triple activation mechanism, which comprises threshold triggering, event triggering and periodic triggering, and after the step of monitoring the real-time operation state of each functional module, the step of monitoring the real-time operation state of each functional module comprises: When the real-time operation state meets any one of the preset threshold trigger condition, the preset event trigger condition and the preset periodic trigger condition, it is determined that the real-time operation state meets the combination optimization trigger condition, and after recording the trigger type and trigger time, the combination optimization process is entered.
5. The method of predictive maintenance of a modular data center of claim 1, wherein, The step of dynamically adjusting the combination relationship between each functional module according to the key association points and the real-time operation state through multi-objective genetic algorithm and constraint screening to obtain a target combination scheme comprises: The functional modules corresponding to the key association points that have a strong association relationship are aggregated to obtain a combination unit; The multi-objective genetic algorithm and the preset operation and maintenance constraint rule are used for iterative calculation and processing of the combination unit based on an optimization objective function, to obtain a candidate combination scheme set, and the candidate combination scheme set is screened based on a preset screening condition, to obtain the target combination scheme. The optimization objective function includes a minimum energy consumption target, a maximum refrigeration efficiency target, and a balanced load distribution target.
6. The method of predictive maintenance of a modular data center of claim 1, wherein, After the step of executing the target combination scheme to perform predictive maintenance on the modular data center, the method further includes: A prediction model constructed by a long short-term memory network model and a gradient boosting tree model is used to perform multi-index evaluation processing on operation data in a preset time interval during the predictive maintenance, to obtain an evaluation result, and the prediction capability is evaluated and optimized based on the evaluation result.
7. The method of predictive maintenance of a modular data center of claim 6, wherein, The step of evaluating and optimizing the prediction capability based on the evaluation result includes: When the evaluation result does not meet a preset evaluation capability standard, fault tree diagnosis processing is performed from a data layer, a model layer, and a hardware layer, respectively, to generate corresponding optimization suggestions.
8. A predictive maintenance device for a modular data center, comprising: The apparatus includes: A system construction module divides types of functional modules according to operation and maintenance requirements of a data center, combines functional modules of associated types into independent closed systems, and collects operation data of each functional module, and determines the operation data as prediction parameters; An association analysis module analyzes an association relationship of the operation data by a graph neural network, determines a key association point according to an association relationship between each functional module and the closed system, and monitors a real-time operation state of each functional module in real time; A combination optimization module dynamically adjusts a combination relationship between each functional module according to the key association point and the real-time operation state when the real-time operation state meets a combination optimization trigger condition, by a multi-objective genetic algorithm and constraint screening, to obtain a target combination scheme, and executes the target combination scheme to perform predictive maintenance on the modular data center.
9. A server, characterized by The computer readable storage medium stores computer executable instructions, and the computer executable instructions, when invoked and executed by the processor, cause the processor to implement the method of any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer executable instructions, and the computer executable instructions, when invoked and executed by the processor, cause the processor to implement the method of any one of claims 1 to 7.