Adaptive control method and system for data center based on data twinning
By collecting multi-dimensional sensing data in the data center and combining it with the multi-physics coupling network of the digital twin space, multi-dimensional feature vectors and risk prediction situation maps are constructed, which solves the problem that the cross-scale heat transport mechanism is not comprehensively considered in traditional methods and achieves more accurate adaptive control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Filing Date
- 2026-05-26
- Publication Date
- 2026-07-17
AI Technical Summary
Traditional adaptive control methods for data centers fail to comprehensively consider the cross-scale heat transport mechanism under the coupling effect of multiple physics fields, resulting in low accuracy of risk prediction situation maps and affecting the accuracy of adaptive control strategies for data centers.
By collecting multiple sensing data from the data center in different dimensions, and combining them with the multi-physics coupling network of the digital twin space to determine the spatial heat conduction path, a multi-dimensional feature vector is constructed to generate a risk prediction situation map. In the process of dynamic identification, multiple risk areas are marked to determine the adaptive control path.
It improves the accuracy of risk prediction situation maps and the accuracy of data center adaptive control strategies, and fully considers multi-dimensional feature vectors and the current working content of devices.
Smart Images

Figure CN122411045A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of digital twins, and more particularly to an adaptive control method and system for a data center based on data twins. Background Technology
[0002] With the explosive growth of high computing power demands from cloud computing, artificial intelligence, and large-scale model training, the power density of single racks in data centers is constantly increasing, and the operating environment of high-density data centers is becoming increasingly complex. In such high-density scenarios, there is a significant physical response lag between the transient load changes of IT equipment and the cooling supply of the cooling system, and there are deep cross-scale coupling relationships between airflow fields, thermal radiation fields, and electromagnetic fields. This poses a huge challenge to real-time thermal management and adaptive control of data centers.
[0003] Traditional prediction methods often rely on single-dimensional temperature thresholds or static heat conduction models, failing to comprehensively consider cross-scale heat transport mechanisms under the coupling of multiple physics fields. They also fail to integrate dynamic heat propagation characteristics such as heat flow vector direction, dynamic distribution of thermal resistance, and heat accumulation rate with IT data load and cooling operation characteristics at the feature level. Because they ignore the dynamic topological evolution and multi-field coupling lag effect in the heat conduction process, the accuracy of risk prediction situation maps is affected, resulting in low accuracy of adaptive control strategies for data centers. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of the prior art. This invention provides an adaptive control method and system for data centers based on data twins.
[0005] This invention provides an adaptive control method for a data center based on data twins, comprising: In this data center, multiple sensing data from different dimensions are collected, and the spatial heat conduction path is determined by combining the multi-physics coupling network of the digital twin space. The dynamic heat propagation characteristics of the data center are determined based on the topological mapping of the spatial heat conduction path. Data load data of data centers is collected, and multi-dimensional feature vectors are constructed by combining dynamic heat propagation characteristics and cooling equipment operation characteristics. Based on the dynamic deduction of multi-dimensional feature vectors, the power consumption drift trend and local heat accumulation risk of data centers in the future within a preset time window are determined, thereby generating a risk prediction situation map. The risk prediction situation map is input into the data twin space, and multiple risk areas are marked during the dynamic identification process. The multiple risk areas are iterated alternately, and multiple control contents are determined by combining the real-time working data of multiple devices in the data center. Multiple control contents are loaded into the distribution map of the data center, and multiple adaptive control paths are determined in combination with the current working sequence of the data center. The multiple adaptive control paths are optimized, and the adaptive control strategy of the data center is constructed in combination with the current working contents of multiple devices.
[0006] This invention provides an adaptive control system for a data twin-based data center, which is applied to the aforementioned adaptive control method for a data twin-based data center. The adaptive control system for the data twin-based data center includes: The digital twin module is used to collect multiple sensing data from different dimensions in the data center, and combine them with the multi-physics coupling network of the digital twin space to determine the spatial heat conduction path. Based on the topological mapping of the spatial heat conduction path, the dynamic heat propagation characteristics of the data center are determined. The risk prediction situation map module is used to collect the data load of the data center and construct a multi-dimensional feature vector by combining dynamic heat propagation characteristics and cooling equipment operation characteristics. Based on the dynamic deduction of the multi-dimensional feature vector, the power consumption drift trend and local heat accumulation risk of the data center in the future preset time window are determined, thereby generating a risk prediction situation map. The control content module is used to input the risk prediction situation map into the data twin space, mark multiple risk areas during the dynamic identification process, iterate over multiple risk areas alternately, and determine multiple control contents by combining the real-time working data of multiple devices in the data center. The adaptive control module is used to load multiple control contents into the distribution map of the data center, determine multiple adaptive control paths in combination with the current working sequence of the data center, optimize the multiple adaptive control paths, and construct the adaptive control strategy of the data center in combination with the current working contents of multiple devices.
[0007] Compared with the prior art, the beneficial effects of the present invention are: (1) In this data center, multiple sensing data of the data center in different dimensions are collected, and the spatial heat conduction path is determined by combining the multi-physics field coupling network of the digital twin space. The dynamic heat propagation characteristics of the data center are determined by the topological mapping of the spatial heat conduction path. The data load of the data center is collected, and a multi-dimensional feature vector is constructed by combining the dynamic heat propagation characteristics and the operating characteristics of the cooling equipment. The power consumption drift trend and local heat accumulation risk of the data center in the future preset time window are determined by the dynamic deduction of the multi-dimensional feature vector, thereby generating a risk prediction situation map. The dynamic heat propagation characteristics of the data center are introduced, and the multi-dimensional feature vector is further controlled, which improves the accuracy of the risk prediction situation map.
[0008] (2) Input the risk prediction situation map into the data twin space, mark multiple risk areas during the dynamic identification process, iterate the multiple risk areas alternately, and determine multiple control contents in combination with the real-time working data of multiple devices in the data center; load the multiple control contents into the distribution map of the data center, determine multiple adaptive control paths in combination with the current working sequence of the data center, optimize the multiple adaptive control paths, and construct the adaptive control strategy of the data center in combination with the current working contents of multiple devices, further control the multiple control contents, fully consider the multiple adaptive control paths and the current working contents of multiple devices, and improve the accuracy of the adaptive control strategy of the data center. Attached Figure Description
[0009] Figure 1 This is a flowchart illustrating the adaptive control method for a data center based on data twins in an embodiment of the present invention. Figure 2 This is a flowchart illustrating step S11 in the adaptive control method for a data center based on data twins in an embodiment of the present invention. Figure 3 This is a flowchart illustrating step S12 in the adaptive control method for a data center based on data twins in an embodiment of the present invention. Figure 4 This is a flowchart illustrating step S13 in the adaptive control method for a data center based on data twins in an embodiment of the present invention. Figure 5 This is a flowchart illustrating step S14 in the adaptive control method for a data center based on data twins in an embodiment of the present invention. Figure 6 This is a schematic diagram of the structural composition of an adaptive control system for a data center based on data twins, as described in an embodiment of the present invention. Detailed Implementation
[0010] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0011] Please see Figures 1 to 6 An adaptive control method for data centers based on data twins, applied to digital twin scenarios; the adaptive control method for data centers based on data twins includes: Step S11: In this data center, collect multiple sensing data from different dimensions of the data center, and combine them with the multi-physics coupling network of the digital twin space to determine the spatial heat conduction path. Based on the topological mapping of the spatial heat conduction path, determine the dynamic heat propagation characteristics of the data center. Step S12: Collect the data load of the data center and construct a multi-dimensional feature vector by combining the dynamic heat propagation characteristics and the operating characteristics of the cooling equipment. Based on the dynamic deduction of the multi-dimensional feature vector, determine the power consumption drift trend and local heat accumulation risk of the data center in the future preset time window, thereby generating a risk prediction situation map. Step S13: Input the risk prediction situation map into the data twin space, mark multiple risk areas during the dynamic identification process, iterate over multiple risk areas alternately, and determine multiple control contents by combining the real-time working data of multiple devices in the data center. Step S14: Load multiple control contents into the distribution map of the data center, determine multiple adaptive control paths in combination with the current working sequence of the data center, optimize the multiple adaptive control paths, and construct the adaptive control strategy of the data center in combination with the current working contents of multiple devices.
[0012] refer to Figure 2 In step S11, the specific steps are as follows: S111: Real-time monitoring data center, collects multiple sensing data from different dimensions through multimodal sensor network, maps multiple sensing data to digital twin space, and performs cross-scale joint analysis of airflow field, thermal radiation field and electromagnetic field through pre-built multi-physics field coupling network in digital twin space to extract spatial heat conduction boundary conditions; S112: Construct dynamically updated spatial heat conduction paths based on spatial heat conduction boundary conditions; perform topological mapping on the spatial heat conduction paths and extract graph structure features; further combine the transient heat source distribution information that evolves over time to determine the dynamic heat propagation characteristics of the data center. The dynamic heat propagation characteristics include the direction of heat flow vector, dynamic distribution of thermal resistance, and heat accumulation rate.
[0013] In the embodiments of this application, a real-time monitoring data center is used to collect multiple sensing data of different dimensions through a multimodal sensing network, and the multiple sensing data are mapped to a digital twin space. Through a pre-constructed multi-physics coupling network in the digital twin space, the airflow field, thermal radiation field and electromagnetic field are jointly analyzed across scales to extract the spatial heat conduction boundary conditions, thus introducing spatial heat conduction boundary conditions.
[0014] At this point, multi-dimensional sensing data is acquired in real time by a heterogeneous sensor array deployed in the physical space of the data center, and the sensing data is mapped to the digital twin space with a unified spatiotemporal reference. The sensing data collected by the multimodal sensor network covers environmental thermodynamic sensing data, IT equipment operating status sensing data, and power supply and distribution system electrical sensing data. Due to the differences in sampling frequency and transmission delay of the above heterogeneous sensors, it is necessary to use interpolation compensation and time alignment based on timestamps and physical space coordinates to perform spatial voxel mapping of the multiple sensing data in the digital twin space. This ensures that each three-dimensional grid cell in the digital twin space is endowed with a multi-dimensional attribute tensor that evolves over time, thereby completing the accurate data feeding from physical sensing to the twin model.
[0015] After the sensing data is mapped, a pre-built multi-physics coupling network in the digital twin space is invoked. The multi-physics coupling network integrates the incompressible Navier-Stokes equations and energy equations for describing airflow and convective heat transfer, the discrete coordinate radiation model for describing radiative heat transfer, and the high-frequency eddy current loss equation for describing electromagnetic heat generation. By performing cross-scale coupled iterative solutions on the airflow field, thermal radiation field, and electromagnetic field, the joint analysis from chip-level micro-heating to data center-level macro-diffusion is achieved. The specific process of the cross-scale joint analysis is as follows: taking the printed circuit boards and chips inside the IT equipment as the microscale and the cabinet layout and channel structure as the macroscale, at the microscale, the high-frequency transient current loss obtained from the electromagnetic field analysis is converted into a Joule heat source, which serves as the internal thermal boundary condition for solving the airflow field and thermal radiation field. At the macroscale, the macroscopic cold transport rate obtained from the airflow field analysis is applied to the equipment surface at the microscale as the convective heat transfer coefficient. At the same time, the thermal radiation field analysis is used to compensate for the radiative heat exchange between the shielded area that the macroscopic airflow cannot reach and the high-temperature cabinet surface. Through the alternating transfer of heat flow and temperature parameters at the microscopic and macroscopic boundaries of the above multiphysics field equations, the joint analysis of cross-scale thermal characteristics is achieved.
[0016] Based on the calculation results of the multi-physics cross-scale joint analysis, boundary conditions for defining spatial heat conduction paths are extracted. These boundary conditions include a first type of temperature boundary condition, a second type of heat flux density boundary condition, and a third type of convective heat transfer boundary condition. The specific extraction logic is as follows: the constant or variable temperature control target at the outlet of the precision air conditioner or liquid-cooled cold plate is calibrated as the first type of temperature boundary; the transient power consumption of the server CPU and GPU caused by dynamic load is converted and extracted as the second type of heat flux density boundary; and the airflow mixing area at the interface of the hot and cold aisles and the surface of the cabinet are extracted as the third type of convective heat transfer boundary. Simultaneously, based on the data deviation fed back in real time by the multimodal sensor network, Kalman filtering or data assimilation is used to dynamically correct and identify the parameters of the extracted boundary conditions to eliminate the accumulated error of steady-state assumptions in the twin model analysis process, ensuring that the extracted spatial heat conduction boundary conditions can truly reflect the real-time physical state of the data center.
[0017] Specifically, in the high-density GPU computing cluster area within the data center that undertakes large model training tasks, the power density of a single rack in this area reaches over 20kW, employing a combination of cold aisle enclosure and row-level air conditioning for cooling. A multimodal sensor network collects in real time the following data in this area: "environmental perception data: such as the supply air temperature at the bottom of the cold aisle is 18℃, the return air temperature at the top of the rack is 35℃, and the supply air velocity is 2.5m / s", "IT equipment perception data: such as the current computing power utilization rate of the GPU cluster is 95%, and the motherboard power consumption is 18.5kW", and "power supply and distribution perception data: such as the transient current fluctuation and busbar temperature measured by the intelligent PDU in a single rack". After time-series alignment, the above heterogeneous data is mapped to the voxel mesh model in the digital twin space corresponding to the physical coordinates of the GPU rack.
[0018] A multiphysics coupling network is triggered to perform cross-scale joint analysis. At the microscale, based on the transient current parameter fed back by the PDU, the Joule heat loss distribution of the high-frequency switching power supply and power cables inside the GPU is calculated through electromagnetic field analysis. This is then coupled as a microscale heat source term into the thermal radiation field and airflow field. At the macroscale, based on the wind speed sensor data in the cold aisle, the airflow field is analyzed to determine the three-dimensional flow field pattern of cold air flowing in from the bottom anti-static floor perforated plate and being drawn upward by the server fan. At the same time, considering the high surface temperature of adjacent high-load racks and the presence of airflow stagnation areas, the thermal radiation field is analyzed to calculate the radiative heat transfer between the side panels of two adjacent racks. This achieves cross-scale joint analysis of electromagnetic heat generation, airflow convection, and radiative heat transfer.
[0019] Based on the above analysis results, the heat conduction boundary conditions of this high-density area are extracted: the current air outlet temperature setpoint of 18℃ for the row-level air conditioner is extracted as the first type of temperature boundary at the bottom of the cabinet; the power consumption of 18.5kW corresponding to 95% utilization of the GPU cluster is extracted as the second type of heat flux density boundary for the heat-generating nodes inside the cabinet; and the cold and hot airflow infiltration zone caused by pressure difference inside and outside the cold aisle is extracted as the third type of convective heat transfer boundary. When the power consumption of the GPU cluster instantly rises to 22kW due to the switching of training tasks, the PDU and motherboard sensor data provide real-time feedback on this transient change. The system dynamically corrects the second type of heat flux density boundary conditions of the cabinet through data assimilation, thereby providing a high-fidelity, time-varying boundary input basis for the construction of dynamic heat conduction paths and the extraction of thermal resistance and heat flux vector features in the subsequent step S112.
[0020] Furthermore, a dynamically updated spatial heat conduction path is constructed based on the spatial heat conduction boundary conditions; the spatial heat conduction path is topologically mapped and graph structure features are extracted; and the dynamic heat propagation characteristics of the data center are determined by combining the transient heat source distribution information that evolves over time. The dynamic heat propagation characteristics include the direction of heat flow vector, dynamic distribution of thermal resistance, and heat accumulation rate, which takes into account the overall consideration of the evolving transient heat source distribution information, ensuring the accuracy of the dynamic heat propagation characteristics of the data center.
[0021] At this point, the spatial heat conduction boundary conditions that evolve over time, extracted in step S111, are used as driving parameters to construct and dynamically update spatial heat conduction paths within the digital twin space. The physical heat sources, cooling equipment, and airflow transition spaces within the data center are discretized into three-dimensional thermodynamic grid nodes. Based on the first type of temperature boundary, the second type of heat flux density boundary, and the third type of convective heat transfer boundary, the thermal conductivity matrix between adjacent grid nodes is calculated. This thermal conductivity matrix is composed of the convective heat transfer coefficient, thermal conductivity, and radiative equivalent heat transfer coefficient between nodes, thereby establishing the heat flow transfer weights between nodes. When the IT load inside the data center changes abruptly or the airflow parameters of the cooling equipment are adjusted, causing a change in the boundary conditions, the thermal conductivity matrix is re-solved, thereby real-time addition, deletion, or modification of the direction and magnitude of the heat flow transfer weights between nodes. This generates dynamically updated spatial heat conduction paths composed of directed edges and weighted nodes. These paths not only characterize the macroscopic physical transmission channel of heat from the heat source to the cold source but also reflect the transient shift in heat flow caused by local airflow turbulence at the microscopic scale.
[0022] After constructing dynamically updated spatial heat conduction paths, the thermodynamic grid of the physical space is mapped to a graph data structure in the topological space, and graph structure features are extracted. Specifically, the discretized three-dimensional thermodynamic grid nodes are abstracted as vertices in the topological graph, the directed heat flow channels between nodes are abstracted as directed edges in the topological graph, the thermal conductivity weights and heat flow directions are quantized as edge attribute features, and the transient temperature and power consumption of nodes are quantized as vertex attribute features, thereby completing the mapping from physical heat conduction paths to the topological graph. Furthermore, graph neural networks or spectral graph theory are used to extract features from the topological graph. The attribute features of each vertex and its adjacent vertices are aggregated through graph convolution operations to obtain a high-dimensional hidden feature vector representing local thermal topological connectivity. The graph structure features can keenly capture local topological distortions in the thermal connectivity graph caused by the mixing of hot and cold air currents, rack layout obstruction, or uneven distribution of cooling capacity within the data center.
[0023] Based on the extracted graph structure features, and further combined with transient heat source distribution information evolving over time, the dynamic heat propagation characteristics of the data center are determined. These dynamic heat propagation characteristics include heat flow vector direction, dynamic thermal resistance distribution, and heat accumulation rate. At this point, the power consumption curves of IT equipment changing in real time with business load are used as transient heat source distribution information, and feature-level fusion is performed with the graph structure features. This data is then input into a pre-built spatiotemporal dynamics decoding network to calculate three major dynamic features: based on the attribute features of directed edges and node pressure differences, the heat flow vector direction is analyzed, and this vector direction characterizes the heat flow in three-dimensional space. The priority of transmission direction and flow intensity is determined; secondly, based on the actual heat flux density between nodes and the temperature difference at both ends, the dynamic distribution of thermal resistance is calculated by inversion. The dynamic distribution of thermal resistance eliminates the interference of static geometric distance and truly reflects the sudden change in local thermal resistance caused by airflow bypass, blind plate absence, or blockage of microchannels in liquid-cooled plates. Based on the difference between the transient heat source input power of the current node and the heat dissipation power flowing out through the heat flow vector, combined with the heat capacity parameter of the medium where the node is located, the heat accumulation rate is calculated. The heat accumulation rate directly quantifies the degree of temperature rise caused by the dynamic imbalance between heat generation and heat dissipation in the local space.
[0024] Specifically, when the GPU cluster starts the iterative calculation phase of the large model training task, causing the power consumption of a single rack to jump from 8kW to 22kW instantly, the second type of heat flux density boundary condition extracted in step S111 changes drastically. The thermal conductivity matrix between the grid node corresponding to the rack and the adjacent cold source node in the digital twin space is triggered to be recalculated. As the server fan speed inside the rack increases with the temperature and exhausts at full load, the pressure difference distribution at the rack's air inlet and outlet changes, causing some of the cold air originally flowing to the adjacent rack to shift to the high-heat rack. At this time, the system dynamically updates the spatial heat conduction path, cuts off the weak heat flow connection between the rack and the adjacent low-load rack in topological logic, and strengthens the strong heat flow directed edge between it and the bottom floor air supply cold source and the top row air conditioner return air vent.
[0025] The physical heat conduction path of the region after the above dynamic update is mapped as a topological graph, in which the 20kW high-heat cabinet is mapped as a high-weight central vertex. The directed edges between it and the row-level air conditioners and the cold source nodes of the floor air outlets are given a large heat flow weight. By extracting features from this local subgraph through a graph convolutional network, it is found that the aggregation weight of the feature vector of the high-heat vertex and the feature vector of the surrounding cold source vertices is abnormally concentrated. The extracted graph structure features accurately represent that a strong "thermal vortex" topological connectivity structure has been formed in this local area, that is, the surrounding cold energy is excessively drawn to the central hot spot.
[0026] The instantaneous 22kW power consumption time-series curve of the GPU cluster was used as input for transient heat source distribution information. This information was then fused with the aforementioned graph structure features to calculate dynamic heat propagation characteristics. The analyzed heat flow vector direction showed that cold air from the bottom of the cold aisle flowed vertically upwards into the upper part of the cabinet with extremely high vector intensity. Due to the instantaneous insufficient cooling supply, some of the heat return vectors pointed towards the weakly cooled area adjacent to the cabinet's rear door. The analyzed dynamic thermal resistance distribution indicated that the extremely large cabinet exhaust volume caused cold air bypassing through the gap at the bottom of the cold aisle's closed door, resulting in a sharp decrease in the thermal resistance of the topological edge corresponding to this gap. Conversely, the thermal resistance at the airflow bottleneck from the cabinet's rear door to the row-level air conditioning return vent significantly increased. The analyzed heat accumulation rate showed that, because the transient heat generation of 22kW far exceeded the dissipation capacity of the current supply cooling capacity, the local heat accumulation rate of the grid containing the GPUs in the upper part of the cabinet reached a critical high-risk value of 5°C per minute. This characteristic parameter directly provided a quantitative spatiotemporal boundary and evolution benchmark for the subsequent step S12 to generate a risk prediction situation map.
[0027] refer to Figure 3 In step S12, the specific steps are as follows: S121: Obtain the operating characteristics of the cooling equipment in the data center, and align the data load with the dynamic heat propagation characteristics and the operating characteristics of the cooling equipment in time and space, so as to construct a multi-dimensional feature vector in the feature-level fusion stage; S122: Input the multi-dimensional feature vector into the time series network and perform dynamic inference to determine the hysteresis coupling effect between the data load and the cooling capacity of the cooling equipment. Then, determine the power consumption drift trend and local heat accumulation risk of the data center in the future preset time window. Based on the gradient change of the power consumption drift trend and the spatial probability distribution of the local heat accumulation risk, generate a risk prediction situation map containing time and spatial dimensions.
[0028] In the embodiments of this application, the operating characteristics of the cooling equipment of the data center are obtained, and the data load situation is spatiotemporally aligned with the dynamic heat propagation characteristics and the operating characteristics of the cooling equipment, thereby constructing a multi-dimensional feature vector in the feature-level fusion stage. This approach is compatible with the overall considerations in the feature-level fusion stage and ensures the accuracy of the multi-dimensional feature vector.
[0029] At this time, the operating characteristics of the cooling equipment are acquired in real time through the communication interface between the Data Center Infrastructure Management System (DCIM) and the controller built into the cooling equipment. The operating characteristics of the cooling equipment cover the operating parameters on the cold source side and the operating parameters on the distribution side. The operating parameters on the cold source side include the supply and return water temperatures of the chiller unit, the compressor frequency, and the coefficient of performance (COP). The operating parameters on the distribution side include the fan speed and frequency of the row-level air conditioning terminal, the set value of the outlet air temperature, the opening degree of the electric water valve, the actual outlet air temperature, and the real-time cooling capacity. The above operating characteristics not only reflect the current cooling capacity and energy consumption status of the cooling system, but also imply the dynamic response lag characteristics and cooling output attenuation boundary of the cooling equipment during the variable operating condition adjustment process, thus providing physical constraint parameters for subsequent evaluation of the dynamic matching degree of cooling and heating supply and demand.
[0030] After acquiring the operating characteristics of the refrigeration equipment, the data load and dynamic heat propagation characteristics obtained in the aforementioned steps are spatiotemporally aligned with the operating characteristics of the refrigeration equipment. Specifically, for the alignment in the time dimension, since the data acquisition frequency of IT equipment load is usually at the second or even millisecond level, while the change cycle of the water valve opening and supply and return water temperature of the refrigeration equipment is mostly at the minute level, there is a significant time scale difference between the two. Therefore, dynamic time warping (DTW) or time displacement compensation based on the physical response lag model is adopted to extrapolate a heat transfer lag time constant backward on the time axis of the refrigeration equipment's cooling output parameters, so that the current refrigeration characteristics are aligned with the IT load characteristics at the historical moment that triggered the refrigeration adjustment in terms of physical causality. For the alignment in the spatial dimension, based on the three-dimensional discrete grid coordinates of the data center physical layout, the IT load data is located to the heat generation grid voxel, the dynamic heat propagation characteristics are located to the grid voxel where the heat flow transport path is located, and the refrigeration equipment operating characteristics are located to the grid voxel where the cold source outlet and cold aisle environment are located, thereby establishing a strict physical correspondence between multi-dimensional parameters under a unified spatiotemporal reference.
[0031] After completing spatiotemporal alignment, a multidimensional feature vector is constructed in the feature-level fusion stage. At this point, the aligned data load, dynamic heat propagation characteristics, and cooling equipment operation characteristics are subjected to feature concatenation and orthogonalization. The data structure of the multidimensional feature vector includes IT load feature sub-vectors, heat propagation feature sub-vectors, and cooling feature sub-vectors. The IT load feature sub-vectors consist of CPU and GPU utilization, dynamic power consumption, and task queue depth. The heat propagation feature sub-vectors consist of quantized components of the local heat flow vector direction, scalar values of dynamic thermal resistance distribution, and heat storage. The cooling feature vector is composed of the row-level air conditioner fan frequency, water valve opening degree, and current effective cooling capacity. Furthermore, in order to eliminate the order-of-magnitude differences between the physical dimensions and extract the implicit coupling correlation across the three domains of cold, heat, and computation, the spliced high-dimensional feature vector is normalized and subjected to principal component analysis (PCA) or autoencoder dimensionality reduction. Finally, a low-dimensional dense multidimensional feature vector is generated that can independently represent the physical state of each domain and contains multi-physics coupling and cooperative information, so as to provide subsequent steps for dynamic inference of power consumption drift trend and local heat accumulation risk.
[0032] Specifically, the operating characteristics of the row-level air conditioning terminals and the bottom-level chiller units corresponding to this high-density area are obtained, including real-time data such as "current operating frequency of row-level air conditioning fans: for example, 45Hz", "opening degree of electric two-way valve for chilled water: for example, 60%", "actual outlet air temperature: for example, 16℃" and "current transient calculated cooling capacity: for example, 25kW". These parameters together constitute the operating characteristics of the refrigeration equipment, which characterize the maximum cooling capacity supply boundary and airflow driving force that the current refrigeration system can provide in this local area.
[0033] When the GPU cluster experiences a sudden load surge due to the start of large model forward propagation computation, the "data load data: rack power consumption suddenly increases to 22kW" and "dynamic heat propagation characteristics: such as heat accumulation rate reaching 5℃ / minute" change abruptly within seconds. However, due to the thermal inertia of the chilled water pipe network and the mechanical action time of the water valve actuator, there is a physical lag of about 2 minutes in the increase of the cooling capacity of the row air conditioner. Therefore, the system adopts a time-series displacement compensation method based on the thermal inertia model to align the cooling characteristics collected at the current moment with the IT load and heat propagation characteristics 2 minutes ago on the time axis. At the same time, in the spatial dimension, the 22kW load characteristics are anchored to the grid voxel of the high-heat rack, and the row air conditioner air outlet characteristics are anchored to the perforated floor and air conditioner air outlet grid voxels adjacent to the rack, ensuring that the cold source supply characteristics and the heat source demand characteristics are strictly registered in three-dimensional space.
[0034] The spatiotemporally aligned heterogeneous parameters are constructed into a multidimensional feature vector. For the specific grid of this high-heat cabinet, the IT load sub-vector [22kW, 95%, 100%] is extracted, corresponding to power consumption, CPU utilization, and GPU utilization, respectively. The heat propagation sub-vector [2.5m / s, 0.15K / W, 5℃ / min] is extracted, corresponding to the Z-axis component of the heat flow vector pointing to the top of the cabinet, the local cold aisle bypass thermal resistance, and the heat accumulation rate, respectively. The cooling sub-vector [45Hz, 60%, 16℃, 25kW] is extracted, corresponding to the fan frequency, water valve opening, outlet air temperature, and effective cooling rate, respectively. The cooling capacity is obtained by concatenating the features of the sub-vectors and performing max-min normalization. The resulting vector is then input into a pre-trained autoencoder network to extract deep coupling features, ultimately generating a 256-dimensional low-dimensional dense multidimensional feature vector. This vector not only accurately encodes the critical physical state of the local area, characterized by "extremely high computing power, rapid heat accumulation, and a gap in cooling capacity," but also implicitly includes the coupling constraint that the heat flow vector is limited by the air conditioner fan frequency and cannot be further accelerated. This provides a high-fidelity, multi-dimensional input tensor for the subsequent step S122 to extrapolate the local heat accumulation risk and power consumption drift trend within the future time window.
[0035] Furthermore, multidimensional feature vectors are input into a time-series network and dynamically extrapolated to determine the hysteresis coupling effect between data load and cooling capacity of the cooling equipment. This allows for the determination of the power consumption drift trend and local heat accumulation risk of the data center within a preset time window. Based on the gradient change of the power consumption drift trend and the spatial probability distribution of local heat accumulation risk, a risk prediction situation map containing both time and spatial dimensions is generated. This approach incorporates the overall consideration of inputting multidimensional feature vectors into the time-series network, ensuring the accuracy of the hysteresis coupling effect between data load and cooling capacity. Simultaneously, the dynamic heat propagation characteristics of the data center are introduced to further control the multidimensional feature vectors, thereby improving the accuracy of the risk prediction situation map.
[0036] At this point, the multidimensional feature vector constructed in step S121 is input into the pre-constructed temporal network for dynamic inference based on time steps. The temporal network adopts a composite architecture of spatiotemporal graph convolutional network or long short-term memory network and graph attention mechanism. Therefore, the temporal network is a composite architecture consisting of a bidirectional long short-term memory network (LSTM) layer and a graph attention mechanism layer connected in series. The network takes the multidimensional feature vector constructed in step S121 as the input sequence, and the input time window length is set to the past ten minutes to infer the state changes within the next fifteen-minute time window. The bidirectional LSTM layer contains sixty-four hidden units, which are used to capture the forward and backward dependencies of the multidimensional feature vector in the time dimension and extract the implicit state features that evolve over time.
[0037] The graph attention mechanism layer receives the hidden state of each time step of the LSTM output and uses the topological graph structure constructed in step S112 as the spatial adjacency prior to perform attention-weighted aggregation of the features between each node, so that the network can dynamically focus on the most influential upstream heat source node and downstream cold source node on the heat propagation path. The output layer of the composite network adopts a fully connected structure, and the output dimension is consistent with the input multidimensional feature vector dimension, which corresponds to the predicted features of each spatial node at each future time step, including predicted power consumption, predicted heat flow vector and predicted heat accumulation rate.
[0038] The training of this time-series network was completed offline in a digital twin space. The training dataset consisted of at least one month of historical operating records from a data center, covering various typical operating conditions, including steady-state low load, sudden peak computing power, cooling equipment rotation switching, and local cooling failures. The dataset was divided into training, validation, and test sets in chronological order, with a ratio of approximately 6:2:2. The loss function used was weighted mean squared error, with higher weights given to the prediction errors of the first five minutes within the next 15 minutes to prioritize the accuracy of short-term risk predictions. During training, an adaptive moment estimation algorithm was used for parameter optimization, with an initial learning rate set to 0.1% and a batch size of 32 time segments. An early stopping mechanism was triggered when the validation set loss did not decrease for 20 consecutive rounds. After training, the LSTM forget gate bias, attention weight matrix, and fully connected layer parameters in the network were all fixed for real-time forward inference.
[0039] Using the multidimensional feature vector as the input sequence, the network continuously fits and predicts the transient jump process of IT equipment data load and the gradual adjustment process of cooling output of refrigeration equipment through the time gating mechanism and state transfer equation within the network. During the extrapolation process, the time-series network extracts the power consumption transient features in the data load as the heat source excitation signal, and simultaneously captures the water valve opening change rate and fan response delay in the refrigeration equipment operation characteristics. By comparing the phase difference and amplitude attenuation of the heat source excitation signal and the cooling response signal of the refrigeration system on the time axis, the hysteresis coupling effect between the heat generation power of the data load and the effective cooling capacity of the refrigeration equipment is quantitatively calculated. The hysteresis coupling effect is characterized by the hysteresis time constant and the overshoot / undershoot attenuation coefficient of cooling capacity, thus establishing the transient energy conservation constraint of the thermodynamic system under the dynamic imbalance boundary.
[0040] Based on the determination of the hysteresis coupling effect, the power consumption drift trend and local heat accumulation risk of the data center within a future preset time window are further deduced. At this time, using the multidimensional feature vector at the current moment as the initial state, the hysteresis coupling effect is embedded as a dynamic constraint into the decoder of the time-series network, and the state parameters of each time step within the future preset time window are predicted by rolling forward along the time axis. The power consumption drift trend is determined by integrating and summing the predicted power consumption sequence of future IT load and the predicted power consumption sequence of auxiliary cooling equipment, and extracting the dynamic change rate of power utilization efficiency, which characterizes the degree of deviation of global energy consumption from load fluctuations. The local heat accumulation risk is based on the predicted heat accumulation rate of each spatial grid node within the future preset time window, combined with the specific heat capacity and safe temperature threshold of the node medium, to calculate the remaining time required for the node temperature to reach the critical shutdown temperature. The reciprocal of this remaining time is defined as the local heat accumulation risk index, thereby quantitatively characterizing the overheating hazard of the local space in the future time series.
[0041] After acquiring the power consumption drift trend and local heat accumulation risk, a risk prediction situation map containing both time and spatial dimensions is generated based on the gradient change of the power consumption drift trend and the spatial probability distribution of the local heat accumulation risk. At this point, the first derivative of the predicted power consumption drift trend time series is calculated to extract the gradient change of the power consumption drift trend, thus identifying the inflection point of sudden energy consumption increases. Simultaneously, for the local heat accumulation risk, the grid node risk index output by the time-series network is treated as a random field. Combined with the uncertainties of the physical boundaries, such as sensor errors and airflow perturbations, a Monte Carlo Dropout method is adopted. The method uses either t or Bayesian approximation to calculate the posterior probability of each grid node exceeding the safe thermal accumulation threshold at each future time step, thereby generating a spatial probability distribution of local thermal accumulation risk. Using the physical three-dimensional topology of the data center as a base, the temporal inflection point of the power consumption drift gradient is marked on the time axis. The spatial probability distribution of thermal accumulation risk is mapped onto the physical space grid using voxel rendering. The evolution trajectory in the time dimension and the probability cloud map in the spatial dimension are fused using tensors to generate a four-dimensional risk prediction situation map. The situation map can intuitively present "when, where, and with what probability" thermal overload and energy efficiency degradation occur.
[0042] Specifically, the "256-dimensional multi-dimensional feature vector corresponding to the GPU cluster region, including parameters such as 22kW instantaneous load, 5℃ / min heat accumulation rate, and 45Hz air conditioning frequency" is input into the time-series network. The network dynamically extrapolates based on the physical decay law learned from historical training. When the large model training task causes the rack power consumption to jump from 8kW to 22kW, the time-series network captures the instantaneous surge in the heat generation power of the IT load. However, due to the thermal inertia of the chilled water pipe network and the action time of the water valve actuator, the effective cooling capacity of the row-level air conditioner needs to undergo a delay of about 120 seconds to gradually climb from 25kW to the target value. By comparing the heat generation power step signal and the cooling capacity ramp response signal, the time-series network quantifies the hysteresis coupling effect between the heating and cooling systems in this high-density area, that is, the hysteresis time constant is 120 seconds, and there is a serious under-adjustment phenomenon in cooling capacity during the hysteresis period.
[0043] By embedding the aforementioned lag time constant and cooling capacity under-adjustment constraint into the decoder, the state evolution within a preset time window of the next 15 minutes was extrapolated. The extrapolation results showed that in the first to third minutes, due to the continuous heating of the 22kW heat source and the insufficient cooling capacity of the row-level air conditioner, the local heat accumulation risk index of the grid where the upper GPU in the rack is located rose sharply to 0.85, close to 1.0, which is the critical shutdown level. At the same time, due to the accelerated response of the row-level air conditioner fan to supplement the cooling capacity, the fan power consumption surged, causing the overall PUE of this area to drift from 1.25 to 1.45 in the fifth minute. This confirmed that the power consumption drift trend was a steep rise followed by a gradual decrease, while the local heat accumulation risk was an extreme deterioration in the first stage and a gradual convergence in the second stage.
[0044] The first-order gradient of the predicted power consumption drift trend time series was calculated, and it was found that the gradient reached its maximum positive value at the 3rd minute in the future. This moment was marked as the extreme inflection point of the power consumption surge. At the same time, for the risk of local heat accumulation, Bayesian inference was used to calculate the spatial probability distribution of the temperature of the high-heat cabinet grid exceeding the 85°C safety threshold in the 2nd minute in the future. The results showed that the probability of thermal runaway in the top third area of the cabinet was as high as 92%, and the probability of thermal runaway in the surrounding adjacent cabinets due to the loss of cold air was also 35%. The system fused the "extreme inflection point of power consumption drift gradient, such as T+3min" with the "the above-mentioned spatial probability distribution of thermal runaway, such as the trend of 92% high probability voxels concentrating in the upper part of the cabinet and spreading to adjacent cabinets" into a four-dimensional tensor, and finally generated a risk prediction situation map. This situation map accurately presents the crisis situation that "in the next 2 to 3 minutes, the top and sides of the GPU cluster cabinet will likely experience local hotspot penetration".
[0045] refer to Figure 4 In step S13, the specific steps are as follows: S131: Input the risk prediction situation map into the data twin space and trigger the dynamic recognition of the risk prediction situation map. In the recognition process, multiple risk areas are extracted and labeled through clustering and topological skeleton. In each risk area, the corresponding regional influence weight matrix is constructed along the energy spillover boundary of each risk area. S132: Multiple risk areas undergo alternating iterations under decoupled control variables based on the regional influence weight matrix. In each iteration, the boundary constraints are dynamically adjusted by combining the real-time working data of multiple devices in the data center, thereby outputting the multi-dimensional control objectives corresponding to each risk area, and then transforming the multi-dimensional control objectives into multiple control contents of the data center.
[0046] In the embodiments of this application, the risk prediction situation map is input into the data twin space and the dynamic recognition of the risk prediction situation map is triggered. During the recognition process, multiple risk areas are extracted and labeled through clustering and topological skeleton. In each risk area, a corresponding regional influence weight matrix is constructed along the energy overflow boundary of each risk area. This takes into account the overall consideration of the energy overflow boundary of each risk area and ensures the accuracy of the corresponding regional influence weight matrix.
[0047] At this point, the risk prediction situation map generated in step S122, which includes time and spatial dimensions, is input into the digital twin space, triggering a dynamic identification process for the risk prediction situation map. Upon receiving the risk prediction situation map, the digital twin space performs rigid registration of its four-dimensional tensor data with the high-fidelity physical geometry model and real-time state model within the twin space, synchronizing the spatial coordinate system and time reference. This ensures that the risk probability cloud map and energy consumption drift gradient accurately cover the corresponding geometric entities of the IT cabinet, cooling equipment, and airflow channels. Based on the dynamic threshold triggering mechanism built into the twin space, when the local thermal accumulation risk probability of a specific voxel in the situation map exceeds the set safety threshold or the power consumption drift gradient exceeds the system steady-state constraint boundary within a preset time window, the dynamic identification module is formally triggered to extract potential disaster-causing topologies and physical boundaries. Optionally, the four-dimensional tensor data can be three-dimensional physical spatial coordinates plus a time dimension.
[0048] Multiple risk regions are marked in the digital twin space using clustering and topology skeleton extraction techniques. Then, a density- and spatial connectivity-based clustering method, such as an improved 3D version of DBSCAN, is employed. The posterior probability value of thermal runaway and the power consumption drift gradient magnitude in the situation map are used as clustering weights to cluster voxel sets in the physical space that continuously exceed risk limits and are spatially adjacent into independent initial risk voxel clusters, completing the initial spatial segmentation of the risk regions. For each initial risk voxel cluster, its topology skeleton is extracted using 3D morphological refinement or distance transformation. This topology skeleton abstracts the irregular 3D risk voxel clusters into centerlines and key topological nodes representing the main direction of heat spread. Based on the connectivity paths represented by the topology skeleton and the spatial contours of the initial risk voxel clusters, multiple risk regions with clear spatial boundaries and core spread skeletons are marked and delineated on the 3D geometric model of the digital twin space. These risk regions not only cover the physical equipment where heat accumulation has occurred but also include potential spread channels driven by heat flow vectors.
[0049] In each risk region, a corresponding regional influence weight matrix is constructed along the energy spillover boundary of each risk region. At this time, for any marked risk region, the outermost voxel set of its three-dimensional spatial contour is extracted as the energy spillover boundary. The energy spillover boundary is the physical interface through which heat and cold are exchanged with the surrounding normal low-temperature region or adjacent risk region. The energy spillover flux of each discrete interface element on the boundary is calculated. The energy spillover flux is jointly determined by the normal component of the heat flow vector and the dynamic distribution parameter of thermal resistance at the element, which characterizes the intensity of heat radiation or convection escape across the boundary. Taking the risk region as the center and the adjacent grids or adjacent risk regions affected by its spillover energy as the target objects, the energy spillover flux of each interface element is projected and integrated according to the spatial topological connection relationship to construct the regional influence weight matrix. The element values in the matrix characterize the degree of thermal interference and cold plunder weight of the risk region to a specific adjacent region, providing quantitative constraint parameters for decoupling cross-regional conflicts in the subsequent iterative process of multiple risk regions.
[0050] Specifically, the digital twin space receives a risk prediction situation map, which shows that in the next 2 to 3 minutes, the probability of thermal runaway in the top third of the rack where the GPU cluster is located is as high as 92%, and the power consumption drift gradient will reach an extreme inflection point in the 3rd minute. The twin space accurately overlays this probability cloud map and gradient curve onto the 3D geometric model of the 20kW high-density rack and the adjacent cold aisle enclosed space. When the system detects that the probability of thermal runaway of the voxel where the GPU node is located at the top of the rack far exceeds the set safety threshold, it immediately triggers the dynamic identification process; the probability of thermal runaway of the voxel where the GPU node is located at the top of the rack is 92%; the set safety threshold is 50%.
[0051] Clustering and skeleton extraction were performed on the voxels exceeding the limit. Using the probability of thermal runaway and the power consumption drift gradient as weights, voxels with continuous high heat at the top of the rack were clustered into "primary risk voxel clusters". At the same time, voxels with a sharp reduction in airflow and abnormal temperature rise at the bottom of adjacent racks due to strong suction of cold air were clustered into "secondary risk voxel clusters". Three-dimensional morphological refinement extraction was performed on the "primary risk voxel clusters" to obtain a topological skeleton extending from the top GPU with the most intense heat along the rear door of the rack to the return air ceiling. Based on this skeleton and the voxel cluster outline, "primary risk area: covering the upper and rear thermal return channel of the GPU rack" and "secondary risk area: covering the lower air inlet side of the adjacent low load rack" were marked in the twin space, clearly defining the core location and spread direction of the crisis.
[0052] For the marked "primary risk area", its outermost voxel is extracted as the energy overflow boundary. Due to the strong negative pressure generated by the full-speed operation of the server fans inside the high-density rack, it is calculated that the interface micro-element on this boundary facing the cold aisle perforated floor has a large heat flux vector normal component and extremely low thermal resistance, resulting in a strong convection suction effect. The interface micro-element facing the rack side panel, on the other hand, exhibits high-temperature radiation overflow flux. The system projects and quantifies the overflow flux in the above different directions onto adjacent areas to construct the regional influence weight matrix of the "primary risk area". The high-weight elements in this matrix clearly point to the cold source area at the bottom of the cold aisle and the adjacent "secondary risk area", indicating that the "primary risk area" not only accumulates its own heat, but also plunders the cold energy that originally belonged to the "secondary risk area" through strong airflow suction. This cross-regional heat plunder weight coefficient quantification provides the necessary conflict focus and decoupling basis for the alternating iterative control of multiple risk areas in the subsequent step S13.
[0053] Furthermore, multiple risk areas undergo alternating iterations under decoupled control variables based on the regional influence weight matrix. In each iteration, the boundary constraints are dynamically adjusted by combining the real-time working data of multiple devices in the data center, thereby outputting multi-dimensional control objectives corresponding to each risk area. This multi-dimensional control objectives are then transformed into multiple control contents of the data center.
[0054] At this point, the multiple risk areas marked in step S131 are solved iteratively under decoupled control variables based on their corresponding regional influence weight matrices. The control variables across the entire data center domain include "IT-side load scheduling variables, such as virtual machine migration ratio and computing power frequency reduction" and "cooling-side cooling capacity supply variables, such as air conditioner fan frequency and chilled water valve opening." Since multiple risk areas are spatially adjacent and have thermal flow coupling, applying control intervention to a certain area will inevitably affect adjacent areas through energy overflow boundaries. Therefore, it is necessary to orthogonalize and decouple the global control variables based on the regional influence weight matrix. This decoupling allocation projects the highly coupled global control space onto the local subspace corresponding to each risk area. Alternating control is then initiated. In the iterative process, in each iteration, the control variables and thermal states of all regions except the current solution region are fixed as known quantities. The control variables of the current region are used as independent variables to be solved. The thermal disturbances transmitted from adjacent regions through the boundary are quantified by the regional influence weight matrix as constant input. The control variable values that reduce the local heat accumulation risk to below the safety threshold and achieve optimal energy consumption are solved independently for the current region. After the current region is solved, the updated control variables and thermal states are passed to the next risk region for the same solution. This process is repeated until the changes in the control variables of all risk regions converge to the preset minimum range, thereby eliminating the conflicts of cold energy competition and hot spot drift caused by parallel control of multiple regions.
[0055] In each iteration of the above alternating process, the boundary constraints need to be dynamically adjusted based on the real-time operating data of multiple devices within the data center. These boundary constraints include the Service Level Agreement (SLA) computing power constraints of the IT devices and the physical security operation constraints of the cooling devices. When iteratively solving for the theoretical values of the control variables, it is essential to ensure that these theoretical values do not exceed the aforementioned boundaries. Simultaneously, due to the transient fluctuations in the operating status of physical devices, the system collects real-time operating data such as the current actual CPU / GPU utilization of the IT devices, virtual machine migration progress, and the current water supply pressure and maximum fan speed of the cooling devices. This real-time operating data is used to calculate the available computing power redundancy on the IT side and the maximum cooling capacity margin that the cooling side can provide at the current moment. This redundancy space and cooling capacity margin are then dynamically updated to the boundary constraints in real-time as upper and lower limits. This ensures that the control variables solved through alternating iterations not only meet thermodynamic defect elimination requirements but are also strictly constrained by the transient executable capabilities of the physical devices, avoiding the output of unexecutable instructions that exceed the physical limits of the devices or violate the business SLA.
[0056] After alternating iterative convergence and dynamic correction of boundary constraints, multi-dimensional control objectives corresponding to each risk area are output, and then these multi-dimensional control objectives are transformed into multiple control contents of the data center. At this time, the multi-dimensional control objectives include target temperature and target heat flux vector distribution in the thermodynamic dimension, target PUE coefficient in the energy consumption dimension, and target load allocation rate in the computing power dimension. The multi-dimensional control objectives are feasible solution sets after decoupling and constraint verification. Through preset control semantic mapping rules, the abstract multi-dimensional control objectives are decoded into control contents that can be executed by specific devices. The control contents include cooling control instruction sets and IT control instruction sets, that is, mapping the target temperature and target PUE to the set air outlet temperature value and fan frequency set value of specific row-level air conditioners, and mapping the target load allocation rate to the instance identifier and quota adjustment instruction for specific virtual machines or containers migrating from the source host to the target host, thereby completing the transformation from twin space mathematical solution to physical space execution instructions.
[0057] Specifically, for the "primary risk area: upper part of the 20kW high-heat cabinet and heat return channel" and "secondary risk area: lower part of adjacent cabinets whose cooling capacity is being robbed" marked in step S131, the system decouples the control variables based on the strong airflow suction coupling relationship revealed by the regional influence weight matrix. The global control variables are decomposed into "local variables of the primary risk area, primary control: increase the cooling capacity of the nearest row-level air conditioner, secondary control: reduce the frequency of GPU computing power" and "local variables of the secondary risk area, primary control: increase the bottom air supply, secondary control: appropriately migrate non-core loads". In the first round of alternating iteration, the system fixes the state of the secondary risk area and solves for the primary risk area to determine the frequency of the nearest row-level air conditioner fan that needs to be adjusted. The theoretical solution was increased to 55Hz. This solution was used as known conditions in the solution of the secondary risk region. It was found that the overall negative pressure of the cold aisle increased due to the speed increase of the main area fan, which further worsened the original cold distribution in the secondary risk region. Therefore, the opening ratio of the bottom perforated plate was significantly increased in the solution of the secondary risk region to compensate for the cold. In the second iteration, the compensation action of the secondary risk region was fed back to the main risk region. The system found that the local hot spots in the main region could be relieved by the cold source added at the bottom. Therefore, the frequency of the main area fan was adjusted down to 50Hz to avoid excessive suction. After three alternating iterations, the rate of change of the control variables in the two regions approached zero and converged, completely solving the control deadlock caused by the airflow competition between the two regions.
[0058] During each iteration of the solution process, the system incorporates real-time operating data from the equipment to correct boundary constraints. For example, when the iterative method calculates that the frequency of the row-level air conditioning fan in the main risk area needs to be increased to 60Hz to quickly suppress hotspots, real-time collected equipment data shows that the current of the air conditioning inverter has reached 95% of its rated limit, and the water supply pipeline pressure is fluctuating slightly. Based on this, the system dynamically corrects the physical safety operation constraint boundary of the refrigeration equipment to 55Hz, forcing the iterative method to find a new solution to compensate for the IT side's computing power reduction under the hard constraint of 55Hz. At the same time, real-time monitoring shows that the GPU cluster in the main risk area is performing the critical gradient synchronization stage of large model training, with a current computing power utilization rate of 98%. Based on this, the system dynamically corrects the IT equipment SLA computing power constraint boundary, prohibiting the execution of frequency reduction or virtual machine migration operations within this time window, compressing the feasible domain of the IT side control variables to zero at this moment, and ensuring the absolute priority of business logic.
[0059] After the iteration converges and real-time boundary verification, the system outputs the multi-dimensional control objectives for the main risk area as follows: the local maximum temperature is reduced to 32℃, the PUE contribution value is reduced by 0.05, and the computing load remains unchanged. The control objectives for the secondary risk area are: the local air intake temperature is maintained at 24℃, and the bottom cold air distribution rate is increased by 20%. The control semantic mapping rules transform the above objectives into specific control content: for the main risk area, a cooling control command is output [lower the air outlet temperature of the row air conditioner at the top of the rack to 15℃, set the fan frequency to 55Hz, and increase the water valve opening to 80%], and for the secondary risk area, a cooling control command is output [adjust the opening of the electric regulating valve of the corresponding floor perforated plate to 100%]. After the critical gradient synchronization phase is completed, an IT control command is added [online migration of the two virtual machines in the main risk area that undertake non-critical inference tasks to the rack in the secondary risk area]. This forms a refined control content set that combines the redistribution of cooling supply and the spatiotemporal staggering of computing load, which is used to generate adaptive control paths in subsequent steps.
[0060] refer to Figure 5 In step S14, the specific steps are as follows: S141: Obtain the distribution map of the data center, load multiple control contents into the distribution map of the data center, construct a spatiotemporal dynamic graph with device nodes as vertices and hot-related paths as edges, and perform spatiotemporal alignment in combination with the current work sequence of the data center to determine the temporal dependency markers of each node in the spatiotemporal dynamic graph, and determine multiple adaptive control paths that satisfy the execution order and resource mutual exclusion constraints by tracing along the temporal dependency markers. S142: Perform global optimization on the multiple adaptive control paths to obtain a Pareto optimal solution set, extract the path with the minimum conflict cost with the current working content of multiple devices from the Pareto optimal solution set, and use it as the target control path. Determine the adaptive control strategy of the data center based on the identification of the target control path.
[0061] In the embodiments of this application, a distribution map of the data center is obtained, multiple control contents are loaded into the distribution map of the data center, a spatiotemporal dynamic graph is constructed with device nodes as vertices and hot-related paths as edges, and spatiotemporal alignment is performed in combination with the current working sequence of the data center to determine the temporal dependency markers of each node in the spatiotemporal dynamic graph. By tracing along the temporal dependency markers, multiple adaptive control paths that satisfy the execution order and resource mutual exclusion constraints are determined. This approach takes into account the overall consideration of temporal dependency markers and ensures the accuracy of multiple adaptive control paths that satisfy the execution order and resource mutual exclusion constraints.
[0062] At this point, the physical logic distribution map of the data center is obtained, and the multiple control contents output in step S132 are loaded onto the distribution map of the data center. A spatiotemporal dynamic graph is constructed with device nodes as vertices and thermally related paths as edges. At this point, the distribution map of the data center includes the spatial coordinates of IT cabinets, the deployment location of cooling equipment, and the physical topology of power supply and chilled water pipelines. The multiple control contents are attached as dynamic attributes to the corresponding physical device entities in the distribution map, thereby transforming the static physical distribution map into a dynamic weighted graph containing control intentions. Further, with the controlled IT equipment and cooling equipment as vertices of the graph, and the heat flow vector direction and thermal resistance coupling relationship characterized by the dynamic heat propagation characteristics established in steps S111 and S112 as directed edges, a spatiotemporal dynamic graph is constructed. The edge weights of the directed edges not only include the thermal conductivity coefficient in the spatial dimension, but also embed the response time difference caused by thermal inertia when the two devices execute control commands, thereby unifying the spatial thermal correlation and temporal evolution logic in the graph structure. The multiple control contents include IT load scheduling commands and cooling parameter adjustment commands.
[0063] After constructing the spatiotemporal dynamic graph, spatiotemporal alignment is performed in conjunction with the current work sequence of the data center to determine the temporal dependency markers of each node in the spatiotemporal dynamic graph. At this time, the current work sequence includes the IT business scheduling, cooling system rotation strategy, and power supply and distribution load balancing strategy of the data center's existing network. The control content attached to each node in the spatiotemporal dynamic graph is compared and inserted for consistency with the current work sequence on the time axis and logical link. Based on the alignment results, the preconditions and postconditions for the control command to take effect are extracted. The temporal causal relationship between nodes is identified using the Directed Acyclic Graph (DAG) parsing technique, such as "cold source increment precondition" or "heat source elimination precondition". Each node in the spatiotemporal dynamic graph is marked with a temporal dependency marker containing the temporal sequence number and logical triggering condition. The temporal dependency marker strictly defines the time phase of the controlled equipment's action to prevent control failure or secondary disasters caused by the reversal of the execution sequence.
[0064] After determining the timing dependency markers, multiple adaptive control paths satisfying the execution order and resource mutual exclusion constraints are determined by tracing the timing dependency markers. At this time, taking the vertices assigned control content in the spatiotemporal dynamic graph as the starting and ending points, path backtracking and forward deduction are performed along the directed edges indicated by the timing dependency markers. During the deduction process, a constraint satisfaction mechanism is introduced, which includes execution order constraints and resource mutual exclusion constraints. The execution order constraints are determined by the logical triggering conditions of the timing dependency markers, ensuring that downstream instructions must be executed after the effects of upstream instructions are manifested. The resource mutual exclusion constraints are based on the exclusive attributes of physical devices, ensuring that the same device cannot respond to conflicting control instructions at the same time, such as simultaneously executing frequency increase and decrease, or simultaneously acting as a virtual machine migration source and migration source. By embedding the dual constraints in the graph search method to prune the feasible solution space, multiple sets of directed paths from control initiation to goal achievement are finally extracted, which are multiple adaptive control paths. These paths not only plan the flow trajectory of control instructions, but also contain strict time waiting and mutual exclusion lock mechanisms.
[0065] Specifically, the system obtains the physical distribution map of the closed cold aisle area of the data center, loads the control content generated in step S132 (for the main risk area: increase the frequency of the row air conditioning fan to 55Hz / open the water valve to 80%; for the secondary risk area: open the perforated plate electric valve to 100% and migrate non-critical virtual machines to the secondary area cabinet) onto the distribution map, and constructs a spatiotemporal dynamic map based on the aforementioned determined dynamic heat propagation characteristics, with the 20kW high-heat cabinet node, adjacent row air conditioning nodes, perforated plate valve nodes and target migration cabinet nodes as vertices. Among them, a directed edge representing strong convection suction is established between the high-heat cabinet node and the row air conditioning node, and the edge weight includes a 2-minute thermal inertia lag time parameter. A directed edge representing cooling supply is also established between the high-heat cabinet node and the perforated plate node.
[0066] The system retrieves the current IT work sequence and cooling rotation sequence of the data center, and aligns the loaded control content with this work sequence in time and space. The alignment reveals that if virtual machine migration is executed immediately, it will disrupt the current gradient synchronization and cause business losses, while fan frequency increase can be executed immediately. Based on this, the nodes of the time-space dynamic graph are marked with time-series dependency tags: row-level air conditioning nodes and perforated plate nodes are marked as [T+0min, execute immediately], "high-heat cabinet nodes, as virtual machine migration sources" are marked as [T+2min, triggered after checkpoint saving], and "target migration cabinet nodes, as migration sources" are marked as [T+2.1min, accepted after migration source release]. This time-series dependency tag clearly solidifies the time-series logic of "first supplying cooling to suppress hotspots, then scheduling computing power to strip away heat sources". Optionally, the current IT work sequence of the data center is a large model training task that is in the critical gradient synchronization period, and will enter an interruptible checkpoint saving period after 2 minutes.
[0067] The system traces and extrapolates along the aforementioned time-dependent markers to generate an adaptive control path. When the extrapolation reaches the virtual machine migration node at time [T+2min], a resource mutual exclusion constraint check is triggered. Since the current work sequence shows that the target rack in the secondary risk area is performing a local storage snapshot backup, its I / O bandwidth and CPU resources are in a mutually exclusive state, and it cannot accept new load as a migration source at the same time. The constraint satisfaction mechanism therefore blocks this direct migration path and forces a search for a bypass solution along the spatiotemporal dynamic graph, ultimately determining an adaptive control path that satisfies dual constraints: "Path 1: Cooling-first path". Row-level air conditioning fan frequency increase (T+0min) → perforated plate valve fully open (T+0.5min, waiting for air pressure balance) → high-heat cabinet hotspot relief; "Path 2: computing power off-peak migration path": high-heat cabinet save checkpoint (T+2min) → target cabinet end snapshot release resources (T+2.5min, satisfy mutual exclusion constraints) → virtual machine online migration (executed after T+2.5min). The above two paths together constitute a set of multi-adaptive control paths that ensure business SLA and equipment safety, strictly adhere to timing and mutual exclusion logic, and provide an executable timing blueprint for the final construction of global control strategy.
[0068] Furthermore, global optimization is performed on the multiple adaptive control paths to obtain a Pareto optimal solution set. The path with the minimum conflict cost with the current working content of multiple devices is extracted from the Pareto optimal solution set and used as the target control path. The adaptive control strategy of the data center is determined based on the identification of the target control path, which is compatible with the overall consideration of the target control path. At the same time, multiple control contents are further controlled, and the current working content of multiple adaptive control paths and multiple devices are fully considered, thereby improving the accuracy of the adaptive control strategy of the data center.
[0069] At this point, global optimization is performed on the multiple adaptive control paths determined in step S141 to obtain a Pareto optimal solution set. These multiple adaptive control paths are then used as inputs to feasible solutions in a multi-objective optimization approach. A multi-objective evaluation function is constructed with the optimization objectives of minimizing global energy consumption in the data center, maximizing hotspot elimination rate, and minimizing IT service performance disturbance. Since the above optimization objectives have mutually constraining game-like relationships in the actual physical system—for example, rapidly eliminating hotspots requires significantly increasing fan frequency, leading to increased global energy consumption, or reducing energy consumption requires load migration, leading to service performance disturbance—a fast non-dominated sorting genetic method with elite retention strategy (NSGA-II) or a multi-objective particle swarm optimization method is used for global iterative search. During the search process, by performing crossover, mutation, and non-dominated sorting on individuals in the population, dominated inferior solutions are eliminated generation by generation. Finally, a set of compromise solutions that cannot further improve the performance of other objectives in the multi-dimensional objective space without compromising any objective is obtained; this is the Pareto optimal solution set. Each path in the Pareto optimal solution set represents an adaptive control scheme under a certain extreme trade-off.
[0070] After obtaining the Pareto optimal solution set, the path with the minimum conflict cost with the current work content of multiple devices is extracted from the solution set and used as the target control path. At this time, for each candidate path in the Pareto optimal solution set, a conflict cost quantification model is constructed. The model compares and verifies the execution sequence of each control instruction and resource occupation requirements in the candidate path with the established operating status of the devices in the current work sequence of the data center step by step. The conflict cost covers "equipment mechanical wear cost, such as fatigue loss caused by frequent start-stop or large-step load changes of compressors and fans", "business SLA breach cost, such as critical task response delay caused by computing power scheduling", and "safety margin exceeding limit cost, such as the risk of water hammer effect in refrigerant pipeline caused by sudden change in water valve opening". By normalizing and weighting the above cost dimensions, the conflict cost index of each candidate path relative to the current network steady state is calculated. The candidate path with the minimum conflict cost index is selected from the Pareto optimal solution set and defined as the target control path that takes into account the optimality of multiple objectives and the lowest intrusion of the current network, ensuring the smooth implementation and extremely high availability of the adaptive control strategy.
[0071] After determining the target control path, an adaptive control strategy for the data center is determined based on the identification of the target control path. At this time, the target control path is parsed and its features are identified, and the timing dependency markers, device action instruction sets, and resource mutex lock states contained in the path are extracted and transformed into a standard strategy script that can be parsed and executed by the data center infrastructure digital twin management and control platform. The adaptive control strategy not only includes a static sequence of control target setpoints, but also embeds dynamic closed-loop feedback adjustment logic and an exception rollback mechanism. That is, in the actual execution process, if the real-time thermal response after a certain step is detected to deviate from the expected threshold, the strategy will automatically trigger local fine-tuning or rollback to the previous steady-state snapshot. In this way, the optimal path obtained by mathematical domain optimization is transformed into an adaptive control strategy with engineering executability, timing rigor, and fault tolerance robustness.
[0072] Specifically, the system inputs the aforementioned multiple adaptive control paths into the NSGA-II method for global optimization, setting the optimization objectives as [minimizing PUE in the next 15 minutes], [minimizing the time for eliminating local hotspots], and [minimizing the decrease in GPU computing power utilization]. After iteration, it outputs a Pareto optimal solution set. This solution set shows that: if the goal is to eliminate hotspots extremely quickly, such as cooling down within 1 minute, then it is necessary to accept the degradation of PUE and a brief drop in GPU computing power, which is the characteristic of the first path; if the goal is to achieve both optimal PUE and computing power, then it is necessary to tolerate the slow decay of hotspots, which is the characteristic of the second path. There is no absolute dominance relationship among all paths in this solution set, which constitutes the Pareto front of a multi-objective game. Optionally, multiple adaptive control paths include the first path: relying on significantly increasing the frequency of row-level air conditioning fans and reducing the outlet temperature of chilled water valves for strong cooling suppression; and the second path: relying on an equal amount of incremental cold source plus online migration of virtual machines to remove heat sources.
[0073] The system calculates the conflict cost for the Pareto optimal solution set. At this time, the system obtains the real-time working content of the current equipment: the row-level air conditioning inverter corresponding to the 20kW high-heat cabinet is operating under high load, and the chiller unit has just completed a load adjustment. The pipeline system is in a period of pressure fluctuation sensitivity. The comparison shows that the first path requires the row-level air conditioning fan frequency to be instantly increased from 45Hz to 60Hz, and at the same time the opening of the chilled water electric valve is increased by 30%. The pipeline water hammer risk cost and inverter overload loss cost caused by this operation are extremely high. At the same time, since the large model is in a critical calculation period, strong cold airflow disturbance is likely to cause the server's internal fan speed to fluctuate drastically, generating additional noise and wear. The second path requires the fan to be gradually increased to 50Hz, and the virtual machine migration is performed after 2 minutes. This path has a small deviation from the current air conditioning system steady-state operation trajectory, and waiting for the checkpoint to be completed before migration avoids the business SLA breach cost. After normalization calculation, the conflict cost index of the second path is much lower than that of the first path. Based on this, the system extracts the second path as the target control path.
[0074] The system identifies the characteristics and instruction sequence of the target control path and compiles them into an adaptive control strategy for the data center. The strategy explicitly outputs: [T+0s] instructs the row-level air conditioning fan frequency to gradually increase to 50Hz and locks the inverter upper limit to prevent overload; [T+30s] instructs the perforated plate electric valve to adjust the opening to 100%; [T+120s] waits for the GPU cluster checkpoint to be saved and then issues the instruction to migrate the virtual machine; [T+150s] dynamically adjusts the fan frequency back to 45Hz based on the temperature drop of the temperature sensor on the top of the rack. At the same time, the strategy has an embedded rollback mechanism: if [T+90s] the cold aisle temperature is detected to rise instead of fall, the waiting logic is immediately stopped and the backup chiller is forcibly triggered to start. In this way, an adaptive control strategy is formed that is deeply adapted to the physical state and business operation rhythm of the current high-density computing power cluster in the data center, smooth and imperceptible, and has a safety net.
[0075] Please see Figure 6 The adaptive control system for a data twin-based data center is applied to the aforementioned adaptive control method for a data twin-based data center; the adaptive control system for a data twin-based data center includes: The digital twin module 21 is used to collect multiple sensing data from different dimensions of the data center, and combine them with the multi-physics coupling network of the digital twin space to determine the spatial heat conduction path, and determine the dynamic heat propagation characteristics of the data center based on the topological mapping of the spatial heat conduction path. The risk prediction situation map module 22 is used to collect the data load of the data center and construct a multi-dimensional feature vector by combining the dynamic heat propagation characteristics and the operating characteristics of the cooling equipment. Based on the dynamic deduction of the multi-dimensional feature vector, the power consumption drift trend and local heat accumulation risk of the data center in the future preset time window are determined, thereby generating a risk prediction situation map. The control content module 23 is used to input the risk prediction situation map into the data twin space, mark multiple risk areas during the dynamic identification process, iterate over multiple risk areas alternately, and determine multiple control contents by combining the real-time working data of multiple devices in the data center. The adaptive control module 24 is used to load multiple control contents into the distribution map of the data center, determine multiple adaptive control paths in combination with the current working sequence of the data center, optimize the multiple adaptive control paths, and construct the adaptive control strategy of the data center in combination with the current working contents of multiple devices.
[0076] It should be noted that although multiple modules are mentioned in the detailed description above, this division is not mandatory; in fact, according to the embodiments of this disclosure, the features and functions of two or more modules or described above can be embodied in one module; conversely, the features and functions of one module described above can be further divided into multiple modules to be embodied.
[0077] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein; this application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein; the specification and embodiments are to be considered exemplary only.
[0078] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. An adaptive control method for a data center based on data twins, characterized in that, include: In this data center, multiple sensing data from different dimensions are collected, and the spatial heat conduction path is determined by combining the multi-physics coupling network of the digital twin space. The dynamic heat propagation characteristics of the data center are determined based on the topological mapping of the spatial heat conduction path. Data load data of data centers is collected, and multi-dimensional feature vectors are constructed by combining dynamic heat propagation characteristics and cooling equipment operation characteristics. Based on the dynamic deduction of multi-dimensional feature vectors, the power consumption drift trend and local heat accumulation risk of data centers in the future within a preset time window are determined, thereby generating a risk prediction situation map. The risk prediction situation map is input into the data twin space, and multiple risk areas are marked during the situation identification process. The multiple risk areas are iterated alternately, and multiple control contents are determined by combining the real-time working data of multiple devices in the data center. Multiple control contents are loaded into the distribution map of the data center, and multiple adaptive control paths are determined in combination with the current working sequence of the data center. The multiple adaptive control paths are optimized, and the adaptive control strategy of the data center is constructed in combination with the current working contents of multiple devices.
2. The adaptive control method for a data center based on data twins according to claim 1, characterized in that, In this data center, multiple sensing data from different dimensions are collected, and the spatial heat conduction path is determined by combining it with the multi-physics coupling network of the digital twin space. Based on the topological mapping of this spatial heat conduction path, the dynamic heat propagation characteristics of the data center are determined, including: The data center is monitored in real time. Multiple sensing data from different dimensions are collected through a multimodal sensor network and mapped to a digital twin space. Through a pre-built multiphysics coupling network in the digital twin space, the airflow field, thermal radiation field and electromagnetic field are jointly analyzed across scales to extract the spatial heat conduction boundary conditions.
3. The adaptive control method for a data center based on data twins according to claim 2, characterized in that, The process of collecting multiple sensing data points across different dimensions within the data center, combining them with a multi-physics coupling network of a digital twin space to determine spatial heat conduction paths, and determining the dynamic heat propagation characteristics of the data center based on the topological mapping of these spatial heat conduction paths, further includes: A dynamically updated spatial heat conduction path is constructed based on the spatial heat conduction boundary conditions; the spatial heat conduction path is topologically mapped and graph structure features are extracted; further, the dynamic heat propagation characteristics of the data center are determined by combining the transient heat source distribution information that evolves over time. The dynamic heat propagation characteristics include the direction of heat flow vector, dynamic distribution of thermal resistance, and heat accumulation rate.
4. The adaptive control method for a data center based on data twins according to claim 1, characterized in that, The system collects data load information from the data center and constructs a multi-dimensional feature vector by combining dynamic heat propagation characteristics and the operating characteristics of cooling equipment. Based on the dynamic deduction of the multi-dimensional feature vector, it determines the power consumption drift trend and local heat accumulation risk of the data center within a preset time window in the future, thereby generating a risk prediction situation map, including: The system acquires the operating characteristics of the data center's cooling equipment, and aligns the data load with the dynamic heat propagation characteristics and the operating characteristics of the cooling equipment in a spatiotemporal manner, thereby constructing a multidimensional feature vector in the feature-level fusion stage.
5. The adaptive control method for a data center based on data twins according to claim 4, characterized in that, The process of collecting data load information from the data center and constructing a multi-dimensional feature vector by combining dynamic heat propagation characteristics and the operating characteristics of cooling equipment, and determining the power consumption drift trend and local heat accumulation risk of the data center within a preset time window based on the dynamic deduction of the multi-dimensional feature vector, thereby generating a risk prediction situation map, also includes: By inputting multidimensional feature vectors into a time-series network and performing dynamic inference, the hysteresis coupling effect between data load and cooling capacity of cooling equipment is determined. This leads to the determination of power consumption drift trend and local heat accumulation risk in the data center within a preset time window. Based on the gradient change of power consumption drift trend and the spatial probability distribution of local heat accumulation risk, a risk prediction situation map containing time and spatial dimensions is generated.
6. The adaptive control method for a data center based on data twins according to claim 1, characterized in that, The process involves inputting a risk prediction situation map into a data twin space, marking multiple risk areas during dynamic identification, iteratively processing these risk areas, and combining this with real-time operational data from multiple devices within the data center to determine multiple control parameters, including: The risk prediction situation map is input into the data twin space, and the dynamic recognition of the risk prediction situation map is triggered. During the recognition process, multiple risk areas are extracted and labeled through clustering and topological skeleton. In each risk area, the corresponding regional influence weight matrix is constructed along the energy spillover boundary of each risk area.
7. The adaptive control method for a data center based on data twins according to claim 6, characterized in that, The process of inputting the risk prediction situation map into the data twin space, marking multiple risk areas during dynamic identification, iterating over multiple risk areas alternately, and determining multiple control contents by combining real-time working data from multiple devices within the data center, also includes: Multiple risk areas are subjected to alternating iterations under decoupled control variables based on the regional influence weight matrix. In each iteration, the boundary constraints are dynamically adjusted by combining the real-time working data of multiple devices in the data center, thereby outputting the multi-dimensional control objectives corresponding to each risk area, and then transforming the multi-dimensional control objectives into multiple control contents of the data center.
8. The adaptive control method for a data center based on data twins according to claim 1, characterized in that, The process involves loading multiple control contents into the data center's distribution map, determining multiple adaptive control paths based on the data center's current work sequence, optimizing these adaptive control paths, and constructing an adaptive control strategy for the data center based on the current work contents of multiple devices. This includes: Obtain the distribution map of the data center, load multiple control contents into the distribution map of the data center, construct a spatiotemporal dynamic graph with device nodes as vertices and hot-related paths as edges, and perform spatiotemporal alignment with the current work sequence of the data center to determine the temporal dependency marker of each node in the spatiotemporal dynamic graph. By tracing along the temporal dependency marker, determine multiple adaptive control paths that satisfy the execution order and resource mutual exclusion constraints.
9. The adaptive control method for a data center based on data twins according to claim 8, characterized in that, The process of loading multiple control contents into the data center's distribution map, determining multiple adaptive control paths based on the data center's current work sequence, optimizing these adaptive control paths, and constructing the data center's adaptive control strategy based on the current work contents of multiple devices further includes: Global optimization is performed on the multiple adaptive control paths to obtain a Pareto optimal solution set. The path with the minimum conflict cost with the current working content of multiple devices in the Pareto optimal solution set is extracted and used as the target control path. The adaptive control strategy of the data center is determined based on the identification of the target control path.
10. An adaptive control system for a data center based on data twins, characterized in that, The adaptive control system for the data twin-based data center is applied to the adaptive control method for the data twin-based data center as described in any one of claims 1-9; the adaptive control system for the data twin-based data center includes: The digital twin module is used to collect multiple sensing data from different dimensions in the data center, and combine them with the multi-physics coupling network of the digital twin space to determine the spatial heat conduction path. Based on the topological mapping of the spatial heat conduction path, the dynamic heat propagation characteristics of the data center are determined. The risk prediction situation map module is used to collect the data load of the data center and construct a multi-dimensional feature vector by combining dynamic heat propagation characteristics and cooling equipment operation characteristics. Based on the dynamic deduction of the multi-dimensional feature vector, the power consumption drift trend and local heat accumulation risk of the data center in the future preset time window are determined, thereby generating a risk prediction situation map. The control content module is used to input the risk prediction situation map into the data twin space, mark multiple risk areas during the dynamic identification process, iterate over multiple risk areas alternately, and determine multiple control contents by combining the real-time working data of multiple devices in the data center. The adaptive control module is used to load multiple control contents into the distribution map of the data center, determine multiple adaptive control paths in combination with the current working sequence of the data center, optimize the multiple adaptive control paths, and construct the adaptive control strategy of the data center in combination with the current working contents of multiple devices.