A management data loading platform based on three-dimensional layout and intelligent optimization algorithm
By constructing a three-dimensional real-time state field and intelligent optimization algorithm, the problems of inaccurate heat generation prediction and incomplete deployment decisions in data centers are solved, precise thermal management and adaptive control under high-density mixed loads are achieved, and the operating stability and energy efficiency of data centers are improved.
Patent Information
- Application Number
- CN202511031700.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-07-25
AI Technical Summary
Existing technologies in data center thermal management and task scheduling suffer from inaccurate heat generation predictions, incomplete deployment decisions, and adaptive control faults. Especially in high-density mixed load scenarios, they are unable to effectively cope with the impact of electrical-thermal coupling and environmental fluctuations.
By constructing a three-dimensional real-time state field and combining it with an intelligent optimization algorithm, the total harmonic distortion rate of the power supply and the power consumption are monitored in real time. The electro-thermal coupling model and principal component analysis method are used to fuse risk indices, achieving closed-loop feedback control and improving the accuracy and adaptability of deployment decisions.
It improves the operational stability and energy efficiency of the data center, achieves high-fidelity heat generation prediction and optimization decision-making, and improves the overall management level of the data center.
Smart Images

Figure CN120523612B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data management technology, and in particular to a management data loading platform based on three-dimensional layout and intelligent optimization algorithm. Background Art
[0002] Supercomputing centers or large financial cloud data centers typically deploy a mix of two distinct workloads: 1) HPC (High-Performance Computing) tasks, such as weather forecasting and gene sequencing, are characterized by long, intensive CPU / GPU usage, creating significant heat and harmonic generation. 2) Low-latency tasks, such as real-time trading risk control and online payments, place extremely stringent requirements on temperature stability and power purity in the operating environment. Existing technologies, when deploying tasks, fail to foresee how the harmonics generated by HPC tasks will degrade power quality, further generating "excessive heat" within the HPC tasks and adjacent servers, leading to unpredictable hotspots and threatening the stability of environmentally sensitive financial tasks. The widespread adoption of intelligent power distribution units (PDUs) enables real-time power consumption monitoring at the server level, providing schedulers with a basis for energy-aware decision-making. However, although the technologies in both thermal management and task scheduling have reached a fairly high level of maturity, existing technical frameworks generally treat the two as independent systems for optimization, and the interaction between them is limited to using the server's computing power consumption as a static heat source input to the thermal management system. This decoupled management paradigm has shown its inherent limitations in the application scenarios of modern high-density mixed-load data centers.
[0003] The existing technology has a publication number of CN114676862A, titled "A Visualized Operation and Maintenance Management Method and System for a Data Center," which includes: obtaining equipment and facility information, pipeline information, and spatial layout information of the data center to obtain a three-dimensional visualization model of the data center; using the three-dimensional visualization model to plan equipment deployment and wiring, and obtaining the data center's operation and maintenance data information; performing energy consumption analysis based on the operation and maintenance data information; and formulating an intelligent control plan based on the energy consumption analysis; classifying the operation and maintenance data, and performing comparative analysis on data of the same type at different time periods to detect data center faults; obtaining the type and location of the fault to generate fault warnings and solutions, and annotating and displaying the fault warnings and solutions in the three-dimensional visualization model. By performing a three-dimensional visualization of the data center's operation and maintenance status and operation and maintenance data, intelligent operation and maintenance of the data center is achieved, human resource allocation is reduced, and operation and maintenance efficiency is improved.
[0004] However, despite the significant progress made in existing technologies, there are still several profound technical bottlenecks that need to be addressed in achieving truly refined and robust intelligent deployment and closed-loop control;
[0005] 1. First, there are fundamental flaws in the accuracy of heat generation predictions. Existing models generally treat the server's computing power consumption (usually obtained from the out-of-band management controller BMC) as equivalent or linearly correlated with heat generation, which ignores a key physical fact: the conversion efficiency of the server's power supply unit (PSU) is not a constant value. The efficiency of the PSU changes dynamically with its load rate and input power quality. In particular, when there is significant total harmonic distortion (THD) in the power grid, the PSU efficiency will drop significantly, causing more electricity to be dissipated in the form of heat inside the server. Existing technologies generally lack the ability to quantitatively model this "electrical-thermal coupling" heat generation effect, resulting in systematic deviations in heat generation predictions from the source, which in turn affects the accuracy of all subsequent decisions;
[0006] Second, deployment decision-making lacks comprehensiveness. Even though some platforms consider multi-dimensional risk factors (such as thermal and power risks), their integration often relies on simple linear methods like static weighted summation. This approach is highly subjective and fails to dynamically reflect the actual contribution and interaction of various risk factors under varying system conditions. For example, aerodynamic coupling risk (such as hot air recirculation) at a location may have minimal impact under low load but dramatically amplify thermal risk under high load. Therefore, the lack of an objective and adaptive metric integration mechanism significantly reduces the context-awareness of deployment decisions.
[0007] 3. Finally, there is a significant gap in post-deployment adaptive control. Currently, deployment and control are typically two independent, open-loop systems. Once a deployment platform achieves its one-time "optimal" placement, its lifecycle ends. Subsequent environmental fluctuations (such as changes in load on adjacent racks) can lead to deviations between the actual and predicted states. This relies on the Building Automation System (BAS) for a crude and delayed response, lacking a closed-loop feedback mechanism to link IT loads with facility controls.
[0008] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the Invention
[0009] The purpose of the present invention is to provide a management data loading platform based on three-dimensional layout and intelligent optimization algorithm to solve the problems raised in the above background technology.
[0010] To achieve the above object, the present invention provides the following technical solutions:
[0011] A management data loading platform based on three-dimensional layout and intelligent optimization algorithm, specifically including:
[0012] 3D real-time state field construction module: This module is used to obtain a 3D physical layout model of the data center. It periodically obtains rack-level total harmonic distortion (THD) from the intelligent power distribution units deployed on the racks, and simultaneously obtains the real-time power consumption of each server from the server out-of-band management controller to construct a 3D real-time state field containing geometric, electrical, and thermal source information.
[0013] The electro-thermal coupling heat generation model calibration module is used to query a preset power efficiency function that characterizes the relationship between server power unit efficiency and harmonic distortion based on the acquired rack-level power supply total harmonic distortion rate to calculate the real-time power efficiency of each server location. The module then uses this real-time power efficiency to correct the server's real-time calculated power consumption, thereby generating a predicted total heat generation for the server.
[0014] Index fusion module: For any server slot to be evaluated, based on the predicted total heat generation and the three-dimensional physical layout model, the module calculates the "local thermal risk index" after the deployment task through computational fluid dynamics simulation. This index is combined with the "power quality impact index" and "aerodynamic coupling risk index" of the rack where the server slot is located. Principal component analysis is used to determine the contribution weights of each index, ultimately fusing them into a "context-aware deployment suitability index."
[0015] Optimal Deployment Decision Module: This module is used to formulate the allocation problem between a set of tasks to be deployed and all available server slots in the data center as an integer linear programming model, thereby solving the optimal allocation plan from tasks to server slots.
[0016] Linkage control module: After completing the task deployment according to the optimal allocation plan, continuously monitor the actual temperature and actual harmonic distortion rate of the deployment location; when the deviation of the actual temperature and actual harmonic distortion rate from the pre-deployment predicted value exceeds the preset threshold, dual-channel linkage control is triggered.
[0017] Compared with the prior art, the present invention has the following beneficial effects:
[0018] 1. By constructing an electro-thermal coupling model that considers the influence of total harmonic distortion, the accuracy of server heat generation prediction is significantly improved, providing high-fidelity input data for subsequent optimization decisions.
[0019] 2. The principal component analysis method is used to objectively integrate multi-dimensional risk indices such as thermal, electrical, and pneumatic. This overcomes the subjectivity and limitations of the traditional static weighting method, making deployment decisions more scientific and comprehensive.
[0020] 3. The innovative dual-channel linkage control module upgrades one-time deployment decisions to a continuous operation and maintenance system with closed-loop feedback and adaptive capabilities, achieving deep collaboration between IT loads and physical facilities at the control level, thereby comprehensively improving the operational stability, reliability and energy efficiency of the data center. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a schematic diagram of the overall system module of the present invention;
[0022] Figure 2 This is a schematic diagram of the server setup on the rack of the present invention. DETAILED DESCRIPTION
[0023] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0024] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0025] With the booming digital economy, demand for compute-intensive applications, such as artificial intelligence, big data analytics, and high-performance computing (HPC), is growing exponentially. To meet this demand, data centers are evolving toward high-density, ultra-large-scale architectures, with cabinet power densities continuously increasing. Against this backdrop, data center management technologies have also undergone profound changes. In the area of thermal management, technologies have evolved from early extensive cooling to sophisticated airflow simulation using computational fluid dynamics (CFD). Combined with a network of temperature sensors throughout the computer room, this enables dynamic and refined control of computer room air conditioning (CRAC) units, striving for the ultimate in power usage effectiveness (PUE). In the area of task scheduling and resource management, advanced scheduling systems such as Kubernetes and Slurm enable automated orchestration and efficient deployment of complex workloads based on the server's CPU, memory, storage, and other computing resource capacities, combined with quality of service (QoS) requirements.
[0026] Example 1:
[0027] See also Figure 1 and Figure 2 , the present invention provides a technical solution:
[0028] A management data loading platform based on three-dimensional layout and intelligent optimization algorithm, specifically including:
[0029] 3D real-time state field construction module: used to obtain the 3D physical layout model of the data center ; Periodically obtain the rack-level power total harmonic distortion rate from the intelligent power distribution unit deployed on the rack , where k is the rack index and t is the timestamp; and synchronously obtain the real-time computing power consumption of each server from the server out-of-band management controller , where i is the server index, to construct a three-dimensional real-time state field containing geometric, electrical and thermal source information.
[0030] The implementation involves gathering the basic data needed to build a digital twin model from the data center infrastructure through standardized interfaces and protocols. The specific operations are as follows: First, a "3D physical layout model" is constructed in one go by importing the Building Information Model (BIM) or using LiDAR scanning technology. , the "3D physical layout model" Accurately describe the physical coordinates and spatial relationships of cabinets, server slots, cooling units, and power distribution units;
[0031] In this embodiment, during the construction of the digital twin model, the quality of the raw data is crucial to the success or failure of all subsequent advanced analysis and decision-making. Existing technologies assume that the collected data is completely accurate and reliable, but this is not the case in reality. Sensors may age, networks may experience delays, and interfaces may experience transient failures, all of which can cause data distortion or failure.
[0032] The following solution introduces a "dynamic data quality credibility assessment" mechanism. Instead of passively receiving data, it actively and in real time assesses the "credibility" of each piece of data and treats it as an intrinsic property of the state field.
[0033] The three-dimensional real-time state field building module is further configured to: obtain the total harmonic distortion rate of the rack-level power supply and real-time computing power consumption of servers At the same time, the "dynamic data quality credibility index" associated with each dynamic data flow is calculated in parallel The dynamic data quality confidence index is a quantitative indicator determined based on the immediacy of data updates, the stability of values, and statistical deviations from historical data. This dynamic data quality confidence index is then bound to the dynamic data and jointly updated into the three-dimensional real-time state field to construct a more reliable state representation that accounts for data source uncertainty, thereby providing confidence-weighted data input for subsequent electro-thermal coupling analysis.
[0034] 3D physical layout model obtained by 3D real-time state field construction module It includes the centimeter-level precision physical coordinates and spatial topology of cabinets, servers, cooling units, and power distribution units; the frequency of dynamic data collection is configured according to system requirements; in this embodiment, during steady-state operation, it is once every 5 seconds, and it is automatically increased to once per second when the task load changes drastically and exceeds the preset value.
[0035] Regarding the Dynamic Data Quality Confidence Index The calculation process includes data collection and preprocessing, sub-dimensional quality index calculation, dynamic data quality credibility index fusion and status field update;
[0036] 1.1) Data collection and preprocessing:
[0037] The platform sends a request to the intelligent PDU of each rack at a set frequency through the network management protocol SNMP protocol or API interface, and obtains the "rack-level power total harmonic distortion rate" at the current time t from the intelligent PDU of rack k. At the same time, read the last recorded time from the memory value;
[0038] The platform queries and summarizes the power consumption readings of core components such as the central processing unit (CPU), graphics processing unit (GPU) and memory through the Redfish-API interface provided by the baseboard management controller (BMC) integrated on the server motherboard, and obtains the current real-time computing power consumption of the server i ;
[0039] 1.2) Calculation of sub-dimensional quality indicators:
[0040] Calculate the "data update immediacy index" and record it as ;Will The calculation logic is as follows: subtract the difference between the current data timestamp and the previous data timestamp from a preset "maximum tolerable data delay", and then divide the difference by the "maximum tolerable data delay". The "data update immediacy index" reflects the real-time degree of the data stream.
[0041] Calculate the "data numerical stability index" and record it as The calculation logic for this indicator is as follows: subtract the absolute difference between the current data value and the previous data value from a preset "maximum allowable instantaneous jump amplitude", and then divide this difference by the "maximum allowable instantaneous jump amplitude". This "data value stability indicator" is used to identify abnormal data changes caused by sensor failure.
[0042] Calculate the "statistical deviation index" and record it as The calculation logic of this indicator is as follows: subtract the absolute difference between the current data value and the moving average value of the data over the past period from 1, and then divide it by three times the moving standard deviation of the data over the same period; this "data statistical deviation indicator" is used to measure whether the current data point deviates from its normal statistical distribution range.
[0043] 1.3) Dynamic Data Quality Credibility Index Fusion:
[0044] In this embodiment, due to the different physical meanings and influences on the final credibility of the three sub-dimensional indicators, the platform uses a subjective-objective weighting model based on the entropy weight method to determine their fusion weights. The calculation logic of this method is as follows: First, domain experts use their experience to assign the three indicators ( 、 、 ) gives a set of subjective weights, which are recorded as , and Then, the information entropy of each indicator in the historical data set is calculated. The smaller the information entropy, the greater the amount of information provided by the indicator and the higher its objective weight. The final fusion weight is the normalized product of the subjective weight and the objective weight. This method combines expert knowledge and utilizes the inherent distribution characteristics of the data, avoiding the one-sidedness of a single weighting method.
[0045] The calculated fusion weight ( , , ) is applied to the corresponding sub-dimensional quality indicators, and the final "dynamic data quality credibility index" is calculated through weighted summation. The calculation logic is as follows: "Dynamic data quality credibility index" is equal to the "data update immediacy index" multiplied by its corresponding fusion weight, plus the "data numerical stability index" multiplied by its corresponding fusion weight, plus the "data statistical deviation index" multiplied by its corresponding fusion weight.
[0046] 1.4) Status field update:
[0047] The collected raw data , and the calculated dynamic data quality confidence index Binding is performed to form a data pair, and the corresponding object properties in the three-dimensional real-time state field are updated accordingly.
[0048] Dynamic Data Quality Confidence Index The quantitative description of the technical effect and algorithm deduction are as follows:
[0049] 1.5) Application-level parameter substitution and calculation:
[0050] Assume that the preset "maximum tolerable data delay" is 5 seconds and the "maximum allowable instantaneous jump amplitude" is At a certain moment, the platform collects the total harmonic distortion rate data of the rack-level power supply of rack k. It is 5.2%, and historical records show that The moving average of the total harmonic distortion rate of the rack-level power supply of this rack is 4.8%, and the moving standard deviation is 0.2%.
[0051] 1.51) Calculation of dimensional indicators:
[0052] Regarding the “Data Update Immediacy Index” The original formula for calculating is a discretized application of Newton's law of cooling, simplified here to a linear relationship. The calculation is: (5 seconds - 1 second) / 5 seconds = 0.8.
[0053] For "Data numerical stationarity index" The original formula for calculating comes from the slope limiter in signal processing. It is calculated as: (10% - |5.2% - 5.1%|) / 10% = (0.1 - 0.001) / 0.1 = 0.99.
[0054] For the “Statistical Deviation Indicator” The original formula is a variation of the Z-score in statistics. It is calculated as: 1-|5.2%-4.8%| / (3*0.2%)=1-0.004 / 0.006≈0.33.
[0055] 1.52) Weight fusion and results:
[0056] Assume that the fusion weight determined by the entropy weight method is: , , .
[0057] The final "dynamic data quality credibility index" The original formula for calculating is the weighted sum model in multi-attribute decision-making theory. The calculation is: 0.8*0.2+0.99*0.3+0.33*0.5=0.16+0.297+0.165=0.622.
[0058] 1.6) Physical dimension and range analysis:
[0059] All sub-dimensional indicators and the final "Dynamic Data Quality Credibility Index" are dimensionless, and their range is strictly designed to be within the closed interval [0,1]. The formula design ensures this by ensuring that the numerator is always less than or equal to the denominator and is non-negative.
[0060] When the parameter is "0" or "1":
[0061] when When it is 0, it means that the data delay has reached the maximum tolerance limit and the data is completely invalid; when it is 1, it means that the data is obtained instantaneously without delay.
[0062] when When it is 0, it means that the data jump amplitude reaches the upper limit, which may be an abnormal point; when it is 1, it means that the data has no jump and is extremely stable.
[0063] when When it approaches 0, it means that the data point deviates seriously from the historical statistical law; when it is 1, it means that the data point is completely consistent with the moving average.
[0064] Dynamic Data Quality Confidence Index The output range is [0,1];
[0065] when The closer the output is to 0, the lower the credibility of the corresponding dynamic data;
[0066] Reasoning: This is due to the poor performance of multiple sub-dimensional indicators at the same time, for example, data update delay Close to the "maximum tolerable data delay", resulting in Approaching 0; at the same time, the data value The greater the deviation from the historical mean, the The closer it is to 0. In this case, even if is not 0, but after weighted summation, the final It will also be small.
[0067] When the platform performs subsequent calculations on the heat generated by the electric-thermal coupling, it will give this low confidence Data is given a relatively low weight, and may even be temporarily replaced with highly reliable data from the previous moment or historical averages. This prevents a single erroneous, unreliable data point (such as a sensor reading error) from causing serious deviations in the entire hotspot prediction and task scheduling decision chain, thereby ensuring system stability and decision-making reliability.
[0068] when The closer the output is to 1, the higher the credibility of the corresponding dynamic data is, and the more reliable the real-time status reflection is.
[0069] Reasoning: This requires that all sub-dimensional indicators perform well. That is, data update delay The smaller, the Approaching 1; data value With the previous moment value Almost the same, is close to 1; at the same time, the data also conforms to its recent statistical laws, making Also approaches 1.
[0070] Technical Effect: The platform uses this data with 100% confidence for all subsequent calculations. This means the platform has greater confidence in its current state perception, and its hotspot predictions, risk assessments, and task scheduling decisions will be based on the most accurate and timely on-site conditions, enabling the most refined energy management and the highest operational efficiency.
[0071] Electric-thermal coupling heat generation model calibration module: used to , query the preset power efficiency function that characterizes the relationship between server power supply unit (PSU) efficiency and harmonic distortion , to calculate the real-time power efficiency of each server location; and using the real-time power efficiency, calculate the real-time power consumption of the server Correction is made to produce a more accurate prediction of the total heat production of the server that takes into account the effects of harmonic heating .
[0072] The specific implementation includes: quantifying the change of power quality as the change of heat generation; the specific implementation process is:
[0073] First, the platform pre-establishes a "power efficiency function" A database, obtained by calibrating various PSU models used in data centers in the laboratory, records the actual operating efficiency of PSUs at different rack-level power supply total harmonic distortion levels;
[0074] When the platform is running, for any server i, the system will obtain the real-time "rack-level power supply total harmonic distortion rate" of the rack k where it is located. ;
[0075] When considering the impact of harmonics on power supply efficiency, existing technologies often overlook an equally critical variable: the load factor of the power supply unit (PSU). A PSU operating at 20% load and a PSU operating at 80% load will have completely different efficiency attenuation curves, even when facing the same harmonic distortion. This improved solution upgrades the original two-dimensional (THD-efficiency) model to a three-dimensional (THD-load-efficiency) model by introducing a "PSU load factor." This improvement makes the efficiency calculation no longer a static function that is only related to external power quality, but a dynamic, accurate model that depends on both external power quality and the internal working state of the server. This improves the accuracy of heat generation predictions;
[0076] The electro-thermal coupling heat generation model calibration module is further configured as follows: based on the total harmonic distortion rate of the rack-level power supply When querying the power efficiency function, synchronously obtain the real-time computing power consumption of server i And combined with the power rating of its power supply unit to calculate a "PSU load factor" ; and using the and stated As dual input, a pre-defined 3D efficiency surface model is queried to determine the “load-aware real-time power efficiency” ; Finally based on this Calculate the predicted total heat generation of the server ;Thereby achieving accurate quantification of harmonic heating effects under different working conditions;
[0077] The preset three-dimensional efficiency surface model is a digital twin model generated by matrix testing the power supply units used in the data center at multiple discrete load points and rack-level power total harmonic distortion injection levels in a controlled experimental environment, and using a three-dimensional interpolation or surface fitting algorithm;
[0078] Set the predicted total heat output of the server The process is sequentially through input parameter acquisition, key factor calculation, core parameter calibration, heat generation calculation and harmonic influence quantification;
[0079] The key factor is the PSU load factor of server i at time t ;
[0080] The harmonic impact quantification is the “harmonic damage factor” of server i at time t , used to characterize the proportion of additional efficiency loss caused by harmonics;
[0081] This embodiment is decomposed into the following calculation process:
[0082] 2.1) Input parameter acquisition:
[0083] Get the "real-time computing power consumption" of the specified server i from the 3D real-time state field module ” and the “rack-level power supply total harmonic distortion rate” of the rack k where it is located .
[0084] Query the PSU rated power capacity of the power supply unit (PSU) configured for server i from the asset database. ;
[0085] 2.2) Calculation of key factors:
[0086] Calculating the "PSU Load Factor" The calculation logic of this factor is as follows: Use "real-time calculation of power consumption ” divided by the “PSU rated power capacity”. This dimensionless factor represents how hard the PSU is currently working.
[0087] 2.3) Core parameter calibration and heat generation calculation:
[0088] Calibrate the "load-aware real-time power efficiency" of server i at time t The calibration process is: using the calculated "PSU load factor" And the obtained "rack-level power total harmonic distortion rate" As coordinates, query or bilinear interpolation is performed in the preset three-dimensional efficiency surface model to obtain the accurate power efficiency under the current comprehensive working conditions;
[0089] Calculate the "predicted total heat generation of the server" The calculation logic of the heat generation is expressed as follows: "Real-time calculation of power consumption ” divided by “Load-aware real-time power efficiency” , minus the "real-time calculation power consumption ". Used to calculate the total power loss of the PSU during the energy conversion process, that is, the total heat generated.
[0090] 2.4) Quantification of harmonic impact:
[0091] To quantify the negative impact of harmonics, the platform further calculates the “harmonic damage factor” of server i at time t This factor is calculated logically as follows: 1 minus the ratio of "Load-Aware Real-Time Power Efficiency" to "Ideal Power Efficiency." Ideal Power Efficiency is the power efficiency at zero rack-level total harmonic distortion (THD) under the same PSU load factor. This value can be directly obtained from a specific section of the 3D efficiency surface model. This factor intuitively reflects the proportion of excess efficiency loss caused by harmonics.
[0092] The technical effect of the electro-thermal coupling heat generation model calibration module is measured by the "harmonic loss factor". and "Predicted total heat generation of the server" The precise calculation of
[0093] The quantitative description of the technical effect and algorithm deduction of the electro-thermal coupling heat generation model calibration module are as follows:
[0094] 2.5) Application-level parameter substitution and calculation: Assume that the PSU rated power capacity of server i is At time t, the "real-time computing power consumption" is obtained from the state field. The total harmonic distortion rate of the rack-level power supply is 400 watts. It is 15%.
[0095] Key factor calculation: "PSU load factor" = L(i,t) = 400 watts / 800 watts = 0.5.
[0096] 2.6) Core parameter calibration and heat generation calculation:
[0097] Query the 3D efficiency surface model:
[0098] When L(i,t)=0.5 and Under the conditions of It is 94%.
[0099] When L(i,t)=0.5 and Under the conditions of It is 91%.
[0100] "Predicted total heat generation of the server" The original formula for the calculation is the law of conservation of energy in physics. The calculation is: (400 watts / 0.91)-400 watts = 439.56 watts-400 watts = 39.56 watts.
[0101] In contrast, in an ideal case without harmonics, the heat generated is (400 W / 0.94) - 400 W = 25.53 W. The presence of harmonics results in an additional heat generation of 14.03 W, an increase of 55%.
[0102] Quantification of harmonic impact:
[0103] “Harmonic Damage Factor” The original formula for calculating is the normalized difference model used in engineering to represent relative performance degradation. The calculation is: 1-(0.91 / 0.94)=1-0.968=0.032.
[0104] The physical dimension of is watt, which is consistent with the dimension of power. , , ,and All are dimensionless pure numbers. The value range of is designed to be in the closed interval [0,1);
[0105] when When it is 0, it means that the "load-aware real-time power efficiency" is equal to the "ideal power efficiency". This means that the impact of harmonics can be ignored under the current working conditions.
[0106] when As the output approaches 0, the additional energy loss caused by harmonics approaches zero, and the negative impact of power quality on the server's heat dissipation becomes smaller.
[0107] reasoning: Approaching 0, it means that the calculation formula and The ratio approaches 1. This is because the input "rack-level power total harmonic distortion rate" The closer it is to zero.
[0108] Technical Effect: This indicates that the cleaner the server's electrical environment, the more efficient the PSU will be under its load. The platform can confirm that the safer the electrical environment at a location, the fewer harmonics must be considered in thermal forecasts, and the more certain its decisions are.
[0109] when As the output approaches 1, the additional energy loss caused by harmonics increases, the efficiency degradation of the PSU increases, and the additional heat generated increases.
[0110] reasoning: Approaching 1 means "load-aware real-time power efficiency" The closer it is to zero. According to the three-dimensional efficiency surface model, this requires the “rack-level power supply total harmonic distortion rate” Reach a higher level.
[0111] Technical Effect: This indicates that the worse the server's electrical environment, the more input electrical energy will be dissipated as heat, and the platform can trigger higher-level alarms accordingly. This factor can sensitively quantify this extreme risk, demonstrating the key role of this invention in ensuring the safe operation of data centers and the rationality of the model design.
[0112] Indicator fusion module: used for any server slot s to be evaluated, based on the predicted total heat generation and 3D physical layout models Through computational fluid dynamics (CFD) simulation, the "local thermal risk index" after the deployment task is calculated; and combined with the "power quality impact index" and "aerodynamic coupling risk index" of the rack where the server slot is located; the contribution weight of each index is determined through principal component analysis, and finally integrated into a "context-aware deployment suitability index."
[0113] The platform uses principal component analysis (PCA) to determine , and Benchmark contribution weight , , The specific algorithm can be expressed as follows: A historical data set containing the three indices and their corresponding system stability results (e.g., whether overheating and frequency reduction occurred) is collected. The covariance matrix of these three indices is subjected to eigenvalue decomposition, and the eigenvector that explains the largest variance in the data (i.e., the load vector of the first principal component) is extracted. The three elements of this vector, after normalization, are defined as the contribution weights of the three risk indices. , , , reflecting the degree of influence of each risk dimension on the overall stability of the system.
[0114] While the principal component analysis (PCA) method used in existing technologies can objectively extract static contribution weights for each risk dimension from historical data, it cannot adapt to the real-time, changing global environment of a data center. For example, during hot summer weather, the actual threat level of heat risk rises dramatically, far exceeding its importance under normal circumstances; and during periods of grid fluctuations, power quality risk becomes the primary concern. This solution introduces a "contextual risk amplifier" to dynamically adjust the static baseline weights derived from PCA in real time. This allows the risk assessment model to evolve from a static model that "summarises after the fact" to a dynamic model that can "adapt on the spot." Its assessment results more accurately reflect the actual deployment risks at a specific point in time and within a specific global environment.
[0115] The indicator fusion module is further configured to: after determining the "baseline contribution weight" of each risk index through principal component analysis, concurrently calculate a "contextual risk amplification factor" based on the real-time global environmental status of the data center; dynamically adjust the baseline contribution weight using the contextual risk amplification factor to generate a set of "dynamic risk weights"; and finally, perform a weighted fusion of the risk indices based on the dynamic risk weights to generate a "context-aware deployment suitability index."
[0116] To ensure real-time performance, the computational fluid dynamics (CFD) simulation process is accelerated by a pre-trained surrogate model, which can predict the thermal characteristics of a specific server slot within one second. The global environmental status includes, but is not limited to, the average cabinet inlet air temperature, outdoor temperature, or grid voltage stability indicators for the entire data center.
[0117] The steps for obtaining the context-aware deployment suitability index include quantifying the basic risk index, determining the static benchmark weight, adjusting the dynamic weight, and integrating the final suitability index.
[0118] Among them, the basic risk index quantification includes the local thermal risk index, power quality impact index and aerodynamic coupling risk index of server slot s;
[0119] The implementation of the context-aware deployment suitability index is broken down into the following calculation process:
[0120] 3.1) Quantification of Basic Risk Index:
[0121] For the server slot s to be evaluated, three standardized basic risk indices are calculated, all with a value range of [0, 1], where 1 represents the optimal value:
[0122] The “local thermal risk index” of server slot s is recorded as The calculation logic is as follows: subtract a quotient from 1. The numerator of the quotient is the difference between the server air inlet temperature predicted after the task is deployed and the optimal operating temperature, and the denominator is the difference between the hardware alarm temperature and the optimal operating temperature.
[0123] The “power quality impact index” of server slot s is recorded as The calculation logic is: subtract the square of a quotient from 1. The numerator of the quotient is the current total harmonic distortion rate of the rack where the server slot is located, and the denominator is the total harmonic distortion rate limit specified by the national standard or industry standard.
[0124] The “aerodynamic coupling risk index” of server slot s is recorded as The calculation logic is as follows: subtract the "hot air recirculation rate" calculated by CFD simulation, which is the air flow from the upstream cabinet outlet back to the air inlet of the server slot, from 1;
[0125] 3.2) Determination of static benchmark weight:
[0126] The platform uses the principal component analysis (PCA) method in statistics to determine the "benchmark contribution weight" of each basic risk index. The determination logic of this method is: collect historical data sets containing the above three indexes and whether the corresponding servers have stability events (such as overheating, frequency reduction, and downtime). The multidimensional data composed of these three indexes are centralized and standardized, and then their covariance matrix is calculated. By performing eigenvalue decomposition on the covariance matrix, the eigenvector corresponding to the maximum eigenvalue is found. The absolute value of each element of the eigenvector, after normalization, is defined as the "benchmark contribution weight" of the three basic risk indices; , and The “benchmark contribution weights” are recorded as , , ;
[0127] 3.3) Dynamic weight adjustment:
[0128] Calculate the "contextual risk magnification factor"; taking thermal risk as an example, its contextual risk magnification factor The calculation method is expressed as: Equal to the base e of the natural logarithm, its exponent is the product of the "thermal risk sensitivity coefficient" and the difference between the "global average inlet air temperature" and the "standard operating temperature". The "thermal risk sensitivity coefficient" is a preset positive constant used to adjust the severity of the amplification effect;
[0129] Will , , The “dynamic risk weights” are recorded as , , The calculation method is as follows: multiply the "benchmark contribution weight" of each underlying risk index by its corresponding "contextual risk magnification factor" to obtain a set of temporary dynamic weights; then normalize this set of temporary weights, that is, divide each temporary weight by the sum of all temporary weights to obtain the final "dynamic risk weight" with a total sum of 1;
[0130] For the "Power Quality Impact Index", its "Contextual Risk Amplification Factor" The calculation method is expressed as: Equal to the base e of the natural logarithm, its exponent is the product of the "power quality sensitivity coefficient" and the difference between the "global grid total harmonic distortion rate" and the "grid harmonic standard limit". The "power quality sensitivity coefficient" is a preset positive constant used to adjust the severity of the amplification effect.
[0131] Parameter explanation: Global Grid Total Harmonic Distortion (THD): This refers to the real-time total harmonic distortion (THD) value of the external grid power supply quality, obtained from the power quality monitoring equipment at the data center's main incoming line. This parameter reflects the overall electrical environment pressure facing the data center.
[0132] Grid harmonic standard limit: refers to the safe upper limit of the total harmonic distortion rate set for the data center power supply system based on industry standards (such as GB / T14549-93 "Power Quality Public Grid Harmonics"). In this embodiment, the initial setting is 5%.
[0133] The power quality sensitivity factor is a dimensionless positive constant, set by operations and maintenance experts based on the sensitivity of the IT equipment within the data center to harmonics. For data centers housing large numbers of sophisticated computing or financial trading servers, this factor is often set higher to significantly amplify risk even with even the slightest deterioration in grid quality.
[0134] For the “gas-heat coupling risk index”, its “context risk amplification factor” The calculation method is expressed as: Equal to the base e of the natural logarithm, its exponent is the product of the "airflow organization sensitivity coefficient" and the percentage by which the "data center total air supply volume" is lower than the "design baseline total air supply volume." The "airflow organization sensitivity coefficient" is a preset positive constant used to adjust the severity of the amplification effect.
[0135] Data Center Total Air Supply: This refers to the sum of the current air supply of all computer room air conditioning (CRAC / CRAH) units, as obtained from the building automation system (BAS) or cooling system controller. This parameter reflects the overall operating intensity and capacity of the data center cooling system.
[0136] Design baseline total airflow: This refers to the theoretical total airflow determined during the data center design phase to meet the cooling requirements of full-load operation. This is a fixed baseline value.
[0137] Airflow sensitivity coefficient: A dimensionless positive constant, set by HVAC experts based on the data center's airflow design characteristics (such as hot and cold aisle containment and cabinet density). In an older data center with chaotic airflow and a tendency for hot air backflow, this coefficient is often set high, significantly amplifying the risk of localized aerodynamic coupling with even the slightest decrease in cooling system capacity.
[0138] 3.4) Final suitability index fusion:
[0139] The “context-aware deployment suitability index” of server slot s is recorded as ; Its calculation method is expressed as: It is equal to the "local thermal risk index" multiplied by its corresponding "dynamic risk weight", plus the "power quality impact index" multiplied by its corresponding "dynamic risk weight", plus the "aerodynamic coupling risk index" multiplied by its corresponding "dynamic risk weight";
[0140] The quantitative description of the technical effect and algorithm deduction are as follows:
[0141] Application-level parameter substitution and calculation: Assume that the three basic risk indices of the slot s to be evaluated are: , , .
[0142] Through principal component analysis (PCA, derived from multivariate statistical analysis), the “benchmark contribution weights” have been determined to be: , , .
[0143] Scenario 1: Normal environment:
[0144] The data center's global average inlet air temperature is 22°C, and the standard operating temperature is 22°C.
[0145] 1. “Contextual Risk Amplification Factor” The original formula for calculating is the exponential gain model in control theory. The result is 1. Assume that the other risk magnification factors are also 1.
[0146] 2. “Dynamic risk weights” are the same as benchmark weights.
[0147] 3. Context-Aware Deployment Suitability Index The original formula for calculating is the weighted sum model (WSM) in multi-attribute decision-making theory. The calculation is: (0.7*0.5)+(0.9*0.3)+(0.8*0.2)=0.35+0.27+0.16=0.78.
[0148] Scenario 2: High temperature environment:
[0149] The data center's Global Average Inlet Temperature rises to 27°C. The Thermal Risk Sensitivity Factor is set to 0.2.
[0150] 1. “Contextual Risk Amplification Factor” The calculation is: e to the power of (0.2*(27-22)) = e to the power of 1 = 2.718. Assume that other risk magnification factors remain at 1.
[0151] 2. The temporary dynamic weight is: (0.5*2.718), (0.3*1), (0.2*1) = (1.359, 0.3, 0.2). The total is 1.859.
[0152] 3. The normalized “dynamic risk weight” is: , , .
[0153] New Context-Aware Deployment Suitability Index The calculation is: (0.7*0.731)+(0.9*0.161)+(0.8*0.108)=0.512+0.145+0.086=0.743.
[0154] Comparative analysis: Although the risk index of the server slot itself remained unchanged, its deployment suitability index dropped from 0.78 to 0.743 due to the increased global thermal risk. The system automatically and quantitatively increased its vigilance against thermal risk.
[0155] All risk indices, weights, sensitivity coefficients and final indices They are all dimensionless pure numbers. The value range of is strictly limited to the closed interval [0,1].
[0156] The effective value range for actual output is (0,1);
[0157] when The closer the output is to 0, the higher the deployment risk of the server slot is in the current global environment, and the less suitable it is for deploying tasks.
[0158] Reasoning: This is because the lower the value of at least one basic risk index, the more significantly its corresponding "dynamic risk weight" is amplified in the current context. For example, the server slot The value is 0.2 (high risk), and the data center is experiencing a heat wave, and the "global average air inlet temperature" is far above the standard. becomes larger, making It dominates all weights. The final weighted sum is seriously dragged down by this low score, approaching 0.
[0159] Technical Effect: The platform accurately identifies deployment points whose conditions could deteriorate from "sub-healthy" to "critical" under specific global catastrophic weather events (such as heat waves and power grid instability). This allows for mandatory avoidance, preventing cascading system failures caused by sudden environmental changes. This demonstrates the model's ability to foresee the potential for "small problems" to evolve into "major disasters" within a broader context.
[0160] when The closer the output is to 1, the more secure and stable the server slot is in the current global environment, and the more suitable it is for deploying tasks.
[0161] reasoning: The closer it is to 1, the condition is that all the basic risk indices are also closer to 1. Because the sum of the weights is 1, any low score will cause the final result to be unable to reach 1.
[0162] Technical effect: The higher-scoring server slots selected by the platform are, the more they can withstand the global environmental pressure currently faced by data centers.
[0163] Optimal deployment decision module: It is used to construct the allocation problem between a set of tasks J to be deployed and all available server slots in the data center as an integer linear programming model; thus solving the optimal allocation plan from tasks to server slots .
[0164] Existing deployment decision models typically treat all tasks to be deployed equally, with the optimization goal being to maximize the sum of all tasks' context-aware deployment suitability indices. However, in real-world operations and maintenance scenarios, different tasks have varying business importance. For example, a core trading system task is far more important than a temporary test task. This solution introduces a "task business priority coefficient" to weight the optimization objective function. This allows the model to prioritize finding the best deployment server slots for high-priority tasks, even if this may result in a slight decrease in the sum of the overall suitability. This improves the commercial rationality and practicality of the decision.
[0165] The optimal deployment decision module is further configured to: associate a preset "task business priority coefficient" with each to-be-deployed task j when constructing an integer linear programming model; and modify the model's optimization objective function to maximize the sum of the product of the "context-aware deployment suitability index" of all deployed tasks and the "task business priority coefficient," with the product being recorded as the weighted suitability. This ensures that the optimal allocation solution obtained can prioritize the allocation of high-importance tasks to the server slots with the most physical security.
[0166] The "task business priority coefficient" is user-configured through a graphical user interface (GUI) or automatically retrieved through API integration with external business management systems (such as CMDBs). This integer linear programming model can be efficiently solved using open source solvers (such as GLPK) or commercial solvers (such as Gurobi).
[0167] Optimal allocation plan The acquisition process is broken down as follows:
[0168] 4.1) Input data preparation and parameter definition:
[0169] Get all available server slots from the indicator fusion module For each task to be deployed Context-Aware Deployment Suitability Index ; Where S represents the total number of available server slots, J represents the total number of tasks to be deployed; s is the index mark of the available server slot, and j is the index mark of the task to be deployed;
[0170] is a dimensionless value with a valid range in the interval (0,1), which represents the overall suitability of deploying task j in slot s;
[0171] Define a "task business priority coefficient" for each task j to be deployed It is a dimensionless positive number greater than zero, representing the business importance of task j. It is determined by a hierarchical mapping table. For example, operations personnel categorize tasks into three levels: "core," "important," and "normal," and map them to priority coefficients of 10, 5, and 1, respectively. This mapping table can be adjusted based on the actual business value of the enterprise.
[0172] Define a binary "deployment decision variable" ;
[0173] When task j is assigned to server slot s, The value of is 1; otherwise it is 0; this is the unknown quantity that needs to be solved by the optimization model;
[0174] Obtain the resource requirements for each task j and the available resource capacity for each server slot s; these parameters represent the number of CPU cores, GB of memory, etc. The task resource requirements are obtained from the task submission request; the slot resource capacity is obtained from the asset management database or real-time monitoring system.
[0175] 4.2) Mathematical optimization model construction:
[0176] Construct a weighted integer linear programming model; this model is derived from the generalized assignment problem in operations research.
[0177] For the weighted integer linear programming model: maximize the sum of the weighted fitnesses of all existing task-slot assignment combinations. The weighted fitness of each combination is equal to the combination's context-aware deployment fitness index multiplied by the task's business priority coefficient multiplied by the combination's deployment decision variable.
[0178] Define constraint 1 (uniqueness of task assignment): For each task to be deployed, the sum of its "deployment decision variables" on all available server slots must be equal to 1;
[0179] Define constraint 2 (slot occupancy uniqueness): For each available server slot, the sum of the "deployment decision variables" of all tasks assigned to it must be less than or equal to 1;
[0180] Define Constraint 3 (Resource Capacity Limit): For each available server slot, the sum of the resource requirements of all tasks assigned to it must be less than or equal to the available resource capacity of that slot. Apply the same constraints to other resources such as memory.
[0181] 4.3) Model solution and solution output:
[0182] Input the constructed weighted integer linear programming model into a standard integer linear programming solver;
[0183] The solver iterates the problem using algorithms such as branch and bound or cutting plane method until it finds the optimal solution that satisfies all constraints and maximizes the objective function value.
[0184] Output the optimal solution, that is, a set of optimal "deployment decision variables" The value of ; this set of variables with a value of 1 constitutes the optimal allocation plan from tasks to server slots.
[0185] 4.4) Quantitative description of the technical effect and algorithm deduction of the task business priority coefficient:
[0186] Application-level parameter substitution and calculation:
[0187] Suppose there are two tasks (core tasks) and (Normal Mission), and two available server slots (Premium Slot) and (normal slot);
[0188] Input data: Context-aware deployment suitability index: , ; , . Task business priority coefficient: , .
[0189] Scenario 1: No "task business priority coefficient" weighting, this is the traditional approach;
[0190] The optimization goal is to maximize .
[0191] Option A1: . Objective function value = 0.9 + 0.6 = 1.5.
[0192] Option B1: . Objective function value = 0.7 + 0.8 = 1.5.
[0193] Conclusion: The objective function values of the two solutions are the same. The model cannot distinguish between the good and the bad. It may randomly select one, such as solution B1, resulting in the core task being deployed to a suboptimal location.
[0194] Scenario 2: The present invention has a weighted “task business priority coefficient”;
[0195] The optimization goal is to maximize .
[0196] Option A2: . Objective function value = (10*0.9)+(1*0.6)=9.0+0.6=9.6.
[0197] Option B2: . Objective function value = (10*0.7)+(1*0.8)=7.0+0.8=7.8.
[0198] Conclusion: The objective function value of Scheme A2 is significantly higher than that of Scheme B2. Therefore, the integer linear programming solver will uniquely and definitively output Scheme A2 as the optimal solution, ensuring the core task Assigned to the best slot .
[0199] Analyze the core component of the optimization goal - "weighted suitability" .
[0200] When the weighted fitness approaches its maximum value, it indicates that important tasks are assigned to a deployment location that is almost perfect under the current environment; for example, for core tasks and premium slots , its maximum value is 10*1=10;
[0201] Reasoning: This requires a "task business priority coefficient" It is high in itself, and the "context-aware deployment suitability index" of this task-slot combination is Approaching 1. It approaches 1, and all the basic risk indices of the server slot are required to approach 1.
[0202] Technical Effect: The optimization model strongly favors this high-scoring combination. This ensures that the data center's most valuable physical resources are used to support core business applications, achieving a precise match between resource value and business value, and maximizing overall business stability and reliability.
[0203] When the "weighted suitability" approaches 0, it means that a task has been assigned to a completely unsuitable deployment location, or the task itself is unimportant and the location is also poor.
[0204] Reasoning: This is due to the Context-Aware Deployment Suitability Index This is caused by the task business priority coefficient approaching 0. High, but multiplied by a value close to 0 , the result is still close to 0.
[0205] Technical Effect: Driven by the goal of maximizing the sum, the optimization model will carefully avoid selecting any allocation combinations whose "weighted suitability" approaches 0. This means that even for low-priority tasks, the model will strive to find a deployment location that is at least "passable," rather than arbitrarily placing them in high-risk areas. This improves the overall operational safety of the data center and demonstrates the comprehensiveness and rationality of the optimization objective design.
[0206] Linkage control module: After completing the task deployment according to the optimal allocation plan, continuously monitor the actual temperature of the deployment location and actual harmonic distortion ; When the actual temperature and actual harmonic distortion rate deviate from the pre-deployment predicted values or When the preset threshold is exceeded, dual-channel linkage control is triggered.
[0207] Further explanation: Dual-channel linkage control adjusts the operating parameters of the computer room air conditioning (CRAC) unit associated with the location to correct the temperature deviation, and generates a harmonic suppression instruction, which includes a recommendation to migrate a task with complementary harmonic characteristics to the location, to achieve closed-loop adaptive environmental control.
[0208] Existing linkage control schemes often use simple threshold triggers and fixed commands (such as "increase wind speed by a specific percentage"). This "open-loop" or "step-based" control approach can easily lead to system overshoot or oscillation, resulting in energy waste from overcooling, or inadequate regulation that fails to resolve the problem. This solution introduces the classic proportional-integral-derivative (PID) control algorithm, upgrading the control logic from a simple "on-off" approach to a sophisticated "regulatory" approach. It dynamically calculates a precise, continuously changing control output based on the magnitude (proportional P), duration (integral I1), and rate of change (differential D) of the deviation. This closed-loop feedback and progressive regulation approach not only eliminates deviations more quickly and stably, but also achieves control objectives in the most economical manner, avoiding energy waste and system instability.
[0209] The linkage control module is further configured to: and actual harmonic distortion When the deviation from the predicted value exceeds a threshold, the proportional-integral-derivative (PID) control algorithm uses the deviation as input to continuously calculate a quantified "cooling adjustment output." Based on this output, progressive control instructions are generated for the computer room air conditioning unit. At the same time, when the harmonic deviation continues to accumulate and exceeds the integral threshold, an "active harmonic cancellation task" based on harmonic spectrum feature matching is triggered, thereby achieving stable, precise, and energy-optimized closed-loop control of the physical environment.
[0210] The three core parameters of the PID controller (proportional coefficient , integral coefficient , differential coefficient ) is initially set using classic tuning methods such as Ziegler-Nichols, and online adaptive optimization is performed during system operation based on historical control performance data. The harmonic spectrum feature matching algorithm performs Fourier analysis on the harmonic spectra of candidate tasks to identify tasks with opposite phases and comparable amplitudes to the target rack's harmonic spectrum at key subharmonics.
[0211] The calculation process of the linkage control module is decomposed as follows:
[0212] 5.1) Continuous monitoring and quantification of deviations:
[0213] After the task is deployed, the linkage control module continuously collects the "actual server inlet air temperature" of the deployed server slot s at a fixed time period. And the actual total harmonic distortion rate of the rack ;
[0214] Obtained through wireless temperature sensors deployed at the air inlet of the server; Obtained through a power quality analyzer installed on the rack PDU;
[0215] Calculate "real-time temperature deviation" and "Real-time Harmonic Deviation" .
[0216] "Real-time temperature deviation" equals "actual server inlet temperature" minus "predicted server inlet temperature" predicted by the model before deployment. The same applies to "real-time harmonic deviation" and is not described in detail here; the predicted value is generated and stored by the previous module when the deployment decision is made.
[0217] 5.2) Temperature channel PID closed loop control:
[0218] Calculate the "cooling control output" at time t This calculation uses the classic PID control algorithm in the industrial control field. Its calculation logic is expressed as: "Cooling control output" is equal to the sum of the three parts;
[0219] The first part is the "proportional term", which is equal to the "proportional coefficient" Multiply by the current "real-time temperature deviation";
[0220] The second part is the "integral term", which is equal to the "integral coefficient" Multiply by the cumulative sum of all "real-time temperature deviations" from the start of trigger control to the current state;
[0221] The third part is the "differential term", which is equal to the "differential coefficient" Multiply by the difference between the current "real-time temperature deviation" and the previous moment's "real-time temperature deviation".
[0222] Proportional coefficient , integral coefficient , differential coefficient These are the core parameters of the PID controller and are pre-tuned based on the thermal inertia and other characteristics of the data center. They are dimensionless constants.
[0223] 5.3) Harmonic channel cancellation control:
[0224] Calculate the "harmonic cumulative risk" at time t The calculation logic is as follows: When the "real-time harmonic deviation" is greater than zero, the "harmonic cumulative risk" is equal to the cumulative sum of all "real-time harmonic deviations" from the start of control triggering to the current state; otherwise, it is zero. This calculation is based on the concept of integral accumulation in control theory.
[0225] Triggering harmonic cancellation: When the "harmonic accumulation risk" When a preset "harmonic intervention threshold" is exceeded, the platform triggers the recommended process of "active harmonic cancellation task".
[0226] The "harmonic intervention threshold" is an empirical value, indicating that the harmonic exceeding the standard problem has persisted for some time and is relatively serious, and needs to be addressed from the source.
[0227] 5.4) Generation and issuance of control instructions:
[0228] Set the "cooling adjustment output" at time t Converted into specific equipment control instructions; the "fan speed adjustment instruction" of the computer room air conditioning CRAC unit is equal to its "base speed" plus the "cooling adjustment output" multiplied by a "speed adjustment coefficient".
[0229] Based on "harmonic cumulative risk" , issuing instructions through the building automation system (BAS) interface and pushing a migration suggestion work order for the "active harmonic cancellation task" to the operation and maintenance system;
[0230] The quantitative description and algorithm deduction of the technical effect of the linkage control module are as follows:
[0231] The technical effect of the linkage control module is to adjust the output through cooling Dynamic calculation and "harmonic cumulative risk" It is reflected by the threshold judgment;
[0232] Perform the following application-level parameter substitution and calculations:
[0233] Assume that the predicted temperature of server slot s is 24°C and the temperature deviation threshold is 1°C. The PID parameters are set to , , The base speed of the CRAC fan is 50%, and the speed adjustment coefficient is 1.
[0234] Time t=1: the actual temperature rises to 25.5°C;
[0235] Real-time temperature deviation If the deviation exceeds 1°C, PID control is triggered.
[0236] Cooling regulation output Calculation (the initial value of the differential term is 0): (0.5*1.5)+(0.1*1.5)+0=0.75+0.15=0.9.
[0237] "Fan speed adjustment command" = 50% + 0.9 * 1 = 50.9%. The fan speed is slightly increased.
[0238] Time t=2: Due to the increase in speed, the temperature rise momentum slows down and the actual temperature is 26°C.
[0239] Real-time temperature deviation .
[0240] Cooling regulation output Calculation: (0.5*2)+(0.1*(1.5+2))+(0.2*(2-1.5))=1+0.35+0.1=1.45.
[0241] "Fan speed adjustment command" = 50% + 1.45 * 1 = 51.45%. The fan speed is further and significantly increased.
[0242] Time t=3: The control effect is apparent and the actual temperature drops to 25°C.
[0243] Real-time temperature deviation .
[0244] Cooling regulation output Calculation: (0.5*1)+(0.1*(1.5+2+1))+(0.2*(1-2))=0.5+0.45-0.2=0.75.
[0245] "Fan speed adjustment command" = 50% + 0.75 * 1 = 50.75%. The adjustment force begins to decrease to avoid overcooling.
[0246] The PID controller dynamically adjusts the control intensity according to the size of the deviation (proportional P), accumulation (integral I1) and change trend (differential D), achieving precise control with fast response and avoiding overshoot.
[0247] when The closer the output is to 0, the more stable the system environment is, and the fewer intervention measures the control system needs;
[0248] reasoning: The closer it is to 0, the more the "real-time temperature deviation" in the PID calculation formula The closer its integral and differential terms are to 0, the closer the actual server inlet air temperature is to the predicted server inlet air temperature.
[0249] Technical Effect: This demonstrates the control system's "silent" capability. When the environment is stable, it avoids unnecessary adjustments, avoiding equipment wear and energy waste caused by frequent starts and stops or fine-tuning. This demonstrates that this closed-loop control achieves both stability and economic efficiency.
[0250] when The closer the output reaches its positive upper limit, the more susceptible the system is to the risk of thermal runaway, and the more intervention measures the control system must take.
[0251] Reasoning: This requires "real-time temperature deviation" The value continues to increase rapidly (resulting in a larger positive value for the differential term). The three terms in the PID formula are superimposed, resulting in a larger positive output value that is saturated by the upper limit of the actuator.
[0252] Technical Results: This demonstrates that even when the system encounters a more severe emergency (such as a CRAC unit failure or an abnormal surge in server load), the PID control system responds with the fastest speed and maximum force, preventing further temperature increases and buying valuable time for human intervention. Furthermore, large control outputs themselves serve as a strong warning signal, demonstrating the algorithm's sensitivity and strength in risk perception and emergency response.
[0253] Compared with the existing technology, this invention fundamentally improves the intelligence and predictability of management by revealing and quantifying the long-neglected electro-thermal coupling effect in the physical field of the data center. First, it constructs the total harmonic distortion rate of the rack-level power supply The heat generation model can accurately quantify the "hidden heat sources" generated by electrical energy pollution, thereby raising the accuracy of hotspot prediction to a new level and achieving a shift from passive response to active prevention. Secondly, the present invention ingeniously integrates the risks of multiple physical fields such as heat, electricity, and airflow into a single quantitative "context-aware deployment suitability index" through objective methods such as principal component analysis, so that task scheduling based on integer linear programming can break away from the limitations of a single dimension and find the truly global optimal solution. More importantly, its closed-loop, dual-channel linkage control mechanism can actively suppress harmonic sources by scheduling tasks at the IT layer while correcting temperature deviations at the facility layer, achieving unprecedented deep coordination between IT loads and physical facilities, and enhancing the operational resilience, stability, and overall energy efficiency of data centers.
[0254] Example 2:
[0255] This embodiment aims to verify the significant advantages of the "linked control module" described in the present invention compared with traditional control strategies in maintaining the stability of the data center microenvironment, improving energy efficiency, and proactively resolving electrical risks.
[0256] Experimental preparation was conducted in a standard high-density computing data center. Two identical server slots, identical in physical location, server hardware configuration, upstream power distribution links, and associated cooling terminals, were selected as test subjects. These slots were designated "Test Server Slot A" (deployed with the linkage control module of the present invention) and "Test Server Slot B" (deployed with a traditional threshold control module). Both slots were equipped with Dell PowerEdge R750 2U rack servers and cooled by the same Emerson Liebert-PEX series precision downflow CRAC unit (with an initial fan speed setting of 50%). The rack PDU was a Schneider AP8853 model, capable of real-time total harmonic distortion (THD) monitoring. Before the experiment began, the prediction module of the present invention predicted the steady-state parameters for both slots after deploying a high-intensity video transcoding task: a server air inlet temperature of 24.0°C and a rack total harmonic distortion of 4.5%. The control parameters for the proposed module are set as follows: a temperature deviation trigger threshold of 1.0°C, PID controller parameters Kp=0.5, Ki=0.1, and Kd=0.2; and a harmonic accumulation risk intervention threshold of 30% / min. The control strategy for the conventional module in comparison is set as follows: when the temperature exceeds 25.0°C, the CRAC fan speed is increased by a fixed 15%; when the THD exceeds 5.5%, only an alarm log is generated.
[0257] Officially starting at time T=0, two identical "high-intensity video transcoding tasks" (with significant nonlinear and perceptual load characteristics) were simultaneously deployed to test server slots A and B. The linkage control module began to continuously monitor and record changes in various environmental parameters on a one-minute cycle and execute the corresponding control logic.
[0258] On the side of slot A of the test server:
[0259] From T=1 to T=2 minutes: Server load rapidly increased, causing the actual inlet air temperature to rise to 25.2°C and the THD to rise to 5.8%. At this point, the temperature deviation (1.2°C) exceeded the threshold, activating the PID control channel. The module calculated the first "cooling control output" using the PID algorithm, fine-tuning the CRAC fan speed to 50.6%. Simultaneously, the harmonic deviation (1.3%) began to be counted towards the "harmonic accumulation risk."
[0260] From T=3 to T=5 minutes: After a slight temperature increase, the PID controller's continuous and gradual adjustments (the fan speed was smoothly increased to 51.8%) stabilized the temperature at 24.9°C at T=5 minutes, achieving rapid, precise, and overshoot-free temperature control. During this period, the "harmonic accumulation risk" continued to accumulate, reaching 32.5%·min at T=5 minutes, successfully exceeding the intervention threshold of 30%·min.
[0261] T = 6 minutes: The harmonic control channel is activated. The module immediately searches the current data center's pool of candidate tasks and, based on harmonic spectrum analysis, identifies an ongoing "Storage Archiving Task" with capacitive load characteristics. The system automatically generates a high-priority migration work order, stating: "Recommend relocating the 'Storage Archiving Task' (ID: SA-081) to rack A07 to offset the 5th and 7th harmonics generated by the 'High-Intensity Video Transcoding Task' (ID: VT-025). This is expected to reduce rack THD by 2.5%."
[0262] On the side of test server slot B (comparison):
[0263] From T=1 to T=2 minutes: The server load also increases. At T=2 minutes, the actual air inlet temperature reaches 25.3°C, exceeding the alarm threshold of 25.0°C. The traditional control module is triggered.
[0264] T = 3 minutes: The module immediately issues a step command, increasing the CRAC fan speed from 50% to 65%. The huge amount of cooling air causes the temperature to drop rapidly, resulting in severe overcooling.
[0265] From T=4 to T=5 minutes, the temperature plummeted to 22.8°C, far below normal operating temperature, resulting in unnecessary energy waste. Because the temperature fell below the 25.0°C threshold, the control module restored the fan speed to 50% at T=5 minutes. Meanwhile, the rack THD reached 6.1% at T=3 minutes, exceeding the 5.5% threshold. The system only generated a text alert in the background stating "Rack B07 THD Exceeds Standard," without any resolution suggestions.
[0266] T = 6 minutes: As the fan speed resumes, the temperature begins to rise rapidly again. The entire temperature control process exhibits a noticeable "oscillation" state, and the environment is extremely unstable. The harmonic problem is completely ignored, posing a persistent safety hazard.
[0267] Through the comparison of the above detailed implementation processes, the present invention demonstrates overwhelming technical advantages in terms of accuracy, stability, energy efficiency and proactive risk management of environmental control.
[0268] Table 1 Study on the effectiveness of linkage control module:
[0269] Parameter name Test server slot A (present invention) Test server slot B (comparison example) Time (minutes) T=0 T=0 Actual air inlet temperature of the server (°C) 22 22 CRAC fan speed adjustment command (%) 50 50 Actual total harmonic distortion rate of the rack (%) 3.5 3.5 Harmonic cumulative risk (% min) 0 0 System response action Task deployment, start monitoring Task deployment, start monitoring Time (minutes) T=2 T=2 Actual air inlet temperature of the server (°C) 25.2 25.3 CRAC fan speed adjustment command (%) 50.6 50 Actual total harmonic distortion rate of the rack (%) 5.8 5.9 Harmonic cumulative risk (% min) 2.5 0 System response action PID control starts, gradual air increase No action, waiting for threshold triggering Time (minutes) T=3 T=3 Actual air inlet temperature of the server (°C) 25.5 22.8 CRAC fan speed adjustment command (%) 51.2 65 Actual total harmonic distortion rate of the rack (%) 6 6.1 Harmonic cumulative risk (% min) 16.5 0 System response action PID continuous adjustment, smooth control Step-up wind, temperature too cold Time (minutes) T=5 T=5 Actual air inlet temperature of the server (°C) 24.9 24.5 CRAC fan speed adjustment command (%) 51.8 50 Actual total harmonic distortion rate of the rack (%) 6.2 6.3 Harmonic cumulative risk (% min) 32.5 0 System response action Temperature stabilization; harmonic intervention threshold triggering The wind boost stops and the temperature starts to rise Time (minutes) T=6 T=6 Actual air inlet temperature of the server (°C) 24.9 25.4 CRAC fan speed adjustment command (%) 51.8 65 Actual total harmonic distortion rate of the rack (%) 6.2 6.4 Harmonic cumulative risk (% min) 34.2 0 System response action Generate harmonic cancellation task migration suggestions Increase the wind speed again and enter oscillation Final Environmental Stability Assessment Stable and optimal energy consumption, with risks actively managed Unstable, oscillating, and risks ignored
[0270] The comparative test data of this embodiment clearly and powerfully demonstrates the significant technical advantages of the "linked control module" of the present invention over the traditional threshold control strategy from the three core dimensions of temperature control accuracy and stability, energy efficiency, and active risk management capabilities.
[0271] 1. Comparative analysis of temperature control accuracy and stability:
[0272] The present invention (test server slot A) demonstrated excellent control accuracy and stability. As can be seen from the data, when the temperature first exceeded the predicted range (25.2°C) at T=2, the present invention's PID control was immediately activated. Its core advantage lies in its progressive regulation: the CRAC fan speed did not increase significantly all at once, but was instead smoothly and subtly adjusted from 50.0% to 50.6%, 51.2%, and finally stabilized at 51.8% at T=5. This fine-tuning based on the size of the deviation, the cumulative amount, and the rate of change allowed the server inlet air temperature to quickly and smoothly return to 24.9°C after a slight peak (25.5°C) and maintain stability. The entire process was free of overcooling, achieving precise stabilization of the microenvironment.
[0273] The comparative example (test server slot B) exposed the inherent flaws of traditional threshold control. At T=3, after the temperature (25.3°C) triggered the threshold, the system implemented step control, sharply increasing the fan speed from 50.0% to 65.0%. This drastic intervention directly led to severe overcooling, with the temperature plummeting to 22.8°C, far below the lower temperature limit required for normal equipment operation. Subsequently, because the temperature fell below the trigger threshold, the system restored the fan speed to 50.0% at T=5, causing the temperature to rise rapidly again to 25.4°C at T=6, triggering another round of increased airflow. This repetitive process caused the server microenvironment to experience severe temperature fluctuations, posing a threat to the long-term reliability of the equipment hardware.
[0274] 2. Comparative analysis of energy efficiency:
[0275] The energy efficiency advantage of this invention (test server slot A) is evident in its on-demand control logic. To maintain a stable temperature of 24.9°C, the system ultimately only needed to maintain a fan speed of 51.8%. This demonstrates that the PID algorithm accurately calculated the minimum cooling capacity required to offset the additional heat load, avoiding any unnecessary energy consumption.
[0276] The control unit (test server slot B) exhibits significant energy waste. First, the overcooling process, which reaches a temperature of 22.8°C, is ineffective cooling output and a pure waste of energy. Second, its control method frequently switches between two states with significantly different energy consumptions: 50.0% and 65.0%. According to the cubic law of fan power and speed, energy consumption at 65.0% speed is much higher than that at 51.8%. Therefore, not only does the control unit waste energy due to overcooling, but its oscillating operating mode also results in higher average operating energy consumption.
[0277] 3. Comparative analysis of proactive risk management capabilities:
[0278] This invention (test server slot A) achieves a leap from passive response to proactive management. The module not only detects excessive harmonic distortion (THD) but also quantifies the persistence and severity of this risk through an innovative "harmonic cumulative risk" parameter. When this cumulative value reaches 32.5%·min at T=5, exceeding the intervention threshold, the system goes beyond the alarm level and automatically generates an actionable solution—recommending the migration of a task with complementary harmonic characteristics. This not only addresses the thermal risk (a temporary solution) through temperature control, but also addresses the root cause of the power quality risk (a fundamental solution) through task scheduling, achieving deep collaboration between IT loads and infrastructure.
[0279] The control unit (test server slot B) was completely passive and isolated in terms of risk management. Although its THD levels continued to exceed the standard (reaching 6.4% at T=6), the system took no further action beyond generating an alarm requiring manual interpretation and resolution. The harmonic risk was completely ignored, persisting and potentially impacting the safe and stable operation of other devices connected to the same PDU. This highlights the inadequacy of traditional control strategies in addressing multi-dimensional, cross-domain risks.
[0280] By introducing PID closed-loop control and a dual-channel linkage mechanism, this system successfully overcomes the oscillation, inefficiency, and risk management blind spots of traditional control strategies. This invention not only achieves a more stable temperature environment with lower energy consumption, but also innovatively incorporates power quality risks into closed-loop control, proactively mitigating these risks through intelligent task scheduling recommendations.
[0281] It should be noted that: All calculation formulas in this application document use regression analysis including but not limited to machine learning algorithms to deeply analyze the relevant parameters collected and identify their natural trends and relationships. Use professional software, such as Python's Scikit-learn library or R language, to automatically generate mathematical models that match the data. Then, objectively evaluate the performance of the model through methods such as cross-validation, and combine continuous feedback and optimization to ensure that the created formula truly reflects the inherent laws of the data, thereby ensuring its effectiveness and accuracy. In all calculation formulas in this application, the parameters in each formula are dimensionally non-dimensionalized within a consistent range to ensure that different physical quantities are compared on the same scale; dimensionless technical means include but are not limited to Min-Max-Normalization and Z-Score standardization;
[0282] The technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random-access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of various embodiments of the present invention.
[0283] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0284] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A management data loading platform based on three-dimensional layout and intelligent optimization algorithm, characterized by: Specifically include: 3D real-time state field construction module: This module is used to obtain a 3D physical layout model of the data center. It periodically obtains rack-level total harmonic distortion (THD) from the intelligent power distribution units deployed on the racks, and simultaneously obtains the real-time power consumption of each server from the server out-of-band management controller to construct a 3D real-time state field containing geometric, electrical, and thermal source information. The electro-thermal coupling heat generation model calibration module is used to query a preset power efficiency function that characterizes the relationship between server power unit efficiency and harmonic distortion based on the acquired rack-level power supply total harmonic distortion rate to calculate the real-time power efficiency of each server location. The module then uses this real-time power efficiency to correct the server's real-time calculated power consumption, thereby generating a predicted total heat generation for the server. The indicator fusion module is used to calculate the "local thermal risk index" after the deployment task for any server slot to be evaluated based on the predicted total heat generation and the three-dimensional physical layout model through computational fluid dynamics simulation. This is combined with the "power quality impact index" and "aerodynamic coupling risk index" of the rack where the server slot is located. The contribution weight of each index is determined through principal component analysis, and finally integrated into a "context-aware deployment suitability index." Optimal Deployment Decision Module: This module is used to formulate the allocation problem between a set of tasks to be deployed and all available server slots in the data center as an integer linear programming model, thereby solving the optimal allocation plan from tasks to server slots. Linkage control module: After completing the task deployment according to the optimal allocation plan, continuously monitor the actual temperature and actual harmonic distortion rate of the deployment location; when the deviation of the actual temperature and actual harmonic distortion rate from the pre-deployment predicted value exceeds the preset threshold, dual-channel linkage control is triggered.
2. The management data loading platform based on three-dimensional layout and intelligent optimization algorithm according to claim 1, characterized in that: Through standardized interfaces and protocols, the basic data required to build digital twin models is collected from the data center infrastructure; The three-dimensional real-time state field construction module is further configured to simultaneously calculate a "dynamic data quality confidence index" associated with each dynamic data stream while obtaining the rack-level power supply total harmonic distortion rate and the real-time calculated power consumption of the server; The 3D physical layout model obtained by the 3D real-time state field construction module includes the centimeter-level precision physical coordinates and spatial topological relationships of the cabinets, servers, cooling units, and power distribution units; The calculation process of the "dynamic data quality credibility index" includes data collection and preprocessing, sub-dimensional quality index calculation, dynamic data quality credibility index fusion, and state field update. The output value range of the "dynamic data quality credibility index" is [0,1]. When the output of "Dynamic Data Quality Credibility Index" is closer to 0, it means that the credibility of the corresponding dynamic data is lower; When the output of "Dynamic Data Quality Credibility Index" is closer to 1, it means that the corresponding dynamic data is more reliable.
3. The management data loading platform based on three-dimensional layout and intelligent optimization algorithm according to claim 2, characterized in that: Quantify the change in power quality as the change in heat generation; the specific implementation process is: First, the platform pre-establishes a "power efficiency function" database, which records the actual operating efficiency of server power supply units at different rack-level power supply total harmonic distortion levels; When the platform is running, for any server, the system will obtain the real-time "rack-level power supply total harmonic distortion rate" of the rack where it is located; The electro-thermal coupling heat generation model calibration module is further configured to: when querying the power efficiency function based on the rack-level power supply total harmonic distortion rate, synchronously obtain the real-time calculated power consumption of the server and combine it with the rated power of its power supply unit to calculate the load factor; and use the rack-level power supply total harmonic distortion rate and the load factor as dual inputs to query a preset three-dimensional efficiency surface model to determine "load-aware real-time power efficiency"; finally, based on the "load-aware real-time power efficiency", calculate the predicted total heat generation of the server.
4. The management data loading platform based on three-dimensional layout and intelligent optimization algorithm according to claim 3 is characterized by: The preset three-dimensional efficiency surface model is a digital twin model generated by matrix testing the power supply units used in the data center at multiple discrete load points and rack-level power total harmonic distortion injection levels in a controlled experimental environment, and using a three-dimensional interpolation or surface fitting algorithm; The predicted total heat production of the server is obtained by input parameter acquisition, key factor calculation, core parameter calibration, heat production calculation and harmonic influence quantification. The key factor is the load factor of the server at the corresponding moment; The harmonic impact is quantified by the server's "harmonic loss factor" at a given moment, which represents the proportion of additional efficiency loss caused by harmonics. The value range of the "harmonic loss factor" is designed to be within the closed interval [0, 1). As the "harmonic loss factor" output approaches 0, the additional energy loss caused by harmonics approaches zero, and the negative impact of power quality on the server's heat dissipation decreases. As the "harmonic loss factor" output approaches 1, the additional energy loss caused by harmonics increases, the efficiency degradation of the server power supply unit increases, and the additional heat generated increases.
5. The management data loading platform based on three-dimensional layout and intelligent optimization algorithm according to claim 4, characterized in that: The platform uses principal component analysis to determine the benchmark contribution weight of each risk index in the indicator fusion module; The indicator fusion module is further configured to: after determining the "baseline contribution weight" of each risk index through principal component analysis, a "contextual risk amplification factor" is calculated in parallel based on the real-time global environmental status of the data center; the contextual risk amplification factor is used to dynamically adjust the baseline contribution weight to generate a set of "dynamic risk weights"; and finally, based on the dynamic risk weights, each risk index is weighted and fused to generate a "context-aware deployment suitability index." 6. The management data loading platform based on three-dimensional layout and intelligent optimization algorithm according to claim 5, characterized in that: The global environmental status includes but is not limited to the average cabinet inlet air temperature, outdoor temperature or grid voltage stability index of the entire data center; The steps for obtaining the context-aware deployment suitability index include quantifying the basic risk index, determining the static benchmark weight, adjusting the dynamic weight, and integrating the final suitability index. Among them, the basic risk index quantification includes the local thermal risk index, power quality impact index and aerodynamic coupling risk index of the server slot; The context-aware deployment suitability index is the actual output with a valid value range of (0,1); As the context-aware deployment suitability index output approaches 0, the server slot is judged to have a higher deployment risk in the current global environment and is less suitable for deploying tasks. As the context-aware deployment suitability index output approaches 1, the server slot is judged to have higher security and stability in the current global environment and is more suitable for deploying tasks.
7. The management data loading platform based on three-dimensional layout and intelligent optimization algorithm according to claim 6, characterized in that: The optimal deployment decision module is further configured to: associate a preset "task business priority coefficient" with each task to be deployed when constructing an integer linear programming model; and modify the model's optimization objective function to maximize the sum of the products of the "context-aware deployment suitability index" of all deployed tasks and the "task business priority coefficient", with the product being recorded as the weighted suitability; The "task business priority coefficient" is configured by the user through a graphical interface; Define a binary "deployment decision variable" ; When task j is assigned to server slot s, The value of is 1; otherwise it is 0; obtain the resource requirements of each task j and the available resource capacity of each server slot s; For the weighted integer linear programming model: maximize the sum of the "weighted fitness" of all existing task-slot assignment combinations; Input the constructed weighted integer linear programming model into a standard integer linear programming solver; Output the optimal solution, that is, a set of optimal "deployment decision variables" The value of ; this set of variable combinations with a value of 1 constitutes the optimal allocation plan from tasks to server slots.
8. The management data loading platform based on three-dimensional layout and intelligent optimization algorithm according to claim 7, characterized in that: The dual-channel linkage control adjusts the operating parameters of the computer room air conditioning unit associated with the server slot to correct the temperature deviation. It also generates a harmonic suppression instruction, which includes a suggestion to migrate a task with complementary harmonic characteristics to that location, to achieve closed-loop adaptive environmental control.
9. The management data loading platform based on three-dimensional layout and intelligent optimization algorithm according to claim 8, characterized in that: The linkage control module is further configured to: when the deviation between the actual temperature and actual harmonic distortion rate and the predicted value exceeds a threshold, based on the proportional-integral-differential control algorithm, use the deviation as input to continuously calculate a quantified "cooling adjustment output"; and based on the output, generate progressive control instructions for the computer room air conditioning unit. At the same time, when the harmonic deviation continues to accumulate and exceeds the integration threshold, it triggers a recommendation for an "active harmonic cancellation task" based on harmonic spectrum feature matching.
10. The management data loading platform based on three-dimensional layout and intelligent optimization algorithm according to claim 9, characterized in that: Calculate the "cooling regulation output" and "harmonic accumulation risk" at the same time; Convert "cooling regulation output" into specific equipment control instructions; Based on the "harmonic accumulation risk level," instructions are issued through the building automation system interface, and a migration recommendation work order for the "active harmonic cancellation task" is pushed to the operation and maintenance system. When the "Cooling Adjustment Output" output is closer to 0, the system environment is more stable and the control system needs fewer intervention measures; As the "cooling regulation output" output approaches its positive upper limit, the system is more likely to experience thermal runaway risk, and the control system will need to take more intervention measures.
Citation Information
Patent Citations
Visual operation and maintenance management method and system for data center
CN114676862A
Circuit board fault diagnosis method and system based on artificial intelligence
CN119203919A
Energy consumption optimization matching system and method for HVAC multi-unit air conditioning system
CN120313176A