AI-based multimodal sensor network cooling management system and method
Through the AI-driven multimodal sensor network cooling management system, cooling parameters are adjusted in real time, solving the shortcomings of traditional systems in heat load adaptability, improving the energy efficiency and reliability of server clusters, and achieving optimal configuration of cooling parameters.
Patent Information
- Application Number
- CN202510294661.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-03-13
AI Technical Summary
Traditional multimodal sensor network cooling management systems lack adaptability to various heat load conditions, resulting in overcooling when the heat load is low and difficulty in responding in a timely manner when the heat load suddenly increases, affecting the energy efficiency and reliability of the equipment.
An AI-based multimodal sensor network cooling management system is adopted. Through the task thermal load classification module, fluid inertia compensation and control module, heat dissipation contribution evaluation module and cooling strategy sharing optimization module, cooling parameters are adjusted in real time, cooling strategies are optimized, and heat dissipation efficiency and energy efficiency are improved.
Targeted cooling measures are implemented according to different task types to improve the overall energy efficiency and reliability of the server. The federated learning mechanism is used to achieve optimal configuration of cooling parameters in the server cluster to ensure overall performance and stability.
Smart Images

Figure CN120161881B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of heat dissipation management technology, and in particular to an AI-based multimodal sensor network cooling management system and method. Background Art
[0002] The field of thermal management technology involves controlling the heat generated during the operation of electronic devices, computing systems, industrial equipment, and other devices to maintain stable operation and improve energy efficiency. This technology covers both passive and active cooling methods, and is optimized by integrating principles such as fluid dynamics, heat transfer, and materials science. In recent years, this field has gradually incorporated technologies such as artificial intelligence, intelligent sensing, and dynamic control to achieve precise temperature management, optimize energy consumption, and improve equipment reliability. Thermal management plays a key role in energy efficiency and equipment lifespan, especially in high-density computing environments.
[0003] The Multimodal Sensor Network Cooling Management System is an intelligent cooling control system based on the fusion of multiple sensor data, specifically designed for server liquid cooling. The system utilizes multi-dimensional sensor data, including temperature, flow rate, and pressure, to analyze coolant status through intelligent algorithms, enabling real-time adjustment of cooling parameters to optimize cooling efficiency, reduce energy consumption, and improve server stability. Primarily used in data centers, cloud computing servers, and high-performance computing environments, it aims to enhance the adaptive control capabilities of liquid cooling systems, ensuring efficient and long-term server operation while reducing energy consumption.
[0004] Traditional management systems rely on a unified cooling strategy and lack adaptability for various heat load conditions. This results in overcooling when the heat load is low, and difficulty responding promptly when the heat load suddenly increases, impacting the energy efficiency and reliability of the equipment. For example, in data centers, the inability to accurately and dynamically assess the heat load of each server can cause some servers to overheat due to insufficient cooling, while other servers overconsume cooling resources under low load. Existing technologies for cooling server clusters lack effective data sharing and coordination mechanisms, preventing adjustments to the overall cooling strategy from being optimized in real time based on the actual operating status of each server, impacting overall cooling efficiency and energy consumption. Summary of the Invention
[0005] The purpose of the present invention is to solve the shortcomings of the prior art and propose an AI-based multimodal sensor network cooling management system and method.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: an AI-based multimodal sensor network cooling management system, the system comprising:
[0007] The task heat load classification module obtains multimodal sensor network data, extracts the task calculation cycle and power consumption accumulation of the server computing task, collects chip temperature change trends, calculates the task heat accumulation rate, and calculates the task heat load information;
[0008] The fluid inertia compensation control module calculates the fluid kinetic energy and flow inertia influence based on the task heat load information, adjusts the coolant flow rate according to the task type, and obtains the optimized cooling flow rate value;
[0009] The heat dissipation contribution evaluation module calculates the temperature drop rate per unit time based on the cooling flow rate optimization value, and calculates the heat dissipation contribution of the server in combination with the energy consumption change before and after the adjustment;
[0010] The cooling strategy sharing optimization module adjusts the server cooling strategy sharing model weight based on the server heat dissipation contribution, selects servers with high contribution as main reference nodes, adjusts the sharing strategy calculation logic, and obtains sharing strategy optimization parameters;
[0011] The cluster thermal management federated computing module constructs a cooling parameter federated learning computing matrix based on the shared strategy optimization parameters, calculates the adaptability of the server cooling execution parameters within the cluster, adjusts the cooling strategy shared parameters, and obtains server cluster cooling control information.
[0012] The present invention has improvements in that the task heat load information specifically includes the calculation cycle heat accumulation ratio, the accumulated power consumption per unit time and the chip temperature change rate; the cooling flow rate optimization value includes the fluid kinetic energy adjustment coefficient, the flow inertia compensation amount and the coolant target flow rate; the server heat dissipation contribution specifically refers to the temperature drop rate per unit time, the cooling energy consumption ratio and the energy consumption change before and after the heat dissipation adjustment; the sharing strategy optimization parameters include the contribution weight adjustment value, the main reference node screening parameters and the sharing strategy calculation correction items; the server cluster cooling control information specifically includes the cooling parameter federation calculation matrix, the server cooling adaptability and the cooling strategy sharing parameters.
[0013] The present invention is improved in that the task heat load classification module includes:
[0014] The task computing feature extraction submodule obtains multimodal sensor network data, extracts the task computing cycle of the server computing task, calculates the cumulative power consumption per unit time, monitors the chip temperature change trend, calculates the chip temperature change rate based on the time series data, and obtains the task computing feature data set;
[0015] The task heat accumulation rate calculation submodule calls the task calculation feature data set and uses the formula:
[0016]
[0017] Calculate and obtain the task heat accumulation rate to obtain the task heat accumulation rate value;
[0018] Among them, R h represents the mission heat accumulation rate, E tot Represents the cumulative power consumption per unit time, T comp Represents the task calculation period, T chip,i represents the chip temperature at the i-th moment, Δt represents the temperature sampling interval, and n represents the length of the time series;
[0019] The task classification determination submodule calls the task heat accumulation rate value, sets the classification threshold, determines the task type based on the ratio, and divides it into short-term high-heat tasks, long-term medium-heat tasks, and low-heat tasks, and establishes task heat load information.
[0020] The present invention is improved in that the fluid inertia compensation control module includes:
[0021] The fluid kinetic energy calculation submodule, based on the task heat load information, calls the server's current coolant flow rate and the specific heat capacity per unit mass of the fluid to calculate the specific kinetic energy per unit mass of the fluid. It also calls the temperature gradient of the heat exchange plate to calculate the energy change during the coolant temperature rise process, and calculates the fluid kinetic energy in combination with the task heat load coefficient.
[0022] The inertia impact analysis submodule calls the fluid kinetic energy to calculate the flow inertia of the coolant in the pipeline, calculates the momentum change rate of the fluid based on the coolant density, flow velocity and fluid cross-sectional area, and combines the temperature gradient of the heat exchange plate to use the formula:
[0023]
[0024] Calculate the influence of flow inertia on the heat transfer process and obtain the flow inertia influence quantity;
[0025] Where I represents the flow inertia influence, m represents the coolant mass per unit volume, v represents the coolant flow rate, ρ represents the coolant density, A represents the cross-sectional area of the coolant pipe, and ΔT p represents the temperature gradient of the heat exchange plate, c p represents the specific heat capacity of the coolant;
[0026] The cooling flow rate control submodule calls the flow inertia influence quantity and adjusts the coolant flow rate according to the task category, including increasing the coolant flow rate for short-term high-heat task servers, periodically adjusting the coolant temperature for long-term medium-heat task servers, and reducing the flow rate for low-heat task servers to obtain the cooling flow rate optimization value.
[0027] The present invention is improved in that the heat dissipation contribution evaluation module includes:
[0028] The temperature change rate calculation submodule obtains the chip temperature data before and after the server cooling system is adjusted based on the cooling flow rate optimization value, calculates the temperature change rate before and after the adjustment, and obtains the chip temperature change rate;
[0029] The cooling energy consumption ratio calculation submodule calls the chip temperature change rate to obtain the server power consumption data before and after the server cooling system is adjusted, using the formula:
[0030]
[0031] Calculate the ratio of the temperature drop rate per unit time to the cooling energy consumption to obtain the cooling energy consumption ratio;
[0032] Among them, R E Represents the cooling energy consumption ratio, ΔT t represents the temperature change rate at time t, E t represents the cooling energy consumption at time t, N represents the total number of data samples, max(ΔT t ) represents the maximum temperature drop rate during the sampling period, max(E t ) represents the maximum cooling energy consumption during the sampling period;
[0033] The heat dissipation contribution calculation submodule calls the cooling energy consumption ratio to obtain the energy consumption data of the server before and after the heat dissipation adjustment, and uses the formula:
[0034]
[0035] Calculate the server heat dissipation contribution;
[0036] Among them, C represents the server heat dissipation contribution, Q before Represents the energy consumption of the cooling system before adjustment, Q after Represents the energy consumption of the cooling system after adjustment, R E Represents the cooling energy consumption ratio.
[0037] The present invention is improved in that the cooling strategy sharing optimization module includes:
[0038] The server contribution weighting submodule obtains the contribution data of each server in the cluster based on the server heat dissipation contribution, calculates the contribution ratio of each server, normalizes the contribution value of each server, calculates the cooling policy impact weight of each server, and obtains the server contribution weight value;
[0039] The shared strategy master node screening submodule calls the server contribution weight value, sets the contribution weight threshold, and screens servers above the threshold as master reference nodes using the formula:
[0040]
[0041] Calculate the influence of the reference node on the overall cooling strategy and obtain the influence value of the main reference node of the shared strategy;
[0042] Among them, W s represents the impact value of the main reference node of the shared strategy, C j represents the heat dissipation contribution of the jth server, P j represents the load power consumption of the jth server, M represents the number of selected main reference servers, Represents the average contribution of the filtering server;
[0043] The sharing strategy optimization calculation submodule calls the influence value of the sharing strategy main reference node, calculates the sharing strategy influence matrix according to the contribution of the main reference node, adjusts the cooling strategy sharing model weight, optimizes the sharing strategy calculation logic, and generates sharing strategy optimization parameters.
[0044] The present invention is improved in that the cluster thermal management federated computing module includes:
[0045] The cooling execution data statistics submodule obtains the cooling execution data of each server in the server cluster based on the shared strategy optimization parameters, records the coolant flow rate, heat exchange plate temperature gradient, server power consumption and temperature drop rate during the execution of the cooling strategy of each server, and establishes a server cooling execution data set;
[0046] The federated computing matrix construction submodule calls the server cooling execution data set to construct a cooling parameter federated learning computing matrix using the formula;
[0047]
[0048] Calculate server cooling suitability;
[0049] Among them, A d Represents the server cooling adaptability, N A Represents the total number of servers participating in the calculation in the server cluster, P A,f represents the power consumption of the fth server, T A,f represents the temperature drop rate of server f, Represents the average power consumption in the server cluster, Represents the average temperature drop rate in the server cluster, represents the standard deviation of power consumption in the server cluster, represents the standard deviation of the temperature drop rate in the server cluster, where e is the base of the natural logarithm;
[0050] The shared parameter optimization submodule calls the server cooling adaptability, adjusts the server cooling strategy shared parameters, sets cluster cooling strategy optimization rules, and obtains server cluster cooling control information.
[0051] An AI-based multimodal sensor network cooling management method is implemented based on the above-mentioned AI-based multimodal sensor network cooling management system, and includes the following steps:
[0052] S1: Acquire multimodal sensor network data, extract the task calculation cycle and power consumption accumulation per unit time of the server computing task, calculate the task heat accumulation rate, and calculate the task heat load information;
[0053] S2: The fluid inertia compensation control module calculates the fluid kinetic energy and flow inertia influence based on the task heat load information, adjusts the coolant flow rate, and obtains the optimized cooling flow rate value;
[0054] S3: The heat dissipation contribution evaluation module calculates the temperature drop rate per unit time based on the cooling flow rate optimization value, and calculates the heat dissipation contribution of the server in combination with the energy consumption change before and after the adjustment;
[0055] S4: The cooling strategy sharing optimization module selects servers with high contribution as primary reference nodes based on the heat dissipation contribution of the servers, adjusts the sharing strategy calculation logic, and obtains sharing strategy optimization parameters;
[0056] S5: The cluster thermal management federated computing module constructs a cooling parameter federated learning computing matrix based on the shared strategy optimization parameters, calculates the adaptability of the server cooling execution parameters within the cluster, adjusts the cooling strategy shared parameters, and obtains server cluster cooling control information.
[0057] Compared with the prior art, the advantages and positive effects of the present invention are:
[0058] In the present invention, by analyzing the ratio of task calculation cycle to cumulative power consumption, the task thermal load is subdivided, so that the cooling parameter adjustment is more in line with the actual thermal load demand, and targeted cooling measures can be implemented according to different task types to improve heat dissipation efficiency. By adjusting the cooling strategy in real time and optimizing the cooling strategy based on the heat dissipation contribution, the overall energy efficiency and reliability of the server are effectively improved. Through the federated computing of cluster thermal management, not only the optimization is achieved within a single server, but also the optimal configuration of cooling parameters is achieved in the entire server cluster through the federated learning mechanism, ensuring the overall performance and stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 is a system flow chart of the present invention;
[0060] Figure 2 This is a flow chart of the task heat load classification module of the present invention;
[0061] Figure 3 This is a flow chart of the fluid inertia compensation control module of the present invention;
[0062] Figure 4 This is a flow chart of the heat dissipation contribution evaluation module of the present invention;
[0063] Figure 5 This is a flow chart of the cooling strategy sharing optimization module of the present invention;
[0064] Figure 6 This is a flow chart of the cluster thermal management federated computing module of the present invention. DETAILED DESCRIPTION
[0065] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0066] In the description of the present invention, it should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings and are only for the convenience of describing the present invention and simplifying the description. They do not indicate or imply that the devices or elements referred to must have a specific direction, be constructed and operate in a specific direction, and therefore should not be understood as limiting the present invention. In addition, in the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0067] See also Figure 1 The present invention provides a technical solution: an AI-based multimodal sensor network cooling management system, the system comprising:
[0068] The task thermal load classification module acquires multimodal sensor network data, extracts the task calculation cycle and cumulative power consumption of server computing tasks, collects chip temperature change trends, extracts task thermal load characteristics, and calculates the task heat accumulation rate based on the ratio of task calculation cycle to cumulative power consumption. The classification threshold is set based on the ratio and tasks are divided into short-term high-heat tasks, long-term medium-heat tasks, and low-heat tasks. Based on the classification results, the task thermal load information is calculated.
[0069] The fluid inertia compensation control module uses the server's current coolant flow rate, heat exchange plate temperature gradient, and fluid unit mass specific heat capacity based on task heat load information to calculate the fluid kinetic energy and flow inertia influence. It then adjusts the coolant flow rate according to the task type to obtain the optimized cooling flow rate value.
[0070] The heat dissipation contribution evaluation module obtains the chip temperature change rate and server power consumption before and after the server cooling system is adjusted based on the cooling flow rate optimization value, calculates the temperature drop rate per unit time, and combines the energy consumption change before and after the server heat dissipation adjustment to calculate the server heat dissipation contribution;
[0071] The cooling strategy sharing optimization module obtains the cooling contribution data of the servers in the cluster based on the server cooling contribution, assigns weights according to the contribution, adjusts the weights of the server cooling strategy sharing model, selects servers with high contribution as the main reference nodes, adjusts the sharing strategy calculation logic, and obtains the sharing strategy optimization parameters;
[0072] The cluster thermal management federated computing module optimizes parameters based on the shared strategy, obtains the cooling execution data of each server in the server cluster, constructs a cooling parameter federated learning calculation matrix, calculates the adaptability of the server cooling execution parameters in the cluster, adjusts the server cooling strategy shared parameters, and obtains the server cluster cooling control information.
[0073] The task thermal load information specifically includes the heat accumulation ratio of the calculation cycle, the accumulated power consumption per unit time, and the chip temperature change rate. The cooling flow rate optimization value includes the fluid kinetic energy adjustment coefficient, the flow inertia compensation amount, and the coolant target flow rate. The server heat dissipation contribution specifically refers to the temperature drop rate per unit time, the cooling energy consumption ratio, and the change in energy consumption before and after the heat dissipation adjustment. The sharing strategy optimization parameters include the contribution weight adjustment value, the main reference node screening parameters, and the sharing strategy calculation correction items. The server cluster cooling control information specifically includes the cooling parameter federation calculation matrix, the server cooling adaptability, and the cooling strategy sharing parameters.
[0074] See also Figure 2 , the task heat load classification module includes:
[0075] The task computing feature extraction submodule obtains multimodal sensor network data, extracts the task computing cycle of the server computing task, calculates the cumulative power consumption per unit time, monitors the chip temperature change trend, calculates the chip temperature change rate based on the time series data, and obtains the task computing feature data set;
[0076] The task computing feature extraction submodule receives data from the multimodal sensor network, including but not limited to real-time power consumption data collected by the power consumption sensor, temperature data obtained by the chip internal temperature sensor, and task scheduling logs on the server side. First, the execution process of the server computing task is analyzed, the task start and end time is read, and the task computing cycle T is calculated. comp As the time difference between the two. For example, if the start time of a task is 12:00:00 and the end time is 12:00:30, then T comp=30s. At the same time, the cumulative power consumption per unit time E in this calculation cycle is calculated. tot , the total power consumption is obtained by accumulating the power consumption measurements per second. For example, if the power consumption of a task in 30 seconds is [8.5W, 8.7W, 8.6W, …, 9.0W], then E tot =∑P i ×Δt, where Δt=1s, and finally we get E tot =258.4J. In addition, the chip temperature change trend is monitored, the temperature values in the time series are read, and the chip temperature change rate is calculated.
[0077] {T chip,0 ,T chip,1 ,...,T chip,n}, calculate the temperature change rate at adjacent moments For example, if the temperature sequence is 50.2℃, 50.5℃, 51.0℃, 51.3℃, the change rate is calculated as [0.3℃ / s, 0.5℃ / s, 0.3℃ / s]. The calculation results are combined to form a task calculation feature data set, including task calculation cycle, power consumption accumulation, temperature change rate and other data items, and stored in the database. In this process, the power consumption accumulation per unit time E tot The measurement standard is based on the design power consumption range of server-side chips, generally selecting 10% to 90% of the chip's peak power consumption as a reasonable measurement range. Taking a common computing chip with a TDP (thermal design power) of 100W as an example, the actual power consumption range in operation is 10W to 90W. Below 10W may indicate idle state, while above 90W may indicate abnormal load. Therefore, if the power consumption data is measured below 10W, it may be under low load or idle state. If it is above 90W, there may be an overload risk, and a comprehensive analysis should be conducted based on the computing cycle.
[0078] The task heat accumulation rate calculation submodule calls the task calculation feature data set and uses the formula:
[0079]
[0080] Calculate and obtain the task heat accumulation rate to obtain the task heat accumulation rate value;
[0081] Among them, R h represents the mission heat accumulation rate, E tot Represents the cumulative power consumption per unit time during the task calculation cycle, T comp Represents the task calculation period, T chip,i represents the chip temperature at the i-th moment, Δt represents the temperature sampling interval, and n represents the length of the time series;
[0082] The task heat accumulation rate calculation submodule receives the task calculation feature data set and uses the formula:
[0083]
[0084] Calculate the power consumption per unit time, that is, Such as E tot =258.4J, T comp =30s, then Next, calculate the sum of the squares of the temperature change rates, such as:
[0085]
[0086] Taking the square root gives Calculate R h =8.613+0.656=9.269, save the result in the database. h When the temperature is increased, the influencing factor of the sum of squares depends on the temperature fluctuation range. If the temperature change rate is large, it indicates that the computing task causes the chip temperature to rise sharply, and the system may have a heat dissipation bottleneck. Therefore, for the temperature change rate calculation part, a correction coefficient α can be set to balance the impact of short-term drastic fluctuations on the overall calculation. The value is usually between 0.5 and 2. The setting of α is based on the chip operating temperature range and the cooling system response time. For example, the normal operating temperature range of a general computing chip is 40°C to 85°C. If the temperature fluctuation exceeds 5°C / s, α is set to 1.5 to increase the temperature change effect on R h If the temperature changes slowly, α is set to 0.8 to reduce the impact of short-term temperature fluctuations. For this calculation, let α = 1.2, then R h The corrected calculation is as follows:
[0087] R h =8.613+1.2×0.656=9.398;
[0088] Final R h =9.398, and the result is stored and used for task classification.
[0089] The task classification determination submodule calls the task heat accumulation rate value, sets the classification threshold, determines the task type based on the ratio, and divides it into short-term high-heat tasks, long-term medium-heat tasks, and low-heat tasks, and establishes task heat load information;
[0090] The task classification determination submodule calls the task heat accumulation rate value R h , classify tasks according to the classification threshold. Set the task classification threshold as follows:
[0091] Short-term high-heat task: R h >10;
[0092] Long-term moderate heat mission: 5≤R h≤10;
[0093] Low Heat Mission: R h <5;
[0094] The classification threshold is set based on the chip's heat dissipation capacity and the type of computing task. Usually, high-heat tasks (R h >10) corresponds to intensive computing loads such as high-performance computing and AI training. Such tasks have high requirements for the cooling system. When scheduling tasks, they should be assigned to computing nodes with active cooling capabilities first. Long-term medium-heat tasks (5≤R h ≤10) are usually ordinary computing tasks, such as database query, video rendering, etc., and the chip temperature is relatively stable. Low-heat tasks (R h <5) are mostly low-power operations such as I / O processing and data access. The classification threshold is adjusted according to chip test data. For example, in the test environment of Intel Xeon processor, the R corresponding to different task loads is h Typical values are: Short-term high-temperature task R h About 12-15, long-term medium heat task R h About 6-9, low heat task R h About 2-4, so the classification threshold is set as above.
[0095] According to the calculation results R h =9.398, the task is judged as a long-term medium-heat task and is recorded in the task heat load information database.
[0096] Table 1: Example of mission heat accumulation rate calculation
[0097]
[0098] As shown in Table 1, the heat accumulation rate calculation process of task number 1 complies with the set formula, and the calculated R h =9.398, and the classification result is a long-term medium-heat mission.
[0099] See also Figure 3 , the fluid inertia compensation control module includes:
[0100] The fluid kinetic energy calculation submodule uses the server's current coolant flow rate and the specific heat capacity per unit mass of the fluid based on the task heat load information to calculate the specific kinetic energy per unit mass of the fluid. It also uses the temperature gradient of the heat exchange plate to calculate the energy change during the coolant temperature rise process, and combines this with the task heat load coefficient to calculate the fluid kinetic energy.
[0101] The fluid kinetic energy calculation submodule calls the server's current coolant flow rate and specific heat capacity per unit mass based on the task heat load information. First, read the current coolant flow rate v. For example, if the flow rate in the coolant pipeline is set to 2.5m / s, the specific heat capacity c is called at the same time. pAs a known physical property of the coolant, for example, 4.18 kJ / (kg·K) (taking water as the coolant), calculate the specific kinetic energy per unit mass of the fluid. Specifically, calculate the kinetic energy The unit mass of the fluid m can be directly taken as 1 kg, and the calculation is:
[0102]
[0103] Next, call the heat exchange plate temperature gradient ΔT p If the temperature gradient of the heat exchange plate is measured to be 15K, calculate the energy change absorbed by the coolant during the temperature rise process, using Q=mc p ΔT p Calculate, where m is set to 1kg, then:
[0104] Q = 1 × 4.18 × 15 = 62.7 J;
[0105] Combined mission heat load coefficient α h The heat load coefficient is used to correct the actual kinetic energy contribution of the coolant. It is set according to the average calculated load ratio of the server. For example, when the server is in full load calculation state, α h Take 1, when the server load is less than 30%, α h Take 0.7, between the load range of 30%-100%, the coefficient is calculated according to the linear relationship with the load change. For example, if the current server load rate is 85%, then calculate Substitute it into the calculation and we get:
[0106] KE′=3.125+0.85×62.7=56.43J;
[0107] The fluid kinetic energy is 56.43 J, and the calculation result is stored in the database.
[0108] The inertia impact analysis submodule uses fluid kinetic energy to calculate the flow inertia of the coolant in the pipeline. It calculates the momentum change rate of the fluid based on the coolant density, flow velocity, and fluid cross-sectional area, and combines it with the temperature gradient of the heat exchange plate using the formula:
[0109]
[0110] Calculate the influence of flow inertia on the heat transfer process and obtain the flow inertia influence quantity;
[0111] Where I represents the flow inertia influence, m represents the coolant mass per unit volume, v represents the coolant flow rate, ρ represents the coolant density, A represents the cross-sectional area of the coolant pipe, and ΔT p represents the temperature gradient of the heat exchange plate, c p represents the specific heat capacity of the coolant;
[0112] The inertia impact analysis submodule calls the fluid kinetic energy to calculate the flow inertia of the coolant in the pipeline. First, the momentum change rate of the coolant is calculated based on the coolant density ρ, flow velocity v and fluid cross-sectional area A. For example, the water density ρ is set to 1000kg / m3, the pipeline cross-sectional area A is set to 0.005m2, and the flow velocity v is set to 2.5m / s.
[0113] Combined with the heat exchange plate temperature gradient ΔT p =15K, using the formula:
[0114]
[0115] Bring in the calculated parameters:
[0116]
[0117] Where, ΔT p The temperature gradient of the heat exchange plate is set based on the temperature difference design benchmark of the cooling system. This benchmark is determined by the maximum operating temperature of the server and the coolant inlet temperature. For example, if the maximum allowable temperature of the server is set to 80°C and the coolant inlet temperature is usually maintained below 65°C, the reasonable heat exchange plate temperature gradient should meet 10K≤ΔT p ≤20K, within this range, ΔT p It can be adjusted according to the heat situation of specific tasks. For example, when the server runs high-load tasks, ΔT p The temperature gradient is set closer to 20 K, and closer to 10 K. In this example, 15 K is used as the actual temperature gradient. Finally, the flow inertia influence I = 48.01 is calculated and stored in the database.
[0118] The cooling flow rate control submodule calls the flow inertia influence variable and adjusts the coolant flow rate according to the task type. This includes increasing the coolant flow rate for servers with short-term high-heat tasks, periodically adjusting the coolant temperature for servers with long-term medium-heat tasks, and reducing the flow rate for servers with low-heat tasks to obtain the optimal cooling flow rate value.
[0119] The cooling flow rate control submodule calls the flow inertia influence and adjusts the coolant flow rate according to the task category. First, the flow rate control rules are set according to the task classification results. The classification standards are as follows: short-term high-temperature tasks (R h >10): Increase the coolant flow rate v′=v+1m / s for long-term medium heat tasks (5≤R h ≤10): Periodically adjust the coolant temperature, such as adjusting ±2K low-heat task (R h <5): Reduce the coolant flow rate v′=v-0.5m / s. Assume that this task is classified as a long-term medium-heat task, then adjust the coolant temperature periodically, adjusting ±2K every 10 minutes. Assume that the current coolant initial temperature T c=20℃, then the coolant temperature in the next cycle is adjusted to T c ′ = 20 ± 2°C, i.e., 18°C or 22°C. During the cooling flow rate control process, the coolant flow rate adjustment threshold is set based on the inertia influence I. The threshold setting range is based on the cooling system's response time to flow rate changes and is generally set within the range of 5 ≤ I ≤ 50. When I > 50, the flow rate adjustment range increases by 20%, and when I < 5, the flow rate adjustment range decreases by 10%. In this example, I = 48.01, close to 50, so the flow rate adjustment threshold takes the maximum adjustment range, that is, the periodic temperature adjustment is maintained at ± 2K without changing the flow rate.
[0120] Table 2: Example of inertia effect calculation
[0121]
[0122] As shown in Table 2, the calculated result of the flow inertia influence of task number 1 is 48.01, which is consistent with the calculation formula. It is finally used for cooling flow rate control. According to the flow rate adjustment threshold, the coolant temperature adjustment range of this task in the next cycle is set to ±2K.
[0123] See also Figure 4 , the heat dissipation contribution evaluation module includes:
[0124] The temperature change rate calculation submodule obtains the chip temperature data before and after the server cooling system is adjusted based on the cooling flow rate optimization value, calculates the temperature change rate before and after the adjustment, and obtains the chip temperature change rate;
[0125] The temperature change rate calculation submodule obtains the chip temperature data before and after the server cooling system is adjusted based on the cooling flow rate optimization value. First, the chip temperature data before adjustment is read. before And the adjusted chip temperature data T after The sampling interval is set to Δt = 5s to ensure the stability of data sampling. For example, the chip temperature before adjustment is [65.2℃, 64.8℃, 64.5℃, 64.0℃], and the temperature after adjustment is [63.0℃, 62.5℃, 62.2℃, 62.0℃]. To calculate the temperature change rate before and after adjustment, use:
[0126]
[0127] Calculation before adjustment:
[0128]
[0129] Adjusted calculation:
[0130]
[0131] The chip temperature change rate is calculated. The average temperature change rate before adjustment is -0.08℃ / s, and after adjustment is -0.066℃ / s. The chip temperature change rate reference value T is set. rate,thresh According to the data center server operating temperature standard, the temperature change rate is generally controlled between [-0.1℃ / s, -0.05℃ / s]. If the temperature change rate is lower than -0.1℃ / s, the coolant may cool down too quickly, affecting the server temperature stability. If the temperature change rate is higher than -0.05℃ / s, the cooling efficiency may be low. Therefore, the reference value T is set. rate,thresh =-0.08℃ / s, which is in line with the temperature change rate before and after adjustment within this range.
[0132] The cooling energy consumption ratio calculation submodule calls the chip temperature change rate to obtain the server power consumption data before and after the server cooling system is adjusted, using the formula:
[0133]
[0134] Calculate the ratio of the temperature drop rate per unit time to the cooling energy consumption to obtain the cooling energy consumption ratio;
[0135] Among them, R E Represents the cooling energy consumption ratio, ΔT t represents the temperature change rate at time t, E t represents the cooling energy consumption at time t, N represents the total number of data samples, max(ΔT t ) represents the maximum temperature drop rate during the sampling period, max(E t ) represents the maximum cooling energy consumption during the sampling period;
[0136] The cooling energy consumption ratio calculation submodule calls the chip temperature change rate to obtain the server power consumption data before and after the server cooling system is adjusted. The total number of data samples is set to N = 3, and the collected cooling energy consumption data E t Assume that before adjustment [30W, 32W, 31W] and after adjustment [29W, 30W, 29.5W], and calculate:
[0137]
[0138] max(ΔT t )=-0.06℃ / s;
[0139] max(E t )=32W;
[0140] Substitute into the formula:
[0141]
[0142] Get the cooling energy consumption ratio RE =0.000048, set cooling energy consumption ratio weight W E According to the energy efficiency ratio of the server cooling system, in general, the energy efficiency ratio of the data center cooling system ranges from 0.00002 to 0.0001, so the cooling energy consumption ratio weight W is set. E =0.000048, this value is used to calculate the energy efficiency of the cooling system. If R E If it is lower than 0.00002, the cooling system efficiency is too low. If it is higher than 0.0001, it may cause the cooling system to over-operate and affect the power consumption balance.
[0143] The heat dissipation contribution calculation submodule calls the cooling energy consumption ratio to obtain the energy consumption data before and after the server heat dissipation adjustment. According to the energy consumption reduction after adjustment, the formula is used:
[0144]
[0145] Calculate the server heat dissipation contribution;
[0146] Among them, C represents the server heat dissipation contribution, Q before Represents the energy consumption of the cooling system before adjustment, Q after Represents the energy consumption of the cooling system after adjustment, R E Represents the cooling energy consumption ratio;
[0147] The heat dissipation contribution calculation submodule calls the cooling energy consumption ratio to obtain the energy consumption data before and after the server heat dissipation adjustment, and sets the cooling system energy consumption Q before adjustment. before =5000J, the energy consumption of the cooling system after adjustment is Q after =4800J, calculate:
[0148]
[0149] Get the server heat dissipation contribution C = 0.00000192, set the heat dissipation contribution coefficient K C According to the cooling system optimization goal, the general setting range is between 0.000001 and 0.00001. If the heat dissipation contribution coefficient is lower than 0.000001, it means that the cooling optimization effect is weak. If it is higher than 0.00001, it may lead to insignificant energy consumption optimization. Therefore, set K C =0.00000192, which meets the reasonable heat dissipation contribution evaluation standard.
[0150] Table 3: Cooling energy consumption ratio calculation example
[0151]
[0152] As shown in Table 3, after the cooling adjustment, the temperature change rate decreased and the cooling energy consumption decreased. The cooling energy consumption ratio was calculated to be 0.000048 and used for the heat dissipation contribution calculation, where the cooling energy consumption ratio weight W E And the heat dissipation contribution coefficient K C Set cooling system evaluation criteria to ensure that the calculation results meet the cooling optimization needs of the data center.
[0153] See also Figure 5 , the cooling strategy sharing optimization module includes:
[0154] The server contribution weighting submodule obtains the contribution data of each server in the cluster based on the server heat dissipation contribution, calculates the contribution ratio of each server, normalizes the contribution value of each server, calculates the cooling policy impact weight of each server, and obtains the server contribution weight value;
[0155] The server contribution weighting submodule obtains the contribution data of each server in the cluster based on the server heat dissipation contribution. First, the heat dissipation contribution C of all servers in the cluster is called. j and load power consumption P j , if the cluster contains M = 4 servers, their heat dissipation contributions are [0.00000192, 0.00000210, 0.00000175, 0.00000225], and their load power consumption is [300W, 320W, 280W, 350W]. Calculate the contribution ratio of each server:
[0156]
[0157] Calculated:
[0158]
[0159]
[0160] Contribution percentage of each server:
[0161]
[0162] Normalize the contribution value of each server, calculate the cooling policy impact weight of the server, obtain the server contribution weight value, and store it in the database. The contribution weight value is set based on the relative influence relationship between the server's heat dissipation contribution and load power consumption. This value fluctuates with the actual load of the server. The calculation method uses the normalization method to ensure that the sum of the contribution weights of all servers is 1. Normalization uses the maximum and minimum normalization method, that is,
[0163]
[0164] Calculated based on server contribution data:
[0165] max(C)=0.00000225;
[0166] min(C)=0.00000175;
[0167]
[0168] Obtain the normalized contribution weight value of each server to ensure that the value is within the range and is not absolutely affected by the contribution value of a single server.
[0169] The shared strategy master node screening submodule calls the server contribution weight value, sets the contribution weight threshold, and screens servers above the threshold as the main reference nodes using the formula:
[0170]
[0171] Calculate the influence of the reference node on the overall cooling strategy and obtain the influence value of the main reference node of the shared strategy;
[0172] Among them, W s represents the impact value of the main reference node of the shared strategy, C j represents the heat dissipation contribution of the jth server, P j represents the load power consumption of the jth server, M represents the number of selected main reference servers, Represents the average contribution of the filtering server;
[0173] The shared strategy master node screening submodule calls the server contribution weight value, sets the contribution weight threshold, and screens servers above the threshold as the main reference nodes. Suppose the contribution weight threshold is set to 0.25, and server 2 (0.2619) and server 4 (0.2803) are screened as the main reference nodes. The formula is:
[0174]
[0175] calculate:
[0176]
[0177] Calculate the contribution variance:
[0178]
[0179] Find the square root:
[0180]
[0181] Calculate the final impact value:
[0182]
[0183] Get the influence value W of the main reference node of the shared strategy s =2.02×10 -6 ;
[0184] The contribution weight threshold is set based on the distribution characteristics of the heat dissipation contribution within the server cluster, and is specifically set using the mean plus standard deviation method of the overall contribution, i.e.
[0185]
[0186] Where σ is the standard deviation, according to the above calculation, Standard deviation σ=5.96×10 -7 ,final
[0187] θ=0.000002005+5.96×10 -7 =0.000002601;
[0188] Servers with contributions higher than this threshold are selected as primary reference nodes.
[0189] The sharing strategy optimization calculation submodule calls the influence value of the sharing strategy master reference node, calculates the sharing strategy influence matrix based on the contribution of the master reference node, adjusts the cooling strategy sharing model weight, optimizes the sharing strategy calculation logic, and generates sharing strategy optimization parameters;
[0190] The shared strategy optimization calculation submodule calls the shared strategy main reference node influence value, calculates the shared strategy influence matrix based on the contribution of the main reference node, and sets the influence matrix element M ij According to the influence relationship between server i and server j, such as:
[0191]
[0192] Adjust the cooling strategy sharing model weights, use the optimized parameter matrix W, and set:
[0193]
[0194] The weight matrix W in the shared strategy optimization calculation is obtained by normalizing the contribution of the main reference node. Each weight value ensures that the sum of the row and column is 1, and is dynamically adjusted with the contribution of the main reference node. The specific value is based on the proportion of the main node's influence on the overall cooling strategy.
[0195] See also Figure 6 ,The cluster thermal management federated computing module includes:
[0196] The cooling execution data statistics submodule obtains the cooling execution data of each server in the server cluster based on the shared strategy optimization parameters, records the coolant flow rate, heat exchange plate temperature gradient, server power consumption and temperature drop rate during the execution of each server's cooling strategy, and establishes a server cooling execution data set;
[0197] The cooling execution data statistics submodule optimizes parameters based on a shared strategy and obtains cooling execution data for each server in the server cluster. First, the coolant flow rate, heat exchange plate temperature gradient, server power consumption, and temperature drop rate of each server in the server cluster are collected. Sampling is set every 5 seconds to ensure data continuity and accuracy. For example, the coolant flow rate of a certain server within 60 seconds is recorded as 2.5, 2.6, 2.4, 2.7, 2.5] m / s, the heat exchange plate temperature gradient is recorded as [12, 13, 11, 14, 12] K, the server power consumption is recorded as [320, 315, 318, 322, 319] W, and the temperature drop rate is recorded as [-0.10, -0.12, -0.08, -0.14, -0.11] °C / s. The data of all servers are integrated to establish a server cooling execution dataset and store it in the database. The sampling period of the coolant flow rate is set according to the weight value of the server heat dissipation contribution. The weight value is calculated by the heat dissipation contribution of each server and converted into a relative weight in the range of 0-1 through normalization calculation, which is then used to allocate the coolant flow rate. The server heat dissipation contribution weight value W is set. C The calculation formula is as follows:
[0198]
[0199] Assume that the heat dissipation contributions of the servers in the cluster are [0.00000192, 0.00000210, 0.00000175, 0.00000225], and calculate their normalized weights:
[0200]
[0201]
[0202] The weight value determines the server cooling policy allocation. For example, server 4 has the highest contribution, and its coolant flow rate allocation weight is higher than that of other servers, which is ultimately used to optimize the allocation of cooling resources.
[0203] The federated computation matrix construction submodule calls the server cooling execution dataset and constructs the cooling parameter federated learning computation matrix using the formula;
[0204]
[0205] Calculate server cooling suitability;
[0206] Among them, A d Represents the server cooling adaptability, N A Represents the total number of servers participating in the calculation in the server cluster, P A,f represents the power consumption of the fth server, T A,f represents the temperature drop rate of server f, Represents the average power consumption in the server cluster, Represents the average temperature drop rate in the server cluster, represents the standard deviation of power consumption in the server cluster, represents the standard deviation of the temperature drop rate in the server cluster, where e is the base of the natural logarithm;
[0207] The federated computing matrix construction submodule calls the server cooling execution dataset to construct the cooling parameter federated learning computing matrix. First, the average power consumption in the server cluster is calculated. and standard deviation Assume that there are N A = 4 servers, their power consumption is recorded as [320, 310, 330, 325] W, and the temperature drop rate is recorded as [-0.10, -0.09, -0.12, -0.11] ℃ / s, then:
[0208]
[0209] Calculate the mean temperature drop rate and standard deviation
[0210]
[0211] Substitute into the formula:
[0212]
[0213] Calculate one by one:
[0214]
[0215]
[0216] Similarly, calculate the values of other servers and find the average, and finally calculate the server cooling adaptability A d =
[0217] 0.689. Average power consumption and standard deviation As a benchmark value, it represents the fluctuation range of server power consumption. Its standard deviation reflects the degree of dispersion of power consumption distribution. This shows that the power consumption of servers in the cluster varies greatly. This indicates that the power consumption is stable. This benchmark value is used to measure the rationality of the cooling adaptability calculation.
[0218] The shared parameter optimization submodule calls the server cooling adaptability, adjusts the server cooling strategy shared parameters, sets cluster cooling strategy optimization rules, and obtains server cluster cooling control information;
[0219] The shared parameter optimization submodule calls the server cooling fitness, adjusts the server cooling strategy shared parameters, sets the cluster cooling strategy optimization rules, and sets the server cooling strategy shared parameter matrix S according to the cooling fitness weight of each server:
[0220]
[0221] The adjusted cooling strategy control parameter S′ is calculated as follows:
[0222]
[0223] The weight values in the shared parameter matrix S are calculated based on the server cooling fitness, and its optimization rule is based on the server cooling fitness A d Scaling is performed, the scaling factor A d =0.689 as the cooling adaptation metric, and proportionally adjust the weight matrix to ensure a reasonable allocation of cooling resources. Finally, the server cluster cooling control information is obtained and stored in the database.
[0224] The AI-based multimodal sensor network cooling management method includes the following steps:
[0225] S1: Acquire multimodal sensor network data, extract the task calculation cycle and cumulative power consumption per unit time of the server computing task, collect chip temperature change trends, calculate the task heat accumulation rate, and calculate the task thermal load information;
[0226] S2: The fluid inertia compensation control module calculates the fluid kinetic energy and flow inertia influence based on the task heat load information, adjusts the coolant flow rate according to the task type, and obtains the optimized cooling flow rate value;
[0227] S3: The heat dissipation contribution evaluation module calculates the temperature drop rate per unit time based on the optimized cooling flow rate. Combined with the energy consumption changes before and after the adjustment, it calculates the server's heat dissipation contribution.
[0228] S4: The cooling strategy sharing optimization module adjusts the server cooling strategy sharing model weights based on the server heat dissipation contribution, selects servers with high contribution as primary reference nodes, adjusts the sharing strategy calculation logic, and obtains the sharing strategy optimization parameters;
[0229] S5: The cluster thermal management federated computing module optimizes parameters based on the shared strategy, constructs a cooling parameter federated learning computing matrix, calculates the adaptability of the server cooling execution parameters within the cluster, adjusts the cooling strategy shared parameters, and obtains the server cluster cooling control information.
[0230] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. AI-based multimodal sensor network cooling management system, characterized by: The system comprises: The task heat load classification module obtains multimodal sensor network data, extracts the task calculation cycle and power consumption accumulation of the server computing task, collects chip temperature change trends, calculates the task heat accumulation rate, and calculates the task heat load information; The fluid inertia compensation control module calculates the fluid kinetic energy and flow inertia influence based on the task heat load information, adjusts the coolant flow rate according to the task type, and obtains the optimized cooling flow rate value; The heat dissipation contribution evaluation module calculates the temperature drop rate per unit time based on the cooling flow rate optimization value, and calculates the heat dissipation contribution of the server in combination with the energy consumption change before and after the adjustment; The cooling strategy sharing optimization module adjusts the server cooling strategy sharing model weight based on the server heat dissipation contribution, selects servers with high contribution as main reference nodes, adjusts the sharing strategy calculation logic, and obtains sharing strategy optimization parameters; The cluster thermal management federated computing module constructs a cooling parameter federated learning computing matrix based on the shared strategy optimization parameters, calculates the adaptability of the server cooling execution parameters within the cluster, adjusts the cooling strategy shared parameters, and obtains server cluster cooling control information.
2. The AI-based multimodal sensor network cooling management system according to claim 1, characterized in that: The task thermal load information specifically includes the calculation cycle heat accumulation ratio, the accumulated power consumption per unit time, and the chip temperature change rate. The cooling flow rate optimization value includes the fluid kinetic energy adjustment coefficient, the flow inertia compensation amount, and the coolant target flow rate. The server heat dissipation contribution specifically refers to the temperature drop rate per unit time, the cooling energy consumption ratio, and the change in energy consumption before and after the heat dissipation adjustment. The sharing strategy optimization parameters include the contribution weight adjustment value, the main reference node screening parameters, and the sharing strategy calculation correction items. The server cluster cooling control information specifically includes the cooling parameter federation calculation matrix, the server cooling adaptability, and the cooling strategy sharing parameters.
3. The AI-based multimodal sensor network cooling management system according to claim 1, characterized in that: The task heat load classification module includes: The task computing feature extraction submodule obtains multimodal sensor network data, extracts the task computing cycle of the server computing task, calculates the cumulative power consumption per unit time, monitors the chip temperature change trend, calculates the chip temperature change rate based on the time series data, and obtains the task computing feature data set; The task heat accumulation rate calculation submodule calls the task calculation feature data set and uses the formula: Calculate and obtain the task heat accumulation rate to obtain the task heat accumulation rate value; Among them, R h represents the mission heat accumulation rate, E tot Represents the cumulative power consumption per unit time during the task calculation cycle, T comp Represents the task calculation period, T chip,i represents the chip temperature at the i-th moment, Δt represents the temperature sampling interval, and n represents the length of the time series; The task classification determination submodule calls the task heat accumulation rate value, sets the classification threshold, determines the task type based on the ratio, and divides it into short-term high-heat tasks, long-term medium-heat tasks, and low-heat tasks, and establishes task heat load information.
4. The AI-based multimodal sensor network cooling management system according to claim 1, characterized in that: The fluid inertia compensation control module includes: The fluid kinetic energy calculation submodule, based on the task heat load information, calls the server's current coolant flow rate and the specific heat capacity per unit mass of the fluid to calculate the specific kinetic energy per unit mass of the fluid. It also calls the temperature gradient of the heat exchange plate to calculate the energy change during the coolant temperature rise process, and calculates the fluid kinetic energy in combination with the task heat load coefficient. The inertia impact analysis submodule calls the fluid kinetic energy to calculate the flow inertia of the coolant in the pipeline, calculates the momentum change rate of the fluid based on the coolant density, flow velocity and fluid cross-sectional area, and combines the temperature gradient of the heat exchange plate to use the formula: Calculate the influence of flow inertia on the heat transfer process and obtain the flow inertia influence quantity; Where I represents the flow inertia influence, m represents the coolant mass per unit volume, v represents the coolant flow rate, ρ represents the coolant density, A represents the cross-sectional area of the coolant pipe, and ΔT p represents the temperature gradient of the heat exchange plate, c p represents the specific heat capacity of the coolant; The cooling flow rate control submodule calls the flow inertia influence quantity and adjusts the coolant flow rate according to the task category, including increasing the coolant flow rate for short-term high-heat task servers, periodically adjusting the coolant temperature for long-term medium-heat task servers, and reducing the flow rate for low-heat task servers to obtain the cooling flow rate optimization value.
5. The AI-based multimodal sensor network cooling management system according to claim 1, characterized in that: The heat dissipation contribution evaluation module includes: The temperature change rate calculation submodule obtains the chip temperature data before and after the server cooling system is adjusted based on the cooling flow rate optimization value, calculates the temperature change rate before and after the adjustment, and obtains the chip temperature change rate; The cooling energy consumption ratio calculation submodule calls the chip temperature change rate to obtain the server power consumption data before and after the server cooling system is adjusted, using the formula: Calculate the ratio of the temperature drop rate per unit time to the cooling energy consumption to obtain the cooling energy consumption ratio; Among them, R E Represents the cooling energy consumption ratio, ΔT t represents the temperature change rate at time t, E t represents the cooling energy consumption at time t, N represents the total number of data samples, max(ΔT t ) represents the maximum temperature drop rate during the sampling period, max(E t ) represents the maximum cooling energy consumption during the sampling period; The heat dissipation contribution calculation submodule calls the cooling energy consumption ratio to obtain the energy consumption data of the server before and after the heat dissipation adjustment, and uses the formula: Calculate the server heat dissipation contribution; Among them, C represents the server heat dissipation contribution, Q before Represents the energy consumption of the cooling system before adjustment, Q after Represents the energy consumption of the cooling system after adjustment, R E Represents the cooling energy consumption ratio.
6. The AI-based multimodal sensor network cooling management system according to claim 1, characterized in that: The cooling strategy sharing optimization module includes: The server contribution weighting submodule obtains the contribution data of each server in the cluster based on the server heat dissipation contribution, calculates the contribution ratio of each server, normalizes the contribution value of each server, calculates the cooling policy impact weight of each server, and obtains the server contribution weight value; The shared strategy master node screening submodule calls the server contribution weight value, sets the contribution weight threshold, and screens servers above the threshold as master reference nodes using the formula: Calculate the influence of the reference node on the overall cooling strategy and obtain the influence value of the main reference node of the shared strategy; Among them, W s represents the impact value of the main reference node of the shared strategy, C j represents the heat dissipation contribution of the jth server, P j represents the load power consumption of the jth server, M represents the number of selected main reference servers, Represents the average contribution of the filtering server; The sharing strategy optimization calculation submodule calls the influence value of the sharing strategy main reference node, calculates the sharing strategy influence matrix according to the contribution of the main reference node, adjusts the cooling strategy sharing model weight, optimizes the sharing strategy calculation logic, and generates sharing strategy optimization parameters.
7. The AI-based multimodal sensor network cooling management system according to claim 1, characterized in that: The cluster thermal management federated computing module includes: The cooling execution data statistics submodule obtains the cooling execution data of each server in the server cluster based on the shared strategy optimization parameters, records the coolant flow rate, heat exchange plate temperature gradient, server power consumption and temperature drop rate during the execution of the cooling strategy of each server, and establishes a server cooling execution data set; The federated computing matrix construction submodule calls the server cooling execution data set to construct a cooling parameter federated learning computing matrix using the formula; Calculate server cooling suitability; Among them, A d Represents the server cooling adaptability, N A Represents the total number of servers participating in the calculation in the server cluster, P A,f represents the power consumption of the fth server, T A,f represents the temperature drop rate of server f, Represents the average power consumption in the server cluster, Represents the average temperature drop rate in the server cluster, represents the standard deviation of power consumption in the server cluster, represents the standard deviation of the temperature drop rate in the server cluster, where e is the base of the natural logarithm; The shared parameter optimization submodule calls the server cooling adaptability, adjusts the server cooling strategy shared parameters, sets cluster cooling strategy optimization rules, and obtains server cluster cooling control information.
8. AI-based multimodal sensor network cooling management method, characterized in that: The AI-based multimodal sensor network cooling management system according to any one of claims 1 to 7 is implemented, comprising the following steps: S1: Acquire multimodal sensor network data, extract the task calculation cycle and power consumption accumulation per unit time of the server computing task, calculate the task heat accumulation rate, and calculate the task heat load information; S2: The fluid inertia compensation control module calculates the fluid kinetic energy and flow inertia influence based on the task heat load information, adjusts the coolant flow rate, and obtains the optimized cooling flow rate value; S3: The heat dissipation contribution evaluation module calculates the temperature drop rate per unit time based on the cooling flow rate optimization value, and calculates the heat dissipation contribution of the server in combination with the energy consumption change before and after the adjustment; S4: The cooling strategy sharing optimization module selects servers with high contribution as primary reference nodes based on the heat dissipation contribution of the servers, adjusts the sharing strategy calculation logic, and obtains sharing strategy optimization parameters; S5: The cluster thermal management federated computing module constructs a cooling parameter federated learning computing matrix based on the shared strategy optimization parameters, calculates the adaptability of the server cooling execution parameters within the cluster, adjusts the cooling strategy shared parameters, and obtains server cluster cooling control information.
Citation Information
Patent Citations
Technologies for assigning workloads based on resource utilization phases
CN109417564A
Heat dissipation control method and device of storage server, equipment and storage medium
CN119440197A