Liquid-cooled cabinet based on double-circulation refrigerating system and use method of liquid-cooled cabinet

By introducing a dual-circulation refrigeration system and a joint precision-cooling capacity optimization controller in the liquid cooling system, combined with graph structure analysis and node importance evaluation, fine cold capacity allocation of non-uniform thermal loads in graph calculation is achieved, solving the problem that system performance and energy efficiency are difficult to be globally optimized in the prior art, and computing performance improvement and energy efficiency optimization are achieved.

CN120018462APending Publication Date: 2025-05-16深圳市前海嘉信科技有限公司
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510237261.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-02
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing liquid cooling system cannot effectively perform fine-tuning cooling capacity distribution for the non-uniform thermal load distribution in the graph calculation, resulting in more uneven thermal load distribution within the system, making it difficult to achieve global optimal system performance and energy efficiency.

Method used

A liquid cooling cabinet based on a dual-circulation refrigeration system is designed. By analyzing the characteristics of the graph structure and evaluating the importance of nodes, a three-dimensional correlation model of calculation accuracy-energy consumption-temperature is constructed to generate an accuracy distribution strategy for graph structure perception, and based on this, a cooling capacity distribution strategy of the liquid cooling system is designed to achieve a joint optimization controller for accuracy-cold capacity to dynamically adapt to graph structure changes.

Benefits of technology

It has achieved computational performance improvement, energy efficiency optimization, thermal safety assurance, dynamic adaptability and scalability, and can effectively coordinate the optimization of computing accuracy and cooling resources in graph computing scenarios to improve the overall performance and energy efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120018462A_ABST
    Figure CN120018462A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of liquid-cooled cabinets, and discloses a liquid-cooled cabinet based on a dual-cycle refrigeration system and a use method thereof, and the method comprises the following steps: analyzing graph structure features and evaluating node importance; constructing a calculation precision-energy consumption-temperature three-dimensional correlation model; generating a precision distribution strategy of graph structure perception based on node importance; designing a cooling capacity distribution strategy of the liquid cooling system based on the precision distribution strategy; a precision-cooling capacity combined optimization controller is realized; designing a dynamic graph structure adaptation mechanism; wherein the precision distribution strategy distributes different calculation precision to different nodes in the graph based on node importance scores, the cooling capacity distribution strategy dynamically adjusts the cooling capacity according to thermal load distribution caused by precision distribution, and the precision-cooling capacity joint optimization controller cooperatively adjusts the calculation precision and the cooling capacity distribution according to a real-time system state. Dynamically adapting to graph structure change; the method has the advantages of calculation performance improvement, energy efficiency optimization, thermal safety guarantee, dynamic adaptive capacity and expandability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of liquid cooling cabinets, and more particularly to a liquid cooling cabinet based on a double-circulation refrigeration system and a use method thereof. Background Art

[0002] With the rapid development of artificial intelligence and big data technologies, graph computing technologies such as graph neural networks and graph algorithms are increasingly widely used in social network analysis, recommendation systems, knowledge graphs, and other fields. However, graph computing, due to its irregular memory access patterns, dynamically changing workloads, and complex dependencies between nodes, places extremely high demands on computing resources, while also bringing significant energy consumption and heat dissipation challenges.

[0003] Existing liquid cooling systems usually adopt a single-loop structure and a unified cooling strategy for all computing units, which makes it impossible to perform refined cooling capacity allocation for the uneven heat load distribution in graph computing.

[0004] In graph computing scenarios, the computational importance of different nodes varies significantly. For example, the computational accuracy of highly central nodes in social networks, key entity nodes in knowledge graphs, or popular item nodes in recommendation systems directly affects the overall algorithm performance. Existing studies have shown that assigning different computational accuracy to nodes of different importance can significantly reduce energy consumption while maintaining algorithm performance, but this non-uniform computational accuracy allocation strategy will lead to a more uneven distribution of thermal load within the system.

[0005] In addition, existing computational optimization methods and cooling system optimization are usually treated as two independent problems and are handled separately. There is a lack of a unified optimization framework to coordinately optimize computational accuracy configuration and cooling capacity distribution, making it difficult to achieve global optimization of system performance and energy efficiency.

[0006] Therefore, how to design an efficient liquid cooling system that can perceive the graph structure and dynamically adjust the cooling distribution according to the characteristics of graph computing, and achieve the coordinated optimization of computing accuracy and cooling resources, is a technical problem that needs to be solved urgently. Summary of the invention

[0007] The present invention provides a liquid cooling cabinet based on a dual-cycle refrigeration system and a method for using the same, which solves the technical problems in the related art and includes the following steps: Analyze graph structure characteristics and evaluate node importance; construct a three-dimensional correlation model of computational accuracy, energy consumption, and temperature; generate a graph structure-aware accuracy allocation strategy based on node importance; design a cooling capacity allocation strategy for a liquid cooling system based on the accuracy allocation strategy; implement an accuracy-cooling capacity joint optimization controller; design a dynamic graph structure adaptation mechanism; Among them, the precision allocation strategy allocates different calculation precisions to different nodes in the graph based on the node importance score. The cooling capacity allocation strategy dynamically adjusts the cooling capacity according to the heat load distribution caused by the precision allocation. The precision-cooling capacity joint optimization controller collaboratively adjusts the calculation precision and cooling capacity allocation according to the real-time system status, and dynamically adapts to changes in the graph structure.

[0008] Furthermore, analyzing the graph structure features and evaluating the importance of nodes includes: Calculate the degree centrality of each node ,in Representation Node The degree, Represents the total number of nodes in the graph; Calculate the PageRank value of each node ,in is the damping coefficient, Representation Node The set of neighbor nodes; Calculate the clustering coefficient for each node ,in Representation Node The number of triangles that exist between the neighbors of Calculate node importance score ,in , and is the weight parameter, satisfying .

[0009] Furthermore, constructing a three-dimensional correlation model includes: defining a set of calculation accuracy levels ,in Indicates the minimum precision, Indicates the highest accuracy; for each accuracy level , define its energy consumption function ,in represents the baseline energy consumption, is the energy consumption coefficient; establish the accuracy-temperature mapping relationship ,in represents the reference temperature, is the temperature coefficient, Representation Node The computational workload is defined as ; represents the computational complexity of the aggregation operation, Indicates the computational complexity of node feature transformation; The three-dimensional correlation model is established through the following mapping function: This model transforms each node At a specific accuracy The computational state is mapped to a point in three-dimensional space.

[0010] Furthermore, the graph structure-aware precision allocation strategy based on node importance includes: dividing the graph nodes into Communities, using Louvain algorithm for community detection; Calculate the importance of each community ; For each community, generate a precision allocation strategy based on node importance distribution and temperature constraints ,satisfy ; Smoothing the precision allocation results ,in Represents the weight of neighbor nodes.

[0011] Furthermore, the design of the cooling capacity allocation strategy of the liquid cooling system based on the precision allocation strategy includes: Establish a physical layout model of the computing chip and map nodes to physical computing units ,in Represents a collection of physical computational units; calculates the heat load for each physical computational unit ; Calculate the cooling capacity distribution of each cooling unit based on the double-cycle liquid cooling system model ,in Assign a balancing factor to the cooling capacity, is the average heat load; Design flow control strategies, including main circulation flow and secondary circulation flow distribution .

[0012] Furthermore, the steps of constructing the precision-cooling capacity joint optimization controller include: Define the system state vector ,in , and Respectively represent the current precision distribution, cooling capacity distribution and temperature distribution; design state transfer model ,in Indicates a control action.

[0013] Furthermore, the network architecture of the precision-cooling capacity joint optimization controller adopts a combination of graph convolutional network and multi-layer perceptron.

[0014] Furthermore, dynamic graph structure adaptation includes: defining a graph structure change detection function; designing a trigger mechanism to re-execute node importance evaluation, precision allocation, and cold allocation strategy generation; and designing a smooth transition mechanism.

[0015] The present invention provides a liquid cooling cabinet based on a double-cycle refrigeration system, which is used to execute the aforementioned method for using the liquid cooling cabinet based on the double-cycle refrigeration system, comprising: The graph structure analysis and node importance assessment module is used to analyze the graph structure characteristics and calculate the node importance score; the three-dimensional correlation model construction module is used to establish the correlation between calculation accuracy, energy consumption and temperature; the accuracy allocation strategy generation module is used to allocate suitable calculation accuracy to different nodes based on node importance; the liquid cooling system control module is used to control the coolant flow distribution according to the accuracy allocation results and heat load distribution; the joint optimization controller is used to coordinate the accuracy configuration and cooling capacity distribution in real time to achieve comprehensive optimization of system performance, energy consumption and temperature; the dynamic graph adaptation module is used to detect graph structure changes and trigger strategy updates so that the system can adapt to dynamically changing graph data.

[0016] Furthermore, it also includes a dual-circulation liquid cooling system, which includes: a main circulation connecting all computing modules, providing basic cooling capacity, and maintaining overall temperature balance; a secondary circulation performs differentiated cooling on specific computing units, and dynamically adjusts the cooling capacity distribution according to the heat load; wherein, the main circulation flow control is adjusted according to the uneven distribution of the heat load in the system, and the greater the heat load difference, the greater the main circulation flow; the secondary circulation flow is finely adjusted according to the calculated cooling capacity distribution ratio.

[0017] The beneficial effects of the present invention are: Improved computing performance, energy efficiency optimization, thermal safety, dynamic adaptability, and scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0019] As an effective tool for processing graph structured data, Graph Neural Network (GNN) faces the following technical challenges in large-scale graph data processing: Unbalanced computing load problem: In large-scale graph data, the node degree distribution usually presents a long-tail distribution, resulting in huge differences in the computational complexity of different nodes, resulting in inefficient utilization of computing resources.

[0020] Hot spot concentration problem: The computing load in the area where high-degree nodes (nodes with high connectivity) are concentrated is too high, forming hot spots on the computing chip, which limits the overall computing performance.

[0021] Accuracy-energy-temperature balance problem: Traditional computing systems cannot flexibly adjust computing accuracy and cooling strategies according to the characteristics of graph data, making it difficult to achieve the optimal balance between computing energy efficiency and accuracy.

[0022] This embodiment aims to propose a method for using a liquid cooling cabinet based on a dual-cycle refrigeration system, and solves the above technical challenges by dynamically adjusting the numerical accuracy and cooling capacity distribution of different calculation areas.

[0023] Embodiment 1: In at least one embodiment of the present invention, a method for using a liquid cooling cabinet based on a dual-cycle refrigeration system is disclosed, such as Figure 1 As shown, the following steps are included: Step 1: Graph structure feature analysis and node importance assessment The input graph data is represented as ,in Represents a collection of nodes. Represents a set of edges.

[0024] 1.1 Calculate the degree centrality of each node: ,in Representation Node The degree of (the number of edges connected to) .

[0025] 1.2 Calculate the PageRank value of each node: in is the damping coefficient (usually set to 0.85), Representation Node The set of neighbor nodes of a node. PageRank is a recursive algorithm used to estimate the importance of a node in a graph, originally used by Google for web page ranking.

[0026] 1.3 Calculate the clustering coefficient of each node: in Representation Node The clustering coefficient measures the degree of interconnection between neighbors of a node and reflects the density of the local network structure.

[0027] 1.4 Calculate the node importance score based on the above indicators: in , and is the weight parameter, satisfying These weights can be used to adjust the influence of different network indicators in node importance evaluation.

[0028] Step 2: Build a 3D correlation model Based on the node importance evaluation results in step 1, a three-dimensional association model is constructed: 2.1 Defining the set of computational accuracy levels ,in Indicates the lowest precision (such as FP16 or INT8), Indicates the highest precision (such as FP64). In actual implementation, the following precision configurations can be selected: INT8 (8-bit integer), FP16 (16-bit floating point), BF16 (16-bit brain floating point number), FP32 (32-bit floating point), FP64 (64-bit floating point).

[0029] 2.2 For each accuracy level , define its energy consumption function: in Indicates the baseline energy consumption (such as energy consumption under FP32 precision), is the energy consumption coefficient, satisfying Based on the energy efficiency characteristics of modern computing chips, a typical energy consumption coefficient may be: (INT8), (FP16), (BF16), (FP32), (FP64).

[0030] 2.3 Establishing the accuracy-temperature mapping relationship: in Indicates the reference temperature (usually the no-load temperature of the chip, such as 35°C). is the temperature coefficient, Representation Node The computational workload is defined as: here, represents the computational complexity of the aggregation operation, Represents the computational complexity of node feature transformation.

[0031] 2.4 Constructing a three-dimensional correlation model of accuracy, energy consumption and temperature: Based on the above definition, construct a three-dimensional correlation model , the model describes the relationship between computing accuracy, energy consumption and temperature: -Accuracy dimension : Indicates the impact of different calculation accuracy levels on the accuracy of calculation results - energy consumption dimension : Represents the energy consumption characteristics at different accuracy levels - temperature dimension : represents the temperature distribution under a specific node importance and accuracy configuration The three-dimensional correlation model is established through the following mapping function: This model transforms each node At a specific accuracy The computing state under the circumstance is mapped to a point in three-dimensional space, realizing the unified representation of the three dimensions of accuracy, energy consumption and temperature.

[0032] 2.5 Define the comprehensive scoring function of the three-dimensional association model: in , and is the weight parameter, satisfying , is the maximum node importance score, and are the minimum and maximum temperature thresholds, respectively. This scoring function takes into account all factors in the three-dimensional association model and is used to subsequently select the optimal accuracy configuration for each node.

[0033] Step 3: Generate a graph structure-aware precision allocation strategy Based on the three-dimensional association model constructed in step 2, the optimal accuracy is assigned to each node in the graph: 3.1 Divide the graph nodes into Communities (sub-graphs): Community detection using Louvain algorithm: in represents the adjacency matrix element, Representation Node The community you belong to, is the Kronecker function, when 1 if the value is 0, otherwise it is 0.

[0034] The Louvain algorithm is a community detection method based on modularity optimization. Its working principle includes two stages: - Stage 1: Initially assign each node to a separate community, and then iteratively move the node to the adjacent community that maximizes the modularity gain. - Stage 2: Aggregate nodes belonging to the same community into "super nodes" to build a new network. These two stages are performed alternately until the modularity no longer increases significantly.

[0035] 3.2 Calculate the importance of each community: Here, the importance of the entire community is evaluated by calculating the average importance score of all nodes in the community, which will be used for subsequent precision strategy allocation.

[0036] 3.3 For each community , according to the node importance distribution and temperature constraints, generate the precision allocation strategy: This step is done by selecting from the precision set defined in step 2 such that the comprehensive score Maximum level of accuracy As a node Accuracy configuration while ensuring that the temperature does not exceed the safety threshold (Usually set to the chip's highest safe operating temperature, such as 85°C).

[0037] 3.4 Smooth the precision allocation results to avoid large precision jumps between adjacent nodes: in Represents the weight of neighbor nodes.

[0038] This smoothing process takes into account the relationship between a node and its neighbors, reduces the spatial discontinuity of the precision configuration, and helps improve computational efficiency and hardware resource utilization.

[0039] Step 4: Generate cooling capacity allocation strategy for liquid cooling system Based on the precision allocation strategy generated in step 3, design the cooling capacity allocation strategy for the liquid cooling system: 4.1 Establish a physical layout model of the computing chip and map nodes to physical computing units: in Represents a collection of physical computation units.

[0040] This mapping relationship allocates the logical nodes in graph data processing to specific processing units on the physical computing architecture, providing a basis for subsequent cooling capacity allocation. In actual implementation, a spatial locality optimization algorithm can be used to map closely connected nodes to physically adjacent computing units to reduce data transmission overhead.

[0041] 4.2 Calculate the heat load of each physical computing unit: This formula accumulates the number of cells assigned to the same physical computation unit. All nodes on The temperature contribution of , using the smoothed precision distribution result generated in step 3 .

[0042] 4.3 Based on the double-cycle liquid cooling system model, calculate the cooling capacity distribution of each cooling unit: in Assign a balance factor to the cooling capacity (between 0 and 1). is the average heat load.

[0043] The dual-loop liquid cooling system is an efficient heat dissipation solution that includes two cooling loops that work together: -Main loop: connects all computing modules, provides basic cooling capacity, and maintains overall temperature balance. -Sub-loop: provides differentiated cooling for specific computing units, and can dynamically adjust cooling capacity distribution according to heat load.

[0044] Cooling capacity distribution balance factor Control the targeted distribution of cooling capacity. When it is close to 1, the cooling capacity distribution tends to be differentiated according to the actual heat load; when When it is close to 0, the cooling capacity is distributed more evenly, which is suitable for situations where the heat load is distributed more evenly.

[0045] 4.4 Design flow control strategy for main loop and sub-loop: Main loop flow: Auxiliary circulation flow distribution: in and are the benchmark flow rates of the main circulation and the secondary circulation respectively, is the flow regulation coefficient.

[0046] The main circulation flow is adjusted according to the uneven distribution of heat load in the system. The greater the difference in heat load, the greater the main circulation flow to ensure the overall temperature balance of the system. The secondary circulation flow is finely adjusted according to the cooling capacity distribution ratio calculated in step 4.3 to ensure that the high heat load area obtains more cooling capacity.

[0047] Step 5: Build a precision-cooling capacity joint optimization controller Based on the strategies of step 3 and step 4, a precision-cooling capacity joint optimization controller is constructed: 5.1 Definition of system state vector ,in , and They respectively represent the accuracy distribution, cooling capacity distribution and temperature distribution at the current moment.

[0048] The system state vector includes the accuracy allocation strategy generated in step 3, the cooling capacity allocation strategy generated in step 4, and the real-time monitored temperature distribution information, which serves as the input of the accuracy-cooling capacity joint optimization controller.

[0049] 5.2 Design state transition model: in Indicates control actions, including precision adjustment actions and cooling capacity adjustment action .

[0050] The state transition model describes the system in its current state. Next, execute the control action Then transfer to the new state In actual implementation, the state transfer function can be constructed by combining the physical model with the data-driven model. .

[0051] 5.3 Define the reward function: in Indicates the calculation accuracy score under the current accuracy configuration. represents the total energy consumption, is the safety temperature threshold, , and is the weight parameter.

[0052] The reward function takes into account three key objectives:- :Evaluate the calculation accuracy of the current accuracy configuration, which can be calculated by weighted average: - : Evaluate the total system energy consumption, including computing energy consumption and cooling energy consumption: -Temperature penalty term: When the temperature of any computing unit exceeds the safety threshold When , a penalty is imposed to ensure the thermal safety of the system.

[0053] 5.4 Training controller policies using deep reinforcement learning : in is the discount factor. The controller network architecture uses a combination of graph convolutional network (GCN) and multi-layer perceptron (MLP): in To add the adjacency matrix after self-connection, is the corresponding degree matrix, For the The weight matrix of the layer, is the activation function.

[0054] Graph Convolutional Network (GCN) is a neural network that specializes in processing graph structure data. Its core idea is to aggregate and transform the features of nodes and their neighbors. The calculation process of the GCN layer can be divided into three steps: 1. Add self-connection: , ensuring that the node's own information is also considered 2. Symmetric normalization: , to avoid too large differences in the degree of information aggregation of nodes with different degrees 3. Feature transformation and nonlinear activation: ; In this controller, GCN is used to process graph structure information, while MLP is used to process non-graph structure information related to the liquid cooling system. The two are combined to form the final control strategy network.

[0055] Step 6: Dynamic graph structure adaptation Based on the joint optimization controller implemented in step 5, dynamic graph structure adaptation is performed: 6.1 Define graph structure change detection function: in represents the graph structure similarity measure, is the time window size.

[0056] Graph structure similarity measurement There are many methods that can be used, for example: -Edge set change rate: ,in Representation symmetric difference - spectral distance: based on the difference in eigenvalues ​​of the graph Laplacian matrix - node feature distribution change: calculate the difference in the statistical distribution of node features 6.2 Design trigger mechanism: in is the trigger threshold.

[0057] Trigger threshold Controls the sensitivity of the system to changes in the graph structure. A smaller value will cause the system to update its strategy even for minor changes, while a larger value will trigger an update only when the graph structure changes significantly. In actual applications, the threshold can be adjusted dynamically based on the characteristics of the application scenario and the limitations of computing resources.

[0058] 6.3 When When the temperature reaches 400 °C, re-execute steps 1-5 to update the precision allocation and cooling capacity allocation strategies.

[0059] When significant changes in the graph structure are detected, the system re-evaluates node importance, updates the 3D association model, regenerates accuracy and cooling allocation strategies, and optimizes controller parameters to adapt to the new graph structure characteristics.

[0060] 6.4 Design a smooth transition mechanism to avoid sudden changes in strategy: in and is the time-varying smoothing factor, and For newly generated policies.

[0061] The smooth transition mechanism is achieved through a time-varying smoothing factor and Control the speed of policy updates to avoid system instability caused by policy mutations. In practical applications, these smoothing factors can be dynamically adjusted according to the magnitude and speed of graph structure changes: - When the graph structure changes greatly, the smoothing factor can be increased to speed up policy updates. - When the graph structure changes slightly, the smoothing factor can be reduced to ensure system stability.

[0062] Embodiment 2: Based on Example 1, further analysis revealed that graph neural networks face the following new technical challenges when processing large-scale graph data: The problem of timing changes in computing load: The computing process of graph neural networks usually includes multiple stages (such as feature aggregation, graph convolution, graph pooling, etc.). There are significant differences in the computing characteristics and thermal load distribution of each stage, and the static optimization strategy in Example 1 cannot fully adapt to such timing changes.

[0063] Cooling system lag problem: There is a time delay from the adjustment of the liquid cooling system to achieving the expected cooling effect. In the scenario where the graph computing load changes rapidly, the cooling effect may not match the actual demand, affecting the system performance and stability.

[0064] Blind resource scheduling problem: The lack of the ability to predict future computing loads results in the precision configuration and cooling allocation strategies being able to respond passively and unable to prepare in advance for the upcoming computing peak.

[0065] The steps of Example 2 are as follows: Step 2-1: Same as step 1 of Example 1 The same as step 1 in Example 1 is used to obtain basic graph structure features and node importance information, which will serve as important input for subsequent workload prediction.

[0066] **Step 2-2: Same as step 2 of Example 1.

[0067] Step 2-3: Identification of graph neural network computational stages and modeling of workload characteristics This step is an innovation of this embodiment and is used to replace step 3 in embodiment 1: 2.3.1 Define the main set of stages of graph neural network calculation: Typical calculation stages include: Input feature processing, Message passing / feature aggregation, Neighbor information aggregation, Node feature updates, Graph representation learning, Task specific output calculations.

[0068] 2.3.2 Constructing the transfer diagram between calculation stages: in represents the transition edge between stages, Represents the transition probability weight. For a given graph structure and computing task, the transition probability of each stage can be determined through historical execution data or model analysis: 2.3.3 Establish workload characteristic model for each computing stage: in , and They represent computational intensity, memory access intensity, and communication intensity, respectively. These characteristics can be quantified as follows: Computational Intensity: in Indicates that in the stage Processing Node The number of floating point operations required. For example, for the graph convolution stage, it can be expressed as: Memory access intensity: in Indicates that in the stage Processing Node The number of memory accesses required.

[0069] Communication intensity: in Indicates that in the stage node and The amount of data transferred between.

[0070] 2.3.4 Establishing a node-level stage-dependent workload model: in , and Is with the node The associated normalization factor reflects the contribution of the node to the overall load: and Similar definition.

[0071] Step 2-4: Build a workload timing prediction model 2.4.1 Constructing a time series graph neural network model for load forecasting: in yes The predicted load after time, is the historical load sequence, are model parameters.

[0072] The Temporal Graph Neural Network (TGNN) combines the GNN’s ability to process graph structures with the time series model’s ability to capture temporal dependencies. Its structure includes: Spatial encoding layer: Time encoding layer: Prediction layer: 2.4.2 Constructing uncertainty estimation model for time series prediction: in represents the variance of the predicted load, is the model parameter.

[0073] 2.4.3 Define the confidence interval for load forecast: in The confidence level is The standard normal quantile at .

[0074] 2.4.4 Based on the calculation phase transition graph and the current execution phase, predict the future phase sequence: By combining the phase forecast and the load forecast, we obtain the phase-dependent load forecast for the future time: Step 2-5: Generate a forward-looking precision-cooling joint scheduling strategy 2.5.1 Based on the workload prediction results, construct the time window sliding optimization problem: in To predict the time window size, is the time discount factor, For the moment The predicted load, and They are the maximum change constraints for accuracy configuration and cooling capacity allocation respectively.

[0075] 2.5.2 Define a robust reward function that takes into account prediction uncertainty: in is the risk aversion factor, and the second term penalizes the reward variance caused by forecast uncertainty.

[0076] 2.5.3 Design a hierarchical preheating mechanism: Based on the predicted future computing load, adjust the cooling system in advance to compensate for its response delay: in is the preheating coefficient, The response delay of the cooling system is is the predicted future heat load.

[0077] 2.5.4 Using the model predictive control (MPC) method to solve the optimization problem: At each control time step , solve the optimization problem to obtain the future control sequence , but only executes the control action at the current moment , and solve the optimization problem again at the next time step.

[0078] The core algorithm steps of the MPC controller: 1. Get the current system status 2. Use time series forecasting models to obtain future load forecasts 3. Solve the optimization problem to obtain the control sequence 4. Application control actions , observe the new state 5. Time stepping, return to step 1 Step 2-6: Learn the adaptive accuracy-cooling response curve and optimize the execution timing 2.6.1 Define the response relationship between accuracy adjustment and temperature change: 2.6.2 Define the response relationship between cooling capacity adjustment and temperature change: 2.6.3 Continuously update the response model using online learning methods: in and is the learning rate, is the actual observed temperature, Predict the temperature for the model.

[0079] 2.6.4 Optimize execution sequence based on learning response curve: Arrange the optimal adjustment sequence by calculating the response time of accuracy / cooling capacity adjustment to ensure that the system has reached the ideal state when calculating load changes: in For the system from the current state Transition to target state Required response time.

[0080] The technical effects of this embodiment are as follows: Improved computing performance: Compared with Example 1, the processing power is further improved by 15% during the multi-stage calculation process of the graph neural network, and the overall improvement is about 50% compared to the traditional method.

[0081] Improved cooling efficiency: Through forward-looking scheduling and hierarchical preheating mechanism, the response delay problem of the cooling system is effectively alleviated, the cooling efficiency is improved by about 20%, and the energy consumption is further reduced by 10%.

[0082] Enhanced system stability: By accurately predicting future load changes, system temperature fluctuations are reduced by 40%, effectively avoiding performance jitter and hardware stress caused by drastic temperature changes.

[0083] Optimized resource utilization: Proactive resource scheduling increases the utilization of computing and cooling resources by 25%, significantly improving the overall energy efficiency of the system.

[0084] Enhanced adaptability: By continuously optimizing the response model through online learning, the system can adapt to different hardware configurations and application scenarios and has stronger generalization capabilities.

[0085] Embodiment 3: comprises the following steps: Step 3-1: Same as step 1 in Example 1 It is consistent with step 1 in Example 1, but is expanded in a distributed environment to provide a basis for subsequent graph partitioning and cross-device precision allocation.

[0086] Step 3-2: Same as step 1 in Example 1 Step 3-3: Generate a heat load-aware graph partitioning strategy This step is the core innovation of this embodiment, replacing step 3 in embodiment 1: 3.3.1 Define the node heat load estimation function: in and are the heat contribution weights of computation and memory access, respectively. and denote the input and output feature dimensions respectively, Represents the node feature dimension.

[0087] 3.3.2 Constructing the thermal balance diagram partition objective function: in Indicates that the figure Divide into The partitioning scheme of the subgraphs, represents the set of edge cuts between partitions, The weight coefficient for balancing thermal load balancing and communication cost.

[0088] 3.3.3 Use multi-level heat balance partition algorithm: Coarsening stage: Recursively merge nodes based on heat load and node similarity to generate a coarse-grained graph sequence Node Merge Scoring: in is the node similarity, is the average heat load.

[0089] Initial partitioning phase: for the coarsest-grained graph Apply spectral partitioning or K-means algorithm to generate initial partitions For K-means partitioning, a heat load-weighted distance metric is used: in Representation Node The spectrum embedding position.

[0090] Refinement stage: From Start by iterating and refining the partitions at each granularity level until the original graph is obtained. Partition At each refinement step, a local optimization algorithm is used to adjust the node assignments to minimize the objective function: in and Respectively represent nodes From the partition Move to Resulting in changes in thermal balance and edge cutting.

[0091] 3.3.4 Partition adaptation for heterogeneous computing clusters: in Indicates The computing power indicator of each computing device allows devices with stronger computing power to be assigned more thermal load.

[0092] Step 3-4: Heterogeneous device-aware precision allocation strategy 3.4.1 Build equipment characteristic model: , where each device The feature vector of contains: here Represents computing power, Indicates the memory capacity. represents the memory bandwidth, Indicates the maximum safe operating temperature. Indicated in precision The energy consumption function.

[0093] 3.4.2 Define the comprehensive scoring model for device feature perception: in Indicates the device For accuracy The advantages of hardware acceleration are as follows: For example, some devices have dedicated acceleration units for INT8 calculations.

[0094] 3.4.3 Consider the precision coordination of nodes across device boundaries: For nodes located at the partition boundary, their precision allocation needs to consider communication efficiency: in is the communication cost weight, and the second term penalizes the extra communication cost caused by the large difference in accuracy with neighboring nodes.

[0095] 3.4.4 Design precision conversion gate: Set precision conversion gates at the boundaries of different precision areas to achieve efficient precision conversion:

[0096] For boundary areas with frequent communication, a precision buffer can be set to cache multi-precision representations to reduce repeated conversion overhead.

[0097] Step 3-5: Coordinated control of multi-zone liquid cooling system 3.5.1 Constructing a multi-region liquid cooling network topology model: in Represents a collection of liquid cooling zones. Represents the cooling medium flow connection between zones.

[0098] 3.5.2 Establish the mapping relationship between regional heat load and equipment distribution: in Indicates cooling area The total heat load, Indicates the device heat load.

[0099] 3.5.3 Design hierarchical cooling capacity distribution strategy: Global cooling capacity allocation: Based on the heat load distribution of each area, determine the cooling capacity allocation ratio of the main cooling source in Assign a balancing factor to the cooling capacity, is the average regional heat load.

[0100] Cooling capacity distribution within the area: In each cooling area, cooling capacity is distributed according to the heat load of the equipment. in For Region The internal cooling capacity distribution balance factor is is the average heat load of the equipment in the area.

[0101] 3.5.4 Design cross-region cooling capacity scheduling protocol: Hot spot area urgent cooling request: when a certain area When the heat load exceeds the safety threshold, an emergency cooling request is sent to the adjacent area. in For Region The heat load safety threshold, For Region The maximum additional cooling capacity that can be provided.

[0102] Cooling response strategy: Adjacent areas determine the response ratio based on their own load conditions and the overall heat distribution of the system. in is the response factor, depending on the region The cooling margin and the area The thermal distribution correlation.

[0103] Cooling transmission path optimization: When cooling needs to be transmitted across multiple areas, the optimal path is selected to minimize energy consumption and cooling loss in represents the energy consumption for pumping water, Indicates the cooling transmission loss.

[0104] Improved computing performance: Through heat load-aware graph partitioning and specialized precision allocation for heterogeneous devices, the distributed processing capability of large-scale graph neural networks is further improved by 25% compared to Example 1, and the overall improvement is more than 60% compared to traditional methods.

[0105] Enhanced system scalability: Supports seamless expansion from a single machine to a cluster of hundreds of nodes. As the number of devices increases, the system acceleration ratio maintains a linear acceleration efficiency of more than 80%.

[0106] Energy efficiency ratio optimization: Multi-zone liquid cooling coordinated control reduces the overall system energy consumption by 35%, a further 7% reduction compared to Example 1. At the same time, performance is improved, and the overall energy efficiency ratio is increased by more than 95%.

[0107] Heterogeneous computing support: It can fully utilize the characteristics of different computing devices (such as CPU, GPU, FPGA, etc.) and automatically allocate the most suitable graph data subset and precision configuration for each type of device.

[0108] The application example of Example 1 is based on the social network analysis scenario, which demonstrates the actual application effect of large-scale graph neural network mixed precision computing and liquid cooling collaborative optimization technology. A social network graph data containing 10 million nodes and 1 billion edges is used for processing. The nodes represent users and the edges represent the social relationships between users. The task is to use GNN to predict user interests.

[0109] Step 1: Graph structure feature analysis and node importance assessment Input data example. The following table shows some node information of the social network dataset:

[0110] Based on the feature analysis results, the importance index of each node is calculated by applying the method of steps 1.1-1.4:

[0111] Taking node 1003 as an example, its importance calculation process is: - Degree centrality: -PageRank: 0.01237 after iterative calculation -Clustering coefficient: There are 44,635 edges between the neighbors of this node, and the maximum possible number of edges is , the clustering coefficient is - Importance score: Step 2: Construction of three-dimensional correlation model of calculation accuracy, energy consumption and temperature Precision configuration definition. To simplify the example, four precision levels are defined:

[0112] Measured data of energy consumption function, energy consumption measurement of different computing operations at different precisions:

[0113] Temperature function measurement data, temperature data measured under different accuracy and energy consumption combinations:

[0114] According to the above measured data, the temperature function is fitted: in is the energy consumption, and this function can predict the average temperature of the system under given energy consumption.

[0115] Step 3: Generate the precision allocation strategy based on node importance and the community division result; Using the Louvain algorithm for community detection, some of the results obtained are:

[0116] Precision allocation strategy, based on community importance and node importance, generates a precision allocation plan:

[0117] Community 5 has the highest average importance, with 87% of nodes assigned FP32 or FP16 precision, while most nodes (76%) in community 1 are assigned INT8 or INT4 precision.

[0118] Step 4: Generate liquid cooling resource allocation plan, physical node mapping, and physical distribution mapping of computing units:

[0119] Hot spot analysis, based on the accuracy allocation in step 3 and the temperature model in step 2, predicts the hot spot area:

[0120] Cooling resource allocation scheme, cooling resource allocation based on hot spot analysis:

[0121] Step 5: Accuracy-cooling joint optimization controller implementation, monitoring indicators, indicators monitored in real time during system operation:

[0122] Adaptive adjustment example. Taking the time point 180s as an example, the system detects that the temperature of hotspot-1 exceeds the threshold (75.0°C), triggering joint optimization adjustment:

[0123] The adjustments performed by the optimization controller include: 1. Increasing the coolant flow rate of hotspot-1 to 5.8L / min 2. Reducing 30% of the FP32 nodes in the hotspot-1 area to FP16 accuracy 3. Dynamically adjusting the flow rate control parameters;

[0124] Step 6: Implementation of dynamic graph structure adaptation mechanism and monitoring of graph structure changes. During a social network data update, the graph structure changes detected are:

[0125] Adaptation process, for the large-scale update of batch 3, the adaptation process records: Importance recalculation:

[0126] Precision configuration adjustment:

[0127] Cooling adjustment:

[0128] The performance test results comparing this embodiment with the traditional method are as follows:

[0129] Energy efficiency ratio improvement, energy efficiency ratio improvement effect under different scale graph data:

[0130] It can be seen from the above real application example data that in the large-scale graph neural network processing scenario, this embodiment successfully solves the technical challenges of unbalanced computing load, hot spot concentration, and computing power-precision-energy consumption-temperature balance through the coordinated optimization of precision calculation and liquid cooling system, and achieves the comprehensive benefits of performance improvement, energy consumption reduction and enhanced system stability.

Claims

1. A hybrid precision computing and liquid cooling collaborative optimization method for large-scale graph neural networks, characterized in that: The following steps are involved: Analyze graph structure characteristics and evaluate node importance; construct a three-dimensional correlation model of calculation accuracy, energy consumption and temperature; Generate a graph structure-aware precision allocation strategy based on node importance; design a cooling capacity allocation strategy for the liquid cooling system based on the precision allocation strategy; Realize precision-cooling capacity joint optimization controller; Design dynamic graph structure adaptation mechanisms; Among them, the precision allocation strategy allocates different calculation precisions to different nodes in the graph based on the node importance score. The cooling capacity allocation strategy dynamically adjusts the cooling capacity according to the heat load distribution caused by the precision allocation. The precision-cooling capacity joint optimization controller collaboratively adjusts the calculation precision and cooling capacity allocation according to the real-time system status, and dynamically adapts to changes in the graph structure.

2. The method for using a liquid cooling cabinet based on a dual-cycle refrigeration system according to claim 1, characterized in that: The analyzing graph structure features and evaluating node importance includes: Calculate the degree centrality of each node ,in Representation Node The degree, Represents the total number of nodes in the graph; Calculate the PageRank value of each node ,in is the damping coefficient, Representation Node The set of neighbor nodes of Calculate the clustering coefficient for each node ,in Representation Node The number of triangles that exist between the neighbors of Calculate node importance score ,in , and is the weight parameter, satisfying .

3. The method for using a liquid cooling cabinet based on a dual-cycle refrigeration system according to claim 1, characterized in that: Building a 3D associative model includes: defining a set of calculation accuracy levels ,in Indicates the minimum precision, Indicates the highest accuracy; for each accuracy level , define its energy consumption function ,in represents the baseline energy consumption, is the energy consumption coefficient; establish the accuracy-temperature mapping relationship ,in represents the reference temperature, is the temperature coefficient, Representation Node The computational workload is defined as ; represents the computational complexity of the aggregation operation, Indicates the computational complexity of node feature transformation; The three-dimensional correlation model is established through the following mapping function: ; This model transforms each node At a specific accuracy The computational state is mapped to a point in three-dimensional space.

4. The method for using a liquid cooling cabinet based on a dual-cycle refrigeration system according to claim 1, characterized in that: The precision allocation strategy based on node importance to generate graph structure awareness includes: dividing the graph nodes into Communities, using Louvain algorithm for community detection; Calculate the importance of each community ; For each community, generate a precision allocation strategy based on node importance distribution and temperature constraints ,satisfy ; Smoothing the precision allocation results ,in Represents the weight of neighbor nodes.

5. The method for using a liquid cooling cabinet based on a dual-cycle refrigeration system according to claim 1, characterized in that: The design of the cooling capacity allocation strategy of the liquid cooling system based on the precision allocation strategy includes: Establish a physical layout model of the computing chip and map nodes to physical computing units ,in Represents a collection of physical computational units; calculates the heat load for each physical computational unit ; Calculate the cooling capacity distribution of each cooling unit based on the double-cycle liquid cooling system model ,in Assign a balancing factor to the cooling capacity, is the average heat load; Design flow control strategies, including main circulation flow and secondary circulation flow distribution .

6. The method for using a liquid cooling cabinet based on a dual-cycle refrigeration system according to claim 1, characterized in that: The steps to construct the precision-cooling capacity joint optimization controller include: Define the system state vector ,in , and Respectively represent the current precision distribution, cooling capacity distribution and temperature distribution; design state transfer model ,in Indicates control actions; Define the reward function: ; Train controller policy using deep reinforcement learning , the optimization goal is .

7. The method for using a liquid cooling cabinet based on a dual-cycle refrigeration system according to claim 6, characterized in that: The network architecture of the precision-cooling capacity joint optimization controller adopts a combination of graph convolutional network and multi-layer perceptron. The calculation process of the graph convolutional network is: ,in To add the adjacency matrix after self-connection, is the corresponding degree matrix, For the The weight matrix of the layer, is the activation function.

8. The method for using a liquid cooling cabinet based on a dual-cycle refrigeration system according to claim 1, characterized in that: Dynamic graph structure adaptation includes: Define graph structure change detection function ,in represents the graph structure similarity measure, is the time window size; Trigger mechanism ,in is the trigger threshold; when When the node importance evaluation, precision allocation and cooling allocation strategy generation are re-executed; smooth transition mechanism and ,in and is the time-varying smoothing factor.

9. A liquid cooling cabinet based on a double-cycle refrigeration system, characterized in that: The method for using a liquid cooling cabinet based on a dual-cycle refrigeration system according to any one of claims 1 to 8 comprises: Graph structure analysis and node importance assessment module, used to analyze graph structure features and calculate node importance scores; A three-dimensional correlation model building module is used to establish the correlation between calculation accuracy, energy consumption and temperature; The precision allocation strategy generation module is used to allocate appropriate computing precision to different nodes based on the importance of the nodes; A liquid cooling system control module, used to control the coolant flow distribution according to the accuracy distribution result and the heat load distribution; Joint optimization controller, used to coordinate precision configuration and cooling capacity allocation in real time to achieve comprehensive optimization of system performance, energy consumption and temperature; The dynamic graph adaptation module is used to detect graph structure changes and trigger strategy updates, so that the system can adapt to dynamically changing graph data.

10. The liquid cooling cabinet based on the dual-cycle refrigeration system according to claim 9, characterized in that: It also includes a dual-circulation liquid cooling system, which includes: The main loop connects all computing modules, provides basic cooling capacity, and maintains overall temperature balance; the secondary loop performs differentiated cooling for specific computing units and dynamically adjusts cooling capacity distribution according to heat load; Among them, the main circulation flow control is adjusted according to the uneven distribution of heat load in the system. The greater the heat load difference, the greater the main circulation flow; the secondary circulation flow is finely adjusted according to the calculated cooling capacity distribution ratio.

Citation Information

Cited By

  • Multi-parameter MPC two-phase liquid cooling energy efficiency optimization distribution control system and method

    CN120916409A

  • A multi-parameter MPC two-phase liquid cooling energy efficiency optimization distribution control system and method

    CN120916409B