Electronic equipment, computing cluster, heat dissipation management method, storage medium
Through the heat dissipation management system with CXL switches as the core, the liquid cooling flow rate and power distribution are dynamically adjusted, which solves the problem of delayed hot and cold load perception in thermal management in the cluster system, realizes dynamic identification of device roles and resource optimization, and improves the performance stability and efficiency of the AI training system.
Patent Information
- Application Number
- CN202510926780.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-07-04
AI Technical Summary
Existing cluster system thermal management technology has coarse granularity in sensing hot and cold loads and delayed response. It is unable to identify the dynamic changes in device roles during the AI training phase and cannot dynamically adjust the resources allocated to the devices, resulting in poor system performance stability.
The heat dissipation management system, which uses the CXL switch as its core, dynamically adjusts the liquid cooling flow rate and power distribution by identifying the computational heat and transmission heat of the computing nodes, ensuring that each device is within the target temperature range. This enables heterogeneous thermal management collaboration across devices and dynamic cooling control driven by data semantics.
It improves the thermal management response speed, resource allocation accuracy, and overall performance stability of the AI training system, avoiding waste of heat dissipation resources and degradation of system performance.
Smart Images

Figure CN120434978B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of heat dissipation management technology, and in particular to an electronic device, a computing cluster, a heat dissipation management method, and a storage medium. Background Art
[0002] In recent years, with the continuous expansion of artificial intelligence (AI) training models, the application of GPU (Graphics Processing Unit) supernode systems has become increasingly widespread in high-performance computing and deep learning. To meet the extreme demands of large-scale AI models for memory capacity, bandwidth, and data transmission efficiency, communication protocols, as a new high-speed interconnect standard, are gradually being applied to the construction of GPU supernode systems. Through a unified memory semantic access mechanism, communication protocols enable efficient interconnection between CPUs (Central Processing Units, or microprocessors), GPUs, and memory devices, becoming a key supporting technology for next-generation heterogeneous computing platforms.
[0003] In this context, ensuring the stable operation of AI training systems requires significantly more complex thermal management and power consumption scheduling. Traditional air-cooling or static liquid-cooling thermal management systems struggle to meet the dynamic cooling requirements of heterogeneous clusters at varying computing stages, data paths, and power consumption states.
[0004] In the research and application of cluster system thermal management, CXL (Compute Express Link) switches have gradually evolved from simple communication nodes to core components capable of protocol parsing, device status awareness, and system-level coordinated control. However, existing cluster system thermal management solutions often fail to fully leverage the key advantages of CXL switches in terms of their awareness and controllability within the cluster topology. In particular, they lack dynamic thermal management that deeply integrates CXL switches with liquid cooling control, power scheduling, and AI training task status. Summary of the Invention
[0005] The present invention provides electronic devices, computing clusters, heat dissipation management methods, and storage media to at least address the problems of existing cluster system thermal management technologies, such as coarse granularity in hot and cold load perception, delayed response, inability to identify dynamic changes in device roles during the AI training phase, and inability to dynamically regulate the resources allocated to devices, resulting in poor system performance stability.
[0006] The present invention provides an electronic device, comprising: a computing node and a switching node, wherein the computing node is connected to the switching node, the computing node includes at least one first device with computing properties and / or at least one second device with transmission properties, at least one first device performs computing tasks in artificial intelligence training, and at least one second device performs transmission tasks in the artificial intelligence training; the switching node includes a control module, which controls the corresponding heat dissipation device through the control module to dissipate heat for each device when the heat dissipation speed of each device is determined based on the computing heat of at least one first device and / or the transmission heat of at least one second device, so that each device is within the corresponding target temperature range.
[0007] The present invention also provides a computing cluster, which includes the above-mentioned electronic device.
[0008] The present invention also provides a heat dissipation management method, which is applied to the above-mentioned computing cluster, wherein the method includes the following steps: identifying the current training stage of each device in the computing cluster, and calculating the computational heat of at least one first device with computational properties and the transmission heat of at least one second device with transmission properties based on the current training stage of each device; calculating the comprehensive thermal load index of each device based on the computational heat of at least one first device and the transmission heat of at least one second device, and determining the current liquid cooling flow rate level of each device based on the comprehensive thermal load index of each device; determining the current power mode based on the current total power of the computing cluster, and dynamically adjusting the liquid cooling flow rate level of each device and / or adjusting the power distribution of each device based on the current power mode, so that each device is in the corresponding target temperature range.
[0009] The present invention also provides a heat dissipation management system, including: an identification module for identifying the current training stage of each device in a computing cluster; a calculation module for calculating the computational heat of at least one first device with computational properties and the transmission heat of at least one second device with transmission properties based on the current training stage of each device; calculating the comprehensive thermal load index of each device based on the computational heat of at least one first device and the transmission heat of at least one second device, and determining the current liquid cooling flow rate level of each device based on the comprehensive thermal load index of each device; a heat dissipation management module for determining the current power mode based on the current total power of the computing cluster, and dynamically adjusting the liquid cooling flow rate level of each device and / or adjusting the power distribution of each device based on the current power mode, so that each device is in the corresponding target temperature range.
[0010] The present invention also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned heat dissipation management method are implemented.
[0011] The present invention also provides a computer program product, comprising a computer program, wherein the computer program implements the above-mentioned heat dissipation management method when executed by a processor.
[0012] Through the present invention, the switching node in the electronic device uses a control module to determine the heat dissipation speed of each device based on the calculation heat of the first device with calculation attributes and / or the transmission heat of the second device with transmission attributes in the computing node. The control module controls the corresponding heat dissipation device to dissipate heat for each device, so that each device is in the corresponding target temperature range. This solves the problem that the existing cluster system thermal management technology has coarse cold and hot load perception granularity and delayed response, cannot identify the dynamic changes of device roles in the AI training stage, cannot dynamically adjust the resources allocated to the device, and leads to poor system performance stability. It comprehensively improves the thermal management response speed, resource allocation accuracy and overall performance stability of the AI training system. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0014] Figure 1 is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention;
[0015] Figure 2 A schematic diagram of a liquid cooling cycle cooling method according to an embodiment of the present invention;
[0016] Figure 3 A schematic diagram of the structure of a switching node according to an embodiment of the present invention;
[0017] Figure 4 A schematic flow chart of a heat dissipation management method according to an embodiment of the present invention;
[0018] Figure 5 A schematic diagram of a computing cluster structure according to an embodiment of the present invention;
[0019] Figure 6 A flowchart of a heat dissipation management method provided according to an embodiment of the present invention;
[0020] Figure 7 FIG. 4 is a schematic diagram of a heat dissipation management system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0022] It should be noted that, in the description of the present invention, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence.
[0023] Before specifically introducing the heat dissipation management method of this application, a brief introduction to current heat dissipation management technologies is provided, which mainly include the following two categories:
[0024] Static liquid cooling systems typically use a simple temperature-based control strategy with a constant flow rate or preset cooling flow rates for each node. This fails to account for the dynamic load changes during AI training. Because they cannot accurately identify the device's current computing phase and the actual source of heat load, they often waste resources or overheat critical nodes.
[0025] The second approach: thermal power management systems based on rack management controllers (RMCs). These solutions often rely on data from temperature and power sensors for basic resource scheduling. This approach suffers from coarse control granularity and high response latency, making it difficult to adapt to the highly dynamic and phased resource demands of AI training tasks. Furthermore, the RMC layer typically lacks in-depth information about the AI task's training configuration and communication patterns, resulting in a lack of targeted control strategies.
[0026] Although existing technologies have made some progress in achieving AI server heat dissipation management, they still have some significant shortcomings:
[0027] (1) Lack of application semantics for control: Traditional systems are unable to identify the differentiated demands for system resources at various stages of AI training (such as parameter transmission, activation calculation, and optimizer update). The control strategy uses "temperature / power exceeding threshold" as the trigger point, which fails to reflect task semantics and results in low cooling efficiency.
[0028] (2) Unable to dynamically distinguish between computing and transmission attributes: GPU devices, HOST devices, and CXL memory devices may assume different functional roles in different training stages, but existing systems usually treat all devices equivalently and fail to allocate resources based on their computing and transmission attributes.
[0029] (3) Unable to achieve coordinated control of power consumption and heat dissipation: The existing thermal management system lacks the ability to globally regulate the total power of the system. When faced with power-constrained scenarios (such as PDU (Power Distribution Unit) capacity bottlenecks), it is unable to dynamically adjust the equipment operating power and liquid cooling resource configuration according to priority, which can easily cause a significant decline in system performance.
[0030] To solve the above problems, an embodiment of the present invention provides an electronic device 10, such as Figure 1 As shown, it includes: a computing node 101 and a switching node 102, the computing node 101 is connected to the switching node 102, the computing node 101 includes at least one first device with computing properties and / or at least one second device with transmission properties, the at least one first device performs computing tasks in artificial intelligence training, and the at least one second device performs transmission tasks in artificial intelligence training; the switching node 102 includes a control module, which controls the corresponding heat dissipation device to dissipate heat for each device through the control module when the heat dissipation speed of each device is determined according to the computing heat of at least one first device and / or the transmission heat of at least one second device, so that each device is in the corresponding target temperature range.
[0031] In this embodiment of the present application, the switching node 102 is a CXL switch, which is connected to the computing node 101. The computing node 101 includes at least one first device with computing attributes and / or at least one second device with transmission attributes. For example, the computing node 101 includes a host device (also called a HOST device) (i.e., at least one first device with computing attributes), a CXL communication device (i.e., at least one second device with transmission attributes), and a CXL communication device. Figure 2 As shown, if the application sets a server host device in the computing node 101 as a computing offloading object, the server host device has computing attributes; if the application sets a HOST DDR (Double Data Rate Synchronous Dynamic Random Access Memory), which refers to a dynamic random access memory directly connected to the server host, as a data offloading object, the server host device has transmission attributes; if the application sets a CXL communication device as a data offloading object, the CXL communication device has transmission attributes.
[0032] In electronic device 10, different devices (HOST devices and CXL communication devices) can be individually preset with a basic liquid cooling flow rate level to ensure that each device maintains its normal operating temperature when idle or under low load. The present invention controls device temperature by adjusting the liquid cooling flow rate level of each device. A higher liquid cooling flow rate level for a device results in a faster liquid cooling flow rate through the device, resulting in faster heat dissipation.
[0033] For devices with significant computing properties (such as GPU devices or HOST devices that undertake computational offloading), the system sets the cooling level based on the computing intensity of their training phase. For devices that primarily undertake data caching tasks (such as CXL communication devices or HOST devices that undertake memory offloading), the cooling intensity is adjusted based on their data flow density. This strategy achieves heterogeneous thermal management collaboration across devices and effectively avoids the waste of cooling resources.
[0034] After determining the liquid cooling flow rate level for at least one first device and at least one second device based on the calculated heat output of at least one first device and / or the transferred heat output of at least one second device, the control module controls the corresponding heat dissipation device (e.g., a fan or liquid cooling system) to dissipate heat for each device, maintaining each device within a corresponding target temperature range (e.g., approximately 25°C). Each liquid cooling flow rate level corresponds to a specific flow rate range. The faster the flow rate, the greater the coolant flow rate through the device per unit time, resulting in greater heat dissipation capability.
[0035] Optionally, in some embodiments, the switching node 102 further includes: a communication module for transmitting data between the computing node 101 and the switching node 102; an application perception module for perceiving the current training stage of each device in the computing node 101 based on the data transmitted by the communication module, and calculating the computing heat of the first device and / or the transmission heat of the second device based on the current training stage of each device; a data congestion perception module for counting the sending key value amount, receiving key value amount and throughput of each device, and identifying congested devices and uncongested devices in the computing node 101 based on the sending key value amount, receiving key value amount and throughput of each device.
[0036] Optionally, in some embodiments, the electronic device 10 further includes: at least one communication protocol, wherein the communication protocol is a cxl protocol or an Ethernet protocol.
[0037] The electronic device 10 supports the CXL protocol or the Ethernet protocol for data transmission, so that users can flexibly select the appropriate communication protocol according to specific application scenarios and needs, and the electronic device can be applicable to a wider range of application scenarios.
[0038] Specifically, the CXL switch consists of the following components: a communication module, a control module, a data congestion awareness module, a power management module, and an application awareness module.
[0039] The communication module implements the CXL protocol, supports data transmission between the computing node 101 and the switching node 102 via CXL semantics, can obtain the data distribution status (cold data volume, hot data volume) of the devices connected to the CXL switch, supports hot plugging, and when a new device is connected to the switch, it will obtain the device type and operating power of the device to form a cluster topology.
[0040] The application awareness module is used to obtain the current training stage of each device in computing node 101 based on data transmitted by the communication module. It also perceives the cluster topology, communication mode, parallel training mode, and computation and data offload locations. The communication mode, parallel training mode, and task offload location are transmitted from the host to the application awareness module, while the cluster topology is transmitted from the communication module to the application awareness module.
[0041] Among them, the method of identifying the current training stage of each device in the present invention is: setting a checkpoint for the key value of each training stage in advance, and the checkpoint is used to mark the training stage; when the device executes to the corresponding checkpoint, it identifies the current training stage where the corresponding checkpoint is located.
[0042] In AI training applications, loading and generating critical data is a core part of the training process. This critical data includes gradients, model parameters, optimizer status, and activation values. Checkpoints act as markers in the training process, dividing the entire training process into several stages. Each checkpoint corresponds to a specific state in the training process, recording the key data and relevant parameters of the model at that moment. Setting checkpoints allows for easy monitoring of the training process.
[0043] When the GPU execution process reaches a preset checkpoint, the notification mechanism is triggered and a notification signal is sent to the application perception module. This notification signal contains the specific stage of the current training.
[0044] Therefore, by setting checkpoints for the key values of each training stage in advance, the current training stage can be accurately identified when the device executes to the corresponding checkpoint. This helps to accurately grasp the operating status of each device in the computing cluster at different training stages, so as to reasonably allocate resources.
[0045] The application perception module determines the computation heat of at least one first device with computation attributes and the transmission heat of at least one second device with transmission attributes according to the training communication mode, parallel mode, computation and memory offloading location set by the application.
[0046] Specifically, the computational heat of at least one first device is calculated based on the model training parameters of the current training phase:
[0047] If the current training phase of at least one first device is the gradient and parameter phase, the calculation heat of at least one first device is:
[0048] ; (1)
[0049] If the current training phase of at least one first device is the optimizer phase, the computational heat of at least one first device is:
[0050] ; (2)
[0051] If the current training phase of the at least one first device is the activation phase, the calculated heat of the at least one first device is:
[0052] . (3)
[0053] The transmission heat of at least one second device is calculated based on the current training phase, specifically, the transmission heat of at least one second device is calculated according to the model training parameters of the current training phase:
[0054] If the current training phase of at least one second device is the gradient and parameter phase, the transferred heat of at least one second device is:
[0055] ; (4)
[0056] If the current training phase of at least one second device is the optimizer phase, the transferred heat of at least one second device is:
[0057] ; (5)
[0058] If the current training phase of the at least one second device is the activation phase, the heat transferred by the at least one second device is:
[0059] . (6)
[0060] Specifically, the application awareness module uses a predefined formula according to the training parameters of the model, that is, uses formula (1-6) to calculate the calculation heat of at least one first device and the transmission heat of at least one second device in different training stages.
[0061] Therefore, by calculating the computing heat and transmission heat of the device in stages, we can more accurately understand the computing heat and transmission heat of the device in different stages, which helps to optimize resource allocation, avoid unnecessary energy waste, and improve the energy utilization efficiency of the system.
[0062] Specifically, the data congestion perception module detects congested devices and uncongested devices in the electronic device 10 in the following manner: the data congestion perception module performs real-time analysis on data packets in the CXL.mem and CXL.cache protocols, counts the amount of key values sent, the amount of key values received, and the throughput of each device, and based on the amount of key values sent, the amount of key values received, and the throughput of each device, establishes a transmission credit mechanism by using the method of establishing a transmission credit mechanism between devices in related technologies, so as to use the transmission credit mechanism to identify congested devices in the electronic device 10 where data transmission congestion occurs.
[0063] Therefore, by counting the sending key value quantity, receiving key value quantity and throughput of each device in the electronic device 10, we can fully grasp the various key indicators of the device during the data transmission process, and establish a transmission credit mechanism through these data to make the evaluation of the device transmission status more accurate and comprehensive.
[0064] Furthermore, in some embodiments, the control module includes: a first control unit, used to calculate the thermal load index of each device based on the calculated heat of the first device and / or the transferred heat of the second device, and determine the first liquid cooling flow rate level of each device based on the thermal load index of each device, and control the liquid cooling system 103 to dissipate heat for each device according to the first liquid cooling flow rate level; a second control unit, used to determine the second liquid cooling flow rate level of the blocked device, and control the liquid cooling system 103 to dissipate heat for the blocked device according to the second liquid cooling flow rate level.
[0065] Specifically, the control module combines the calculation heat of at least one computing device calculated by the application perception module and the transmission heat of at least one transmission device according to a certain weight coefficient to obtain a comprehensive heat load index of each device, thereby determining the first liquid cooling flow rate level of each device.
[0066] The control module adjusts the liquid cooling flow rate level of the blocked device to a second liquid cooling flow rate level according to the detected blocked device, and controls the liquid cooling system 103 to dissipate heat for the blocked device according to the second liquid cooling flow rate level.
[0067] Each liquid cooling flow rate level corresponds to a specific flow rate range. The faster the flow rate, the greater the coolant flow through the device per unit time, and the stronger the heat dissipation capacity. The coefficients a and b can be set based on experimental debugging results before deployment to ensure that the device temperature is controlled within the target temperature range (for example, around 25°C) under different application configurations.
[0068] For example, if the calculated heat of a device is Qc and the transmitted heat is Qt, the comprehensive heat load index Qtotal of the device can be calculated by the following formula:
[0069] Qtotal=a×Qc+b×Qt;
[0070] Here, a + b = 1, and 0 ≤ a ≤ 1, 0 ≤ b ≤ 1. In this way, the calculated heat and the transmitted heat are combined according to the set weights to obtain an indicator that can comprehensively reflect the thermal load of the equipment.
[0071] When the device is in a low heat load state, the liquid cooling flow rate is set to a low speed level. Because under low heat load conditions, a lower flow rate can meet the device's heat dissipation needs and also reduce the energy consumption of the liquid cooling system.
[0072] When the device is under medium heat load, the liquid cooling flow rate is set to medium speed level. Medium speed flow rate can provide stronger heat dissipation capacity to cope with the increase of device heat load.
[0073] When the device is under high heat load, the liquid cooling flow rate is set to high speed. The high flow rate can quickly remove the heat generated by the device and prevent the device from overheating.
[0074] Through the above technical solution, the comprehensive consideration of the two types of heat can more comprehensively reflect the actual thermal state of the device, thereby more accurately determining the first liquid cooling flow rate level for each device. Because the thermal load of the device changes dynamically with the workload, the control module can dynamically determine the first liquid cooling flow rate level for each device based on the real-time calculated thermal load index and control the liquid cooling system to adjust the heat dissipation flow rate accordingly. This dynamic adaptability enables the liquid cooling system to respond promptly to changes in the device's heating conditions, ensuring that the device always operates in an appropriate temperature environment, effectively addressing heat dissipation issues caused by device blockage, ensuring the normal operation of the device, and optimizing the allocation of heat dissipation resources.
[0075] Optionally, in some embodiments, the switching node 102 further includes: a power management module, configured to determine a current power mode of the electronic device 10 based on the total cluster power of the computing cluster in which the electronic device 10 is located, wherein the current power mode includes a power-limited mode and a power-sufficient mode.
[0076] The total cluster power of the computing cluster where the electronic device 10 is located is the sum of the current power of the switch, the current power of each device, and the liquid cooling power.
[0077] The power management module operates in two modes: sufficient power mode and limited power mode. Preset power thresholds are set within the module. The module continuously monitors the total cluster power of the computing cluster in which electronic device 10 resides, including the power of the CXL switch, the power of each host device and CXL communication equipment, and the liquid cooling power of each device. The module compares the total cluster power of the computing cluster in which electronic device 10 resides with the preset power threshold. The liquid cooling power of a device is proportional to the liquid cooling flow rate.
[0078] When the total cluster power of the computing cluster where the electronic device 10 is located is less than or equal to the preset power threshold, the electronic device 10 is determined to be in sufficient power mode and can support full power operation of the device. When the total cluster power of the computing cluster where the electronic device is located is greater than the preset power threshold, the system enters power-limited mode.
[0079] The preset power threshold may be a threshold pre-set by a user, a threshold obtained through a limited number of experiments, or a threshold obtained through a limited number of computer simulations, which is not specifically limited here.
[0080] By monitoring the total power of the computing cluster housing the electronic devices in real time and comparing it with a preset power threshold, it's possible to promptly detect whether the system's power is approaching or exceeding its capacity limit. When the total power exceeds the threshold, power-limited mode is activated, and appropriate measures can be taken to prevent system failure or damage due to overload, ensuring stable operation.
[0081] When the current power mode is the sufficient power mode, once a transmission congestion is detected between a pair of sending and receiving devices in the electronic device 10, the control module will timely control the liquid cooling system to increase the liquid cooling flow rate level of the receiving device to provide it with additional heat dissipation capacity, thereby maintaining the optimal performance of the device.
[0082] In power-limited mode, the power management module identifies the blocked device based on the data congestion perception module. If there is a blocked device with data transmission congestion, the system will prioritize allocating more power to the blocked device and increase the liquid cooling flow rate level of the receiving device in the blocked device, thereby preventing the key communication link from becoming a performance bottleneck.
[0083] Therefore, in the power-limited mode, by detecting whether there is a device group with data transmission congestion in the electronic device 10 and preferentially allocating higher power to the blocked devices, the data transmission congestion can be effectively alleviated, and the power of the receiving device in the blocked device group can be increased, thereby improving its data processing and transmission capabilities, ensuring smooth data transmission of key communication links, and avoiding problems such as data loss and increased delay caused by congestion; and this solution combines power distribution and liquid cooling, regulates the power of the equipment and adjusts the liquid cooling flow rate level, thereby achieving coordinated optimization of heat dissipation and power distribution, while ensuring the normal operation of the equipment, it also improves the overall heat dissipation efficiency of the system.
[0084] For unblocked devices where no data transmission congestion occurs, the power management module will obtain the throughput of the devices with computing attributes among the unblocked devices through the data congestion perception module, and give priority to allocating normal operating power to the first device with high throughput among the devices with computing attributes. The power of the second device with low throughput among the devices with computing attributes will be limited. At the same time, the control module controls the liquid cooling system to reduce the liquid cooling flow rate level of the second device.
[0085] Through the above technical solution, devices with high throughput are allocated more power than their current power. Increasing power can speed up the computing speed of these devices and improve their processing capabilities, thereby improving the computing performance and response speed of the entire system, ensuring that critical tasks can be completed efficiently.
[0086] For devices with low throughput, their demand for computing resources is relatively low in current business scenarios. Appropriately reducing their power can avoid power waste and allocate more limited power resources to high-throughput devices, thereby improving power utilization efficiency.
[0087] While reducing the power of equipment with low throughput, the liquid cooling flow rate level is also reduced. As the power is reduced, the heat generated by the equipment will also be reduced accordingly. At this time, lowering the liquid cooling flow rate level can reduce the energy consumption and operating costs of the liquid cooling system while meeting the heat dissipation requirements of the equipment.
[0088] Furthermore, the power management module will preferentially allocate the amount of hot and cold data of the transmission attribute device among the unblocked devices with no data transmission congestion obtained by the communication module to the normal working power of the third device with a high amount of hot data among the transmission attribute devices, while the power of the fourth device with a low amount of hot data among the transmission attribute devices will be limited, and the liquid cooling flow rate level of the fourth device will be reduced at the same time.
[0089] The above technical solution prioritizes devices with high hot and cold data transmission capacity among transmission devices, allocating higher power to them. This ensures rapid processing and transmission of critical business data, improving the overall system response speed and business processing efficiency.
[0090] For devices with small amounts of hot and cold data, appropriately reducing their power can avoid power waste and allocate more limited power resources to devices that process hot data, thus achieving reasonable allocation and efficient utilization of power resources.
[0091] While appropriately reducing the power of devices with small amounts of hot and cold data, the liquid cooling flow rate level can also be reduced. As the power is reduced, the heat generated by the equipment will also be reduced accordingly. At this time, lowering the liquid cooling flow rate level can reduce the energy consumption and operating costs of the liquid cooling system while meeting the heat dissipation requirements of the equipment.
[0092] Optionally, in some embodiments, the electronic device 10 further includes: a liquid cooling system 103, the liquid cooling system 103 is connected to the switching node 102 and the computing node 101 respectively, and the cooling water in the liquid cooling system 103 passes through the switching node 102 and the computing node 101 in turn and then flows back into the liquid cooling system 103.
[0093] The electronic device 10 adopts a closed liquid cooling cycle cooling solution as shown in the attached Figure 2 As shown:
[0094] After flowing out of the Coolant Distribution Unit (CDU) in the liquid cooling system 103 , the cooling water flows through the CXL switch and the computing node 101 in sequence and finally flows back to the CDU, completing the cooling cycle.
[0095] This technical solution allows cooling water to directly reach the key heat-generating components of the switching and computing nodes, enabling precise cooling of these components. This precise cooling ensures timely heat dissipation, prevents local overheating, and ensures equipment operates in a stable temperature environment.
[0096] The specific implementation process of the present invention is described below. Figure 4 shown.
[0097] Step S401: System startup: The control module sets the basic liquid cooling flow rate level for each device and initializes the power management module operating mode. When the AI training task begins, the application perception module collects the topology and communication parameters of the training task, and the power management module monitors the total power of the electronic equipment and the thermal load of each device in the electronic equipment in real time:
[0098] Step S402: Identify the functional attributes of each device connected to the CXL switch through the communication module, and establish a CXL device topology relationship based on each device;
[0099] Step S403: Loading the AI training application;
[0100] Step S404: Identify the current training stage of each device;
[0101] Step S405: Calculate the calculation heat of the device with calculation attribute and the transmission heat of the device with transmission attribute;
[0102] Step S406: Identify blocked devices and unblocked devices among the electronic devices;
[0103] Step S407: Calculate the comprehensive heat load index of each device, and assign a liquid cooling flow rate level to each device according to the comprehensive heat load index of each device;
[0104] Step S408: confirming the current power mode of the computing cluster where the electronic device is located;
[0105] Step S409: Dynamically adjust the liquid cooling flow rate levels and power distribution of the blocked device and the unblocked device based on the current power mode.
[0106] In summary, the electronic device proposed in the present invention has the following significant beneficial effects compared to the prior art:
[0107] (1) Build a global collaborative heat dissipation management system centered on the CXL switch to implement differentiated cooling strategies based on the functional attributes of the equipment:
[0108] This invention breaks through the static configuration limitations of traditional liquid cooling solutions in terms of heat dissipation paths and control strategies, and for the first time uses CXL switches as the core of cluster-level liquid cooling control. The system dynamically identifies the actual role of the device in the current AI training task and the source of the heat load based on the functional attributes of the device (i.e., computing attributes and transmission attributes). For devices with significant computing attributes (such as GPU nodes or HOST nodes that undertake computing offload), the system sets the cooling level according to the computing intensity of their training phase; for devices that mainly undertake data caching tasks (such as CXL memory or HOST nodes that undertake memory offload), the cooling intensity is adjusted according to their data flow density. This strategy realizes heterogeneous collaboration of thermal management across devices, effectively avoiding the waste of cooling resources.
[0109] (2) Introducing a data semantics-driven dynamic cooling control mechanism to achieve phased thermal management based on key training data:
[0110] This invention innovatively integrates the generation and use of key AI training data types (such as activations, gradients, parameters, and optimizer states) into the thermal management control logic. Based on the memory read / write pressure and computational intensity generated by data at different training stages, the system meticulously calculates the thermal load indicators for each device and pre-configures device cooling priorities by stage. This close connection between data and resources avoids the energy redundancy inherent in traditional systems, which rely on high-flow cooling at all times.
[0111] (3) Implementing priority-based joint scheduling of power and resources in power-constrained scenarios to ensure that the performance of key equipment is not affected:
[0112] Taking into account that AI clusters are often subject to total power budget constraints (such as computer room PDU capacity, electricity cost, etc.) in actual deployment, the present invention introduces a power management module to achieve coupled control of power and thermal management. In power-limited mode, the system prioritizes identifying transmission bottleneck paths or computing / transmission high-load nodes, and ensures their stable operation through dynamic power allocation and resource tilting. This process builds a priority queue based on real-time throughput (computing equipment) and hot and cold data flow density (transmission equipment), compresses the power consumption and cooling intensity of secondary task equipment, and thus maximizes training efficiency without exceeding the total power budget. Compared with the traditional average power reduction strategy, this method is more fine-grained and task-aware, effectively avoiding performance degradation of key links.
[0113] (4) Realize deep linkage between the liquid cooling system, power scheduling system and AI task configuration to enhance the system’s intelligent management capabilities:
[0114] Through the collaborative design of the communication module, application awareness module, data congestion awareness module, and power management module, a complete closed loop has been established: "Device topology awareness → Application behavior understanding → Thermal power modeling → Cooling / power coordinated control." Compared to traditional thermal management solutions dominated by the operating system or BMC (Baseboard Management Controller) control layer, this system has stronger vertical penetration (directly linked to application layer task characteristics) and horizontal coordination capabilities (dynamic scheduling across devices), supporting adaptive cooling and thermal power management in complex, heterogeneous, and high-load AI cluster environments.
[0115] According to the electronic device proposed in the embodiment of the present invention, the switching node uses a control module to control the corresponding heat dissipation device to dissipate heat for each device based on the calculation heat of the first device with calculation properties and / or the transmission heat of the second device with transmission properties in the computing node, so that each device is in the corresponding target temperature range. This solves the problem that the existing cluster system thermal management technology has coarse granularity in sensing hot and cold loads, delayed response, inability to identify dynamic changes in device roles during the AI training phase, inability to dynamically adjust the resources allocated to the device, resulting in poor system performance stability. This comprehensively improves the thermal management response speed, resource allocation accuracy, and overall performance stability of the AI training system.
[0116] The embodiment of the present invention also provides a computing cluster, such as Figure 5 As shown, the computing cluster 20 includes the electronic device 10 described above.
[0117] Secondly, if Figure 6As shown, an embodiment of the present invention further provides a heat dissipation management method, which is applied to the above-mentioned computing cluster, wherein the method includes the following steps:
[0118] Step S601 : Identify the current training phase of each device in the computing cluster.
[0119] Optionally, in some embodiments, identifying the current training stage of each device in the computing cluster includes: setting a checkpoint for the key value of each training stage in advance, the checkpoint is used to mark the training stage; when the device executes to the corresponding checkpoint, identifying the current training stage where the corresponding checkpoint is located.
[0120] As you can understand, each device in the computing cluster has a checkpoint. Each checkpoint corresponds to a specific state in the training process, recording the key data and relevant parameters of the model at that moment. By setting checkpoints, you can easily monitor the training process.
[0121] When the execution process of the device reaches a preset checkpoint, the notification mechanism will be triggered to send a notification signal to the application perception module. This notification signal contains the current training stage of the current training.
[0122] Step S602: Calculate the computational heat of at least one first device with computational properties and the transmission heat of at least one second device with transmission properties based on the current training stage of each device; calculate the comprehensive thermal load index of each device based on the computational heat of at least one first device and the transmission heat of at least one second device; and determine the current liquid cooling flow rate level of each device based on the comprehensive thermal load index of each device.
[0123] Optionally, in some embodiments, calculating the computational heat of at least one first device and the transmission heat of at least one second device based on the current training phase includes: calculating the computational heat of at least one first device according to model training parameters of the current training phase, wherein:
[0124] If the current training phase is the gradient and parameter phase, the calculation heat of at least one first device is:
[0125] ;
[0126] If the current training phase is the optimizer phase, the computational heat of at least one first device is:
[0127] ;
[0128] If the current training phase is the activation phase, the calculated heat of at least one first device is:
[0129] .
[0130] Optionally, in some embodiments, calculating the transferred heat of at least one second device based on the current training phase further includes: calculating the transferred heat of at least one second device based on the model training parameters of the current training phase, wherein:
[0131] If the current training phase is the gradient and parameter phase, the amount of heat transferred by at least one second device is:
[0132] ;
[0133] If the current training phase is the optimizer phase, the amount of heat transferred by the at least one second device is:
[0134] ;
[0135] If the current training phase is the activation phase, the amount of heat transferred by the at least one second device is:
[0136] .
[0137] Specifically, the application awareness module calculates the computational heat of at least one device with computational attributes and the transmission heat of at least one device with transmission attributes in the computing cluster at different training stages according to the training parameters of the model.
[0138] The application perception module combines the calculated heat and the transmitted heat according to certain weight coefficients a and b to obtain a comprehensive heat load index, thereby obtaining the current liquid cooling flow rate level of each device.
[0139] Furthermore, the current power mode is determined based on the current total power of the computing cluster, and the liquid cooling flow rate level of each device and / or the power distribution of each device are dynamically adjusted according to the current power mode so that each device is in the corresponding target temperature range.
[0140] Step S603: determine the current power mode based on the current total power of the computing cluster, and dynamically adjust the liquid cooling flow rate level of each device and / or adjust the power allocation of each device according to the current power mode so that each device is in the corresponding target temperature range.
[0141] Optionally, in some embodiments, the current total power of the computing cluster is the sum of the current power of the switch node, the current power of each device, and the liquid cooling power.
[0142] The target temperature range may be a threshold value pre-set by the user, a threshold value obtained through a limited number of experiments, or a threshold value obtained through a limited number of computer simulations, which is not specifically limited here.
[0143] Optionally, in some embodiments, the current power mode is determined based on the current total power of the computing cluster, including: monitoring the current total power of the computing cluster; determining whether the current total power is greater than a preset power threshold; if the current total power is greater than the preset power threshold, determining that the current power mode is a power-limited mode, otherwise, determining that the current power mode is a power-sufficient mode.
[0144] The preset power threshold may be a threshold pre-set by a user, a threshold obtained through a limited number of experiments, or a threshold obtained through a limited number of computer simulations, which is not specifically limited here.
[0145] It should be understood that when the current total power of the computing cluster is less than or equal to the preset power threshold, the computing cluster is determined to be in sufficient power mode and can support full power operation of the equipment. When the current total power of the computing cluster is greater than the preset power threshold, the computing cluster enters power-limited mode.
[0146] Optionally, in some embodiments, the liquid cooling flow rate level of each device and / or the power allocation of each device are dynamically adjusted according to the current power mode, including: if the current power mode is a power-limited mode, detecting whether there is a blocked device in the computing cluster where data transmission is blocked; if there is a blocked device where data transmission is blocked, preferentially allocating the first power to the blocked device, and increasing the liquid cooling flow rate level of the receiving device in the blocked device.
[0147] It should be understood that after the computing cluster enters the power-limited mode, the power management module confirms whether there is a blocked device with data transmission congestion in the computing cluster based on the data congestion perception module. If there is a blocked device with data transmission congestion in the computing cluster, the computing cluster prioritizes allocating more power to the blocked device with data transmission congestion, that is, allocating the first power to the blocked device. The power of the blocked device after the first power is allocated is greater than the power before allocation, and the liquid cooling flow rate level of the receiving device in the blocked device is increased.
[0148] Optionally, in some embodiments, after determining whether there is a blocked device in the computing cluster where data transmission is blocked, the method further includes: determining unblocked devices in the computing cluster where no data transmission is blocked, and obtaining the throughput of devices with computing attributes among the unblocked devices; determining a first device with a throughput greater than a first preset value and a second device with a throughput less than or equal to the first preset value among the devices with computing attributes; after allocating a second power to the first device among the devices with computing attributes, allocating a third power to the second device among the devices with computing attributes, and simultaneously reducing the liquid cooling flow rate level of the second device.
[0149] The first preset value may be a threshold value pre-set by a user, a threshold value obtained through a limited number of experiments, or a threshold value obtained through a limited number of computer simulations, which is not specifically limited here.
[0150] It should be understood that for unblocked devices in the computing cluster where no data transmission congestion occurs, the power management module will obtain the throughput of the devices with computing attributes in the unblocked devices based on the data congestion perception module, and give priority to allocating normal operating power (i.e., allocating the second power) to the first device with high throughput. The power of the second device with low throughput will be limited, that is, the third power will be allocated to the second device, and at the same time, the liquid cooling flow rate level of the second device will be reduced. The power of the second device after the third power is allocated is less than the power before allocation.
[0151] Optionally, in some embodiments, after determining the unblocked devices in the computing cluster where no data transmission congestion occurs, it also includes: obtaining the amount of hot and cold data of the devices with transmission properties in the unblocked devices; determining a third device with an amount of hot and cold data greater than a second preset value and a fourth device with an amount of hot and cold data less than or equal to the second preset value among the devices with transmission properties; after allocating the fourth power to the third device among the devices with transmission properties, allocating the fifth power to the fourth device among the devices with transmission properties, and at the same time reducing the liquid cooling flow rate level of the fourth device.
[0152] The second preset value may be a threshold value pre-set by the user, a threshold value obtained through a limited number of experiments, or a threshold value obtained through a limited number of computer simulations, which is not specifically limited here.
[0153] It should be understood that the power management module preferentially allocates normal operating power to the third device with high hot data volume among the devices with transmission attributes based on the amount of hot and cold data of the devices with transmission attributes among the unblocked devices with no data transmission congestion obtained by the communication module, that is, the fourth power is allocated to the third device, and the power of the third device after the fourth power is allocated is greater than the power before allocation, while the power of the fourth device with low hot data volume will be limited, that is, the fifth power is allocated to the fourth device, and the liquid cooling flow rate level of the fourth device is reduced at the same time, and the power of the fourth device after the fifth power is allocated is less than the power before allocation.
[0154] Optionally, in some embodiments, when the current power mode is a sufficient power mode, it includes: determining whether there is a target device in each device where data transmission is blocked; if there is a target device where data transmission is blocked, increasing the liquid cooling flow rate level of the receiving device in the target device.
[0155] It is understandable that when the current power mode is sufficient power mode, once a data transmission congestion is detected between a pair of sending and receiving devices (i.e., target devices), the computing cluster will timely increase the liquid cooling flow rate level of the receiving device to provide it with additional heat dissipation capacity, thereby maintaining the optimal performance of the device.
[0156] Optionally, in some embodiments, detecting whether there is a blocked device where data transmission is blocked in the computing cluster also includes: counting the sending key value amount, receiving key value amount and throughput of each device in the computing cluster; establishing a transmission credit mechanism between devices based on the sending key value amount, receiving key value amount and throughput of each device, so as to utilize the transmission credit mechanism between devices to identify the blocked device where data transmission is blocked.
[0157] It can be understood that the data congestion perception module counts the sending key value quantity, receiving key value quantity and throughput of each device in the computing cluster, and based on the sending key value quantity, receiving key value quantity and throughput of each device, establishes a transmission credit mechanism by using the method of establishing a transmission credit mechanism between devices in related technologies, so as to use the transmission credit mechanism to identify blocked devices and non-blocked devices in the computing cluster where data transmission congestion occurs.
[0158] According to the heat dissipation management method proposed in an embodiment of the present invention, the current training stage of each device in the computing cluster is identified, and based on the current training stage of each device, the computational heat of at least one first device with computational properties and the transmission heat of at least one second device with transmission properties are calculated; the comprehensive thermal load index of each device is calculated based on the computational heat of at least one first device and the transmission heat of at least one second device, and the current liquid cooling flow rate level of each device is determined based on the comprehensive thermal load index of each device; the current power mode is determined based on the current total power of the computing cluster, and the liquid cooling flow rate level of each device and / or the power allocation of each device are dynamically adjusted according to the current power mode, so that each device is in the corresponding target temperature range. In this way, the existing cluster system thermal management technology solves the problems of coarse granularity of hot and cold load perception, delayed response, inability to identify dynamic changes in device roles in the AI training stage, inability to dynamically adjust the resources allocated to the device, resulting in poor system performance stability, and comprehensively improves the thermal management response speed, resource allocation accuracy and overall performance stability of the AI training system.
[0159] Figure 7 FIG. 4 is a schematic diagram of a heat dissipation management system according to an embodiment of the present invention.
[0160] like Figure 7 As shown, the heat dissipation management system 30 includes: an identification module 100 , a calculation module 200 and a heat dissipation management module 300 .
[0161] Among them, the identification module 100 is used to identify the current training stage of each device in the computing cluster; the computing module 200 is used to calculate the computing heat of at least one first device with computing properties and the transmission heat of at least one second device with transmission properties based on the current training stage of each device, calculate the comprehensive thermal load index of each device according to the computing heat of the first device and the transmission heat of the second device, and determine the current liquid cooling flow rate level of each device according to the comprehensive thermal load index of each device; the heat dissipation management module 300 is used to determine the current power mode based on the current total power of the computing cluster, and dynamically adjust the liquid cooling flow rate level of each device and / or adjust the power distribution of each device according to the current power mode, so that each device is in the corresponding target temperature range.
[0162] Optionally, in some embodiments, the calculation module 200 is further configured to calculate the calculation heat of the first device according to the model training parameters of the current training phase, wherein if the current training phase is the gradient and parameter phase, the calculation heat of the first device is:
[0163] ;
[0164] If the current training phase is the optimizer phase, the computational heat of the first device is:
[0165] ;
[0166] If the current training phase is the activation phase, the calculated heat of the first device is:
[0167] .
[0168] Optionally, in some embodiments, the calculation module 200 is further configured to calculate the transferred heat of the second device according to the model training parameters of the current training phase, wherein:
[0169] If the current training phase is the gradient and parameter phase, the transferred heat of the second device is:
[0170] ;
[0171] If the current training phase is the optimizer phase, the transferred heat of the second device is:
[0172] ;
[0173] If the current training phase is the activation phase, the transferred heat of the second device is:
[0174] .
[0175] Optionally, in some embodiments, the heat dissipation management module 300 is further used to: monitor the current total power of the computing cluster; determine whether the current total power is greater than a preset power threshold; if the current total power is greater than the preset power threshold, determine that the current power mode is a power-limited mode; otherwise, determine that the current power mode is a power-sufficient mode.
[0176] Optionally, in some embodiments, the heat dissipation management module 300 is further used to: if the current power mode is a power-limited mode, detect whether there is a blocked device in the computing cluster where data transmission is blocked; if there is a blocked device where data transmission is blocked, preferentially allocate the first power to the blocked device, and increase the liquid cooling flow rate level of the receiving end device in the blocked device.
[0177] Optionally, in some embodiments, after determining whether there is a blocked device in the computing cluster where data transmission is blocked, the heat dissipation management module 300 is further used to: determine unblocked devices in the computing cluster where no data transmission is blocked, and obtain the throughput of devices with computing attributes among the unblocked devices; determine a first device with a throughput greater than a first preset value and a second device with a throughput less than or equal to the first preset value among the devices with computing attributes; after allocating a second power to the first device among the devices with computing attributes, allocate a third power to the second device among the devices with computing attributes, and simultaneously reduce the liquid cooling flow rate level of the second device.
[0178] Optionally, in some embodiments, after determining unblocked devices in the computing cluster where no data transmission congestion has occurred, the heat dissipation management module 300 is further used to: obtain the amount of hot and cold data of the devices with transmission properties in the unblocked devices; determine a third device with an amount of hot and cold data greater than a second preset value and a fourth device with an amount of hot and cold data less than or equal to the second preset value among the devices with transmission properties; after allocating a fourth power to the third device among the devices with transmission properties, allocate a fifth power to the fourth device among the devices with transmission properties, and simultaneously reduce the liquid cooling flow rate level of the fourth device.
[0179] Optionally, in some embodiments, when the current power mode is the sufficient power mode, the heat dissipation management module 300 is also used to: determine whether there is a target device in each device where data transmission is blocked; if there is a target device where data transmission is blocked, increase the liquid cooling flow rate level of the receiving device in the target device.
[0180] Optionally, in some embodiments, to detect whether there is a blocked device in the computing cluster where data transmission is blocked, the heat dissipation management module 300 is also used to: count the sending key value amount, receiving key value amount and throughput of each device in the computing cluster; establish a transmission credit mechanism between devices based on the sending key value amount, receiving key value amount and throughput of each device, so as to utilize the transmission credit mechanism between devices to identify the blocked device where data transmission is blocked.
[0181] Optionally, in some embodiments, the current total power of the computing cluster is the sum of the current power of the switch node, the current power of each device, and the liquid cooling power.
[0182] Optionally, in some embodiments, the identification module 100 is further used to: pre-set checkpoints for key values of each training stage, where the checkpoints are used to mark the training stages; and identify the current training stage where the corresponding checkpoint is located when the device executes to the corresponding checkpoint.
[0183] According to the heat dissipation management system proposed in an embodiment of the present invention, the current training stage of each device in the computing cluster is identified, and based on the current training stage of each device, the computational heat of at least one first device with computational properties and the transmission heat of at least one second device with transmission properties are calculated; the comprehensive thermal load index of each device is calculated based on the computational heat of at least one first device and the transmission heat of at least one second device, and the current liquid cooling flow rate level of each device is determined based on the comprehensive thermal load index of each device; the current power mode is determined based on the current total power of the computing cluster, and the liquid cooling flow rate level of each device and / or the power allocation of each device are dynamically adjusted according to the current power mode, so that each device is in the corresponding target temperature range. In this way, the existing cluster system thermal management technology solves the problems of coarse granularity of hot and cold load perception, delayed response, inability to identify dynamic changes in device roles in the AI training stage, inability to dynamically adjust the resources allocated to the device, resulting in poor system performance stability, and comprehensively improves the thermal management response speed, resource allocation accuracy and overall performance stability of the AI training system.
[0184] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned heat dissipation management method embodiments when running.
[0185] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0186] An embodiment of the present invention further provides a computer program product, including a computer program, which implements the above-mentioned heat dissipation management method when executed by a processor.
[0187] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0188] The electronic device, computing cluster, heat dissipation management method, and storage medium provided by the present invention are introduced in detail above. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, the present invention can also be improved and modified in several ways, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.
Claims
1. An electronic device comprising: Computing nodes and switching nodes, characterized in that, The computing node is connected to the switching node, and the computing node includes at least one first device with computing properties and / or at least one second device with transmission properties, wherein at least one first device performs computing tasks in artificial intelligence training, and at least one second device performs transmission tasks in the artificial intelligence training; The switching node includes a control module, which controls a corresponding heat dissipation device to dissipate heat for each device so that each device is within a corresponding target temperature range, when a heat dissipation rate of each device is determined based on a calculated heat amount of at least one first device and / or a transmitted heat amount of at least one second device. The switching node further includes: a communication module, configured to transmit data between the computing node and the switching node; an application sensing module, configured to sense a current training stage of each device in the computing node based on data transmitted by the communication module, and calculate a computational heat of at least one of the first devices and / or a transmission heat of at least one of the second devices based on the current training stage of each device; The data congestion perception module is used to count the sending key value amount, receiving key value amount and throughput of each device, and identify the congested devices and non-congested devices in the computing node based on the sending key value amount, receiving key value amount and throughput of each device.
2. The electronic device according to claim 1, wherein Also includes: A liquid cooling system is connected to the switching node and the computing node respectively, and cooling water in the liquid cooling system flows through the switching node and the computing node in sequence and then flows back into the liquid cooling system.
3. The electronic device according to claim 1, wherein The control module includes: a first control unit, configured to calculate a heat load index of each of the devices based on the calculated heat of at least one of the first devices and / or the transferred heat of at least one of the second devices, determine a first liquid cooling flow rate level for each device based on the heat load index of each device, and control the liquid cooling system to dissipate heat for each device according to the first liquid cooling flow rate level; The second control unit is used to determine a second liquid cooling flow rate level of the blocked device and control the liquid cooling system to dissipate heat for the blocked device according to the second liquid cooling flow rate level.
4. The electronic device according to claim 1, wherein: The switching node further includes: The power management module is configured to determine a current power mode of the electronic device according to the total cluster power of the computing cluster in which the electronic device is located, wherein the current power mode includes a power-limited mode and a power-sufficient mode.
5. The electronic device according to claim 1, wherein Also includes: At least one communication protocol, wherein the communication protocol is a CXL protocol or an Ethernet protocol.
6. A computing cluster, characterized in that: include: At least one electronic device according to any one of claims 1 to 5.
7. A heat dissipation management method, characterized in that: Applied to the computing cluster according to claim 6, wherein the method comprises the following steps: Identify the current training phase of each device in the computing cluster; calculating, based on a current training phase of each device, a computational heat of at least one first device having a computational attribute and a transferred heat of at least one second device having a transferred attribute, calculating a comprehensive heat load index of each device based on the computational heat of the at least one first device and the transferred heat of the at least one second device, and determining a current liquid cooling flow rate level of each device based on the comprehensive heat load index of each device; A current power mode is determined based on the current total power of the computing cluster, and the liquid cooling flow rate level of each device and / or the power distribution of each device are dynamically adjusted according to the current power mode so that each device is in a corresponding target temperature range.
8. The heat dissipation management method according to claim 7, characterized in that: Calculating the computational heat of at least one first device and the transmission heat of at least one second device based on the current training phase includes: Calculating the computational heat of at least one of the first devices according to the model training parameters of the current training phase, wherein: If the current training phase is the gradient and parameter phase, the calculation heat of at least one of the first devices is: ; If the current training phase is the optimizer phase, the computational heat of at least one of the first devices is: ; If the current training phase is an activation phase, the calculated heat of at least one of the first devices is: 。 9. The heat dissipation management method according to claim 8, characterized in that: Calculating the transferred heat of at least one second device based on the current training phase further includes: Calculating the transferred heat of at least one of the second devices according to the model training parameters of the current training phase, wherein: If the current training phase is the gradient and parameter phase, the amount of heat transferred by at least one of the second devices is: ; If the current training phase is the optimizer phase, the amount of heat transferred by at least one of the second devices is: ; If the current training phase is the activation phase, the amount of heat transferred by at least one of the second devices is: 。 10. The heat dissipation management method according to claim 7, characterized in that: The determining the current power mode based on the current total power of the computing cluster includes: monitoring the current total power of the computing cluster; Determining whether the current total power is greater than a preset power threshold; If the current total power is greater than the preset power threshold, the current power mode is determined to be a power-limited mode; otherwise, the current power mode is determined to be a power-sufficient mode.
11. The heat dissipation management method according to claim 10, characterized in that: Dynamically adjusting the liquid cooling flow rate level of each device and / or adjusting the power allocation of each device according to the current power mode includes: If the current power mode is the power limited mode, detecting whether there is a blocked device in the computing cluster where data transmission is blocked; If there is a blocked device where data transmission is blocked, the first power is preferentially allocated to the blocked device, and the liquid cooling flow rate level of the receiving end device in the blocked device is increased.
12. The heat dissipation management method according to claim 11, characterized in that: After detecting whether there is a blocked device in the computing cluster where data transmission is blocked, the method further includes: Determine unblocked devices in the computing cluster that are not blocked in data transmission, and obtain the throughput of devices with computing attributes among the unblocked devices; Determine, among the devices having computing attributes, a first device whose throughput is greater than a first preset value and a second device whose throughput is less than or equal to the first preset value; After allocating the second power to the first device among the devices with computing attributes, allocating the third power to the second device among the devices with computing attributes, and reducing the liquid cooling flow rate level of the second device.
13. The heat dissipation management method according to claim 12, wherein: After determining an unblocked device in the computing cluster where no data transmission blockage occurs, the method further includes: Obtaining the amount of hot and cold data of the device with transmission attributes in the unblocked device; Determine, among the devices having the transmission attribute, a third device whose amount of hot and cold data is greater than a second preset value and a fourth device whose amount of hot and cold data is less than or equal to the second preset value; After allocating the fourth power to the third device among the devices with transmission properties, allocating the fifth power to the fourth device among the devices with transmission properties, and reducing the liquid cooling flow rate level of the fourth device.
14. The heat dissipation management method according to claim 10, characterized in that: When the current power mode is the sufficient power mode, the method includes: Determining whether there is a target device with data transmission congestion among each of the devices; If there is a target device where data transmission is blocked, the liquid cooling flow rate level of the receiving end device in the target device is increased.
15. The heat dissipation management method according to claim 11, characterized in that: The detecting whether there is a blocked device in the computing cluster where data transmission is blocked also includes: Counting the key value sent, key value received, and throughput of each device in the computing cluster; A transmission credit mechanism between devices is established based on the sending key value amount, receiving key value amount and throughput of each device, so as to identify the blocked device where data transmission is blocked by utilizing the transmission credit mechanism between devices.
16. The heat dissipation management method according to claim 7, characterized in that: The current total power of the computing cluster is the sum of the current power of the switch node, the current power of each device, and the liquid cooling power.
17. The heat dissipation management method according to claim 7, characterized in that: The identification computing cluster includes the following steps: Setting a checkpoint for a key value of each training stage in advance, wherein the checkpoint is used to mark the training stage; When the device reaches a corresponding checkpoint, it identifies the current training phase where the corresponding checkpoint is located.
18. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the heat dissipation management method according to any one of claims 7 to 17.
Citation Information
Patent Citations
Distributed training system, method, device and equipment and readable storage medium
CN115310566A
Server and thermal management system, method, product, equipment and medium thereof
CN120066222A