Multi-heat-source partition heat pipe array packaging system and method applied to AI server
By adopting a multi-heat source partitioned heat pipe array packaging system on the AI server, combined with technical means such as real-time temperature data acquisition, power consumption distribution map acquisition and three-dimensional coupling model construction, the problem of low heat dissipation efficiency of AI servers is solved, and more efficient and accurate heat dissipation effects are achieved.
Patent Information
- Application Number
- CN202510548171.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-28
AI Technical Summary
During operation, due to the increase in computing performance and chip integration, the heat caused by the increase in AI servers, it is difficult for existing heat dissipation technologies to accurately locate high-temperature areas and efficient heat dissipation, resulting in low heat dissipation efficiency.
A multi-heat source partitioned heat pipe array packaging system is adopted. The system includes a real-time temperature data acquisition module, a power consumption distribution map acquisition module, a full-domain thermal network acquisition module, a three-dimensional coupled model construction module, an embedded heat dissipation module generation module and an AI server temperature control module. Through dynamic temperature sensing networks, adaptive non-uniform grids, cluster analysis and three-dimensional coupling models and other technical means, it realizes accurate monitoring and control of the internal temperature and power consumption of the AI server.
It improves the accuracy and efficiency of the thermal dissipation of AI servers, can accurately locate high-temperature areas and take targeted thermal dissipation measures, significantly improving the stable operation and performance of the AI server.
Smart Images

Figure CN120076278A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of chip heat dissipation, and particularly to a multi-heat-source partition heat pipe array packaging system and method applied to an AI server. Background Art
[0002] During the operation of an AI server, the heat dissipation problem is of crucial importance. With the rapid development of AI technology, the computing performance of AI servers has been continuously improved, and the chip integration has become higher and higher, which has caused a significant increase in the heat generated by the server during operation. When dealing with the multi-heat-source heat dissipation of AI servers, the existing heat dissipation technologies have obvious deficiencies in terms of accuracy. Traditional temperature monitoring methods often use a small number of sensors at fixed positions, making it difficult to comprehensively and accurately obtain the complex temperature distribution inside the server, and unable to accurately locate high-temperature areas, resulting in weak pertinence of heat dissipation measures.
[0003] In terms of heat dissipation effect, the existing heat dissipation technologies are also difficult to meet the requirements. The common air-cooled heat dissipation method only relies on forced convection of fans. In the case of multiple heat sources with uneven heat source distribution, it is unable to efficiently carry away heat. For some key chips with high heat generation, the heat dissipation capacity of air cooling is limited, which is likely to cause local overheating and affect the performance and lifespan of the chips. Although liquid cooling technology has relatively high heat dissipation efficiency, it has problems such as high cost and complex maintenance, and there are also certain challenges in adapting to the internal structure of AI servers, making it unable to be widely applied. These problems comprehensively lead to low efficiency of the existing technology in AI server heat dissipation, seriously restricting the stable operation and performance of AI servers. Summary of the Invention
[0004] The present invention provides a multi-heat-source partition heat pipe array packaging system and method applied to an AI server, and its main purpose is to solve the problem of low heat dissipation efficiency of AI servers.
[0005] To achieve the above object, a multi-heat-source partition heat pipe array packaging system applied to an AI server provided by the present invention is characterized in that the system includes a real-time temperature data acquisition module, a power consumption distribution map acquisition module, a global heat network acquisition module, a three-dimensional coupling model construction module, an embedded heat dissipation module generation module, and an AI server temperature control module, wherein: The real-time temperature data acquisition module is used to dynamically configure the sensor acquisition period for the AI server embedded with a distributed thermocouple array to obtain the dynamic temperature sensing network of the AI server, and based on the dynamic temperature sensing network, obtain the real-time temperature data of the AI server; The power consumption distribution map acquisition module is used to create an adaptive non-uniform grid of the AI server based on pre-acquired real-time power consumption data, fit the mapping relationship between current and heat generation in the AI server based on the adaptive non-uniform grid, and generate the power consumption distribution map of the AI server based on the mapping relationship and the dynamic reduction algorithm; The global heat network acquisition module is used to perform clustering analysis on the heat source distribution of the AI server to obtain the global heat network of the AI server; The three-dimensional coupling model construction module is used to construct a three-dimensional coupling model of the AI server based on the real-time temperature data, the power consumption distribution map, and the global heat network; The embedded heat dissipation module generation module is used to construct a multi-heat source partition heat pipe array of the AI server based on the global heat network, and perform interface packaging on the chips of the AI server and the multi-heat source partition heat pipe array to obtain the embedded heat dissipation module of the AI server; The AI server temperature control module is used to perform temperature control of the AI server on the embedded heat dissipation module based on the three-dimensional coupling model.
[0006] In a preferred embodiment, when the real-time temperature data acquisition module executes the dynamic configuration of the sensor acquisition period for the AI server with the embedded distributed thermocouple array to obtain the dynamic temperature sensing network of the AI server, it is specifically used for: Embed miniaturized K-type thermocouples arranged in a hexagonal close-packed pattern on the packaging substrate of the AI server to obtain the embedded distributed thermocouple array of the AI server; Dynamically configure the acquisition period for the embedded distributed thermocouple array to obtain the 3D temperature field model of the AI server; Superimpose the 3D temperature field model on the digital twin of the physical substrate to obtain the dynamic temperature sensing network of the AI server.
[0007] In a preferred embodiment, when the power consumption distribution map acquisition module executes the creation of the adaptive non-uniform grid of the AI server based on pre-acquired real-time power consumption data, it is specifically used for: Read the real-time power consumption data of the AI server through the SVID protocol; Calculate the global interpolation field of the AI server based on the real-time power consumption data, where the calculation of the global interpolation field is as follows: In the formula, is the adaptive non-uniform grid, is the three-dimensional coordinate point of the chip of the AI server, is the total number of interpolations of the full-field interpolation field, is the interpolation number of the full-field interpolation field, is the weight of the interpolation of the full-field interpolation field affected by temperature, is the Gaussian kernel decay function, is the real-time power consumption impact factor in the real-time power consumption data, is the spatial distance between two interpolations of the full-field interpolation field, is the three-dimensional coordinate point where the embedded distributed thermocouple array is deployed, is the attenuation factor of heat transfer; Based on the global interpolation field, mesh division is performed on the AI server to obtain the adaptive non-uniform grid of the AI server.
[0008] In a preferred embodiment, when the power consumption distribution map acquisition module executes fitting the mapping relationship between the current and the heat generation amount in the AI server based on the adaptive non-uniform grid and generating the power consumption distribution map of the AI server based on the mapping relationship and the dynamic order reduction algorithm, it is specifically used for: Extracting the busbar layer current density distribution of the AI server based on the adaptive non-uniform grid; Generating the Joule heat increment field of the AI server based on the adaptive non-uniform grid; Performing electromagnetic-thermal coupling on the busbar layer current density distribution and the Joule heat increment field to obtain the electromagnetic-thermal force correlation of the AI server; Generating the mapping relationship between the current and the heat generation amount in the AI server based on the electromagnetic-thermal force correlation; Constructing a 2000-order state space model of the AI server based on the mapping relationship; Based on the dynamic order reduction algorithm, performing hierarchical feature truncation and order reduction on the 2000-order state space model to obtain the hierarchical feature vector of the AI server, where the dynamic order reduction algorithm is as follows: In the formula, is the hierarchical feature vector obtained after order reduction and compression of the 2000-order state space model, is the 2000-order state space model, is the inverted heat conduction matrix, is the heat conduction matrix, is the matrix inversion flag, is the hierarchical feature vector, is the reconstructed temperature field after reducing the error, is the balance parameter of the thermophysical constraint and the statistical characteristics, is the potential feature distribution of the temperature field in the 2000-order state space model, is the regularization constraint on the potential features of the temperature field in the 2000-order state space model, The global KL constraint ignores the importance of local hotspots. is the potential feature of the temperature field. is the original temperature field in the 2000-order state space model. is the reconstructed temperature field in the 2000-order state space model.
[0009] Perform power consumption feature fusion on the hierarchical feature vector to obtain the power consumption distribution map of the AI server.
[0010] In a preferred embodiment, when the global heat network acquisition module performs clustering analysis on the heat source distribution of the AI server to obtain the global heat network of the AI server, it specifically is used for: Extract thermodynamic features from the heat source distribution of the AI server to obtain the thermodynamic data of the heat source distribution; Perform clustering analysis on the thermodynamic data to obtain the high-temperature cluster data of the AI server; Perform grid division on the heat source distribution to obtain the grid of the heat source distribution; Mark the grid based on the high-temperature cluster data: If the high-temperature cluster data of the grid exceeds the preset high-temperature threshold, mark it as a temperature-sensitive grid; If the high-temperature cluster data of the grid is lower than the preset high-temperature threshold, mark it as a temperature-insensitive grid; Collect the temperature-sensitive grids and the temperature-insensitive grids as the global heat network of the AI server.
[0011] In a preferred embodiment, when the three-dimensional coupling model construction module constructs the three-dimensional coupling model of the AI server based on the real-time temperature data, the power consumption distribution map, and the global heat network, it specifically is used for: Perform time series alignment processing on the power consumption distribution map and the real-time temperature data to obtain the standard data of the AI server; Construct an unstructured tetrahedral grid model of the physical structure of the AI server based on the global heat network to obtain the thermal gradient region model of the AI server; Assign multi-physical field coupling nodes to the high current density area in the power consumption distribution map to obtain the initial model of the AI server; Perform electro-thermal coupling analysis on the AI server based on the multi-physical field coupling nodes of the initial model and the standard data to obtain the electro-thermal coupling model of the AI server; Construct the three-dimensional coupling model of the AI server based on the electro-thermal coupling model and the thermal gradient region model.
[0012] In a preferred embodiment, when constructing the three-dimensional coupling model of the AI server based on the electro-thermal coupling model and the thermal gradient region model, the three-dimensional coupling model construction module is specifically configured to: Define a coupling interface in the key thermal gradient area of the thermal gradient region model to obtain the overlapping domain coupling interface of the thermal gradient region model; Extract the global temperature residual during the iteration of the thermal gradient region model; Extract the heat flux residual during the iteration of the electro-thermal coupling model; When the global temperature residual is less than the thermal gradient threshold and the heat flux residual is less than the electro-thermal threshold, obtain a converged thermal gradient region model and a converged electro-thermal coupling model; Perform thermal-electric-force cross-scale coupling on the converged thermal gradient region model and the converged electro-thermal coupling model based on the overlapping domain coupling interface to obtain the three-dimensional coupling model of the AI server.
[0013] In a preferred embodiment, when constructing the multi-source partition heat pipe array of the AI server based on the global thermal network and performing interface encapsulation on the chip of the AI server and the multi-source partition heat pipe array to obtain the embedded heat dissipation module of the AI server, the embedded heat dissipation module generation module is specifically configured to: Extract the heat source characteristics of the global thermal network to obtain the heat source characteristics of the AI server; Arrange heat pipes for the AI server based on the heat source characteristics to obtain the multi-source partition heat pipe array of the AI server; Perform gradient brazing on the multi-source partition heat pipe array and the chip of the AI server to obtain the initial heat dissipation module of the AI server; Perform multi-layer hermetic packaging on the initial heat dissipation module to obtain the embedded heat dissipation module of the AI server.
[0014] In a preferred embodiment, when controlling the temperature of the AI server for the embedded heat dissipation module based on the three-dimensional coupling model, the AI server temperature control module is specifically configured to: Perform short-time domain rolling prediction on the AI server based on the three-dimensional coupling model; When the prediction result is that the AI server has no overheating risk, the embedded heat dissipation module maintains the current working state; When the prediction result is that the AI server has a slight overheating risk, the embedded heat dissipation module increases the rotation speed of the condensation end fan of the partition heat pipe; When the prediction result is that the AI server has a moderate overheating risk, the embedded heat dissipation module activates the thermoelectric cooler to perform local point cooling on the AI server; When the prediction result indicates a severe overheating risk for the AI server, the embedded heat dissipation module activates dynamic voltage and frequency adjustment, increases the power of the thermoelectric cooler, and increases the rotational speed of the condensation end fan of the partition heat pipe to perform global cooling on the AI server.
[0015] To solve the above problems, the present invention also provides a multi-heat source partition heat pipe array packaging method applied to an AI server, and the method includes: S1. Dynamically configure the sensor acquisition period for the AI server with an embedded distributed thermocouple array to obtain the dynamic temperature sensing network of the AI server, and obtain the real-time temperature data of the AI server based on the dynamic temperature sensing network; S2. Create an adaptive non-uniform grid for the AI server based on the pre-acquired real-time power consumption data, fit the mapping relationship between current and heat generation in the AI server based on the adaptive non-uniform grid, and generate a power consumption distribution map of the AI server based on the mapping relationship and the dynamic reduction algorithm; S3. Perform clustering analysis on the heat source distribution of the AI server to obtain the global heat network of the AI server; S4. Construct a three-dimensional coupling model of the AI server based on the real-time temperature data, the power consumption distribution map, and the global heat network; S5. Construct a multi-heat source partition heat pipe array for the AI server based on the global heat network, and perform interface packaging on the chips of the AI server and the multi-heat source partition heat pipe array to obtain the embedded heat dissipation module of the AI server; S6. Perform temperature control of the AI server on the embedded heat dissipation module based on the three-dimensional coupling model.
[0016] Compared with the prior art, the present invention has the following beneficial effects: 1. Through the real-time temperature data acquisition module, the present invention uses miniaturized K-type thermocouples arranged in a hexagonal close-packed pattern when embedding a distributed thermocouple array, and cooperates with dynamically configuring the sensor acquisition period to construct a dynamic temperature sensing network. This network can accurately collect the 3D temperature field model of the AI server, and superimpose it on the digital twin of the physical substrate to form a digital twin thermal mirror, so as to grasp the internal temperature distribution of the server in real time, comprehensively and accurately, accurately locate the global heat network domain, provide a reliable basis for the precise implementation of subsequent heat dissipation measures, and greatly improve the accuracy of heat dissipation.
[0017] 2. The power consumption distribution map acquisition module in the present invention fits the mapping relationship between current and heat generation in the AI server based on an adaptive non-uniform grid, generating an accurate power consumption distribution map, enabling the heat dissipation design to more specifically match the heat generation conditions of each area; meanwhile, the embedded heat dissipation module generation module constructs a multi-heat source partition heat pipe array based on the global heat network, and adopts gradient brazing and multi-layer hermetic packaging technologies to enhance the heat dissipation capacity; in addition, the AI server temperature control module performs short-time domain rolling prediction based on a three-dimensional coupling model, and takes corresponding measures according to different overheating risk levels, such as increasing the rotation speed of the condenser fan, activating the semiconductor refrigeration chip, performing dynamic voltage and frequency scaling, etc., to achieve precise and efficient heat dissipation, effectively solving the multi-heat source heat dissipation problem of the AI server and comprehensively improving the heat dissipation effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 FIG. is a system architecture diagram of a multi-heat source partition heat pipe array packaging system applied to an AI server provided by an embodiment of the present invention; Figure 2 FIG. is a flowchart of a method for packaging a multi-heat source partition heat pipe array applied to an AI server provided by an embodiment of the present invention.
[0019] The realization, functional features and advantages of the objectives of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0021] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms of "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. "Plurality" generally includes at least two.
[0022] Depending on the context, the words "if" or "when" as used herein may be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detecting (stated condition or event)" may be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".
[0023] In addition, the step timings in the following method embodiments are only examples, not strictly limited.
[0024] In fact, the server devices deployed in the multi-heat-source partition heat pipe array packaging system applied to the AI server may be composed of one or more devices. The multi-heat-source partition heat pipe array packaging system applied to the AI server described above can be implemented as: a service instance, a virtual machine, or a hardware device. For example, the multi-heat-source partition heat pipe array packaging system applied to the AI server can be implemented as a service instance deployed on one or more devices in a cloud node. Simply put, the multi-heat-source partition heat pipe array packaging system applied to the AI server can be understood as a software deployed on a cloud node, used to provide the multi-heat-source partition heat pipe array packaging system applied to the AI server for each client. Alternatively, the multi-heat-source partition heat pipe array packaging system applied to the AI server can also be implemented as a virtual machine deployed on one or more devices in a cloud node. An application software for managing each client is installed in the virtual machine. Or, the multi-heat-source partition heat pipe array packaging system applied to the AI server can also be implemented as a server composed of many identical or different types of hardware devices, with one or more hardware devices set to provide the multi-heat-source partition heat pipe array packaging system applied to the AI server for each client.
[0025] In terms of implementation form, the multi-heat-source partition heat pipe array packaging system applied to the AI server and the client adapt to each other. That is, if the multi-heat-source partition heat pipe array packaging system applied to the AI server is an application installed on a cloud service platform, then the client is a client that establishes a communication connection with this application; or if the multi-heat-source partition heat pipe array packaging system applied to the AI server is implemented as a website, then the client is implemented as a web page; or if the multi-heat-source partition heat pipe array packaging system applied to the AI server is implemented as a cloud service platform, then the client is implemented as a small program in an instant messaging application.
[0026] As Figure 1 shown, it is the system architecture diagram of the multi-heat-source partition heat pipe array packaging system applied to the AI server provided by an embodiment of the present invention.
[0027] The multi-heat-source partition heat pipe array packaging system 100 applied to an AI server according to the present invention can be set in a cloud server. In terms of implementation form, it can be used as one or more service devices, or can be used as an application installed on the cloud (such as the server of a mobile service operator, a server cluster, etc.), or can also be developed into a website. According to the implemented functions, the multi-heat-source partition heat pipe array packaging system 100 applied to an AI server can include a real-time temperature data acquisition module 101, a power consumption distribution map acquisition module 102, a global heat network acquisition module 103, a three-dimensional coupling model construction module 104, an embedded heat dissipation module generation module 105, and an AI server temperature control module 106. The modules in the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by an electronic device processor and can complete fixed functions, and are stored in the memory of the electronic device.
[0028] In an embodiment of the present invention, in the multi-heat-source partition heat pipe array packaging system applied to an AI server, each of the above modules can be independently implemented and called with other modules. Here, the call can be understood as that a certain module can be connected to multiple modules of another type and provide corresponding services for the multiple modules it is connected to. In the multi-heat-source partition heat pipe array packaging system applied to an AI server provided by the embodiment of the present invention, without modifying the program code, the applicable range of the architecture of the multi-heat-source partition heat pipe array packaging system applied to an AI server can be adjusted by adding modules and directly calling, so as to achieve cluster-level horizontal expansion, so as to achieve the purpose of quickly and flexibly expanding the multi-heat-source partition heat pipe array packaging system applied to an AI server. In practical applications, the above modules can be set in the same device or different devices, or can also be set in virtual devices, such as service instances in a cloud server.
[0029] The following will combine specific embodiments to separately illustrate the respective components and specific working processes of the multi-heat-source partition heat pipe array packaging system applied to an AI server: The real-time temperature data acquisition module 101 is used to dynamically configure the sensor acquisition period for an AI server embedded with a distributed thermocouple array, obtain the dynamic temperature sensing network of the AI server, and acquire the real-time temperature data of the AI server based on the dynamic temperature sensing network; In an embodiment of the present invention, when the real-time temperature data acquisition module executes the operation of dynamically configuring the sensor acquisition period for an AI server embedded with a distributed thermocouple array to obtain the dynamic temperature sensing network of the AI server, it specifically is used for: Embedding miniaturized K-type thermocouples arranged in a hexagonal close-packed pattern on the packaging substrate of the AI server to obtain the embedded distributed thermocouple array of the AI server; Dynamically configure the acquisition period for the embedded distributed thermocouple array to obtain the 3D temperature field model of the AI server; Overlay the 3D temperature field model onto the digital twin of the physical substrate to obtain the dynamic temperature sensing network of the AI server.
[0030] Specifically, the packaging substrate is a core component in the field of electronic packaging, specifically referring to a support platform that bears chips and integrates an interconnect structure. The packaging substrate is usually made of high thermal conductivity ceramic composite materials (such as aluminum nitride) and metal lamination structures (such as copper - polyimide).
[0031] Specifically, the miniaturized K - type thermocouple is a super - small temperature sensor designed for high - density electronic packaging. It belongs to the K - type (nickel - chromium - nickel - aluminum) thermocouple family, and the volume of the traditional thermocouple is compressed to the sub - millimeter level through micro - manufacturing technology.
[0032] Specifically, the physical substrate is the core thermal management carrier of the AI server packaging system. It uses a copper - ceramic composite substrate, with a hexagonal close - packed heat pipe array embedded, and realizes global heat conduction through multi - layer metallized channels.
[0033] Specifically, design a hexagonal close - packed layout on the packaging substrate of the AI server to ensure uniform distribution of thermocouples and coverage of key heat - generating areas.
[0034] Furthermore, select miniaturized K - type thermocouples, embed them on the substrate surface or in the interlayer through a high - precision mounting process, and optimize the wiring to avoid signal interference.
[0035] Furthermore, the hexagonal arrangement of thermocouples can maximize the space utilization rate and improve the reliability of temperature data through redundant measurements of adjacent nodes.
[0036] Furthermore, conduct electrical tests after completion to ensure stable signal transmission for each thermocouple.
[0037] Furthermore, through an embedded controller or FPGA, monitor the temperature change rate of each area in real - time and dynamically adjust the acquisition frequency of thermocouples: high - temperature or large - fluctuation areas use high - frequency sampling, while low - temperature and stable areas are reduced to 1Hz to save resources.
[0038] Furthermore, use machine learning algorithms to identify temperature change patterns and automatically optimize the period configuration strategy.
[0039] Furthermore, after data is synchronized by timestamp, combined with the spatial coordinates of thermocouples, construct a spatio - temporally continuous original temperature data set.
[0040] Furthermore, based on the physical position coordinates and real - time data of thermocouples, use interpolation algorithms to convert discrete points into a continuous 3D temperature field.
[0041] Furthermore, the finite element analysis is used to simulate the heat conduction process and correct the error caused by insufficient thermocouple density.
[0042] Furthermore, the model output is a voxelized grid, and each voxel contains a temperature value and a confidence parameter, and can be updated in real time.
[0043] Furthermore, a dynamic temperature cloud map is rendered through a visualization tool, supporting slice viewing and hot spot tracking.
[0044] Furthermore, the CAD model of the physical substrate is imported into the digital twin platform for coordinate alignment and mesh matching with the 3D temperature field model.
[0045] Furthermore, real-time data stream transmission is achieved through the API interface, enabling the digital twin to synchronously display the temperature field changes.
[0046] Furthermore, on this basis, an alarm module and a prediction function are integrated, and finally a closed-loop dynamic temperature sensing network is formed, supporting remote monitoring and active thermal management strategy optimization.
[0047] Generally speaking, miniaturized K-type thermocouples arranged in a hexagonal close-packed pattern are embedded in the packaging substrate of the AI server to form an embedded distributed thermocouple array. This layout can sense the temperature inside the server more comprehensively compared to traditional sensors with a small number of fixed positions.
[0048] Generally speaking, by dynamically configuring the sensor acquisition period, a 3D temperature field model of the AI server can be obtained. This model can accurately reflect the temperature distribution of the server at different times and positions, providing accurate data support for subsequent heat dissipation measures.
[0049] Generally speaking, by superimposing the obtained 3D temperature field model on the digital twin of the physical substrate, a dynamic temperature sensing network of the AI server is obtained. This network is equivalent to constructing a digital twin thermal mirror of the server's temperature distribution, enabling operators to intuitively grasp the internal temperature changes of the server in real time.
[0050] Generally speaking, through the analysis of the digital twin thermal mirror, the global thermal network domain can be accurately located, that is, the temperature-sensitive and insensitive areas in the server can be determined, providing a reliable basis for subsequent targeted design and implementation of heat dissipation solutions, and greatly improving the accuracy of heat dissipation.
[0051] Generally speaking, accurate and comprehensive temperature data and the digital twin thermal mirror can help the system better understand the internal heat generation situation of the server.
[0052] For example, when the power consumption distribution map acquisition module, the embedded heat dissipation module generation module, and the AI server temperature control module are working, more targeted heat dissipation strategies can be designed based on this data.
[0053] Generally speaking, in the power consumption distribution map acquisition module, fitting the mapping relationship between current and heat generation amount based on accurate temperature data will be more accurate; the embedded heat dissipation module generation module can reasonably arrange heat pipes according to temperature-sensitive areas; the AI server temperature control module can adjust the heat dissipation method in a timely manner according to temperature changes, thereby improving the overall heat dissipation effect and solving the problem of multi-heat source heat dissipation of AI servers.
[0054] The power consumption distribution map acquisition module 102 is used to create an adaptive non-uniform grid of the AI server based on pre-acquired real-time power consumption data, fit the mapping relationship between current and heat generation amount in the AI server based on the adaptive non-uniform grid, and generate the power consumption distribution map of the AI server based on the mapping relationship and the dynamic reduction algorithm; In the embodiment of the present invention, when the power consumption distribution map acquisition module executes creating the adaptive non-uniform grid of the AI server based on pre-acquired real-time power consumption data, it specifically is used for: Reading the real-time power consumption data of the AI server through the SVID protocol; Calculating the global interpolation field of the AI server based on the real-time power consumption data, where the calculation of the global interpolation field is as follows: In the formula, is the adaptive non-uniform grid, is the three-dimensional coordinate point of the chip of the AI server, is the total number of interpolations of the full-field interpolation field, is the interpolation number of the full-field interpolation field, is the weight of the interpolation of the full-field interpolation field affected by temperature, is the Gaussian kernel attenuation function, is the real-time power consumption influence factor in the real-time power consumption data, is the spatial distance between two interpolations of the full-field interpolation field, is the three-dimensional coordinate point where the embedded distributed thermocouple array has been deployed, is the attenuation factor of heat transfer; Based on the global interpolation field, performing grid division on the AI server to obtain the adaptive non-uniform grid of the AI server.
[0055] When the power consumption distribution map acquisition module executes fitting the mapping relationship between current and heat generation amount in the AI server based on the adaptive non-uniform grid, and generating the power consumption distribution map of the AI server based on the mapping relationship and the dynamic reduction algorithm, it specifically is used for: Extract the busbar layer current density distribution of the AI server based on the adaptive non-uniform grid; Generate the Joule heat increment field of the AI server based on the adaptive non-uniform grid; Perform electromagnetic-thermal coupling on the busbar layer current density distribution and the Joule heat increment field to obtain the electromagnetic-thermal-force correlation of the AI server; Generate the mapping relationship between current and heat generation in the AI server based on the electromagnetic-thermal-force correlation; Construct a 2000-order state space model of the AI server based on the mapping relationship; Based on the dynamic order reduction algorithm, perform hierarchical feature truncation and order reduction on the 2000-order state space model to obtain the hierarchical feature vector of the AI server, where the dynamic order reduction algorithm is as follows: In the formula, is the hierarchical feature vector obtained after order reduction and compression of the 2000-order state space model, is the 2000-order state space model, is the inverted heat conduction matrix, is the heat conduction matrix, is the matrix inversion flag, is the hierarchical feature vector, is the reconstructed temperature field after reducing errors, is the balance parameter of thermal physical constraints and statistical characteristics, is the potential feature distribution of the temperature field in the 2000-order state space model, is the regularization constraint on the potential features of the temperature field in the 2000-order state space model, is that the global KL constraint ignores the importance of local hotspots, is the potential feature of the temperature field, is the original temperature field in the 2000-order state space model, is the reconstructed temperature field in the 2000-order state space model.
[0056] Perform power consumption feature fusion on the hierarchical feature vector to obtain the power consumption distribution map of the AI server.
[0057] Specifically, the SVID protocol is a serial voltage identification protocol, a serial protocol for communication between x86 architecture CPUs and voltage regulation modules (VRMs), which transmits power management information through three signal lines and is used to accurately regulate CPU voltage and power status.
[0058] Specifically, the temperature field in the 2000 - order state - space model is a mathematical abstraction of the temperature distribution within the AI server, covering temperature information in each region, used to analyze the temperature change pattern, assist in constructing the power consumption distribution map, and achieve precise heat dissipation control.
[0059] Specifically, to establish a communication connection with the server, it is necessary to follow the specifications of the SVID protocol, clarify the server's address and port number, and use an appropriate network programming library (such as Python's socket library) to construct a communication socket.
[0060] Furthermore, construct and send a request message according to the protocol requirements to obtain real - time power consumption data. During the communication process, pay attention to the data format and encoding to ensure that the received data can be correctly parsed. Further, to enhance the communication stability, it is also necessary to handle possible network exceptions, such as connection timeouts, data loss, etc.
[0061] Furthermore, after successfully receiving the real - time power consumption data returned by the AI server, it is necessary to parse it.
[0062] Furthermore, convert the received raw data into a numerical form that can be used for subsequent calculations according to the data format specified by the SVID protocol.
[0063] Furthermore, after parsing, store these real - time power consumption data into a suitable data structure, such as a Python list or a NumPy array.
[0064] Furthermore, this makes it convenient to access and process the data subsequently.
[0065] Furthermore, for the convenience of subsequent analysis, corresponding timestamps or server identification information can be added to each power consumption data.
[0066] Furthermore, based on the stored real - time power consumption data, start calculating the global interpolation field of the AI server.
[0067] Furthermore, then, use a suitable mathematical library to implement the selected interpolation algorithm.
[0068] Furthermore, during the calculation process, consider the physical structure and layout of the server, and associate the power consumption data with the spatial location of the server.
[0069] Furthermore, the obtained global interpolation field can more accurately reflect the overall power consumption distribution of the server.
[0070] Furthermore, after obtaining the global interpolation field, perform grid division on the AI server according to this field.
[0071] Furthermore, the purpose of grid division is to divide the physical space of the server into multiple small regions for a more refined analysis of power consumption distribution.
[0072] Furthermore, when dividing, the characteristics of the global interpolation field should be considered so that the grid can adaptively reflect the changes in power consumption.
[0073] Furthermore, the grid density can be adjusted according to the gradient or rate of change of power consumption. A finer grid is used in areas with a larger change in power consumption, while a sparser grid is used in areas with a smaller change in power consumption.
[0074] Furthermore, in this way, an adaptive non-uniform grid of the AI server is finally obtained.
[0075] Specifically, in the calculation of the global interpolation field, data measured by distributed thermocouples deployed on the AI server chip are collected. These data correspond to the situation at specific positions on the chip, and at the same time, the total number of these known measurement points is determined.
[0076] Furthermore, two important values need to be determined. One is the influence degree value of each known measurement point on the final calculation result, and the other is the value related to the real-time power consumption that affects the attenuation rate of heat transfer. These two values need to be determined based on the actual physical model or past experience.
[0077] Furthermore, for the target position where the estimated value is to be calculated, calculate the spatial distance between it and each known measurement point.
[0078] Furthermore, based on the value related to the real-time power consumption determined previously and the calculated spatial distance, with the help of a function that can reflect the attenuation characteristic of heat transfer with the increase of distance, calculate the attenuation degree of the influence of each known measurement point on the target position.
[0079] Furthermore, multiply the influence degree value of each known measurement point by the corresponding attenuation degree of influence, and then add up all the multiplied results to obtain the estimated value of the target position, that is, the value of the global interpolation field at this position.
[0080] Furthermore, from the perspective of distance, the farther the distance between the target position and the known measurement point, the smaller the influence of this known measurement point on the estimated value of the target position.
[0081] Furthermore, this is because as the distance increases, the function that reflects the attenuation characteristic of heat transfer will cause the influence of the known measurement point to decrease rapidly.
[0082] Furthermore, the value related to the real-time power consumption also has an obvious influence on the result.
[0083] Furthermore, if this value is large, then as the distance increases, the influence of known measurement points on the target position will rapidly decrease, and the effect of distant measurement points on the estimated target position will quickly become negligible; conversely, if this value is small, the influence of distant measurement points on the target position will decrease more slowly, and relatively speaking, their effect on the estimated target position will also be greater.
[0084] Furthermore, the main function of this formula is to use the known thermocouple measurement data distributed at different positions on the chip to estimate the physical quantities, such as temperature, at any position on the AI server chip.
[0085] Furthermore, by considering the spatial distance between data points and the attenuation of heat transfer with distance, the estimation results are made to better conform to the actual physical phenomena.
[0086] Furthermore, the global interpolation field of the entire AI server chip can be obtained, providing important basic data for subsequent operations such as grid division and thermal management of the chip.
[0087] Specifically, it is necessary to obtain the adaptive non-uniform grid data of the AI server obtained previously.
[0088] Furthermore, this grid data contains the relevant physical information of different positions of the server.
[0089] Furthermore, based on the principles of electromagnetism and the physical characteristics of the busbar layer, a suitable algorithm is used to analyze the grid data.
[0090] Furthermore, during the analysis process, factors such as the material properties, geometric shape, and current conduction law of the busbar layer are considered.
[0091] Furthermore, by calculating and evaluating each grid cell, the current density values at various positions of the busbar layer are extracted.
[0092] Furthermore, these current density values are sorted and recorded according to the positions of the grids, thereby obtaining the current density distribution of the busbar layer of the AI server.
[0093] Furthermore, based on the adaptive non-uniform grid, the Joule heat increment field of the AI server is generated according to Joule's law.
[0094] Furthermore, Joule's law states that the heat generated by an electric current passing through a conductor is proportional to the square of the current, the resistance of the conductor, and the time of power supply.
[0095] Furthermore, for each grid cell, the resistance of the cell is calculated by combining the previously extracted current density information.
[0096] Furthermore, the calculation of resistance needs to consider factors such as the material properties and dimensions of the grid cells. Then, the increment of Joule heat generated by the grid cell is calculated based on the current density and resistance.
[0097] Furthermore, the same calculation is performed on all grid cells, and the increments of Joule heat of each grid cell obtained are arranged according to their positions in the adaptive non-uniform grid, and finally the Joule heat increment field of the AI server is generated.
[0098] Furthermore, the previously obtained current density distribution of the busbar layer and the Joule heat increment field are combined to carry out electromagnetic-thermal coupling analysis.
[0099] Furthermore, this process needs to apply relevant theories and equations of electromagnetics and thermotics, considering factors such as the magnetic field generated by the current, the influence of the magnetic field on the current, and heat transfer and distribution.
[0100] Furthermore, by establishing a suitable mathematical model, such as a coupled model of Maxwell's equations and heat conduction equations, the interaction between the current density distribution and the Joule heat increment field is simulated and calculated.
[0101] Furthermore, during the calculation process, the electromagnetic and thermal properties of the material are considered to vary with temperature and magnetic field. Through analysis and calculation, the correlation between electromagnetic phenomena and thermal phenomena in the AI server is obtained, that is, electromagnetic-thermal correlation.
[0102] Furthermore, based on the previously obtained electromagnetic-thermal correlation, a corresponding mapping relationship is generated. This mapping relationship can describe the corresponding relationship between electromagnetic parameters and thermal parameters in the AI server.
[0103] Furthermore, methods such as function fitting and data statistical analysis can be used to process the electromagnetic-thermal correlation data.
[0104] Furthermore, by analyzing a large number of data points, a mathematical expression or data table that can accurately describe the relationship between the two is found.
[0105] Furthermore, this mapping relationship will provide an important basis for the subsequent construction of the state space model.
[0106] Furthermore, using the previously generated mapping relationship, a 2000-order state space model of the AI server is constructed.
[0107] Furthermore, the state space model is a mathematical model used to describe the dynamic characteristics of a system, which represents the relationship between the state variables, input variables and output variables of the system with a set of differential equations or difference equations.
[0108] Further, determine the state variables, input variables, and output variables of the AI server system. The state variables may include physical quantities such as current density and temperature, the input variables may be power input, etc., and the output variables may be the temperature or current at certain key positions, etc.
[0109] Further, establish the state equation and output equation according to the mapping relationship and the physical characteristics of the system.
[0110] Further, by reasonably selecting the state variables and parameters, construct a 2000-order state space model to accurately describe the dynamic characteristics of the AI server.
[0111] Further, in order to simplify the model and improve the calculation efficiency, adopt a dynamic order reduction algorithm to perform hierarchical feature truncation and order reduction on the 2000-order state space model.
[0112] Further, analyze the characteristics of the 2000-order state space model to determine which characteristics have a greater impact on the dynamic characteristics of the system and which have a smaller impact. Then, according to the analysis results, stratify the characteristics of the model.
[0113] Further, for the characteristic layer with a smaller impact, perform truncation processing, that is, ignore the impact of these characteristics on the system. In this way, reduce the order of the model and the complexity of the model.
[0114] Further, during the order reduction process, it is necessary to ensure that the reduced-order model can still better reflect the main dynamic characteristics of the original system.
[0115] Further, perform power consumption characteristic fusion on the hierarchical feature vectors obtained after hierarchical feature truncation and order reduction.
[0116] Further, power consumption characteristic fusion is to integrate various characteristic information related to power consumption to more comprehensively describe the power consumption distribution of the AI server.
[0117] Further, methods such as weighted average and principal component analysis can be used to process the hierarchical feature vectors. During the fusion process, consider the influence degree of different characteristics on power consumption and assign different weights.
[0118] Further, through the fusion operation, merge multiple hierarchical feature vectors into a comprehensive power consumption characteristic vector.
[0119] Further, according to this comprehensive power consumption characteristic vector, combined with the physical structure of the AI server and the adaptive non-uniform grid information, generate the power consumption distribution map of the AI server.
[0120] Further, this map can intuitively display the power consumption distribution of the AI server at different positions.
[0121] Specifically, in the calculation of the dynamic order reduction algorithm, to reduce the order of a 2000-order state space model, it is first necessary to obtain the heat conduction matrix and its inverse matrix.
[0122] Furthermore, the heat conduction matrix reflects the relevant characteristics of heat conduction in the model. Multiply the inverted heat conduction matrix by the 2000-order state space model, and then multiply by the heat conduction matrix again. Through such a matrix operation process, the reduced-order and compressed hierarchical feature vector can be obtained.
[0123] Furthermore, this process is actually to simplify the original model, removing some information with less impact on the whole, and only retaining the key features to make the model more concise and efficient.
[0124] Furthermore, first calculate the reconstructed temperature field after reducing the error, which reflects the difference between the reconstructed temperature field and the original temperature field. The smaller this difference, the better the reconstruction effect.
[0125] Furthermore, it is necessary to consider the balance between thermophysical constraints and statistical characteristics. There is a balance parameter here, which plays a role in adjusting the weights of the two.
[0126] Furthermore, it is also necessary to calculate the relationship between the potential feature distribution of the temperature field in the 2000-order state space model and the regularization constraint on this potential feature. This relationship is measured by the global KL constraint, but it should be noted that the global KL constraint may ignore the importance of local hot spots.
[0127] Furthermore, add the reconstructed temperature field after reducing the error to the result of multiplying the balance parameter by the global KL constraint. The obtained value is the hierarchical feature vector.
[0128] Furthermore, if the element values in the heat conduction matrix change significantly, it may mean that there are significant changes in the characteristics of heat conduction.
[0129] Furthermore, when such a change occurs, the reduced-order and compressed hierarchical feature vector will also change accordingly.
[0130] For example, if heat conduction is enhanced, the reduced-order model may highlight the key features related to heat conduction and weaken other less relevant features.
[0131] Furthermore, when the value of the reconstructed temperature field after reducing the error increases, it means that the difference between the reconstructed temperature field and the original temperature field is increasing, which may lead to an increase in the value of the hierarchical feature vector.
[0132] Furthermore, the change in the balance parameter will directly affect the weights of thermophysical constraints and statistical characteristics in the final result.
[0133] Furthermore, if the balance parameter increases, the influence of the relationship between thermophysical constraints and statistical characteristics on the hierarchical eigenvector will be more significant; conversely, if the balance parameter decreases, the influence of the reconstructed temperature field after reducing the error on the hierarchical eigenvector will be relatively greater.
[0134] Furthermore, the main function of these two formulas is to perform order reduction processing and feature extraction on the 2000-order state space model.
[0135] Furthermore, the order reduction compression formula can simplify the complex 2000-order model into a hierarchical eigenvector, reduce the complexity of the model, improve the calculation efficiency, and at the same time retain the key feature information of the model.
[0136] Furthermore, the formula for calculating the hierarchical eigenvector comprehensively considers the error of the reconstructed temperature field, the balance between thermophysical constraints and statistical characteristics, and the distribution of potential features of the temperature field, so that the obtained hierarchical eigenvector can more comprehensively and accurately reflect the characteristics of the model, providing a more effective basis for subsequent analysis and processing of the AI server.
[0137] Generally speaking, creating an adaptive non-uniform grid through real-time power consumption data can flexibly divide regions according to the actual power consumption distribution of the server, accurately reflecting the heat generation differences at different positions.
[0138] Generally speaking, compared with the traditional uniform grid, it is more suitable for the characteristics of multi-heat sources and uneven heat source distribution of the AI server, providing a basis for subsequent accurate analysis.
[0139] Generally speaking, based on the adaptive non-uniform grid to fit the mapping relationship between current and heat generation, the correlation between current and heat generation is clarified.
[0140] Generally speaking, this enables the heat dissipation design to closely focus on the actual heat generation situation of each region. For example, heat dissipation components are key-layout in regions with high heat generation, improving the utilization efficiency of heat dissipation resources and enhancing the overall heat dissipation capacity.
[0141] Generally speaking, the power consumption distribution map generated by using the mapping relationship and the dynamic order reduction algorithm visually presents the power consumption distribution of each part of the server.
[0142] Generally speaking, this provides a key reference for the embedded heat dissipation module generation module, enabling it to reasonably construct a multi-heat source partition heat pipe array; it also provides a decision-making basis for the temperature control module of the AI server. According to the power consumption and heat generation risks in different regions, corresponding temperature control strategies are adopted to achieve precise heat dissipation.
[0143] The global heat network acquisition module 103 is used to perform clustering analysis on the heat source distribution of the AI server to obtain the global heat network of the AI server; In the embodiments of the present invention, when the global heat network acquisition module performs clustering analysis on the heat source distribution of the AI server to obtain the global heat network of the AI server, it specifically is used for: Extract thermodynamic characteristics from the heat source distribution of the AI server to obtain the thermodynamic data of the heat source distribution; Perform clustering analysis on the thermodynamic data to obtain the high-temperature cluster data of the AI server; Perform grid division on the heat source distribution to obtain the grid of the heat source distribution; Mark the grid based on the high-temperature cluster data: If the high-temperature cluster data of the grid exceeds a preset high-temperature threshold, mark it as a temperature-sensitive grid; If the high-temperature cluster data of the grid is lower than a preset high-temperature threshold, mark it as a temperature-insensitive grid; Collect the temperature-sensitive grids and the temperature-insensitive grids as the global heat network of the AI server.
[0144] Specifically, collect the thermodynamic data of the AI server, which may include information such as the temperature and heat flux at different positions of the server.
[0145] Furthermore, select a suitable clustering algorithm, such as K-means clustering, hierarchical clustering, etc. The clustering algorithm will divide the thermodynamic data into different clusters according to the similarity between the data.
[0146] Furthermore, during the clustering process, the algorithm will automatically identify the clusters corresponding to the high-temperature regions in the data. After clustering analysis, select the cluster data representing the high-temperature regions from all the clusters. These data are the high-temperature cluster data of the AI server, which reflect the characteristics of the parts with higher temperature in the server.
[0147] Furthermore, in order to analyze the heat source distribution of the AI server in more detail, it is necessary to perform grid division on it.
[0148] Furthermore, determine the rules and accuracy of grid division, such as the size and shape of the grid. These parameters can be set according to the physical structure and actual requirements of the server.
[0149] Furthermore, divide the heat source distribution area of the AI server into small grid units according to the set rules.
[0150] Furthermore, each grid unit has its specific position and range. In this way, the continuous heat source distribution can be discretized, which is convenient for subsequent analysis and processing, and finally the grid of the heat source distribution is obtained.
[0151] Furthermore, using the high-temperature cluster data obtained previously, the grids of the heat source distribution are marked.
[0152] Furthermore, for each grid, it is compared and analyzed with the high-temperature cluster data. The relationship between the high-temperature cluster data of this grid and the preset high-temperature threshold is judged.
[0153] Furthermore, the preset high-temperature threshold is a temperature limit preset according to factors such as the performance requirements and safety standards of the AI server.
[0154] Furthermore, through comparison, it is determined whether this grid belongs to the temperature-sensitive or temperature-insensitive type.
[0155] Furthermore, when the high-temperature cluster data of a certain grid exceeds the preset high-temperature threshold, it indicates that the temperature of this grid area is relatively high, which may have a greater impact on the performance and stability of the server. Therefore, it is marked as a temperature-sensitive grid.
[0156] Furthermore, the marking method can be to add a specific identifier in the grid data for subsequent identification and processing.
[0157] Furthermore, doing so helps to focus on these temperature-sensitive areas and take corresponding heat dissipation or optimization measures.
[0158] Furthermore, if the high-temperature cluster data of a certain grid is lower than the preset high-temperature threshold, it indicates that the temperature of this grid area is relatively low and has a small impact on the server performance. It is marked as a temperature-insensitive grid.
[0159] Furthermore, corresponding identifiers are added in the grid data for distinction. Marking the temperature-insensitive grids can help us clarify which areas do not require special heat dissipation treatment, so as to reasonably allocate resources.
[0160] Furthermore, all the marked temperature-sensitive grids and temperature-insensitive grids are summarized and integrated.
[0161] Furthermore, taking these grids as basic units, a complete network structure is constructed, and this network is the global thermal network of the AI server.
[0162] Furthermore, the global thermal network comprehensively reflects the heat distribution of the AI server, including temperature-sensitive areas and temperature-insensitive areas.
[0163] Furthermore, by analyzing the global thermal network, a clearer understanding of the overall thermal state of the server can be obtained, providing a basis for subsequent thermal management and optimization.
[0164] Generally speaking, by performing thermodynamic feature extraction and clustering analysis on the heat source distribution, high-temperature cluster data can be found.
[0165] Generally speaking, based on this, temperature-sensitive grids and temperature-insensitive grids are marked out, and the areas prone to high-temperature problems in the AI server are accurately located, providing a basis for subsequent key heat dissipation.
[0166] For example, in actual operation, heat dissipation measures can be preferentially strengthened for these temperature-sensitive areas to avoid local overheating from affecting the performance and lifespan of the chips.
[0167] Generally speaking, the global thermal network provides key information for the embedded heat dissipation module generation module.
[0168] Generally speaking, by extracting heat source characteristics based on the global thermal network, heat pipes can be reasonably arranged according to the heat source characteristics of different regions to construct a multi-heat-source partitioned heat pipe array.
[0169] Generally speaking, increasing the number of heat pipes or adjusting the heat pipe layout in temperature-sensitive areas can improve the heat dissipation efficiency, making the heat dissipation module more in line with the actual heat dissipation requirements of the server and enhancing the overall heat dissipation capacity.
[0170] Generally speaking, based on the global thermal network and the three-dimensional coupling model, the temperature control module of the AI server can more accurately monitor and predict the temperatures of different regions.
[0171] Generally speaking, when different overheating risks are predicted, more targeted measures can be taken for temperature-sensitive areas to ensure the stable operation of the server.
[0172] The three-dimensional coupling model construction module 104 is used to construct the three-dimensional coupling model of the AI server based on the real-time temperature data, the power consumption distribution map, and the global thermal network; In the embodiment of the present invention, when the three-dimensional coupling model construction module executes constructing the three-dimensional coupling model of the AI server based on the real-time temperature data, the power consumption distribution map, and the global thermal network, it is specifically used for: Performing time series alignment processing on the power consumption distribution map and the real-time temperature data to obtain the standard data of the AI server; Constructing an unstructured tetrahedral grid model of the physical structure of the AI server based on the global thermal network to obtain the thermal gradient region model of the AI server; Allocating multi-physical-field coupling nodes to the high current density area in the power consumption distribution map to obtain the initial model of the AI server; Performing electro-thermal coupling analysis on the AI server based on the multi-physical-field coupling nodes of the initial model and the standard data to obtain the electro-thermal coupling model of the AI server; Constructing the three-dimensional coupling model of the AI server based on the electro-thermal coupling model and the thermal gradient region model.
[0173] When the three-dimensional coupling model building module builds the three-dimensional coupling model of the AI server based on the electrothermal coupling model and the thermal gradient region model, it is specifically used to: Defining a coupling interface in a thermal gradient key region in the thermal gradient region model to obtain an overlapping domain coupling interface of the thermal gradient region model; Extracting the global temperature residual when the thermal gradient regional model is iterated; Extracting the heat flow residual when the electrothermal coupling model is iterated; When the global temperature residual is less than a thermal gradient threshold and the heat flow residual is less than an electrothermal threshold, a converged thermal gradient regional model and a converged electrothermal coupling model are obtained; Based on the overlapping domain coupling interface, the converged thermal gradient region model and the converged electrothermal coupling model are thermally-electrically-mechanically coupled across scales to obtain a three-dimensional coupling model of the AI server.
[0174] Specifically, the global thermal network of the AI server obtained previously is obtained, which shows in detail the distribution of temperature-sensitive and insensitive areas within the server.
[0175] Furthermore, based on the information of the global thermal network, the physical structure of the AI server is analyzed in detail to determine the shape, position and relationship of each component.
[0176] Furthermore, an unstructured tetrahedral meshing method is used to decompose the physical space of the AI server into numerous irregular tetrahedral mesh units.
[0177] Furthermore, in the division process, the heat distribution characteristics reflected by the global thermal network are fully considered. In areas with drastic temperature changes, that is, places with large thermal gradients, finer grid divisions are used to accurately capture thermal changes; while in areas with relatively stable temperatures, sparser grid divisions are used to improve calculation efficiency.
[0178] Furthermore, an unstructured tetrahedral mesh model is constructed in this way, which can clearly show the thermal gradient distribution within the AI server, thereby obtaining a thermal gradient regional model of the AI server.
[0179] Furthermore, we first obtain the power consumption distribution map of the AI server, which intuitively shows the current density distribution at different locations in the server.
[0180] Furthermore, high current density areas are accurately identified from the map. These areas are usually key areas with high heat generation and significant impact on server performance.
[0181] Furthermore, multiphysics coupling nodes are assigned in these high current density regions.
[0182] Furthermore, the multi-physical field coupling nodes can comprehensively consider the interactions between multiple physical fields such as the electric field and the thermal field.
[0183] Furthermore, according to the size, shape, and distribution characteristics of the high current density region, reasonably determine the number and positions of the nodes to ensure that the nodes can effectively reflect the changes in the physical fields within this region.
[0184] Furthermore, by allocating multi-physical field coupling nodes to the high current density region and combining these nodes with the physical structure of the server, the initial model of the AI server is obtained.
[0185] Furthermore, obtain the initial model of the AI server and the relevant standard data, where the standard data includes information such as the electrical and thermal property parameters of the material, boundary conditions, etc.
[0186] Furthermore, use professional multi-physical field analysis software or numerical calculation methods to perform electro-thermal coupling analysis on the AI server based on the multi-physical field coupling nodes and the standard data in the initial model.
[0187] Furthermore, during the analysis process, fully consider the mutual influence between the electric field and the thermal field. For example, current passing through a conductor generates heat, and the change in temperature will in turn affect the resistance and current distribution of the conductor.
[0188] Furthermore, by solving a series of electrical and thermal equations, simulate the electro-thermal behavior of the AI server during actual operation.
[0189] Furthermore, according to the analysis results, optimize and adjust the initial model to finally obtain an electro-thermal coupling model that can accurately reflect the electro-thermal interaction of the AI server.
[0190] Furthermore, this model can be used to predict the performance of the server, evaluate the effectiveness of the heat dissipation scheme, etc.
[0191] Specifically, start the iterative calculation process of the thermal gradient region model.
[0192] Furthermore, in each iteration, update the temperature values of each grid node in the model according to the physical rules and algorithms inside the model.
[0193] Furthermore, after the update is completed, compare the temperature values of each node obtained in this iteration with the corresponding node temperature values in the previous iteration.
[0194] Furthermore, for each grid node, calculate the difference between the temperature values of the two iterations.
[0195] Further, by using appropriate mathematical methods, such as taking the square root of the sum of the squares of the temperature differences of all nodes or calculating the sum of the absolute values of the temperature differences of all nodes, etc., these differences are aggregated to obtain a single value that can represent the degree of temperature change difference in the entire thermal gradient region model. This value is the global temperature residual.
[0196] Further, iterative calculations are performed on the electro-thermal coupling model. During the iteration process, according to the electrical and thermal principles followed by the model, the heat flux values of each part in the model are updated.
[0197] Further, the heat flux value reflects the flow of heat within the model.
[0198] Further, the heat flux distribution obtained from the current iteration is carefully compared with the heat flux distribution of the previous iteration, and for each position in the model, the difference in the heat flux value is calculated.
[0199] Further, similar to the treatment of the global temperature residual, appropriate mathematical means are used to integrate the heat flux differences at all positions, and finally a value that can reflect the degree of heat flux change difference in the entire electro-thermal coupling model is obtained, namely the heat flux residual.
[0200] Further, in advance, according to the performance indicators, physical characteristics, and calculation accuracy requirements of the AI server, the thermal gradient threshold and the electro-thermal threshold are set.
[0201] Further, after each iteration calculation of the thermal gradient region model and the electro-thermal coupling model is completed, and the global temperature residual and the heat flux residual are respectively extracted, the global temperature residual is compared with the thermal gradient threshold, and at the same time the heat flux residual is compared with the electro-thermal threshold.
[0202] Further, when the global temperature residual is less than the thermal gradient threshold and the heat flux residual is less than the electro-thermal threshold, this indicates that the thermal gradient region model and the electro-thermal coupling model have reached sufficient convergence accuracy during the iteration process, and at this time the converged thermal gradient region model and the converged electro-thermal coupling model are obtained.
[0203] Further, after obtaining the converged thermal gradient region model and the converged electro-thermal coupling model, an overlapping domain coupling interface is introduced.
[0204] Further, the overlapping domain coupling interface can achieve effective connection and information transfer between different scales and different physical fields.
[0205] Further, the converged thermal gradient region model and the converged electro-thermal coupling model are analyzed to determine the parts and overlapping regions that need to be coupled between them.
[0206] Furthermore, through the overlapping domain coupling interface, the information of the thermal field, electric field, and force field is fused and interacted. During the coupling process, the interactions and influences between different physical fields are fully considered, such as the influence of the electric field on the thermal field, the influence of the thermal field on the force field, etc.
[0207] Furthermore, after coupling, the numerical values of the thermal, electrical, and force physical quantities are obtained and integrated according to the three-dimensional physical structure of the AI server.
[0208] Furthermore, using three-dimensional modeling software or programming tools, with these physical quantities as parameters, a three-dimensional model that can intuitively present the thermal, electrical, and force characteristics of the AI server is constructed.
[0209] Furthermore, in the model, through different visualization methods such as colors, textures, or isosurfaces, the distribution of each physical quantity in space is displayed, and finally a three-dimensional coupling model that comprehensively reflects the thermal, electrical, and force characteristics of the AI server is formed.
[0210] Generally speaking, the real-time temperature data reflects the current actual temperature state of the server, the power consumption distribution map shows the heat generation situation in each area, and the global thermal network clarifies the heat source distribution characteristics.
[0211] Generally speaking, the three-dimensional coupling model integrates these key factors to achieve a comprehensive analysis of the multi-physical fields of heat-electricity-force of the AI server, comprehensively master the complex physical processes inside the server, and avoid the limitations of single-factor analysis.
[0212] Generally speaking, through this model, the heat transfer and temperature change laws in different regions can be clearly understood, providing accurate guidance for the generation module of the embedded heat dissipation module.
[0213] For example, according to the model results, the heat pipe layout is reasonably adjusted and the heat dissipation material configuration is optimized in the area with a large thermal gradient, so that the multi-source partition heat pipe array is more suitable for the actual heat dissipation requirements of the server, improving the heat dissipation efficiency.
[0214] The embedded heat dissipation module generation module 105 is used to construct the multi-source partition heat pipe array of the AI server based on the global thermal network, and perform interface encapsulation on the chip of the AI server and the multi-source partition heat pipe array to obtain the embedded heat dissipation module of the AI server; In the embodiment of the present invention, when the embedded heat dissipation module generation module executes constructing the multi-source partition heat pipe array of the AI server based on the global thermal network, performing interface encapsulation on the chip of the AI server and the multi-source partition heat pipe array to obtain the embedded heat dissipation module of the AI server, it is specifically used for: Performing heat source feature extraction on the global thermal network to obtain the heat source features of the AI server; Based on the heat source characteristics, arrange heat pipes for the AI server to obtain a multi-heat-source partition heat pipe array of the AI server; Perform gradient brazing on the multi-heat-source partition heat pipe array and the chips of the AI server to obtain an initial heat dissipation module of the AI server; Perform multi-layer hermetic packaging on the initial heat dissipation module to obtain an embedded heat dissipation module of the AI server.
[0215] Specifically, use data analysis and processing technologies to deeply mine the global heat network. By analyzing information such as the temperature change trend, temperature distribution, and heat flow direction in each area of the heat network, identify the heat source positions within the server.
[0216] Furthermore, extract the characteristics of these heat sources, such as the heat generation power, heat generation stability, and heat generation frequency of the heat sources.
[0217] Furthermore, organize and summarize the extracted heat source position and characteristic information to obtain the heat source characteristics of the AI server. These characteristics are crucial for subsequent heat dissipation design and can provide an accurate basis for heat pipe arrangement.
[0218] Furthermore, based on the heat source characteristics of the AI server obtained in the previous step, start the heat pipe arrangement design.
[0219] Furthermore, divide the server into multiple different areas according to the positions and heat generation characteristics of the heat sources, with each area corresponding to one or more heat sources.
[0220] Furthermore, for each partition, select heat pipes with appropriate specifications and quantities. The selection of heat pipes should consider parameters such as their heat transfer capacity, length, and diameter to ensure that the heat generated by the heat sources within the partition can be effectively transferred out.
[0221] Furthermore, install the heat pipes in the corresponding partitions according to the designed layout, making the heat pipes in close contact with the heat sources to form a multi-heat-source partition heat pipe array.
[0222] Furthermore, such an arrangement method can perform precise heat dissipation for the heat source characteristics of different partitions and improve the heat dissipation efficiency.
[0223] Furthermore, after obtaining the multi-heat-source partition heat pipe array, connect it to the chips of the AI server.
[0224] Furthermore, adopt the method of gradient brazing. First, clean and pre-treat the connection parts of the chips and heat pipes to remove surface impurities and oxides to ensure good welding effects.
[0225] Furthermore, according to the material properties of the chip and the heat pipe, appropriate brazing materials and process parameters are selected. During the welding process, the gradient heating method is used to gradually increase the temperature of the welding area to avoid damage to the chip or heat pipe due to too rapid temperature changes.
[0226] Furthermore, by precisely controlling the brazing process, the multi-source partitioned heat pipe array is firmly connected to the chip to form an integrated heat dissipation structure, that is, the initial heat dissipation module of the AI server.
[0227] Furthermore, in order to protect the initial heat dissipation module and improve its stability and reliability, it is necessary to perform multi-layer hermetic packaging on it.
[0228] Furthermore, appropriate packaging materials are selected, such as metal shells, ceramic materials, etc. These materials should have good heat insulation, moisture-proof and protection properties.
[0229] Furthermore, the initial heat dissipation module is placed inside the packaging structure, and the multi-layer packaging method is adopted, that is, a sealing layer is set between different layers to enhance the sealing effect.
[0230] Furthermore, during the packaging process, it is necessary to ensure that a relatively stable gas environment is formed inside to avoid the entry of impurities such as external air, moisture and dust.
[0231] Furthermore, through precise packaging technology, the initial heat dissipation module is completely sealed inside the packaging structure, and finally the embedded heat dissipation module of the AI server is obtained. This module can better adapt to the working environment of the server and improve the heat dissipation performance.
[0232] Generally speaking, the global heat network accurately reflects the heat source distribution of the AI server. Based on this, constructing a multi-source partitioned heat pipe array can make the heat pipe layout highly adaptable to the heat source distribution.
[0233] Generally speaking, in the temperature-sensitive area, the heat pipes can be densely arranged to ensure that the heat can be quickly conducted and dissipated, greatly improving the heat dissipation efficiency, avoiding the occurrence of local overheating, and ensuring the stable operation of the chip in a suitable temperature environment.
[0234] Generally speaking, by performing interface packaging on the chip of the AI server and the multi-source partitioned heat pipe array and adopting the gradient brazing technology, a good heat conduction connection can be formed between the chip and the heat pipe, reducing the thermal resistance and improving the heat transfer efficiency.
[0235] Generally speaking, the multi-layer hermetic packaging further protects the internal structure, prevents external factors such as dust and moisture from affecting the heat dissipation performance, and enhances the stability and reliability of the entire embedded heat dissipation module.
[0236] The AI server temperature control module 106 is used to perform temperature control of the AI server on the embedded heat dissipation module based on the three-dimensional coupling model.
[0237] In an embodiment of the present invention, when the AI server temperature control module performs temperature control of the AI server on the embedded heat dissipation module based on the three-dimensional coupling model, it is specifically used for: Performing short-time domain rolling prediction on the AI server based on the three-dimensional coupling model; When the prediction result is that there is no overheating risk for the AI server, the embedded heat dissipation module maintains the current working state; When the prediction result is that there is a slight overheating risk for the AI server, the embedded heat dissipation module increases the rotational speed of the condensation end fan of the partition heat pipe; When the prediction result is that there is a medium overheating risk for the AI server, the embedded heat dissipation module activates the thermoelectric cooler to perform local point cooling on the AI server; When the prediction result is that there is a severe overheating risk for the AI server, the embedded heat dissipation module starts dynamic voltage and frequency adjustment, increases the power of the thermoelectric cooler, and increases the rotational speed of the condensation end fan of the partition heat pipe to perform global cooling on the AI server.
[0238] Specifically, set the time step and prediction duration of the short-time domain. The time step needs to be reasonably selected according to the change frequency of the server operation data to ensure that the dynamic changes of the server physical state can be captured in a timely manner.
[0239] Furthermore, taking the operation parameters of the current server as the initial conditions, based on the interaction relationship between heat, electricity, and force in the three-dimensional coupling model, using numerical calculation methods, calculate physical quantities such as temperature, heat flux, and stress of the server in the future short-time domain step by step in time.
[0240] Furthermore, after each calculation is completed, take the latest calculation result as the initial condition for the next prediction, and continuously roll and update the prediction process, so as to realize continuous prediction of the physical state change of the AI server in the short-time domain.
[0241] Furthermore, obtain the result of the short-time domain rolling prediction, and extract key indicators related to temperature in the prediction data, such as the surface temperature of each chip, the temperature of the hot spot area, etc.
[0242] Furthermore, compare these temperature indicators with the pre-set safety temperature threshold. If all temperature indicators are lower than the safety threshold, it indicates that there is no overheating risk for the AI server during the prediction period. At this time, the embedded heat dissipation module does not need to make any adjustments, maintains the current working state, and maintains the normal heat dissipation power and operation mode to ensure the stable operation of the server while reducing energy consumption.
[0243] Furthermore, if the prediction result shows that the temperature index of some areas exceeds the safety threshold but does not reach the severe level, it is determined that the AI server has a mild overheating risk.
[0244] Furthermore, in response to this situation, control the condensation end fan of the partition heat pipe in the embedded heat dissipation module.
[0245] Furthermore, by adjusting the control circuit of the fan, increase the fan speed so that the fan can generate a greater air volume and accelerate the heat dissipation at the condensation end of the heat pipe.
[0246] Furthermore, since the mildly overheated area is relatively small, by only increasing the fan speed and enhancing the air convection heat dissipation method, the server temperature can be effectively reduced and brought back to the safe range.
[0247] Furthermore, when the prediction result indicates that the overheating situation of the server is relatively obvious and reaches the moderate overheating risk level, that is, there are multiple areas with excessive temperatures and the temperatures of some key chips are close to or exceed the danger threshold, start the thermoelectric cooler in the embedded heat dissipation module.
[0248] Furthermore, according to the distribution of the overheated areas, accurately control the starting position and working intensity of the thermoelectric cooler to perform targeted refrigeration on local chips or areas with too high temperatures.
[0249] Furthermore, through the Peltier effect, the thermoelectric cooler can quickly absorb the surrounding heat after being powered on, achieve local point refrigeration, effectively reduce the temperature of key parts, and prevent the overheating problem from deteriorating further.
[0250] Furthermore, if the prediction result shows that the AI server faces a severe overheating risk, that is, the temperatures of a large number of areas far exceed the safety threshold and may cause irreversible damage to the server hardware, take multiple heat dissipation measures to work together.
[0251] Furthermore, start dynamic voltage and frequency adjustment through the server's power management system to reduce the operating voltage and frequency of the server, thereby reducing the power consumption and heat generation of the server.
[0252] Furthermore, significantly increase the working power of the thermoelectric cooler to enhance its refrigeration capacity and perform more powerful refrigeration on the entire server.
[0253] Furthermore, further increase the speed of the condensation end fan of the partition heat pipe to the maximum to give full play to the heat dissipation efficiency of the heat pipe and air convection heat dissipation.
[0254] Furthermore, through this series of combined operations, perform global refrigeration on the AI server, quickly reduce the overall temperature of the server, and ensure the safe operation of the server hardware.
[0255] Generally speaking, the three-dimensional coupling model integrates real-time temperature data, power consumption distribution maps, and global heat network information, comprehensively reflecting the complex relationships among heat, electricity, and power within the AI server.
[0256] Generally speaking, through short-time domain rolling prediction of the AI server using this model, it is possible to accurately judge the temperature change trends in different regions of the server, know in advance whether there is an overheating risk and the degree of overheating, providing a reliable basis for temperature control.
[0257] Generally speaking, based on the prediction results, the system can flexibly adjust the embedded heat dissipation module for different overheating risk levels.
[0258] Generally speaking, when there is no overheating risk, maintain the current state of the module to avoid unnecessary energy consumption; when slightly overheated, increase the rotational speed of the fan at the condensation end of the partitioned heat pipe to enhance air-cooled heat dissipation; when moderately overheated, activate the thermoelectric cooler for local point cooling to accurately reduce the temperature; when severely overheated, start dynamic voltage and frequency adjustment, increase the power of the thermoelectric cooler and accelerate the rotational speed of the fan for global cooling.
[0259] Generally speaking, this hierarchical control strategy makes the heat dissipation measures more targeted and effectively responds to various heating situations.
[0260] Generally speaking, precise temperature control can ensure that each chip of the AI server operates within an appropriate temperature range, avoiding problems such as performance degradation, shortened lifespan, and even hardware damage caused by overheating.
[0261] Generally speaking, through intelligent adjustment of the embedded heat dissipation module, maintain a stable operating environment for the server, improve the reliability and stability of the server, and ensure the continuous and efficient progress of AI computing tasks.
[0262] Refer to Figure 2 As shown, it is a schematic flow chart of a multi-heat-source partitioned heat pipe array packaging method applied to an AI server provided by an embodiment of the present invention. In this embodiment, the multi-heat-source partitioned heat pipe array packaging method applied to an AI server includes: S1. Dynamically configure the sensor acquisition period for the AI server with an embedded distributed thermocouple array to obtain the dynamic temperature sensing network of the AI server, and obtain the real-time temperature data of the AI server based on the dynamic temperature sensing network; S2. Create an adaptive non-uniform grid for the AI server based on pre-acquired real-time power consumption data, fit the mapping relationship between current and heat generation amount in the AI server based on the adaptive non-uniform grid, and generate the power consumption distribution map of the AI server based on the mapping relationship and the dynamic reduction algorithm; S3. Conduct cluster analysis on the heat source distribution of the AI server to obtain the global heat network of the AI server; S4. Construct a three-dimensional coupling model of the AI server based on the real-time temperature data, the power consumption distribution map, and the global heat network; S5. Construct a multi-heat-source partitioned heat pipe array of the AI server based on the global heat network, and perform interface packaging on the chips of the AI server and the multi-heat-source partitioned heat pipe array to obtain the embedded heat dissipation module of the AI server; S6. Perform temperature control of the AI server on the embedded heat dissipation module based on the three-dimensional coupling model.
[0263] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.
[0264] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0265] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A multi-heat source zoned heat pipe array packaging system applied to AI servers, characterized in that: The system includes a real-time temperature data acquisition module, a power consumption distribution map acquisition module, a global heat network acquisition module, a three-dimensional coupling model construction module, an embedded heat dissipation module generation module, and an AI server temperature control module, wherein: A real-time temperature data acquisition module, used to dynamically configure a sensor acquisition cycle for an AI server with an embedded distributed thermocouple array, obtain a dynamic temperature sensing network of the AI server, and obtain real-time temperature data of the AI server based on the dynamic temperature sensing network; A power consumption distribution map acquisition module, used to create an adaptive non-uniform grid of the AI server based on pre-acquired real-time power consumption data, fit a mapping relationship between current and heat generation in the AI server based on the adaptive non-uniform grid, and generate a power consumption distribution map of the AI server based on the mapping relationship and a dynamic order reduction algorithm; A global heat network acquisition module, used to perform cluster analysis on the heat source distribution of the AI server to obtain the global heat network of the AI server; A three-dimensional coupling model building module, used to build a three-dimensional coupling model of the AI server based on the real-time temperature data, the power consumption distribution map and the global heat network; An embedded heat dissipation module generation module is used to construct a multi-heat source partitioned heat pipe array of the AI server based on the global heat network, and perform interface packaging between the chip of the AI server and the multi-heat source partitioned heat pipe array to obtain an embedded heat dissipation module of the AI server; The AI server temperature control module is used to control the temperature of the AI server by the embedded heat dissipation module based on the three-dimensional coupling model.
2. The multi-heat source zoned heat pipe array packaging system for AI server according to claim 1, characterized in that: When the real-time temperature data acquisition module dynamically configures the sensor acquisition cycle of the AI server embedded with the distributed thermocouple array to obtain the dynamic temperature sensing network of the AI server, it is specifically used to: Embedding miniaturized K-type thermocouples densely packed in a hexagon on a packaging substrate of the AI server to obtain an embedded distributed thermocouple array of the AI server; Dynamically configure a collection period for the embedded distributed thermocouple array to obtain a 3D temperature field model of the AI server; The 3D temperature field model is superimposed on the digital twin of the physical substrate to obtain a dynamic temperature sensing network of the AI server.
3. The multi-heat source zoned heat pipe array packaging system for AI server according to claim 1, characterized in that: When the power consumption distribution map acquisition module creates an adaptive non-uniform grid of the AI server based on pre-acquired real-time power consumption data, it is specifically used to: Reading real-time power consumption data of the AI server through the SVID protocol; The global interpolation field of the AI server is calculated based on the real-time power consumption data, wherein the global interpolation field is calculated as follows: In the formula, is the adaptive non-uniform grid, is the three-dimensional coordinate point of the chip of the AI server, is the total number of interpolation values of the full-field interpolation field, is the interpolation number of the full-field interpolation field, is the weight of the interpolation of the full field interpolation field affected by temperature, is the Gaussian kernel decay function, is the real-time power consumption influencing factor in the real-time power consumption data, is the spatial distance between two interpolations of the full-field interpolation field, are the three-dimensional coordinate points where the embedded distributed thermocouple array has been deployed, is the attenuation factor of heat transfer; The AI server is gridded based on the global interpolation field to obtain an adaptive non-uniform grid of the AI server.
4. The multi-heat source zoned heat pipe array packaging system for AI server according to claim 1, characterized in that: When the power consumption distribution map acquisition module performs the mapping relationship between the current and the heat generation in the AI server based on the adaptive non-uniform grid fitting and generates the power consumption distribution map of the AI server based on the mapping relationship and the dynamic order reduction algorithm, it is specifically used to: Extracting busbar layer current density distribution of the AI server based on the adaptive non-uniform grid; Generating a Joule heat increment field of the AI server based on the adaptive non-uniform grid; Performing electromagnetic-thermal coupling on the busbar layer current density distribution and the Joule heat increment field to obtain the electromagnetic-thermal correlation of the AI server; Generate a mapping relationship between current and heat generation in the AI server based on the electromagnetic-thermal correlation; Constructing a 2000-order state space model of the AI server based on the mapping relationship; The 2000-order state space model is subjected to hierarchical feature truncation and order reduction based on a dynamic order reduction algorithm to obtain a hierarchical feature vector of the AI server, wherein the dynamic order reduction algorithm is as follows: In the formula, is the hierarchical feature vector obtained after reducing the order of the 2000-order state space model, is the 2000-order state-space model, is the inverted heat conduction matrix, is the heat conduction matrix, is the matrix inversion flag, is the hierarchical feature vector, To reduce the error of the reconstructed temperature field, is the balance parameter between thermophysical constraints and statistical properties, is the potential characteristic distribution of the temperature field in the 2000-order state-space model, is a regularization constraint on the potential characteristics of the temperature field in the 2000-order state-space model, The global KL constraint ignores the importance of local hot spots. is the potential characteristic of the temperature field, is the original temperature field in the 2000-order state-space model, is the reconstructed temperature field in the 2000-order state-space model; The power consumption features of the hierarchical feature vectors are fused to obtain a power consumption distribution map of the AI server.
5. The multi-heat source zoned heat pipe array packaging system for AI server according to claim 1, characterized in that: When the global heat network acquisition module performs cluster analysis on the heat source distribution of the AI server to obtain the global heat network of the AI server, it is specifically used to: Extracting thermodynamic features of the heat source distribution of the AI server to obtain thermodynamic data of the heat source distribution; Performing cluster analysis on the thermodynamic data to obtain high temperature cluster data of the AI server; Performing grid division on the heat source distribution to obtain a grid of the heat source distribution; The grid is labeled based on the high temperature cluster data: If the high temperature cluster data of the grid exceeds a preset high temperature threshold, it is marked as a temperature sensitive grid; If the high temperature cluster data of the grid is lower than a preset high temperature threshold, it is marked as a temperature insensitive grid; The temperature-sensitive grids and the temperature-insensitive grids are aggregated into a global heat network for the AI server.
6. The multi-heat source zoned heat pipe array packaging system for AI server according to claim 1, characterized in that: When the three-dimensional coupling model building module builds the three-dimensional coupling model of the AI server based on the real-time temperature data, the power consumption distribution map and the global heat network, it is specifically used to: Performing time series alignment processing on the power consumption distribution map and the real-time temperature data to obtain standard data of the AI server; Based on the global thermal network, an unstructured tetrahedral mesh model is constructed for the physical structure of the AI server to obtain a thermal gradient regional model of the AI server; Allocating a multi-physics field coupling node to a high current density area in the power consumption distribution map to obtain an initial model of the AI server; Performing electrothermal coupling analysis on the AI server based on the multi-physics field coupling nodes of the initial model and the standard data to obtain an electrothermal coupling model of the AI server; A three-dimensional coupling model of the AI server is constructed based on the electrothermal coupling model and the thermal gradient region model.
7. The multi-heat source zoned heat pipe array packaging system for AI server according to claim 6, characterized in that: When the three-dimensional coupling model building module builds the three-dimensional coupling model of the AI server based on the electrothermal coupling model and the thermal gradient region model, it is specifically used to: Defining a coupling interface in a thermal gradient key region in the thermal gradient region model to obtain an overlapping domain coupling interface of the thermal gradient region model; Extracting the global temperature residual when the thermal gradient regional model is iterated; Extracting the heat flow residual when the electrothermal coupling model is iterated; When the global temperature residual is less than a thermal gradient threshold and the heat flow residual is less than an electrothermal threshold, a converged thermal gradient regional model and a converged electrothermal coupling model are obtained; Based on the overlapping domain coupling interface, the converged thermal gradient region model and the converged electrothermal coupling model are thermally-electrically-mechanically coupled across scales to obtain a three-dimensional coupling model of the AI server.
8. The multi-heat source zoned heat pipe array packaging system for AI server according to claim 1, characterized in that: When the embedded heat dissipation module generation module executes the multi-heat source partitioned heat pipe array of the AI server based on the global heat network, performs interface packaging between the chip of the AI server and the multi-heat source partitioned heat pipe array, and obtains the embedded heat dissipation module of the AI server, it is specifically used to: Extracting heat source features from the global heat network to obtain heat source features of the AI server; Arranging heat pipes of the AI server based on the heat source characteristics to obtain a multi-heat source partitioned heat pipe array of the AI server; Gradient brazing the multi-heat source partitioned heat pipe array and the chip of the AI server to obtain an initial heat dissipation module of the AI server; The initial heat dissipation module is hermetically packaged in multiple layers to obtain an embedded heat dissipation module of the AI server.
9. The multi-heat source zoned heat pipe array packaging system for AI server according to claim 1, characterized in that: When the AI server temperature control module performs temperature control of the AI server on the embedded heat dissipation module based on the three-dimensional coupling model, the AI server temperature control module is specifically used to: Performing short-term rolling prediction on the AI server based on a three-dimensional coupling model; When the prediction result shows that the AI server has no overheating risk, the embedded heat dissipation module maintains the current working state; When the prediction result indicates that the AI server has a slight overheating risk, the embedded heat dissipation module increases the fan speed of the condensing end of the zoned heat pipe; When the prediction result indicates that the AI server has a moderate overheating risk, the embedded heat dissipation module activates the semiconductor cooling sheet to perform local point cooling on the AI server; When the prediction result is that the AI server has a serious overheating risk, the embedded heat dissipation module starts dynamic voltage and frequency adjustment, increases the power of the semiconductor refrigeration plate, increases the fan speed of the condensing end of the zoned heat pipe, and performs global cooling on the AI server.
10. A multi-heat source partitioned heat pipe array packaging method applied to an AI server, characterized in that: The method comprises: S1. Dynamically configure a sensor collection cycle for an AI server with an embedded distributed thermocouple array to obtain a dynamic temperature sensing network of the AI server, and obtain real-time temperature data of the AI server based on the dynamic temperature sensing network; S2. Creating an adaptive non-uniform grid of the AI server based on pre-acquired real-time power consumption data, fitting a mapping relationship between current and heat generation in the AI server based on the adaptive non-uniform grid, and generating a power consumption distribution map of the AI server based on the mapping relationship and a dynamic order reduction algorithm; S3. Performing cluster analysis on the heat source distribution of the AI server to obtain the global heat network of the AI server; S4. Building a three-dimensional coupling model of the AI server based on the real-time temperature data, the power consumption distribution map, and the global heat network; S5. Construct a multi-heat source partitioned heat pipe array of the AI server based on the global heat network, perform interface packaging on the chip of the AI server and the multi-heat source partitioned heat pipe array, and obtain an embedded heat dissipation module of the AI server; S6. Control the temperature of the AI server on the embedded heat dissipation module based on the three-dimensional coupling model.
Citation Information
Patent Citations
Modeling subsurface processes on unstructured grid
CN101903803A
Method for positioning resonance-induced subsynchronous oscillation source in offshore wind plant
CN113612237A
Embedded digital twinning structure damage quantification method, equipment and medium
CN119089326A
Method and system for measuring temperature and power distributions of a device in a package
US20060039114A1
Cited By
Electronic component heat distribution self-adaptive manifold distribution system and liquid distribution method
CN120387424A