A multi-heat source zoned heat pipe array packaging system and method for AI servers

By building a multi-heat source partitioned heat pipe array packaging system in an AI server, using real-time temperature data and power consumption distribution map acquisition module, and combining a three-dimensional coupling model for temperature control, the problem of low heat dissipation efficiency of AI servers is solved, and precise positioning of high-temperature areas and efficient heat dissipation is achieved.

CN120076278BActive Publication Date: 2025-08-19SUZHOU HUASHENGYUAN ELECTROMECHANICAL CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510548171.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-19
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

When the heat dissipation technology of existing AI servers is difficult to accurately locate high-temperature areas when dissipating heat from multiple heat sources, resulting in poor targeted cooling measures and the existing technology is difficult to meet the needs of efficient heat dissipation, which affects the stable operation and performance of the server.

Method used

The real-time temperature data acquisition module, power consumption distribution map acquisition module, full-domain thermal network acquisition module, three-dimensional coupled model construction module and embedded heat dissipation module generation module are used to build a multi-heat source partitioned heat pipe array for packaging and temperature control through embedded distributed thermocouple array, adaptive non-uniform grid fitting, clustering analysis and three-dimensional coupling model.

Benefits of technology

It realizes accurate monitoring and efficient heat dissipation of the internal temperature distribution of AI servers, can accurately locate high-temperature areas, improve heat dissipation accuracy and efficiency, and ensure the stable operation and performance of the server.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120076278B_ABST
    Figure CN120076278B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of chip heat dissipation technology, and discloses a multi-heat source zoned heat pipe array packaging system and method for AI servers. The system includes a real-time temperature data acquisition module, a power consumption distribution map acquisition module, a global heat network acquisition module, a three-dimensional coupling model construction module, an embedded heat dissipation module generation module, and an AI server temperature control module. The system configures a cycle for the embedded distributed thermocouple array to obtain a dynamic temperature sensing network, and acquires data based on the dynamic temperature sensing network. An adaptive non-uniform grid is created based on real-time power consumption data, and a mapping relationship between current and heat generation is fitted based on the adaptive non-uniform grid. A power consumption distribution map is generated based on the mapping relationship. The heat source distribution is analyzed to obtain a global heat network. A three-dimensional coupling model is constructed. The chip is packaged with the multi-heat source zoned heat pipe array to obtain an embedded heat dissipation module. Temperature control is performed based on the three-dimensional coupling model. The present invention can improve the heat dissipation efficiency of AI servers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of chip heat dissipation technology, and in particular to a multi-heat source partitioned heat pipe array packaging system and method for AI servers. Background Art

[0002] Heat dissipation is crucial for the operation of AI servers. With the rapid development of AI technology, the computing performance of AI servers continues to improve, and chip integration is becoming increasingly dense. This significantly increases the heat generated by these servers during operation. Existing heat dissipation technologies lack accuracy when addressing the multiple heat sources in AI servers. Traditional temperature monitoring methods often use a small number of fixed-position sensors, making it difficult to comprehensively and accurately capture the complex temperature distribution within the server. They are unable to precisely locate high-temperature areas, resulting in limited targeted heat dissipation measures.

[0003] Existing cooling technologies also struggle to meet demand for heat dissipation. Common air-cooling methods rely solely on forced convection through fans, which cannot efficiently remove heat from multiple heat sources with uneven distribution. For some critical chips with high heat generation, air cooling has limited heat dissipation capacity, which can easily cause local overheating and affect chip performance and lifespan. While liquid cooling technology offers relatively high heat dissipation efficiency, it suffers from high costs and complex maintenance. Furthermore, it faces challenges in adapting to the internal structure of AI servers, preventing widespread application. These combined issues result in inefficient cooling technologies for AI servers, severely restricting their stable operation and performance. Summary of the Invention

[0004] The present invention provides a multi-heat source partitioned heat pipe array packaging system and method for AI servers, the main purpose of which is to solve the problem of low heat dissipation efficiency of AI servers.

[0005] To achieve the above objectives, the present invention provides a multi-heat source zoned heat pipe array packaging system for AI servers, characterized in that the system includes a real-time temperature data acquisition module, a power consumption distribution map acquisition module, a global thermal network acquisition module, a three-dimensional coupling model construction module, an embedded heat dissipation module generation module, and an AI server temperature control module, wherein:

[0006] A real-time temperature data acquisition module, configured to dynamically configure a sensor acquisition cycle for an AI server embedded with a distributed thermocouple array, obtain a dynamic temperature sensing network for the AI server, and obtain real-time temperature data of the AI server based on the dynamic temperature sensing network;

[0007] a power consumption distribution map acquisition module, configured to create an adaptive non-uniform grid for the AI server based on pre-acquired real-time power consumption data, fit a mapping relationship between current and heat generation in the AI server based on the adaptive non-uniform grid, and generate a power consumption distribution map for the AI server based on the mapping relationship and a dynamic order reduction algorithm;

[0008] A global heat network acquisition module, configured to perform cluster analysis on the heat source distribution of the AI server to obtain the global heat network of the AI server;

[0009] a three-dimensional coupling model construction module, configured to construct a three-dimensional coupling model of the AI server based on the real-time temperature data, the power consumption distribution map, and the global thermal network;

[0010] An embedded heat dissipation module generation module is used to construct a multi-heat source zoned heat pipe array for the AI server based on the global heat network, and perform interface packaging between the chip of the AI server and the multi-heat source zoned heat pipe array to obtain an embedded heat dissipation module for the AI server;

[0011] The AI server temperature control module is used to control the temperature of the AI server by the embedded heat dissipation module based on the three-dimensional coupling model.

[0012] In a preferred embodiment, the real-time temperature data acquisition module is specifically configured to:

[0013] Embedding miniaturized K-type thermocouples densely packed in a hexagonal pattern on a packaging substrate of the AI server to obtain an embedded distributed thermocouple array of the AI server;

[0014] Dynamically configuring an acquisition cycle for the embedded distributed thermocouple array to obtain a 3D temperature field model of the AI server;

[0015] The 3D temperature field model is superimposed on the digital twin of the physical substrate to obtain a dynamic temperature sensing network of the AI server.

[0016] In a preferred embodiment, when the power consumption distribution map acquisition module creates the adaptive non-uniform grid of the AI server based on the pre-acquired real-time power consumption data, it is specifically configured to:

[0017] Reading real-time power consumption data of the AI server through the SVID protocol;

[0018] A global interpolation field of the AI server is calculated based on the real-time power consumption data, wherein the global interpolation field is calculated as follows: Where, is the adaptive non-uniform grid, is the three-dimensional coordinate point of the chip of the AI server, is the total number of interpolation values of the full-field interpolation field, is the interpolation number of the full-field interpolation field, is the weight of the interpolation of the full field interpolation field affected by temperature, is the Gaussian kernel decay function, is the real-time power consumption influencing factor in the real-time power consumption data, is the spatial distance between two interpolations of the full-field interpolation field, are the three-dimensional coordinate points where the embedded distributed thermocouple array has been deployed, is the attenuation factor of heat transfer;

[0019] The AI server is grid-divided based on the global interpolation field to obtain an adaptive non-uniform grid of the AI server.

[0020] In a preferred embodiment, when the power consumption distribution map acquisition module performs the mapping relationship between the current and the heat generation in the AI server based on the adaptive non-uniform grid fitting and generates the power consumption distribution map of the AI server based on the mapping relationship and the dynamic order reduction algorithm, it is specifically used to:

[0021] Extracting a busbar layer current density distribution of the AI server based on the adaptive non-uniform grid;

[0022] generating a Joule heat increment field of the AI server based on the adaptive non-uniform grid;

[0023] Performing electromagnetic-thermal coupling on the busbar layer current density distribution and the Joule heat increment field to obtain an electromagnetic-thermal correlation of the AI server;

[0024] generating a mapping relationship between current and heat generation in the AI server based on the electromagnetic-thermal correlation;

[0025] Constructing a 2000-order state space model of the AI server based on the mapping relationship;

[0026] The 2000-order state space model is subjected to hierarchical feature truncation and order reduction based on a dynamic order reduction algorithm to obtain a hierarchical feature vector of the AI server, wherein the dynamic order reduction algorithm is as follows: Where, is the hierarchical feature vector obtained after reducing the order of the 2000-order state space model, is the 2000-order state-space model, is the inverted heat conduction matrix, is the heat conduction matrix, is the matrix inversion flag, is the hierarchical feature vector, To reduce the error of the reconstructed temperature field, is the balance parameter between thermophysical constraints and statistical properties, is the potential characteristic distribution of the temperature field in the 2000-order state-space model, is a regularization constraint on the potential characteristics of the temperature field in the 2000-order state-space model, The global KL constraint ignores the importance of local hot spots. is the potential characteristic of the temperature field, is the original temperature field in the 2000-order state-space model, is the reconstructed temperature field in the 2000-order state-space model.

[0027] Power consumption features are fused on the hierarchical feature vectors to obtain a power consumption distribution map of the AI server.

[0028] In a preferred embodiment, when performing cluster analysis on the heat source distribution of the AI server to obtain the global heat network of the AI server, the global heat network acquisition module is specifically configured to:

[0029] performing thermodynamic feature extraction on the heat source distribution of the AI server to obtain thermodynamic data of the heat source distribution;

[0030] Performing cluster analysis on the thermodynamic data to obtain high-temperature cluster data of the AI server;

[0031] Performing grid division on the heat source distribution to obtain a grid of the heat source distribution;

[0032] The grid is labeled based on the high temperature cluster data:

[0033] If the high temperature cluster data of the grid exceeds a preset high temperature threshold, it is marked as a temperature sensitive grid;

[0034] If the high temperature cluster data of the grid is lower than the preset high temperature threshold, it is marked as a temperature insensitive grid;

[0035] The temperature-sensitive grids and the temperature-insensitive grids are aggregated into a global heat network for the AI server.

[0036] In a preferred embodiment, when the three-dimensional coupling model construction module executes the construction of the three-dimensional coupling model of the AI server based on the real-time temperature data, the power consumption distribution map, and the global thermal network, it is specifically used to:

[0037] Performing time series alignment processing on the power consumption distribution map and the real-time temperature data to obtain standard data of the AI server;

[0038] Constructing an unstructured tetrahedral mesh model of the physical structure of the AI server based on the global thermal network to obtain a thermal gradient regional model of the AI server;

[0039] Allocating a multi-physics field coupling node to a high current density area in the power consumption distribution map to obtain an initial model of the AI server;

[0040] performing an electrothermal coupling analysis on the AI server based on the multi-physics field coupling nodes of the initial model and the standard data to obtain an electrothermal coupling model of the AI server;

[0041] A three-dimensional coupling model of the AI server is constructed based on the electrothermal coupling model and the thermal gradient region model.

[0042] In a preferred embodiment, when the three-dimensional coupling model construction module constructs the three-dimensional coupling model of the AI server based on the electrothermal coupling model and the thermal gradient region model, it is specifically used to:

[0043] Defining a coupling interface in a thermal gradient key region of the thermal gradient region model to obtain an overlapping domain coupling interface of the thermal gradient region model;

[0044] extracting a global temperature residual when the thermal gradient regional model is iterated;

[0045] extracting the heat flow residual when the electrothermal coupling model is iterated;

[0046] When the global temperature residual is less than a thermal gradient threshold and the heat flow residual is less than an electrothermal threshold, a converged thermal gradient regional model and a converged electrothermal coupling model are obtained;

[0047] Based on the overlapping domain coupling interface, the converged thermal gradient region model and the converged electrothermal coupling model are thermal-electrical-mechanical cross-scale coupled to obtain a three-dimensional coupling model of the AI server.

[0048] In a preferred embodiment, when the embedded heat dissipation module generation module executes the multi-heat source zoned heat pipe array construction of the AI server based on the global heat network, and performs interface packaging between the chip of the AI server and the multi-heat source zoned heat pipe array to obtain the embedded heat dissipation module of the AI server, it is specifically used to:

[0049] Extracting heat source features from the global heat network to obtain heat source features of the AI server;

[0050] Arranging heat pipes of the AI server based on the heat source characteristics to obtain a multi-heat source zoned heat pipe array of the AI server;

[0051] Gradient brazing is performed on the multi-heat source zoned heat pipe array and the chip of the AI server to obtain an initial heat dissipation module of the AI server;

[0052] The initial heat dissipation module is hermetically packaged in multiple layers to obtain the embedded heat dissipation module of the AI server.

[0053] In a preferred embodiment, when the AI server temperature control module performs temperature control of the AI server on the embedded heat dissipation module based on the three-dimensional coupling model, it is specifically configured to:

[0054] Performing short-term rolling prediction on the AI server based on a three-dimensional coupling model;

[0055] When the prediction result indicates that the AI server has no overheating risk, the embedded heat dissipation module maintains the current working state;

[0056] When the prediction result indicates that the AI server is at a slight overheating risk, the embedded heat dissipation module increases the fan speed of the condensing end of the zoned heat pipe;

[0057] When the prediction result indicates that the AI server has a moderate overheating risk, the embedded heat dissipation module activates the semiconductor cooling chip to perform local point cooling on the AI server;

[0058] When the prediction result indicates that the AI server is at serious risk of overheating, the embedded heat dissipation module starts dynamic voltage and frequency adjustment, increases the power of the semiconductor refrigeration plate, increases the fan speed of the condensing end of the zoned heat pipe, and performs global cooling on the AI server.

[0059] In order to solve the above problems, the present invention also provides a multi-heat source zoned heat pipe array packaging method for an AI server, the method comprising:

[0060] S1. Dynamically configure a sensor acquisition cycle for an AI server with an embedded distributed thermocouple array to obtain a dynamic temperature sensing network of the AI server, and acquire real-time temperature data of the AI server based on the dynamic temperature sensing network.

[0061] S2. Creating an adaptive non-uniform grid for the AI server based on pre-acquired real-time power consumption data, fitting a mapping relationship between current and heat generation in the AI server based on the adaptive non-uniform grid, and generating a power consumption distribution map of the AI server based on the mapping relationship and a dynamic order reduction algorithm.

[0062] S3. Perform cluster analysis on the heat source distribution of the AI server to obtain a global heat network of the AI server;

[0063] S4. Constructing a three-dimensional coupling model of the AI server based on the real-time temperature data, the power consumption distribution map, and the global thermal network;

[0064] S5. Construct a multi-heat source zoned heat pipe array for the AI server based on the global heat network, and perform interface packaging between the chip of the AI server and the multi-heat source zoned heat pipe array to obtain an embedded heat dissipation module for the AI server.

[0065] S6. Control the temperature of the AI server on the embedded heat dissipation module based on the three-dimensional coupling model.

[0066] Compared with the prior art, the present invention has the following beneficial effects:

[0067] 1. This invention utilizes a real-time temperature data acquisition module, employs hexagonally densely packed miniaturized K-type thermocouples within an embedded distributed thermocouple array, and dynamically configures the sensor acquisition cycle to construct a dynamic temperature sensing network. This network accurately captures the 3D temperature field model of the AI server and overlays it onto the digital twin of the physical substrate to form a digital twin thermal image. This allows for a comprehensive and precise understanding of the server's internal temperature distribution in real time, accurately locating the global thermal domain, and providing a reliable basis for the precise implementation of subsequent heat dissipation measures, significantly improving heat dissipation accuracy.

[0068] 2. The power consumption distribution map acquisition module in the present invention generates an accurate power consumption distribution map based on the mapping relationship between the current and heat generation in the AI server through adaptive non-uniform grid fitting, so that the heat dissipation design can more specifically match the heat generation conditions in each area; at the same time, the embedded heat dissipation module generation module constructs a multi-heat source partitioned heat pipe array based on the global heat network, and adopts gradient brazing and multi-layer airtight packaging technology to enhance the heat dissipation capacity; in addition, the AI server temperature control module performs short-term rolling prediction based on a three-dimensional coupling model, and takes corresponding measures according to different overheating risk levels, such as increasing the fan speed at the condensing end, activating semiconductor refrigeration plates, and performing dynamic voltage and frequency scaling, etc., to achieve accurate and efficient heat dissipation, effectively solving the multi-heat source heat dissipation problem of the AI server and comprehensively improving the heat dissipation effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 This is a system architecture diagram of a multi-heat source zoned heat pipe array packaging system for an AI server provided by one embodiment of the present invention;

[0070] Figure 2 A schematic flow chart of a method for packaging a multi-heat source zoned heat pipe array for an AI server, provided in accordance with one embodiment of the present invention.

[0071] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0072] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments belong to some embodiments of the present invention, but not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0073] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise, and "a plurality" generally includes at least two.

[0074] As used herein, the words “if” or “when” may be interpreted as “at the time of” or “when” or “in response to determining” or “in response to detecting,” depending on the context. Similarly, the phrases “if it is determined” or “if (stated condition or event) is detected” may be interpreted as “when it is determined” or “in response to the determination” or “when detecting (stated condition or event)” or “in response to detecting (stated condition or event),” depending on the context.

[0075] In addition, the step sequence in the following method embodiments is only an example and not a strict limitation.

[0076] In practice, the server-side device deployed in the multi-heat source partitioned heat pipe array packaging system for AI servers may be composed of one or more devices. The multi-heat source partitioned heat pipe array packaging system for AI servers can be implemented as a service instance, a virtual machine, or a hardware device. For example, the multi-heat source partitioned heat pipe array packaging system for AI servers can be implemented as a service instance deployed on one or more devices in a cloud node. Simply put, the multi-heat source partitioned heat pipe array packaging system for AI servers can be understood as software deployed on a cloud node, used to provide the multi-heat source partitioned heat pipe array packaging system for AI servers to each client. Alternatively, the multi-heat source partitioned heat pipe array packaging system for AI servers can be implemented as a virtual machine deployed on one or more devices in a cloud node. This virtual machine contains application software for managing each client. Alternatively, the multi-heat source partitioned heat pipe array packaging system for AI servers can be implemented as a server-side device composed of numerous hardware devices of the same or different types, with one or more hardware devices configured to provide the multi-heat source partitioned heat pipe array packaging system for AI servers to each client.

[0077] In terms of implementation, the multi-heat source partitioned heat pipe array packaging system applied to the AI server and the user end are mutually adapted. That is, the multi-heat source partitioned heat pipe array packaging system applied to the AI server is an application installed on the cloud service platform, and the user end is the client that establishes a communication connection with the application; or the multi-heat source partitioned heat pipe array packaging system applied to the AI server is implemented as a website, and the user end is implemented as a web page; or the multi-heat source partitioned heat pipe array packaging system applied to the AI server is implemented as a cloud service platform, and the user end is implemented as a small program in the instant messaging application.

[0078] like Figure 1 1 is a system architecture diagram of a multi-heat source partitioned heat pipe array packaging system for an AI server provided by one embodiment of the present invention.

[0079] The multi-heat source partitioned heat pipe array packaging system 100 for AI servers described in the present invention can be set in a cloud server. In terms of implementation, it can be used as one or more service devices, or as an application installed on the cloud (such as a mobile service operator's server, server cluster, etc.), or it can also be developed as a website. According to the functions implemented, the multi-heat source partitioned heat pipe array packaging system 100 for AI servers can include a real-time temperature data acquisition module 101, a power consumption distribution map acquisition module 102, a global thermal network acquisition module 103, a three-dimensional coupling model construction module 104, an embedded heat dissipation module generation module 105, and an AI server temperature control module 106. The module described in the present invention can also be called a unit, which refers to a series of computer program segments that can be executed by an electronic device processor and can complete fixed functions, which are stored in the memory of the electronic device.

[0080] In an embodiment of the present invention, in a multi-heat source partitioned heat pipe array packaging system applied to an AI server, each of the above modules can be implemented independently and called with other modules. The call here can be understood as a module that can connect to multiple modules of another type and provide corresponding services to the multiple modules connected to it. In the multi-heat source partitioned heat pipe array packaging system applied to an AI server provided by an embodiment of the present invention, the scope of application of the multi-heat source partitioned heat pipe array packaging system architecture applied to the AI server can be adjusted by adding modules and directly calling them without modifying the program code, thereby realizing cluster-type horizontal expansion, so as to achieve the purpose of quickly and flexibly expanding the multi-heat source partitioned heat pipe array packaging system applied to the AI server. In actual applications, the above modules can be set in the same device or different devices, or in a virtual device, such as a service instance in a cloud server.

[0081] The following describes the various components and specific workflows of the multi-heat source zoned heat pipe array packaging system for AI servers in conjunction with specific embodiments.

[0082] A real-time temperature data acquisition module 101 is configured to dynamically configure a sensor acquisition cycle for an AI server embedded with a distributed thermocouple array, obtain a dynamic temperature sensing network for the AI server, and obtain real-time temperature data of the AI server based on the dynamic temperature sensing network;

[0083] In an embodiment of the present invention, when dynamically configuring a sensor acquisition cycle for an AI server embedded with a distributed thermocouple array to obtain a dynamic temperature sensing network of the AI server, the real-time temperature data acquisition module is specifically configured to:

[0084] Embedding miniaturized K-type thermocouples densely packed in a hexagonal pattern on a packaging substrate of the AI server to obtain an embedded distributed thermocouple array of the AI server;

[0085] Dynamically configuring an acquisition cycle for the embedded distributed thermocouple array to obtain a 3D temperature field model of the AI server;

[0086] The 3D temperature field model is superimposed on the digital twin of the physical substrate to obtain a dynamic temperature sensing network of the AI server.

[0087] Specifically, the packaging substrate is a core component in the field of electronic packaging, specifically a supporting platform that carries chips and integrates interconnection structures. The packaging substrate is usually made of high thermal conductivity ceramic composite materials (such as aluminum nitride) and metal laminate structures (such as copper-polyimide).

[0088] Specifically, the miniaturized K-type thermocouple is an ultra-small temperature sensor designed specifically for high-density electronic packaging. It belongs to the K-type (nickel-chromium-nickel-aluminum) thermocouple family and uses micro-processing technology to compress the volume of traditional thermocouples to the sub-millimeter level.

[0089] Specifically, the physical substrate is the core thermal management carrier of the AI server packaging system. It adopts a copper-ceramic composite substrate, embedded with a hexagonal close-packed heat pipe array, and realizes full-area heat flow through multi-layer metallized channels.

[0090] Specifically, a hexagonal close-packed layout is designed on the packaging substrate of the AI server to ensure that the thermocouples are evenly distributed and cover key heating areas.

[0091] Furthermore, miniaturized K-type thermocouples are selected and embedded on the substrate surface or in the interlayer through a high-precision mounting process, while the wiring is optimized to avoid signal interference.

[0092] Furthermore, the hexagonal arrangement of thermocouples can maximize space utilization and improve the reliability of temperature data through redundant measurements of adjacent nodes.

[0093] Furthermore, electrical testing is performed upon completion to ensure stable signal transmission from each thermocouple.

[0094] Furthermore, the temperature change rate of each area is monitored in real time through an embedded controller or FPGA, and the acquisition frequency of the thermocouple is dynamically adjusted: high-frequency sampling is used in areas with high temperature or large fluctuations, while the low-temperature stable area is reduced to 1Hz to save resources.

[0095] Furthermore, machine learning algorithms are used to identify temperature change patterns and automatically optimize cycle configuration strategies.

[0096] Furthermore, after the data are synchronized by timestamps, they are combined with the spatial coordinates of the thermocouples to construct a spatiotemporally continuous raw temperature dataset.

[0097] Furthermore, based on the physical position coordinates and real-time data of the thermocouples, an interpolation algorithm is used to convert the discrete points into a continuous 3D temperature field.

[0098] Furthermore, the heat conduction process was simulated by finite element analysis to correct the error caused by insufficient thermocouple density.

[0099] Furthermore, the model output is a voxelized grid, where each voxel contains a temperature value and a confidence parameter, and can be updated in real time.

[0100] Furthermore, dynamic temperature cloud maps are rendered through visualization tools, supporting slice viewing and hotspot tracking.

[0101] Furthermore, the CAD model of the physical substrate is imported into the digital twin platform for coordinate alignment and grid matching with the 3D temperature field model.

[0102] Furthermore, real-time data streaming is achieved through the API interface, allowing the digital twin to synchronously display temperature field changes.

[0103] Furthermore, on this basis, alarm modules and prediction functions are integrated to eventually form a closed-loop dynamic temperature sensing network that supports remote monitoring and active thermal management strategy optimization.

[0104] In general, miniaturized K-type thermocouples with dense hexagonal arrangements are embedded in the packaging substrate of the AI server to form an embedded distributed thermocouple array. Compared with traditional sensors with a small number of fixed positions, this layout can more comprehensively sense the internal temperature of the server.

[0105] In general, by dynamically configuring the sensor acquisition cycle, we can obtain a 3D temperature field model of the AI server. This model can accurately reflect the temperature distribution of the server at different times and locations, providing accurate data support for subsequent heat dissipation measures.

[0106] In general, the acquired 3D temperature field model is superimposed on the digital twin of the physical substrate to obtain the dynamic temperature sensing network of the AI server. This network is equivalent to building a digital twin thermal mirror of the server's temperature distribution, allowing operators to grasp the internal temperature changes of the server in real time and intuitively.

[0107] In general, through the analysis of the digital twin thermal image, the global thermal domain can be accurately located, that is, the temperature-sensitive and insensitive areas in the server can be determined, providing a reliable basis for the subsequent targeted design of cooling solutions and implementation of cooling measures, greatly improving the accuracy of cooling.

[0108] In general, accurate and comprehensive temperature data and digital twin thermal images can help the system understand the heat conditions inside the server more accurately.

[0109] For example, when the power consumption distribution map acquisition module, the embedded heat dissipation module generation module, and the AI server temperature control module are working, a more targeted heat dissipation strategy can be designed based on these data.

[0110] In general, in the power consumption distribution map acquisition module, fitting the mapping relationship between current and heat generation based on accurate temperature data will be more accurate; the embedded heat dissipation module generation module can reasonably arrange heat pipes according to temperature-sensitive areas; the AI server temperature control module can adjust the heat dissipation method in time according to temperature changes, thereby improving the overall heat dissipation effect and solving the heat dissipation problem of multiple heat sources in the AI server.

[0111] a power consumption distribution map acquisition module 102 for creating an adaptive non-uniform grid for the AI server based on pre-acquired real-time power consumption data, fitting a mapping relationship between current and heat generation in the AI server based on the adaptive non-uniform grid, and generating a power consumption distribution map for the AI server based on the mapping relationship and a dynamic order reduction algorithm;

[0112] In an embodiment of the present invention, when the power consumption distribution map acquisition module creates an adaptive non-uniform grid of the AI server based on pre-acquired real-time power consumption data, it is specifically configured to:

[0113] Reading real-time power consumption data of the AI server through the SVID protocol;

[0114] A global interpolation field of the AI server is calculated based on the real-time power consumption data, wherein the global interpolation field is calculated as follows: Where, is the adaptive non-uniform grid, is the three-dimensional coordinate point of the chip of the AI server, is the total number of interpolation values of the full-field interpolation field, is the interpolation number of the full-field interpolation field, is the weight of the interpolation of the full field interpolation field affected by temperature, is the Gaussian kernel decay function, is the real-time power consumption influencing factor in the real-time power consumption data, is the spatial distance between two interpolations of the full-field interpolation field, are the three-dimensional coordinate points where the embedded distributed thermocouple array has been deployed, is the attenuation factor of heat transfer;

[0115] The AI server is grid-divided based on the global interpolation field to obtain an adaptive non-uniform grid of the AI server.

[0116] When the power consumption distribution map acquisition module performs fitting of the mapping relationship between current and heat generation in the AI server based on the adaptive non-uniform grid and generates the power consumption distribution map of the AI server based on the mapping relationship and the dynamic order reduction algorithm, the module is specifically configured to:

[0117] Extracting a busbar layer current density distribution of the AI server based on the adaptive non-uniform grid;

[0118] generating a Joule heat increment field of the AI server based on the adaptive non-uniform grid;

[0119] Performing electromagnetic-thermal coupling on the busbar layer current density distribution and the Joule heat increment field to obtain an electromagnetic-thermal correlation of the AI server;

[0120] generating a mapping relationship between current and heat generation in the AI server based on the electromagnetic-thermal correlation;

[0121] Constructing a 2000-order state space model of the AI server based on the mapping relationship;

[0122] The 2000-order state space model is subjected to hierarchical feature truncation and order reduction based on a dynamic order reduction algorithm to obtain a hierarchical feature vector of the AI server, wherein the dynamic order reduction algorithm is as follows: Where, is the hierarchical feature vector obtained after reducing the order of the 2000-order state space model, is the 2000-order state-space model, is the inverted heat conduction matrix, is the heat conduction matrix, is the matrix inversion flag, is the hierarchical feature vector, To reduce the error of the reconstructed temperature field, is the balance parameter between thermophysical constraints and statistical properties, is the potential characteristic distribution of the temperature field in the 2000-order state-space model, is a regularization constraint on the potential characteristics of the temperature field in the 2000-order state-space model, The global KL constraint ignores the importance of local hot spots. is the potential characteristic of the temperature field, is the original temperature field in the 2000-order state-space model, is the reconstructed temperature field in the 2000-order state-space model.

[0123] Power consumption features are fused on the hierarchical feature vectors to obtain a power consumption distribution map of the AI server.

[0124] Specifically, the SVID protocol is a serial voltage identification protocol, which is a serial protocol for communication between the x86 architecture CPU and the voltage regulation module VRM. It transmits power management information through three signal lines and is used to accurately control the CPU voltage and power status.

[0125] Specifically, the temperature field in the 2000-order state-space model is a mathematical abstraction of the temperature distribution within the AI server, covering temperature information in each area. It is used to analyze temperature variation patterns, assist in constructing power consumption distribution maps, and achieve precise heat dissipation control.

[0126] Specifically, to establish a communication connection with the server, you need to specify the server's address and port number according to the SVID protocol specifications, and use an appropriate network programming library (such as Python's socket library) to build a communication socket.

[0127] Furthermore, a request message is constructed and sent according to the protocol requirements to obtain real-time power consumption data. During communication, attention must be paid to the data format and encoding to ensure that the received data can be correctly parsed. Furthermore, to enhance communication stability, it is necessary to handle possible network anomalies such as connection timeouts and data loss.

[0128] Furthermore, after successfully receiving the real-time power consumption data returned by the AI server, it needs to be parsed.

[0129] Furthermore, according to the data format specified by the SVID protocol, the received raw data is converted into a numerical form that can be used for subsequent calculations.

[0130] Furthermore, after the parsing is completed, the real-time power consumption data is stored in a suitable data structure, such as a Python list or a NumPy array.

[0131] Furthermore, this facilitates subsequent access and processing of the data.

[0132] Furthermore, to facilitate subsequent analysis, corresponding timestamps or server identification information may be added to each power consumption data.

[0133] Furthermore, based on the stored real-time power consumption data, the global interpolation field of the AI server is calculated.

[0134] Furthermore, then, a suitable mathematical library is used to implement the selected interpolation algorithm.

[0135] Furthermore, during the calculation process, the physical structure and layout of the server should be considered, and the power consumption data should be associated with the spatial location of the server.

[0136] Furthermore, the global interpolation field obtained in this way can more accurately reflect the overall power consumption distribution of the server.

[0137] Furthermore, after obtaining the global interpolation field, the AI server is gridded according to the field.

[0138] Furthermore, the purpose of grid partitioning is to divide the physical space of the server into multiple small areas to enable a more detailed analysis of power consumption distribution.

[0139] Furthermore, when dividing, the characteristics of the global interpolation field should be considered so that the grid can adaptively reflect the changes in power consumption.

[0140] Furthermore, the density of the grid can be adjusted according to the gradient or change rate of power consumption, using a finer grid in areas with larger power consumption changes and a sparser grid in areas with smaller power consumption changes.

[0141] Furthermore, in this way, an adaptive non-uniform grid of the AI server is finally obtained.

[0142] Specifically, in the calculation of the global interpolation field, data measured by distributed thermocouples deployed on the AI server chip are collected. These data correspond to the conditions at specific locations on the chip, and the total number of these known measurement points is determined.

[0143] Furthermore, two important values need to be determined: one is the degree of influence of each known measurement point on the final calculation result, and the other is a value related to real-time power consumption that affects the heat transfer attenuation rate. These two values need to be determined based on actual physical models or past experience.

[0144] Furthermore, for the target position whose estimated value is to be calculated, the spatial distance between it and each known measurement point is calculated.

[0145] Furthermore, based on the previously determined value related to real-time power consumption and the calculated spatial distance, the attenuation degree of the influence of each known measurement point on the target position is calculated with the help of a function that can reflect the attenuation characteristics of heat transfer with increasing distance.

[0146] Furthermore, the influence degree value of each known measurement point is multiplied by the corresponding influence attenuation degree, and then all the multiplied results are added together, so that the estimated value of the target position can be obtained, that is, the value of the global interpolation field at that position.

[0147] Furthermore, from the perspective of distance, the greater the distance between the target position and the known measurement point, the smaller the influence of the known measurement point on the estimated value of the target position.

[0148] Furthermore, this is because as the distance increases, the function that reflects the heat transfer attenuation characteristics will cause the influence of the known measurement point to decrease rapidly.

[0149] Furthermore, the value related to real-time power consumption also has a significant impact on the results.

[0150] Furthermore, if this value is large, then as the distance increases, the influence of the known measurement points on the target position will decrease rapidly, and the effect of the distant measurement points on the target position estimate will soon become insignificant; conversely, if this value is small, the influence of the distant measurement points on the target position will decrease more slowly, and relatively speaking, the effect on the target position estimate will be greater.

[0151] Furthermore, the main function of this formula is to use known thermocouple measurement data distributed at different locations on the chip to estimate physical quantities, such as temperature, at any location on the AI server chip.

[0152] Furthermore, it makes the estimation results more consistent with actual physical phenomena by considering the spatial distance between data points and the attenuation of heat transfer with distance.

[0153] Furthermore, the global interpolation field of the entire AI server chip can be obtained, providing important basic data for subsequent operations such as chip meshing and thermal management.

[0154] Specifically, it is necessary to obtain the adaptive non-uniform grid data of the AI server obtained previously.

[0155] Furthermore, these raster data contain relevant physical information of different locations of the server.

[0156] Furthermore, based on electromagnetic principles and the physical properties of the busbar layer, the grid data is analyzed using appropriate algorithms.

[0157] Furthermore, during the analysis process, factors such as the material properties, geometric shape, and current conduction law of the busbar layer are considered.

[0158] Furthermore, by calculating and evaluating each grid unit, the current density value of the busbar layer at each position is extracted.

[0159] Furthermore, these current density values are sorted and recorded according to the grid positions to obtain the current density distribution of the AI server bus layer.

[0160] Furthermore, based on the adaptive non-uniform grid, the Joule heat increment field of the AI server is generated according to Joule's law.

[0161] Furthermore, Joule's law states that the heat generated by current passing through a conductor is proportional to the square of the current, the resistance of the conductor, and the time the current is passed through the conductor.

[0162] Furthermore, for each grid cell, the resistance of the cell is calculated in combination with the current density information extracted previously.

[0163] Furthermore, the calculation of resistance needs to take into account factors such as the material properties and size of the grid unit. Then, the Joule heat increment generated by the grid unit is calculated based on the current density and resistance.

[0164] Furthermore, the same calculation is performed on all grid cells, and the obtained Joule heat increments of each grid cell are arranged according to their positions in the adaptive non-uniform grid, and finally the Joule heat increment field of the AI server is generated.

[0165] Furthermore, the previously obtained busbar layer current density distribution and Joule heat increment field are combined to carry out electromagnetic-thermal coupling analysis.

[0166] Furthermore, this process requires the application of relevant theories and equations of electromagnetism and thermodynamics, taking into account factors such as the magnetic field generated by the current, the impact of the magnetic field on the current, and heat transfer and distribution.

[0167] Furthermore, by establishing a suitable mathematical model, such as a coupled model of Maxwell's equations and the heat conduction equation, the interaction between the current density distribution and the Joule heat increment field is simulated and calculated.

[0168] Furthermore, the calculation process considers how the electromagnetic and thermal properties of the material change with temperature and magnetic field. Through analysis and calculation, the correlation between electromagnetic and thermal phenomena in the AI server is obtained, namely the electromagnetic-thermal correlation.

[0169] Furthermore, based on the electromagnetic-thermal correlation obtained above, a corresponding mapping relationship is generated. This mapping relationship can describe the correspondence between electromagnetic parameters and thermal parameters in the AI server.

[0170] Furthermore, methods such as function fitting and data statistical analysis can be used to process electromagnetic-thermal correlation data.

[0171] Furthermore, by analyzing a large number of data points, a mathematical expression or data table that can accurately describe the relationship between the two is found.

[0172] Furthermore, this mapping relationship will provide an important basis for the subsequent construction of the state space model.

[0173] Furthermore, the previously generated mapping relationship is used to construct a 2000-order state space model of the AI server.

[0174] Furthermore, a state-space model is a mathematical model used to describe the dynamic characteristics of a system, which expresses the relationship between the state variables, input variables, and output variables of the system using a set of differential equations or difference equations.

[0175] Furthermore, the state variables, input variables, and output variables of the AI server system are determined. State variables can include physical quantities such as current density and temperature, input variables can include power input, and output variables can include temperature or current at certain key locations.

[0176] Furthermore, the state equation and output equation are established according to the mapping relationship and the physical characteristics of the system.

[0177] Furthermore, by rationally selecting state variables and parameters, a 2000-order state space model is constructed to accurately describe the dynamic characteristics of the AI server.

[0178] Furthermore, in order to simplify the model and improve computational efficiency, a dynamic order reduction algorithm is used to perform hierarchical feature truncation and order reduction on the 2000-order state space model.

[0179] Furthermore, the characteristics of the 2000-order state-space model are analyzed to determine which characteristics have a greater impact on the system's dynamic characteristics and which have a smaller impact. Then, based on the analysis results, the model characteristics are stratified.

[0180] Furthermore, for feature layers with less influence, truncation is performed, that is, the influence of these features on the system is ignored. In this way, the order of the model is reduced and the complexity of the model is reduced.

[0181] Furthermore, during the order reduction process, it is necessary to ensure that the reduced model can still better reflect the main dynamic characteristics of the original system.

[0182] Furthermore, power consumption feature fusion is performed on the hierarchical feature vectors obtained after hierarchical feature truncation and order reduction.

[0183] Furthermore, power consumption feature fusion is to integrate various feature information related to power consumption to more comprehensively describe the power consumption distribution of AI servers.

[0184] Furthermore, weighted averaging, principal component analysis, and other methods can be used to process the hierarchical feature vectors. During the fusion process, different features are given different weights based on their impact on power consumption.

[0185] Furthermore, multiple hierarchical feature vectors are merged into a comprehensive power consumption feature vector through a fusion operation.

[0186] Furthermore, based on this comprehensive power consumption feature vector, combined with the physical structure and adaptive non-uniform grid information of the AI server, a power consumption distribution map of the AI server is generated.

[0187] Furthermore, the map can intuitively display the power consumption distribution of AI servers in different locations.

[0188] Specifically, in the calculation of the dynamic order reduction algorithm, in order to reduce the order of the 2000-order state space model, it is necessary to first obtain the heat conduction matrix and its inverse matrix.

[0189] Furthermore, the heat conduction matrix reflects the relevant characteristics of heat conduction in the model. By multiplying the inverted heat conduction matrix with the 2000-order state-space model and then multiplying it with the heat conduction matrix again, through this matrix operation process, we can obtain the reduced-order compressed hierarchical feature vector.

[0190] Furthermore, this process actually simplifies the original model, removes some information that has little impact on the overall situation, and retains only key features, making the model more concise and efficient.

[0191] Furthermore, the reconstructed temperature field after reducing the error is calculated first, which reflects the degree of difference between the reconstructed temperature field and the original temperature field. The smaller the difference, the better the reconstruction effect.

[0192] Furthermore, it is necessary to consider the balance between thermophysical constraints and statistical characteristics. There is a balance parameter that plays the role of adjusting the weight of the two.

[0193] Furthermore, we need to calculate the relationship between the potential feature distribution of the temperature field in the 2000-order state-space model and the regularization constraint on the potential feature. This relationship is measured using the global KL constraint, but it should be noted that the global KL constraint may ignore the importance of local hot spots.

[0194] Furthermore, the reconstructed temperature field after reducing the error is added to the result of multiplying the equilibrium parameter by the global KL constraint, and the resulting value is the hierarchical eigenvector.

[0195] Furthermore, if the values of the elements in the heat conduction matrix vary significantly, it may mean that the characteristics of heat conduction have changed significantly.

[0196] Furthermore, when such a change occurs, the hierarchical feature vectors after rank reduction and compression will also change accordingly.

[0197] For example, if heat conduction is enhanced, the reduced-order model may emphasize key features related to heat conduction while deemphasizing other less relevant features.

[0198] Furthermore, when the value of the reconstructed temperature field after reducing the error increases, it means that the difference between the reconstructed temperature field and the original temperature field is increasing, which may cause the value of the hierarchical eigenvector to increase.

[0199] Furthermore, changes in equilibrium parameters will directly affect the weights of thermophysical constraints and statistical properties in the final results.

[0200] Furthermore, if the balance parameter increases, the relationship between thermophysical constraints and statistical characteristics will have a more significant impact on the hierarchical eigenvector; conversely, if the balance parameter decreases, the reconstructed temperature field with reduced errors will have a relatively greater impact on the hierarchical eigenvector.

[0201] Furthermore, the main function of these two formulas is to reduce the order and extract features of the 2000-order state space model.

[0202] Furthermore, the order reduction and compression formula can simplify the complex 2000-order model into a hierarchical feature vector, reducing the complexity of the model and improving computational efficiency while retaining the key feature information of the model.

[0203] Furthermore, the formula for calculating the hierarchical eigenvector comprehensively considers the error in reconstructing the temperature field, the balance between thermal physical constraints and statistical characteristics, and the distribution of the potential characteristics of the temperature field, so that the obtained hierarchical eigenvector can more comprehensively and accurately reflect the characteristics of the model, providing a more effective basis for subsequent analysis and processing of the AI server.

[0204] In general, by creating an adaptive non-uniform grid through real-time power consumption data, we can flexibly divide the area according to the actual power consumption distribution of the server and accurately reflect the heat differences in different locations.

[0205] In general, compared with traditional uniform grids, it is more in line with the characteristics of multiple heat sources and uneven heat source distribution of AI servers, providing a basis for subsequent accurate analysis.

[0206] In summary, the mapping relationship between current and heat generation is clearly established based on the adaptive non-uniform grid fitting.

[0207] In general, this allows the heat dissipation design to be closely based on the actual heat generation conditions in each area. For example, heat dissipation components can be focused on areas with high heat generation, thereby improving the utilization efficiency of heat dissipation resources and enhancing overall heat dissipation capabilities.

[0208] In general, the power consumption distribution map generated by the mapping relationship and dynamic order reduction algorithm can intuitively present the power consumption distribution of various parts of the server.

[0209] In general, this provides a key reference for the generation module of the embedded heat dissipation module, enabling it to reasonably construct a multi-heat source partitioned heat pipe array; it also provides a decision-making basis for the AI server temperature control module, and adopts corresponding temperature control strategies according to the power consumption and heat risk of different areas to achieve precise heat dissipation.

[0210] A global heat network acquisition module 103 is configured to perform cluster analysis on the heat source distribution of the AI server to obtain the global heat network of the AI server;

[0211] In an embodiment of the present invention, when performing cluster analysis on the heat source distribution of the AI server to obtain the global heat network of the AI server, the global heat network acquisition module is specifically configured to:

[0212] performing thermodynamic feature extraction on the heat source distribution of the AI server to obtain thermodynamic data of the heat source distribution;

[0213] Performing cluster analysis on the thermodynamic data to obtain high-temperature cluster data of the AI server;

[0214] Performing grid division on the heat source distribution to obtain a grid of the heat source distribution;

[0215] The grid is labeled based on the high temperature cluster data:

[0216] If the high temperature cluster data of the grid exceeds a preset high temperature threshold, it is marked as a temperature sensitive grid;

[0217] If the high temperature cluster data of the grid is lower than the preset high temperature threshold, it is marked as a temperature insensitive grid;

[0218] The temperature-sensitive grids and the temperature-insensitive grids are aggregated into a global heat network for the AI server.

[0219] Specifically, the thermodynamic data of the AI server is collected, which may include information such as temperature and heat flux at different locations of the server.

[0220] Furthermore, a suitable clustering algorithm is selected, such as K-means clustering, hierarchical clustering, etc. The clustering algorithm will divide the thermodynamic data into different clusters based on the similarity between the data.

[0221] Furthermore, during the clustering process, the algorithm automatically identifies clusters corresponding to high-temperature areas in the data. After cluster analysis, cluster data representing high-temperature areas is selected from all clusters. This data is the high-temperature cluster data of the AI server, reflecting the characteristics of the hottest parts of the server.

[0222] Furthermore, in order to analyze the heat source distribution of the AI server in more detail, it is necessary to divide it into grids.

[0223] Furthermore, the rules and accuracy of grid division are determined, such as the size and shape of the grid. These parameters can be set according to the physical structure of the server and actual needs.

[0224] Furthermore, the heat source distribution area of the AI server is divided into small grid units according to the set rules.

[0225] Furthermore, each grid unit has its specific position and range. In this way, the continuous heat source distribution can be discretized to facilitate subsequent analysis and processing, and finally a grid of heat source distribution is obtained.

[0226] Furthermore, the high temperature cluster data obtained above is used to mark the grid of heat source distribution.

[0227] Furthermore, for each grid, a comparison analysis is performed between the grid and the high temperature cluster data to determine the magnitude relationship between the high temperature cluster data of the grid and the preset high temperature threshold.

[0228] Furthermore, the preset high temperature threshold is a temperature limit set in advance based on factors such as the performance requirements and safety standards of the AI server.

[0229] Furthermore, through comparison, it is determined whether the grid is temperature sensitive or temperature insensitive.

[0230] Furthermore, when the high temperature cluster data of a certain grid exceeds the preset high temperature threshold, it indicates that the temperature of the grid area is high, which may have a significant impact on the performance and stability of the server, so it is marked as a temperature-sensitive grid.

[0231] Furthermore, the marking method may be to add a specific identifier to the grid data for subsequent identification and processing.

[0232] Furthermore, this helps focus on these temperature-sensitive areas and take corresponding heat dissipation or optimization measures.

[0233] Furthermore, if the high temperature cluster data of a certain grid is lower than the preset high temperature threshold, it indicates that the temperature of the grid area is relatively low and has little impact on server performance, and it is marked as a temperature-insensitive grid.

[0234] Furthermore, adding corresponding identifiers to the grid data to distinguish them can help us identify which areas do not require special heat dissipation treatment, thereby allocating resources rationally.

[0235] Furthermore, all marked temperature-sensitive grids and temperature-insensitive grids are summarized and integrated.

[0236] Furthermore, a complete network structure is constructed using these grids as basic units. This network is the global heat network of the AI server.

[0237] Furthermore, the global thermal network comprehensively reflects the thermal distribution of AI servers, including temperature-sensitive areas and temperature-insensitive areas.

[0238] Furthermore, by analyzing the global thermal network, we can have a clearer understanding of the overall thermal status of the server, providing a basis for subsequent thermal management and optimization.

[0239] In general, high-temperature cluster data can be found by extracting thermodynamic features and performing cluster analysis on heat source distribution.

[0240] In general, based on this, temperature-sensitive grids and temperature-insensitive grids are marked, and areas in the AI server that are prone to high temperature problems are accurately located, providing a basis for subsequent focused heat dissipation.

[0241] For example, in actual operation, priority can be given to strengthening heat dissipation measures in these temperature-sensitive areas to avoid local overheating that affects chip performance and life.

[0242] In general, the global thermal network provides key information for embedded heat dissipation module generation modules.

[0243] In general, by extracting heat source characteristics based on the global thermal network, heat pipes can be reasonably arranged according to the heat source characteristics of different regions, and a multi-heat source partitioned heat pipe array can be constructed.

[0244] In general, increasing the number of heat pipes or adjusting the heat pipe layout in temperature-sensitive areas can improve heat dissipation efficiency, make the heat dissipation module more in line with the actual heat dissipation needs of the server, and enhance the overall heat dissipation capacity.

[0245] In general, the AI server temperature control module can monitor and predict the temperature of different areas more accurately based on the global thermal network and three-dimensional coupling model.

[0246] In general, when different overheating risks are predicted, more targeted measures can be taken for temperature-sensitive areas to ensure stable server operation.

[0247] A three-dimensional coupling model construction module 104 is configured to construct a three-dimensional coupling model of the AI server based on the real-time temperature data, the power consumption distribution map, and the global thermal network;

[0248] In an embodiment of the present invention, when the three-dimensional coupling model construction module constructs the three-dimensional coupling model of the AI server based on the real-time temperature data, the power consumption distribution map, and the global thermal network, it is specifically used to:

[0249] Performing time series alignment processing on the power consumption distribution map and the real-time temperature data to obtain standard data of the AI server;

[0250] Constructing an unstructured tetrahedral mesh model of the physical structure of the AI server based on the global thermal network to obtain a thermal gradient regional model of the AI server;

[0251] Allocating a multi-physics field coupling node to a high current density area in the power consumption distribution map to obtain an initial model of the AI server;

[0252] performing an electrothermal coupling analysis on the AI server based on the multi-physics field coupling nodes of the initial model and the standard data to obtain an electrothermal coupling model of the AI server;

[0253] A three-dimensional coupling model of the AI server is constructed based on the electrothermal coupling model and the thermal gradient region model.

[0254] When the three-dimensional coupling model construction module constructs the three-dimensional coupling model of the AI server based on the electrothermal coupling model and the thermal gradient region model, the three-dimensional coupling model construction module is specifically used to:

[0255] Defining a coupling interface in a thermal gradient key region of the thermal gradient region model to obtain an overlapping domain coupling interface of the thermal gradient region model;

[0256] extracting a global temperature residual when the thermal gradient regional model is iterated;

[0257] extracting the heat flow residual when the electrothermal coupling model is iterated;

[0258] When the global temperature residual is less than a thermal gradient threshold and the heat flow residual is less than an electrothermal threshold, a converged thermal gradient regional model and a converged electrothermal coupling model are obtained;

[0259] Based on the overlapping domain coupling interface, the converged thermal gradient region model and the converged electrothermal coupling model are thermal-electrical-mechanical cross-scale coupled to obtain a three-dimensional coupling model of the AI server.

[0260] Specifically, the global thermal network of the AI server obtained previously is obtained, which details the distribution of temperature-sensitive and insensitive areas within the server.

[0261] Furthermore, based on the information of the global thermal network, the physical structure of the AI server is analyzed in detail to determine the shape, position and relationship of each component.

[0262] Furthermore, an unstructured tetrahedral meshing method is used to decompose the physical space of the AI server into numerous irregular tetrahedral mesh units.

[0263] Furthermore, during the division process, the thermal distribution characteristics reflected by the global thermal network are fully considered. In areas with drastic temperature changes, that is, where the thermal gradient is large, a finer grid division is used to accurately capture the thermal changes; while in areas with relatively stable temperatures, a sparser grid division is used to improve computational efficiency.

[0264] Furthermore, an unstructured tetrahedral mesh model was constructed in this way, which can clearly show the thermal gradient distribution within the AI server, thereby obtaining a thermal gradient area model of the AI server.

[0265] Furthermore, we first obtain the power consumption distribution map of the AI server, which intuitively shows the current density distribution at different locations within the server.

[0266] Furthermore, high current density areas are accurately identified from the map. These areas are usually key areas with high heat generation and significant impact on server performance.

[0267] Furthermore, multiphysics coupling nodes are assigned in these high current density regions.

[0268] Furthermore, the multi-physics coupling node can comprehensively consider the interactions between multiple physical fields such as electric fields and thermal fields.

[0269] Furthermore, the number and location of nodes are reasonably determined according to the size, shape and distribution characteristics of the high current density area to ensure that the nodes can effectively reflect the changes in the physical field in the area.

[0270] Furthermore, by assigning multi-physics coupling nodes to high current density areas and combining these nodes with the physical structure of the server, an initial model of the AI server is obtained.

[0271] Furthermore, the initial model of the AI server and related standard data are obtained. The standard data includes information such as the material's electrical and thermal properties, boundary conditions, etc.

[0272] Furthermore, professional multi-physics field analysis software or numerical calculation methods are used to perform electrothermal coupling analysis on the AI server based on the multi-physics field coupling nodes and standard data in the initial model.

[0273] Furthermore, during the analysis process, the mutual influence between the electric field and the thermal field is fully considered. For example, the current passing through the conductor generates heat, and the temperature change affects the resistance and current distribution of the conductor.

[0274] Furthermore, by solving a series of electrical and thermal equations, the electrical and thermal behavior of the AI server during actual operation is simulated.

[0275] Furthermore, based on the analysis results, the initial model was optimized and adjusted, and finally an electrothermal coupling model was obtained that can accurately reflect the electrothermal interaction of the AI server.

[0276] Furthermore, this model can be used to predict server performance, evaluate the effectiveness of cooling solutions, etc.

[0277] Specifically, an iterative calculation process of the thermal gradient region model is started.

[0278] Furthermore, in each iteration, the temperature value of each grid node in the model is updated according to the physical rules and algorithms within the model.

[0279] Furthermore, after the update is completed, the temperature value of each node obtained in this iteration is compared with the temperature value of the corresponding node in the previous iteration.

[0280] Furthermore, for each grid node, the difference between the two iterative temperature values is calculated.

[0281] Furthermore, appropriate mathematical methods are used, such as taking the square root of the sum of the temperature differences of all nodes, or calculating the sum of the absolute values of the temperature differences of all nodes, to summarize these differences and obtain a single value that can represent the degree of temperature change difference in the entire thermal gradient area model. This value is the global temperature residual.

[0282] Furthermore, an iterative calculation is performed on the electrothermal coupling model. During the iteration process, the heat flow values of each part of the model are updated according to the electrical and thermal principles followed by the model.

[0283] Furthermore, the heat flux value reflects the flow of heat inside the model.

[0284] Furthermore, the heat flux distribution obtained in the current iteration is carefully compared with the heat flux distribution in the previous iteration, and the difference in heat flux values is calculated for each location in the model.

[0285] Furthermore, just like processing the global temperature residual, appropriate mathematical methods are used to integrate the heat flow differences at all locations, and finally a value that can reflect the degree of difference in heat flow changes in the entire electrothermal coupling model is obtained, namely the heat flow residual.

[0286] Furthermore, the thermal gradient threshold and electrothermal threshold are set in advance based on the performance indicators, physical characteristics, and computing accuracy requirements of the AI server.

[0287] Furthermore, after completing the iterative calculation of the thermal gradient regional model and the electrothermal coupling model each time and extracting the global temperature residual and heat flow residual respectively, the global temperature residual is compared with the thermal gradient threshold, and the heat flow residual is compared with the electrothermal threshold.

[0288] Furthermore, when the global temperature residual is less than the thermal gradient threshold and the heat flow residual is less than the electrothermal threshold, this indicates that the thermal gradient regional model and the electrothermal coupling model have achieved sufficient convergence accuracy during the iteration process. At this time, a converged thermal gradient regional model and a converged electrothermal coupling model are obtained.

[0289] Furthermore, after obtaining a converged thermal gradient region model and a converged electrothermal coupling model, the Overlapping Domain Coupling interface is introduced.

[0290] Furthermore, the overlapping domain coupling interface enables effective connection and information transfer between different scales and different physical fields.

[0291] Furthermore, the converged thermal gradient region model and the converged electrothermal coupling model are analyzed to determine the parts and overlapping regions that need to be coupled between them.

[0292] Furthermore, the overlapping domain coupling interface allows for the fusion and interaction of thermal, electric, and force field information. During the coupling process, the interactions and influences between different physical fields are fully considered, such as the influence of the electric field on the thermal field, and vice versa.

[0293] Furthermore, through coupling, the distribution values of thermal, electrical and force physical quantities are obtained and integrated according to the three-dimensional physical structure of the AI server.

[0294] Furthermore, using 3D modeling software or programming tools, with these physical quantities as parameters, a 3D model is constructed that can intuitively present the thermal, electrical, and mechanical characteristics of the AI server.

[0295] Furthermore, in the model, the distribution of various physical quantities in space is displayed through visualization methods such as different colors, textures or isosurfaces, ultimately forming a three-dimensional coupling model that comprehensively reflects the thermal, electrical and mechanical characteristics of the AI server.

[0296] In general, real-time temperature data reflects the actual temperature status of the server, the power consumption distribution map shows the heat generation conditions in each area, and the global heating network clearly defines the heat source distribution characteristics.

[0297] In general, the three-dimensional coupling model integrates these key factors to achieve a comprehensive analysis of the thermal, electrical, and mechanical multi-physics fields of AI servers, fully grasp the complex physical processes inside the server, and avoid the limitations of single-factor analysis.

[0298] In general, this model can provide a clear understanding of the heat transfer and temperature change patterns in different areas, providing precise guidance for generating modules for embedded heat dissipation modules.

[0299] For example, based on the model results, the heat pipe layout is reasonably adjusted and the heat dissipation material configuration is optimized in areas with large thermal gradients, so that the multi-heat source partitioned heat pipe array can better meet the actual heat dissipation needs of the server and improve the heat dissipation efficiency.

[0300] An embedded heat dissipation module generation module 105 is configured to construct a multi-heat source zoned heat pipe array for the AI server based on the global heat network, and perform interface packaging between the chip of the AI server and the multi-heat source zoned heat pipe array to obtain an embedded heat dissipation module for the AI server;

[0301] In an embodiment of the present invention, when the embedded heat dissipation module generates a multi-heat source zoned heat pipe array for the AI server based on the global heat network and interfaces the chip of the AI server with the multi-heat source zoned heat pipe array to obtain the embedded heat dissipation module of the AI server, the module is specifically configured to:

[0302] Extracting heat source features from the global heat network to obtain heat source features of the AI server;

[0303] Arranging heat pipes of the AI server based on the heat source characteristics to obtain a multi-heat source zoned heat pipe array of the AI server;

[0304] Gradient brazing is performed on the multi-heat source zoned heat pipe array and the chip of the AI server to obtain an initial heat dissipation module of the AI server;

[0305] The initial heat dissipation module is hermetically packaged in multiple layers to obtain the embedded heat dissipation module of the AI server.

[0306] Specifically, data analysis and processing technologies are used to conduct in-depth research on the global heating network. By analyzing temperature trends, temperature distribution, heat flow direction, and other information in each area of the heating network, the location of heat sources within the server can be identified.

[0307] Furthermore, the characteristics of these heat sources are extracted, such as the heating power, heating stability, heating frequency, etc. of the heat sources.

[0308] Furthermore, by organizing and summarizing the extracted heat source locations and characteristics, we obtain the AI server's heat source characteristics. These characteristics are crucial for subsequent heat dissipation design and provide an accurate basis for heat pipe layout.

[0309] Furthermore, based on the heat source characteristics of the AI server obtained in the previous step, the heat pipe layout design was started.

[0310] Furthermore, the server is divided into multiple different areas according to the location and heating characteristics of the heat source, and each area corresponds to one or more heat sources.

[0311] Furthermore, for each zone, heat pipes of appropriate specifications and quantity are selected. The selection of heat pipes should take into account parameters such as heat transfer capacity, length, and diameter to ensure that the heat generated by the heat source in the zone can be effectively transferred out.

[0312] Furthermore, the heat pipes are installed in the corresponding partitions according to the designed layout, so that the heat pipes are in close contact with the heat sources, forming a multi-heat source partitioned heat pipe array.

[0313] Furthermore, this arrangement can precisely dissipate heat according to the heat source characteristics of different partitions, thereby improving heat dissipation efficiency.

[0314] Furthermore, after obtaining the multi-heat source partitioned heat pipe array, it is necessary to connect it to the chip of the AI server.

[0315] Furthermore, a gradient brazing method is adopted, and the connection parts of the chip and the heat pipe are first cleaned and pretreated to remove surface impurities and oxides to ensure a good welding effect.

[0316] Furthermore, appropriate brazing materials and process parameters are selected based on the material properties of the chip and heat pipe. During the soldering process, gradient heating is used to gradually increase the temperature of the soldering area to avoid damage to the chip or heat pipe due to rapid temperature changes.

[0317] Furthermore, by precisely controlling the brazing process, the multi-heat source partitioned heat pipe array is firmly connected to the chip to form an overall heat dissipation structure, which is the initial heat dissipation module of the AI server.

[0318] Furthermore, in order to protect the initial heat dissipation module and improve its stability and reliability, it needs to be hermetically sealed with multiple layers.

[0319] Furthermore, suitable packaging materials are selected, such as metal shells, ceramic materials, etc. These materials should have good heat insulation, moisture resistance and protective properties.

[0320] Furthermore, the initial heat dissipation module is placed inside the packaging structure, and a multi-layer packaging method is adopted, that is, a sealing layer is set between different layers to enhance the sealing effect.

[0321] Furthermore, during the packaging process, it is necessary to ensure that a relatively stable gas environment is formed inside to prevent the entry of impurities such as external air, moisture, and dust.

[0322] Furthermore, through a precise packaging process, the initial heat dissipation module is completely sealed within the packaging structure, ultimately obtaining an embedded heat dissipation module for the AI server. This module can better adapt to the server's working environment and improve heat dissipation performance.

[0323] In general, the global thermal network accurately reflects the heat source distribution of the AI server. Based on this, building a multi-heat source partitioned heat pipe array can make the heat pipe layout highly adaptable to the heat source distribution.

[0324] In general, heat pipes can be densely arranged in temperature-sensitive areas to ensure that heat can be quickly conducted and dissipated, greatly improving the heat dissipation efficiency, avoiding local overheating, and ensuring that the chip works stably in a suitable temperature environment.

[0325] In general, the interface packaging of the AI server chip with the multi-heat source partitioned heat pipe array and the use of gradient brazing technology can form a good thermal conduction connection between the chip and the heat pipe, reduce thermal resistance, and improve heat transfer efficiency.

[0326] In general, the multi-layer airtight packaging further protects the internal structure, prevents external factors such as dust and moisture from affecting the heat dissipation performance, and enhances the stability and reliability of the entire embedded heat dissipation module.

[0327] The AI server temperature control module 106 is used to control the temperature of the AI server based on the embedded heat dissipation module based on the three-dimensional coupling model.

[0328] In an embodiment of the present invention, when the AI server temperature control module performs temperature control of the AI server on the embedded heat dissipation module based on the three-dimensional coupling model, it is specifically configured to:

[0329] Performing short-term rolling prediction on the AI server based on a three-dimensional coupling model;

[0330] When the prediction result indicates that the AI server has no overheating risk, the embedded heat dissipation module maintains the current working state;

[0331] When the prediction result indicates that the AI server is at a slight overheating risk, the embedded heat dissipation module increases the fan speed of the condensing end of the zoned heat pipe;

[0332] When the prediction result indicates that the AI server has a moderate overheating risk, the embedded heat dissipation module activates the semiconductor cooling chip to perform local point cooling on the AI server;

[0333] When the prediction result indicates that the AI server is at serious risk of overheating, the embedded heat dissipation module starts dynamic voltage and frequency adjustment, increases the power of the semiconductor refrigeration plate, increases the fan speed of the condensing end of the zoned heat pipe, and performs global cooling on the AI server.

[0334] Specifically, the time step and prediction duration of the short time domain are set. The time step needs to be reasonably selected according to the frequency of changes in the server operation data to ensure that the dynamic changes in the server's physical state can be captured in a timely manner.

[0335] Furthermore, taking the current operating parameters of the server as the initial conditions and based on the interaction between heat, electricity and force in the three-dimensional coupling model, numerical calculation methods are used to calculate the temperature, heat flow, stress and other physical quantities of the server in the future short time domain step by time.

[0336] Furthermore, after each calculation is completed, the latest calculation result is used as the initial condition for the next prediction, and the prediction process is continuously updated, thereby realizing continuous prediction of the physical state changes of the AI server in the short-term domain.

[0337] Furthermore, the results of short-term rolling prediction are obtained, and key temperature-related indicators in the prediction data are extracted, such as the surface temperature of each chip and the temperature of the hot spot area.

[0338] Furthermore, these temperature indicators are compared with the pre-set safety temperature thresholds. If all temperature indicators are lower than the safety thresholds, it indicates that the AI server has no overheating risk during the predicted period. At this time, the embedded heat dissipation module does not need to be adjusted, and it maintains the current working state, maintaining normal heat dissipation power and operating mode to ensure stable operation of the server while reducing energy consumption.

[0339] Furthermore, if the prediction results show that the temperature indicators in some areas exceed the safety threshold but do not reach a serious level, it is judged that the AI server is at risk of mild overheating.

[0340] Furthermore, in response to this situation, the condensation end fans of the zoned heat pipes in the embedded heat dissipation module are controlled.

[0341] Furthermore, by adjusting the control circuit of the fan, the fan speed is increased, so that the fan can generate a larger air volume, thereby accelerating the heat dissipation at the condensation end of the heat pipe.

[0342] Furthermore, since the mild overheating area is relatively small, the server temperature can be effectively reduced to a safe range simply by increasing the fan speed and enhancing air convection heat dissipation.

[0343] Furthermore, when the prediction results show that the server is overheating significantly and reaches a moderate overheating risk level, that is, when the temperatures in multiple areas exceed the standard and the temperatures of some key chips are close to or exceed the dangerous threshold, the semiconductor cooling plate in the embedded heat dissipation module is started.

[0344] Furthermore, based on the distribution of overheated areas, the starting position and working intensity of the semiconductor refrigeration plate are precisely controlled to carry out targeted cooling of local chips or areas with excessively high temperatures.

[0345] Furthermore, through the Peltier effect, the semiconductor refrigeration chip can quickly absorb the surrounding heat after power is turned on, achieving local point cooling, effectively reducing the temperature of key parts, and preventing the overheating problem from further deteriorating.

[0346] Furthermore, if the prediction results show that the AI server faces a serious overheating risk, that is, the temperature in a large number of areas far exceeds the safety threshold and may cause irreversible damage to the server hardware, multiple cooling measures will be taken to work together.

[0347] Furthermore, dynamic voltage and frequency adjustment is enabled through the server's power management system to reduce the server's operating voltage and frequency, thereby reducing the server's power consumption and heat generation.

[0348] Furthermore, the working power of the semiconductor refrigeration chip is greatly improved, its cooling capacity is enhanced, and the entire server is cooled more powerfully.

[0349] Furthermore, the speed of the fan at the condensing end of the zoned heat pipe is further increased to the maximum, so as to give full play to the heat dissipation efficiency of the heat pipe and air convection.

[0350] Furthermore, through this series of combined operations, the AI server can be globally cooled, quickly reducing the overall temperature of the server and ensuring the safe operation of the server hardware.

[0351] In general, the three-dimensional coupling model integrates real-time temperature data, power consumption distribution maps and global thermal network information to fully reflect the complex relationship between heat, electricity and power within the AI server.

[0352] In general, by using this model to perform short-term rolling predictions on AI servers, we can accurately determine the temperature change trends in different areas of the server, know in advance whether there is an overheating risk and the degree of overheating, and provide a reliable basis for temperature control.

[0353] In general, based on the prediction results, the system can flexibly adjust the embedded cooling module for different overheating risk levels.

[0354] In general, when there is no risk of overheating, maintain the current state of the module to avoid unnecessary energy consumption; when it is slightly overheated, increase the fan speed of the zoned heat pipe condensing end to enhance air cooling and heat dissipation; when it is moderately overheated, activate the semiconductor refrigeration chip for local point cooling and precise temperature reduction; when it is severely overheated, start dynamic voltage and frequency adjustment, increase the power of the semiconductor refrigeration chip and increase the fan speed for global cooling.

[0355] In general, this hierarchical control strategy makes heat dissipation measures more targeted and effectively addresses various heating situations.

[0356] In general, precise temperature control can ensure that each chip in the AI server operates within an appropriate temperature range, avoiding problems such as performance degradation, shortened lifespan, and even hardware damage caused by overheating.

[0357] In general, through intelligent regulation of the embedded heat dissipation module, a stable operating environment of the server can be maintained, the reliability and stability of the server can be improved, and AI computing tasks can be carried out continuously and efficiently.

[0358] Reference Figure 2 FIG2 is a flow chart of a method for packaging a multi-heat source zoned heat pipe array for an AI server according to an embodiment of the present invention. In this embodiment, the method for packaging a multi-heat source zoned heat pipe array for an AI server includes:

[0359] S1. Dynamically configure a sensor acquisition cycle for an AI server with an embedded distributed thermocouple array to obtain a dynamic temperature sensing network of the AI server, and acquire real-time temperature data of the AI server based on the dynamic temperature sensing network.

[0360] S2. Creating an adaptive non-uniform grid for the AI server based on pre-acquired real-time power consumption data, fitting a mapping relationship between current and heat generation in the AI server based on the adaptive non-uniform grid, and generating a power consumption distribution map of the AI server based on the mapping relationship and a dynamic order reduction algorithm.

[0361] S3. Perform cluster analysis on the heat source distribution of the AI server to obtain a global heat network of the AI server;

[0362] S4. Constructing a three-dimensional coupling model of the AI server based on the real-time temperature data, the power consumption distribution map, and the global thermal network;

[0363] S5. Construct a multi-heat source zoned heat pipe array for the AI server based on the global heat network, and perform interface packaging between the chip of the AI server and the multi-heat source zoned heat pipe array to obtain an embedded heat dissipation module for the AI server.

[0364] S6. Control the temperature of the AI server on the embedded heat dissipation module based on the three-dimensional coupling model.

[0365] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0366] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0367] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A multi-heat source zoned heat pipe array packaging system for AI servers, characterized by: The system includes a real-time temperature data acquisition module, a power consumption distribution map acquisition module, a global thermal network acquisition module, a three-dimensional coupling model construction module, an embedded heat dissipation module generation module, and an AI server temperature control module, wherein: A real-time temperature data acquisition module, configured to dynamically configure a sensor acquisition cycle for an AI server embedded with a distributed thermocouple array, obtain a dynamic temperature sensing network for the AI server, and obtain real-time temperature data of the AI server based on the dynamic temperature sensing network; A power consumption distribution map acquisition module, configured to read the real-time power consumption data of the AI server through the SVID protocol; A global interpolation field of the AI server is calculated based on the real-time power consumption data, wherein the global interpolation field is calculated as follows: Where, is an adaptive non-uniform grid, is the three-dimensional coordinate point of the chip of the AI server, is the total number of interpolation values of the full-field interpolation field, is the interpolation number of the full-field interpolation field, is the weight of the interpolation of the full field interpolation field affected by temperature, is the Gaussian kernel decay function, is the real-time power consumption influencing factor in the real-time power consumption data, is the spatial distance between two interpolations of the full-field interpolation field, are the three-dimensional coordinate points where the embedded distributed thermocouple array has been deployed, is the attenuation factor of heat transfer; Meshing the AI server based on the global interpolation field to obtain an adaptive non-uniform grid of the AI server; Extracting a busbar layer current density distribution of the AI server based on the adaptive non-uniform grid; generating a Joule heat increment field of the AI server based on the adaptive non-uniform grid; Performing electromagnetic-thermal coupling on the busbar layer current density distribution and the Joule heat increment field to obtain an electromagnetic-thermal correlation of the AI server; generating a mapping relationship between current and heat generation in the AI server based on the electromagnetic-thermal correlation; Constructing a 2000-order state space model of the AI server based on the mapping relationship; The 2000-order state space model is subjected to hierarchical feature truncation and order reduction based on a dynamic order reduction algorithm to obtain a hierarchical feature vector of the AI server, wherein the dynamic order reduction algorithm is as follows: Where, is the hierarchical feature vector obtained after reducing the order of the 2000-order state space model, is the 2000-order state-space model, is the inverted heat conduction matrix, is the heat conduction matrix, is the matrix inversion flag, is the hierarchical feature vector, To reduce the error of the reconstructed temperature field, is the balance parameter between thermophysical constraints and statistical properties, is the potential characteristic distribution of the temperature field in the 2000-order state-space model, is a regularization constraint on the potential characteristics of the temperature field in the 2000-order state-space model, The global KL constraint ignores the importance of local hot spots. is the potential characteristic of the temperature field, is the original temperature field in the 2000-order state-space model, is the reconstructed temperature field in the 2000-order state-space model; Performing power consumption feature fusion on the hierarchical feature vectors to obtain a power consumption distribution map of the AI server; A global heat network acquisition module, configured to perform cluster analysis on the heat source distribution of the AI server to obtain the global heat network of the AI server; a three-dimensional coupling model construction module, configured to construct a three-dimensional coupling model of the AI server based on the real-time temperature data, the power consumption distribution map, and the global thermal network; An embedded heat dissipation module generation module is used to construct a multi-heat source zoned heat pipe array for the AI server based on the global heat network, and perform interface packaging between the chip of the AI server and the multi-heat source zoned heat pipe array to obtain an embedded heat dissipation module for the AI server; The AI server temperature control module is used to control the temperature of the AI server by the embedded heat dissipation module based on the three-dimensional coupling model.

2. The multi-heat source zoned heat pipe array packaging system for AI server according to claim 1, characterized in that: The real-time temperature data acquisition module is specifically used to: Embedding miniaturized K-type thermocouples densely packed in a hexagonal pattern on a packaging substrate of the AI server to obtain an embedded distributed thermocouple array of the AI server; Dynamically configuring an acquisition cycle for the embedded distributed thermocouple array to obtain a 3D temperature field model of the AI server; The 3D temperature field model is superimposed on the digital twin of the physical substrate to obtain a dynamic temperature sensing network of the AI server.

3. The multi-heat source zoned heat pipe array packaging system for AI server according to claim 1, characterized in that: When the global heat network acquisition module performs cluster analysis on the heat source distribution of the AI server to obtain the global heat network of the AI server, it is specifically used to: performing thermodynamic feature extraction on the heat source distribution of the AI server to obtain thermodynamic data of the heat source distribution; Performing cluster analysis on the thermodynamic data to obtain high-temperature cluster data of the AI server; Performing grid division on the heat source distribution to obtain a grid of the heat source distribution; The grid is labeled based on the high temperature cluster data: If the high temperature cluster data of the grid exceeds a preset high temperature threshold, it is marked as a temperature sensitive grid; If the high temperature cluster data of the grid is lower than the preset high temperature threshold, it is marked as a temperature insensitive grid; The temperature-sensitive grids and the temperature-insensitive grids are aggregated into a global heat network for the AI server.

4. The multi-heat source zoned heat pipe array packaging system for AI server according to claim 1, characterized in that: When the three-dimensional coupling model construction module constructs the three-dimensional coupling model of the AI server based on the real-time temperature data, the power consumption distribution map, and the global thermal network, it is specifically used to: Performing time series alignment processing on the power consumption distribution map and the real-time temperature data to obtain standard data of the AI server; Constructing an unstructured tetrahedral mesh model of the physical structure of the AI server based on the global thermal network to obtain a thermal gradient regional model of the AI server; Allocating a multi-physics field coupling node to a high current density area in the power consumption distribution map to obtain an initial model of the AI server; performing an electrothermal coupling analysis on the AI server based on the multi-physics field coupling nodes of the initial model and the standard data to obtain an electrothermal coupling model of the AI server; A three-dimensional coupling model of the AI server is constructed based on the electrothermal coupling model and the thermal gradient region model.

5. The multi-heat source zoned heat pipe array packaging system for AI server according to claim 4, characterized in that: When the three-dimensional coupling model construction module constructs the three-dimensional coupling model of the AI server based on the electrothermal coupling model and the thermal gradient region model, the three-dimensional coupling model construction module is specifically used to: Defining a coupling interface in a thermal gradient key region of the thermal gradient region model to obtain an overlapping domain coupling interface of the thermal gradient region model; extracting a global temperature residual when the thermal gradient regional model is iterated; extracting the heat flow residual when the electrothermal coupling model is iterated; When the global temperature residual is less than a thermal gradient threshold and the heat flow residual is less than an electrothermal threshold, a converged thermal gradient regional model and a converged electrothermal coupling model are obtained; Based on the overlapping domain coupling interface, the converged thermal gradient region model and the converged electrothermal coupling model are thermal-electrical-mechanical cross-scale coupled to obtain a three-dimensional coupling model of the AI server.

6. The multi-heat source zoned heat pipe array packaging system for AI server according to claim 1, characterized in that: The embedded heat dissipation module generation module, when executing the construction of the multi-heat source zoned heat pipe array of the AI server based on the global heat network, and performing interface packaging between the chip of the AI server and the multi-heat source zoned heat pipe array to obtain the embedded heat dissipation module of the AI server, is specifically used to: Extracting heat source features from the global heat network to obtain heat source features of the AI server; Arranging heat pipes of the AI server based on the heat source characteristics to obtain a multi-heat source zoned heat pipe array of the AI server; Gradient brazing is performed on the multi-heat source zoned heat pipe array and the chip of the AI server to obtain an initial heat dissipation module of the AI server; The initial heat dissipation module is hermetically packaged in multiple layers to obtain the embedded heat dissipation module of the AI server.

7. The multi-heat source zoned heat pipe array packaging system for AI server according to claim 1, characterized in that: When the AI server temperature control module performs temperature control of the AI server on the embedded heat dissipation module based on the three-dimensional coupling model, the AI server temperature control module is specifically configured to: Performing short-term rolling prediction on the AI server based on a three-dimensional coupling model; When the prediction result indicates that the AI server has no overheating risk, the embedded heat dissipation module maintains the current working state; When the prediction result indicates that the AI server is at a slight overheating risk, the embedded heat dissipation module increases the fan speed of the condensing end of the zoned heat pipe; When the prediction result indicates that the AI server has a moderate overheating risk, the embedded heat dissipation module activates the semiconductor cooling chip to perform local point cooling on the AI server; When the prediction result indicates that the AI server is at serious risk of overheating, the embedded heat dissipation module starts dynamic voltage and frequency adjustment, increases the power of the semiconductor refrigeration plate, increases the fan speed of the condensing end of the zoned heat pipe, and performs global cooling on the AI server.

8. A multi-heat source partitioned heat pipe array packaging method for AI servers, characterized in that: The method comprises: S1. Dynamically configure a sensor acquisition cycle for an AI server with an embedded distributed thermocouple array to obtain a dynamic temperature sensing network of the AI server, and acquire real-time temperature data of the AI server based on the dynamic temperature sensing network. S2. Read the real-time power consumption data of the AI server through the SVID protocol; A global interpolation field of the AI server is calculated based on the real-time power consumption data, wherein the global interpolation field is calculated as follows: Where, is an adaptive non-uniform grid, is the three-dimensional coordinate point of the chip of the AI server, is the total number of interpolation values of the full-field interpolation field, is the interpolation number of the full-field interpolation field, is the weight of the interpolation of the full field interpolation field affected by temperature, is the Gaussian kernel decay function, is the real-time power consumption influencing factor in the real-time power consumption data, is the spatial distance between two interpolations of the full-field interpolation field, are the three-dimensional coordinate points where the embedded distributed thermocouple array has been deployed, is the attenuation factor of heat transfer; Meshing the AI server based on the global interpolation field to obtain an adaptive non-uniform grid of the AI server; Extracting a busbar layer current density distribution of the AI server based on the adaptive non-uniform grid; generating a Joule heat increment field of the AI server based on the adaptive non-uniform grid; Performing electromagnetic-thermal coupling on the busbar layer current density distribution and the Joule heat increment field to obtain an electromagnetic-thermal correlation of the AI server; generating a mapping relationship between current and heat generation in the AI server based on the electromagnetic-thermal correlation; Constructing a 2000-order state space model of the AI server based on the mapping relationship; The 2000-order state space model is subjected to hierarchical feature truncation and order reduction based on a dynamic order reduction algorithm to obtain a hierarchical feature vector of the AI server, wherein the dynamic order reduction algorithm is as follows: Where, is the hierarchical feature vector obtained after reducing the order of the 2000-order state space model, is the 2000-order state-space model, is the inverted heat conduction matrix, is the heat conduction matrix, is the matrix inversion flag, is the hierarchical feature vector, To reduce the error of the reconstructed temperature field, is the balance parameter between thermophysical constraints and statistical properties, is the potential characteristic distribution of the temperature field in the 2000-order state-space model, is a regularization constraint on the potential characteristics of the temperature field in the 2000-order state-space model, The global KL constraint ignores the importance of local hot spots. is the potential characteristic of the temperature field, is the original temperature field in the 2000-order state-space model, is the reconstructed temperature field in the 2000-order state-space model; Performing power consumption feature fusion on the hierarchical feature vectors to obtain a power consumption distribution map of the AI server; S3. Perform cluster analysis on the heat source distribution of the AI server to obtain a global heat network of the AI server; S4. Constructing a three-dimensional coupling model of the AI server based on the real-time temperature data, the power consumption distribution map, and the global thermal network; S5. Construct a multi-heat source zoned heat pipe array for the AI server based on the global heat network, and perform interface packaging between the chip of the AI server and the multi-heat source zoned heat pipe array to obtain an embedded heat dissipation module for the AI server. S6. Control the temperature of the AI server on the embedded heat dissipation module based on the three-dimensional coupling model.

Citation Information

Patent Citations

  • Modeling subsurface processes on unstructured grid

    CN101903803A

  • Method and system for measuring temperature and power distributions of a device in a package

    US20060039114A1