Microfluidic network heat dissipation system based on intelligent predictive control and control method thereof

CN122803157APending Publication Date: 2026-09-22TIANFU YONGXING LAB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611274117.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-21
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

现有热管理方案主要分为两类:一类是在芯片顶部外表面加装液冷冷板,通过导热硅脂和均温盖板将热量传递至冷却液,但该方案在3D堆叠架构下面临多层串联热阻的制约,热量需依次穿透上层芯片、导热界面材料和盖板方能到达冷板,导致底层发热源结温过高;另一类是在玻璃基板内部集成玻璃金属过孔(TGV)与中空微流体通道(TGC),将冷却液引入封装内部以缩短导热路径,但该方案在控制层面仍采用恒定流量或简单开环控制,冷却液流量不随芯片实际发热状态动态调节

Benefits of technology

[0057]1、第一,本发明通过在玻璃基板内部嵌设玻璃毛细管散热网络,并将逻辑裸片和/或高带宽内存堆叠体直接倒装于玻璃基板上方,使发热体与冷却液之间的距离缩短至数十微米量级,能够大幅地缩短导热路径,有效地降低芯片结温。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122803157A_ABST
    Figure CN122803157A_ABST
Patent Text Reader

Abstract

The application provides a microfluid network heat dissipation system based on intelligent predictive control and a control method thereof, and relates to the technical field of chip thermal management. The system comprises a glass substrate, a laser fusion micro-manifold, a distributed temperature sensing array, a flow regulation actuator array, an intelligent control unit, and a logic die and / or a high-bandwidth memory stack. The logic die and / or the high-bandwidth memory stack are inverted above the glass substrate. The laser fusion micro-manifold is bonded to the bottom or side edge of the glass substrate. The distributed temperature sensing array is embedded in the glass substrate. The intelligent control unit is signal-connected with the distributed temperature sensing array and the flow regulation actuator array. The intelligent control unit comprises a digital twin thermal model module and a model predictive control decision module. The application changes the passive response after the temperature exceeds the standard to active intervention before the temperature exceeds the standard, effectively eliminates the control lag and temperature overshoot, and can run in real time on a low-cost embedded platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of chip thermal management technology, specifically to a microfluidic network heat dissipation system and its control method based on intelligent predictive control. Background Technology

[0002] With the exponential growth in computing power demands for high-performance computing and large-scale AI model training, chip power consumption has climbed to the 1000W or even 1500W level. 2.5D / 3D heterogeneous packaging technology has become an important path to continue Moore's Law. Existing thermal management solutions are mainly divided into two categories: one is to install a liquid cooling plate on the top outer surface of the chip, transferring heat to the coolant through thermal grease and a vapor chamber cover. However, this solution faces the constraint of multi-layer series thermal resistance in 3D stacked architectures. Heat must penetrate through the upper chip, thermal interface material, and cover plate in sequence to reach the cold plate, resulting in excessively high junction temperatures of the bottom heat source. The other is to integrate glass-metal vias (TGV) and hollow microfluidic channels (TGC) inside the glass substrate to introduce coolant into the package to shorten the heat conduction path. However, this solution still uses constant flow or simple open-loop control at the control level, and the coolant flow rate is not dynamically adjusted according to the actual heating state of the chip.

[0003] Existing microfluidic cooling systems suffer from at least the following technical deficiencies at the control level: First, they employ passive response control strategies, increasing cooling flow only after the temperature exceeds the limit, resulting in inherent control lag and temperature overshoot. Second, they lack the ability to predict chip temperature evolution trends, making it impossible to implement proactive thermal management for millisecond-level power fluctuations in AI chips. Third, in 3D heterogeneous packaging, multiple heat sources coexist, such as logic dies and high-bandwidth memory stacks, with varying heat dissipation characteristics, making it difficult for existing single-loop control architectures to coordinate the thermal coupling effects between multiple heat sources. Fourth, existing model predictive control algorithms have high computational complexity, making it difficult to achieve millisecond-level real-time decision-making on low-cost embedded microcontrollers.

[0004] The aforementioned defects result in poor heat dissipation efficiency and poor chip temperature uniformity in existing solutions when facing dynamic loads, which restricts the reliability of 3D heterogeneous packaging in high power density scenarios. Summary of the Invention

[0005] To address the problems in related technologies, this invention provides a microfluidic network heat dissipation system and its control method based on intelligent predictive control.

[0006] To achieve the above objectives, the technical solution adopted by the present invention includes:

[0007] According to a first aspect of the present invention, a microfluidic network heat dissipation system based on intelligent predictive control is provided, comprising:

[0008] A glass substrate, which has a solid glass-metal via array and a glass capillary heat dissipation network embedded inside, the glass capillary heat dissipation network including multiple hollow microfluidic channels;

[0009] At least one logic die and / or a high-bandwidth memory stack is flip-chipped over the glass substrate;

[0010] Laser-fused micromanifolds are bonded to the bottom or side edge of the glass substrate and communicate with the glass capillary for distributing coolant.

[0011] A distributed temperature sensing array is embedded in different areas inside the glass substrate to monitor the temperature distribution of each area of ​​the chip in real time.

[0012] An array of flow regulating actuators is integrated at the outlet of each branch channel of the micromanifold to independently regulate the flow rate of coolant entering each glass capillary branch.

[0013] An intelligent control unit is signal-connected to the distributed temperature sensor array and the flow regulation actuator array. The intelligent control unit includes a digital twin thermal model module and a model predictive control decision module. The intelligent control unit is configured to: predict future temperature evolution trends based on real-time temperature measurements from the distributed temperature sensor array using the digital twin thermal model module, and solve for the optimal coolant flow distribution command using the model predictive control decision module, and output the command to the flow regulation actuator array.

[0014] Optionally, the distributed temperature sensing array is an M×N two-dimensional temperature sensing array, where M≥5, N≥5, the sensor spacing is no greater than 2mm, and the time resolution is no greater than 1ms.

[0015] Optionally, the digital twin thermal model module uses discrete-time state-space equations to describe the temperature evolution dynamics of the chip's thermal system:

[0016]

[0017]

[0018] In the formula, for The system state vector at any given time includes the temperature and rate of temperature change of each region. for The system state vector at any given time. for The input vector, i.e., the coolant flow rate of each glass capillary branch, is controlled at all times. for Time-process noise vector, Here is the state transition matrix. To control the input matrix, for The output vector at any given time is the temperature measured by the sensor. For the output matrix, for Measure the noise vector at any time.

[0019] Optionally, the model predictive control decision module performs the following optimization in each control cycle:

[0020] The objective function is minimized based on the following criteria:

[0021]

[0022] In the formula, To predict the number of time-domain steps, To control the number of time-domain steps, In order to be in Time prediction The output vector at any given time is the predicted temperature value. For the target temperature vector, In order to be in Time prediction The rate of change of the control quantity at all times , In order to be in Time prediction The constant-time control input vector, i.e., the predicted coolant flow rate of each glass capillary branch, This is the weighted matrix for temperature tracking deviation. To control the rate of change penalty weighting matrix, The weighting matrix is ​​used to control the amplitude penalty.

[0023] And it satisfies the following constraints:

[0024] Control amplitude constraints: ;

[0025] Control quantity change rate constraint: ;

[0026] Hard constraint on upper temperature limit: ;

[0027] In the formula, Let this be the lower limit vector of the flow rate for each glass capillary branch. Let the upper limit vector of the flow rate of each glass capillary branch be denoted as . This represents the upper limit vector of the flow rate change rate for each glass capillary branch. This is the upper limit vector of the chip's rated operating temperature.

[0028] Optionally, the model predictive control decision module adopts an offline-online hybrid computing architecture:

[0029] In the offline phase, the solution to the optimization problem is pre-computed as a piecewise affine function:

[0030]

[0031] In the formula, This is the optimal control quantity. For piecewise affine functions, the state vector and target temperature The mapping is to the optimal control quantity; the piecewise affine function is stored in the read-only memory of the intelligent control unit in the form of a lookup table or decision tree;

[0032] During the online phase, the intelligent control unit queries the piecewise affine function stored in the read-only controller based on the current sensor readings to obtain the optimal control quantity.

[0033] Optionally, the piecewise affine function is specifically:

[0034]

[0035] In the formula, This is the optimal control quantity. For augmented state vectors and The units of its components are based on and The components are determined; This represents the total number of state space partitions. This is the affine gain matrix for the first partition. This is the affine gain matrix for the second partition. For the first Affine gain matrix for each partition; Let be the affine bias vector for the first partition. This is the affine bias vector for the second partition. For the first Affine bias vectors for each partition; , , ..., Partition the state space.

[0036] Optionally, the model predictive control decision module employs an adaptive predictive time-domain reduction method, the first... Prediction time steps for each temperature monitoring area Determined by the following function:

[0037]

[0038] In the formula, For adaptive prediction time-domain mapping function, For the first The thermal time constant of each temperature monitoring area For the first The deviation between the current temperature and the target temperature in each temperature monitoring area For the first The current temperature of each temperature monitoring area For the first The target temperature of each temperature monitoring area For the first The current rate of change norm of the temperature monitoring zone control input.

[0039] According to a second aspect of the present invention, a control method for a microfluidic network heat dissipation system based on intelligent predictive control is also provided, applied to the microfluidic network heat dissipation system based on intelligent predictive control described in any of the technical solutions of the first aspect of the present invention, the method comprising the following steps:

[0040] Step S1: Real-time acquisition of current temperature distribution data in various regions of the chip through a distributed temperature sensing array embedded inside the glass substrate;

[0041] Step S2: Input the temperature distribution data into the intelligent control unit and update the state vector of the digital twin thermal model through a Kalman filter;

[0042] Step S3: Based on the updated state vector, use the digital twin thermal model to predict the temperature evolution trajectory of each region in the future prediction time domain;

[0043] Step S4: With the optimization objectives of minimizing temperature tracking deviation, minimizing the rate of change of control variable, and minimizing total coolant flow rate, and under the premise of satisfying the control variable amplitude constraint, control variable rate of change constraint, and temperature upper limit constraint, solve the optimal coolant flow rate distribution sequence in the control time domain.

[0044] Step S5: Output the control command for the current moment in the optimal coolant flow distribution sequence to the flow regulating actuator array to adjust the coolant flow of each glass capillary branch.

[0045] Optionally, the method for solving the optimal coolant flow distribution sequence in step S4 is as follows: a hybrid calculation method combining offline pre-calculation and online table lookup is adopted. In the offline stage, the solution to the optimization problem is pre-calculated as a piecewise affine control law on the state space partition through multi-parameter quadratic programming and stored as a lookup table. In the online stage, the optimal control quantity is obtained by looking up the table based on the current state vector and the target temperature. When the system state falls in the boundary region of the lookup table entry or encounters a working condition not covered in the offline stage, the online quadratic programming solver is started to solve the problem.

[0046] Optionally, the control method further includes:

[0047] During chip operation, control input data and temperature sensor output data are continuously collected. A recursive least squares algorithm with a forgetting factor is used to update the state transition matrix and control input matrix of the digital twin thermal model online.

[0048]

[0049]

[0050]

[0051] In the formula, for The estimated value of the parameter vector to be identified at time step. for The estimated value of the parameter vector to be identified at time step. for Time-regression vector, for transpose, for Kalman gain vector at time step for Output vector at each time step for Time-varying covariance matrix for Time-varying covariance matrix Forgetting factor and This is used to control the rate at which historical data decays.

[0052] Calculate the relative change between the current model parameters and the calibration model parameters. :

[0053]

[0054] In the formula, This is the current online identification state transition matrix. This is the initially calibrated state transition matrix. It is the Frobenius norm;

[0055] when When the preset threshold is exceeded, the model re-identification process is triggered to generate an updated digital twin thermal model, and the piecewise affine control law is adjusted according to the updated model parameters.

[0056] Beneficial effects:

[0057] 1. First, the present invention embeds a glass capillary heat dissipation network inside the glass substrate and directly flip-chips logic dies and / or high-bandwidth memory stacks on top of the glass substrate, thereby shortening the distance between the heat source and the coolant to the order of tens of micrometers, which can significantly shorten the heat conduction path and effectively reduce the chip junction temperature.

[0058] Secondly, this invention uses a distributed temperature sensing array to sense the chip temperature field in real time and combines a digital twin model to predict future temperature evolution trends. The model predicts and controls the decision-making module to output the optimal coolant flow distribution command in advance before the temperature exceeds the limit. This can fundamentally solve the passive response defect of the traditional solution of "temperature rise first, then regulation" and eliminate the problems of control lag and temperature overshoot.

[0059] Third, this invention independently adjusts the coolant flow rate of each glass capillary branch by means of a flow regulating actuator array, and combines it with the multi-input multi-output optimization solution of the intelligent control unit to achieve differentiated cooling for different heat sources such as logic dies and high-bandwidth memory stacks, effectively reducing the temperature non-uniformity on the chip surface.

[0060] Fourth, this invention adopts a hybrid computing architecture that combines offline pre-computation of piecewise affine control laws with online table lookup, reducing the computational load of a single decision from traditional MPC online solution of nonlinear optimization to the level of table lookup operations. This enables model predictive control algorithms to achieve millisecond-level real-time decision-making on low-cost embedded microcontrollers, demonstrating feasibility for industrial-grade deployment and engineering economics.

[0061] Fifth, this invention continuously updates the parameters of the digital twin model online through a recursive least squares algorithm with a forgetting factor, and triggers model re-identification and control law adjustment by monitoring changes in model parameters. This enables the control system to adapt to slow time-varying factors such as chip aging and coolant property drift, ensuring the control accuracy and reliability of the heat dissipation system throughout its entire life cycle.

[0062] 2. Other beneficial effects or advantages of the present invention will be described in detail in the specific embodiments. Attached Figure Description

[0063] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0064] in:

[0065] Figure 1 This is a schematic diagram of liquid cooling solutions in existing related technologies;

[0066] Figure 2 This is a schematic diagram of the arrangement of a microfluidic network heat dissipation system based on intelligent predictive control, provided in an exemplary embodiment of the present invention.

[0067] Figure 3 This is a schematic diagram of microchannel distribution provided in an exemplary embodiment of the present invention. Figure 3 In the diagram, 'a' represents the overall size of the microchannel structure, and 'p' represents the spacing between the microchannels.

[0068] Figure 4 This is a schematic diagram comparing the thermal resistance of the present invention with that of traditional heat dissipation methods;

[0069] Figure 5 This is a schematic diagram of a controller offline calibration and online adaptive update method provided in an exemplary embodiment of the present invention;

[0070] Figure 6 This is a flowchart illustrating step 1 of an exemplary embodiment of the present invention;

[0071] Figure 7 This is a flowchart illustrating step 2 of an exemplary embodiment of the present invention;

[0072] Figure 8 This is a flowchart illustrating step 3 of an exemplary embodiment of the present invention;

[0073] Figure 9 This is a flowchart illustrating step 4 of an exemplary embodiment of the present invention;

[0074] Figure 10 This is a flowchart illustrating step 5 of an exemplary embodiment of the present invention;

[0075] Figure 11 This is a flowchart illustrating step 6 of an exemplary embodiment of the present invention;

[0076] Figure 12 This is a flowchart illustrating step 7 of an exemplary embodiment of the present invention.

[0077] Explanation of the labels in the attached drawings:

[0078] 100 - Glass substrate; 101 - Logic die; 102 - High-bandwidth memory stack; 103 - Server motherboard; 201 - Solid glass-metal via; 202 - Hollow microfluidic channel; 300 - Laser-welded micromanifold; Detailed Implementation

[0079] To facilitate a clearer and more accurate understanding of the technical solutions of this invention by those skilled in the art, the existing related technologies and their technical problems will be described in more detail below.

[0080] The demand for computing power in training large models, exemplified by ChatGPT, is growing exponentially, giving rise to AI chips with ultra-high heat flux, consuming up to 1000W or even 1500W. In order to accommodate more transistors and high-speed memory within a limited area, the industry widely adopts 2.5D / 3D heterogeneous packaging architecture, stacking logic computing chips and high-bandwidth memory (HBM) side by side or vertically through silicon interposers.

[0081] However, the resulting three-dimensional thermal barrier has become the biggest bottleneck to improving computing power. For example... Figure 1 As shown, existing liquid cooling solutions mainly focus on heat dissipation on the top outer surface of the chip, i.e., pressing a cold plate onto the top cover plate of the chip. For 3D stacked chips, heat must penetrate from the bottom heat source through the upper chip, multiple layers of thermal grease, a homogeneous copper cover plate, and a second layer of thermal grease before finally exchanging heat with the coolant. This layered series thermal resistance results in extremely high internal temperatures of the bottom silicon wafer, causing not only severe thermal throttling and frequency reduction but also exacerbating thermal stress warping and even deformation between heterogeneous materials, leading to the failure of the cold plate's seal.

[0082] In recent years, research has proposed integrating microchannels and through-glass vias (TGVs) within glass substrates to achieve in-package liquid cooling, shortening the heat conduction path at the hardware level. However, existing solutions still suffer from the following common drawbacks at the control level:

[0083] First, passive response control. Existing solutions mostly use constant flow or simple open-loop control, where the coolant flow rate is not dynamically adjusted according to the actual heat generation of the chip. When the chip's workload suddenly increases, the flow rate is only passively increased after the temperature has already exceeded the limit, resulting in serious control lag and temperature overshoot problems.

[0084] Second, it lacks the ability to predict thermal evolution. Existing solutions only perform feedback control based on the current temperature and cannot predict temperature change trends in the next few seconds. For the millisecond-level power fluctuations of AI chips, traditional feedback control struggles to achieve forward-looking thermal management.

[0085] Third, there is a lack of multi-heat source coupling control. In 3D heterogeneous packaging, there are multiple heat sources such as logic dies and HBM stacks. The heating characteristics, time constants and spatial distributions of each heat source are different. Existing single-loop control is difficult to coordinate the thermal coupling effect between multiple heat sources.

[0086] Fourth, the control algorithm has low computational efficiency. Existing model predictive control methods face problems such as high computational complexity, large memory consumption, and poor real-time performance when implemented on microcontrollers, making it difficult to meet the requirements for millisecond-level heat dissipation response.

[0087] In view of this, the present invention provides a novel solution: the microfluidic network heat dissipation system and its control method based on intelligent predictive control. Under the premise of keeping the TGV-TGC hardware architecture embedded in the glass substrate unchanged, the present invention achieves advanced prediction and dynamic optimal allocation of coolant flow by deploying a distributed temperature sensing network in the package, establishing a digital twin model of the chip thermal system, and running a lightweight MPC algorithm. This fundamentally solves the problems of lag, passivity and low energy efficiency in heat dissipation control in 3D heterogeneous packaging.

[0088] The technical concept of this invention lies in addressing the passive response deficiency of existing microfluidic heat dissipation systems in 3D heterogeneous packaging, which rely on "temperature rise first, then regulation." It proposes an integrated intelligent heat dissipation technology solution encompassing "sensing-prediction-decision-execution." Specifically, a distributed temperature sensor array is deployed within the glass substrate to achieve real-time sensing of the chip's three-dimensional temperature field. A digital twin model is used to predict the thermal evolution trends of various regions of the chip in advance. A lightweight model predictive control (MPC) decision module, under multiple constraints such as flow rate amplitude, rate of change, and temperature upper limit, continuously solves for the optimal coolant distribution scheme for each glass capillary branch. The core of this invention is to upgrade traditional "feedback cooling" to "predictive thermal management," that is, deploying cooling resources in advance based on prediction results before the temperature exceeds the limit, fundamentally eliminating control lag and temperature overshoot problems. Meanwhile, to enable the MPC algorithm to run in real time on a low-cost embedded platform, this invention further adopts a hybrid computing architecture that combines offline pre-computation of piecewise affine control laws with online table lookup. While ensuring control performance, this reduces the computational load of a single decision from the online QP solution of traditional MPC to the level of table lookup operations, thus making this invention feasible for industrial-grade deployment and economically viable.

[0089] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.

[0090] like Figure 2 As shown, this embodiment provides a microfluidic network heat dissipation system based on intelligent predictive control according to a first aspect of the present invention, comprising:

[0091] A glass substrate, which has a solid glass-metal via array and a glass capillary heat dissipation network embedded inside. The glass capillary heat dissipation network includes multiple hollow microfluidic channels.

[0092] At least one logic die and / or a high-bandwidth memory stack is flip-chip mounted on a glass substrate;

[0093] Laser-fused micromanifolds are bonded to the bottom or side edge of a glass substrate and connected to glass capillaries for distributing coolant.

[0094] A distributed temperature sensing array is embedded in different areas inside the glass substrate to monitor the temperature distribution in different areas of the chip in real time.

[0095] An array of flow regulating actuators is integrated at the outlet of each branch channel of the micromanifold to independently regulate the flow rate of coolant entering each glass capillary branch.

[0096] The intelligent control unit is connected to the distributed temperature sensor array and the flow regulation actuator array. The intelligent control unit includes a digital twin thermal model module and a model predictive control decision module. The intelligent control unit is configured to: predict the future temperature evolution trend based on the real-time temperature measurement value of the distributed temperature sensor array through the digital twin thermal model module, and solve the optimal coolant flow distribution command by the model predictive control decision module, and output it to the flow regulation actuator array.

[0097] First, this invention effectively shortens the heat conduction path and reduces the chip junction temperature. Specifically, in existing top liquid cooling solutions, heat must sequentially penetrate the upper chip, multiple layers of thermal grease, and a vapor chamber before reaching the cold plate. This series thermal resistance is extremely high, resulting in the internal temperature of the bottom heat source being much higher than the coolant temperature. This invention, however, by embedding a glass capillary heat dissipation network inside the glass substrate and directly flip-chipping the logic die and / or high-bandwidth memory stack onto the glass substrate, shortens the distance between the heat source and the coolant to the tens of micrometers, thus significantly reducing the heat conduction path. Simultaneously, by combining laser-welded micromanifolds to introduce coolant into the branches of the glass capillary, close-range heat exchange between the coolant and the heat source can be achieved, thereby effectively reducing the chip junction temperature.

[0098] Secondly, this invention enables proactive thermal management, eliminating control lag and temperature overshoot. Specifically, existing microfluidic cooling systems often employ constant flow or simple open-loop control, where the coolant flow rate is not dynamically adjusted according to the chip's real-time heating status. When the chip load suddenly increases, the flow rate is passively increased only after the temperature has already exceeded the limit, resulting in severe control lag and temperature overshoot. In this invention, a distributed temperature sensor array embedded in different regions within the glass substrate enables real-time sensing of the temperature distribution in various areas of the chip. An intelligent control unit containing a digital twin thermal model module can predict future temperature trends based on current temperature measurements. Furthermore, a model predictive control decision module can proactively make decisions and output optimal coolant flow distribution commands before the temperature exceeds the limit. Thus, through the coordinated operation of the sensor array, the twin model, and the MPC decision-making, the system of this invention possesses the ability to "sense the present, predict the future, and make proactive decisions," fundamentally solving the control lag problem of traditional solutions that rely on "temperature rise first, then adjustment."

[0099] Third, this invention enables independent and precise control of multiple heat sources, improving temperature uniformity. Specifically, 3D heterogeneous packaging contains multiple heat sources, including logic dies and high-bandwidth memory stacks, each with different heating characteristics, time constants, and spatial distributions. This invention, by integrating a flow regulating actuator array at the outlet of each branch channel of the micromanifold, can independently regulate the coolant flow rate entering each glass capillary branch. Combined with an intelligent control unit that solves for the optimal coolant flow rate allocation command, differentiated cooling for different regions and heat sources can be achieved. Thus, the system of this invention can precisely allocate more cooling flow to high heat flux density areas and less flow to low heat flux density areas based on the real-time temperature status and predicted evolution trend of each region, thereby effectively reducing temperature non-uniformity on the chip surface.

[0100] Fourth, this invention can form a complete closed-loop control chain of perception-prediction-decision-execution, realizing on-demand cooling and energy efficiency optimization. Specifically, in existing solutions, sensing, decision-making, and execution are isolated from each other, lacking a complete closed loop from temperature sensing to flow regulation. This invention, however, constructs a complete closed-loop control chain from "real-time temperature monitoring" to "future trend prediction," then to "optimal flow solution," and finally to "independent execution of each branch" through the signal connection relationship between a distributed temperature sensor array, an intelligent control unit (including a digital twin model module and a model predictive control decision module), and a flow regulation actuator array. In this way, through this closed-loop mechanism, the coolant flow rate can be dynamically adjusted according to the actual heat dissipation state of the chip, achieving "on-demand cooling" rather than "constant cooling," thus meeting the chip's heat dissipation needs while avoiding unnecessary coolant pump power consumption.

[0101] In one embodiment of the present invention, the distributed temperature sensing array of the present invention can be a two-dimensional temperature sensing array of M×N, wherein M≥5, N≥5, the sensor spacing is not greater than 2mm, and the time resolution is not greater than 1ms.

[0102] In existing temperature monitoring solutions, the sparse sensor layout makes it difficult to capture subtle changes in the chip's temperature field, leading to distorted temperature data used for control decisions and affecting the accuracy of regulation. In this implementation, a distributed temperature sensor array (an M×N two-dimensional temperature sensor array, M≥5, N≥5, with a sensor spacing of no more than 2mm) ensures that the sensor network has sufficient spatial resolution for the temperature distribution in different areas of the chip. Simultaneously, the temporal resolution is no greater than 1ms, enabling the sensor network to capture millisecond-level dynamic changes in the chip's temperature field. This provides real-time, high-fidelity input data for accurate state updates of the digital twin model and advanced predictions by the MPC decision module, thereby improving the overall regulation accuracy of the control system.

[0103] In one embodiment of the present invention, the digital twin thermal model module of the present invention uses discrete-time state-space equations to describe the temperature evolution dynamics of the chip thermal system:

[0104]

[0105]

[0106] In the formula, for The system state vector at any given time includes the temperature and rate of temperature change of each region. for The system state vector at any given time. for The input vector, i.e., the coolant flow rate of each glass capillary branch, is controlled at all times. for Time-process noise vector, Here is the state transition matrix. To control the input matrix, for The output vector at any given time is the temperature measured by the sensor. For the output matrix, for Measure the noise vector at any time.

[0107] In this implementation, through the state transition matrix Describe the inherent dynamic characteristics and control input matrix of a thermal system. Describing the relationship between coolant flow rate and temperature enables a concise mathematical characterization of the dynamic behavior of a multi-region thermally coupled chip system, providing a quantitative computational basis for the MPC decision module to be used for prediction and optimization.

[0108] In one embodiment of the present invention, the model predictive control decision module performs the following optimization in each control cycle:

[0109] The objective function is minimized based on the following criteria:

[0110]

[0111] In the formula, To predict the number of time-domain steps, To control the number of time-domain steps, In order to be in Time prediction The output vector at any given time is the predicted temperature value. For the target temperature vector, In order to be in Time prediction The rate of change of the control quantity at all times , In order to be in Time prediction The constant-time control input vector, i.e., the predicted coolant flow rate of each glass capillary branch, This is the weighted matrix for temperature tracking deviation. To control the rate of change penalty weighting matrix, The weighting matrix is ​​used to control the amplitude penalty.

[0112] And it satisfies the following constraints:

[0113] Control amplitude constraints: ;

[0114] Control quantity change rate constraint: ;

[0115] Hard constraint on upper temperature limit: ;

[0116] In the formula, Let this be the lower limit vector of the flow rate for each glass capillary branch. Let the upper limit vector of the flow rate of each glass capillary branch be denoted as . This represents the upper limit vector of the flow rate change rate for each glass capillary branch. This is the upper limit vector of the chip's rated operating temperature.

[0117] In this implementation, the MPC decision module (Model Predictive Control Decision Module) is defined to minimize the objective function, which includes a weighted sum of three terms—temperature tracking deviation, control variable rate of change, and control variable amplitude—in each control cycle, while simultaneously satisfying three constraints: control variable amplitude constraint, control variable rate of change constraint, and a hard constraint on the upper temperature limit. Specifically, the temperature tracking deviation term drives the system to approach the target temperature as close as possible; the control variable rate of change penalty term suppresses drastic flow fluctuations to protect the actuator; and the control variable amplitude penalty term automatically minimizes the total flow to achieve energy savings while meeting heat dissipation requirements. Thus, through the synergy of these three terms, multi-objective optimization of "precise temperature control, stable adjustment, and on-demand cooling" can be achieved. The three constraints ensure that the instructions do not exceed the physical capabilities of the actuator and that the chip temperature does not exceed safe limits.

[0118] In one embodiment of the present invention, the model predictive control decision module adopts an offline-online hybrid computing architecture:

[0119] In the offline phase, the solution to the optimization problem is pre-computed as a piecewise affine function:

[0120]

[0121] In the formula, This is the optimal control quantity. For piecewise affine functions, the state vector and target temperature The mapping is to the optimal control quantity; the piecewise affine function is stored in the read-only memory of the intelligent control unit in the form of a lookup table or decision tree;

[0122] During the online phase, the intelligent control unit queries the piecewise affine function stored in the read-only controller based on the current sensor readings to obtain the optimal control quantity.

[0123] Existing MPC algorithms suffer from high computational complexity when solving constrained quadratic programming problems online, making it difficult to achieve millisecond-level real-time decision-making on low-cost embedded platforms. In this implementation, the solution to the MPC optimization problem is pre-calculated as a piecewise affine function and stored in the form of a lookup table or decision tree during the offline phase. This transforms the high-dimensional optimization problem, which originally required online solution, into pre-calculated data in read-only memory. During the online phase, the optimal control input can be obtained simply by querying the stored piecewise affine function based on the current sensor readings. This reduces the online computational load from nonlinear optimization to simple table lookup or tree traversal operations, enabling the MPC algorithm to achieve millisecond-level real-time decision-making on low-cost embedded microcontrollers.

[0124] In one embodiment of the present invention, the piecewise affine function of the present invention is specifically:

[0125]

[0126] In the formula, This is the optimal control quantity. For augmented state vectors and The units of its components are based on and The components are determined; This represents the total number of state space partitions. This is the affine gain matrix for the first partition. This is the affine gain matrix for the second partition. For the first Affine gain matrix for each partition; Let be the affine bias vector for the first partition. This is the affine bias vector for the second partition. For the first Affine bias vectors for each partition; , , ..., Partition the state space.

[0127] In this implementation, the augmented state vector system state vector and target temperature They are all included in the space of independent variables. Each partition completely covers the entire possible working state space, and each partition corresponds to an affine expression. Among them, the number of partitions It can be flexibly configured according to the model complexity, avoiding the shortcomings of a single linear control law that cannot adapt to all operating conditions. At the same time, the partitioned structure can ensure the integrity of the state space partitioning, so that only judgment is needed in the online phase. Once a region falls into a certain partition, the corresponding affine expression can be directly substituted to calculate the optimal control quantity.

[0128] In one embodiment of the present invention, the model predictive control decision module employs an adaptive predictive time-domain reduction method. Prediction time steps for each temperature monitoring area Determined by the following function:

[0129]

[0130] In the formula, For adaptive prediction time-domain mapping function, For the first The thermal time constant of each temperature monitoring area For the first The deviation between the current temperature and the target temperature in each temperature monitoring area For the first The current temperature of each temperature monitoring area For the first The target temperature of each temperature monitoring area For the first The current rate of change norm of the temperature monitoring zone control input.

[0131] Existing MPC solutions employ a fixed prediction time domain. However, in 3D heterogeneous packaging, the thermal time constant of the logic die core region is short (milliseconds), while the thermal time constant of the HBM stack region is long (hundreds of milliseconds). A fixed prediction time domain makes it difficult to simultaneously meet the dual requirements of fast response and advanced planning. In this implementation, the thermal time constant of each region... This reflects the rate of thermal response in the region and the deviation between the current temperature and the target temperature. This reflects the severity of the deviation from the expected value in the region, and controls the current rate of change norm of the input. This reflects the level of activity in the control behavior; all three factors serve as input, and are processed by an adaptive mapping function. Dynamically determine the prediction time-domain steps for each region In this way, regions with short thermal time constants can use shorter prediction time domains to achieve fast response, while regions with long thermal time constants can use longer prediction time domains to achieve advanced planning. This allows for minimizing overall computational overhead while ensuring control performance in each region, and avoiding unnecessary long-time domain calculations in fast-response regions.

[0132] According to a second aspect of the present invention, a control method for a microfluidic network heat dissipation system based on intelligent predictive control is also provided, applicable to the microfluidic network heat dissipation system based on intelligent predictive control according to any one of the technical solutions in the first aspect of the present invention, the method comprising the following steps:

[0133] Step S1: Real-time acquisition of current temperature distribution data in various regions of the chip through a distributed temperature sensing array embedded inside the glass substrate;

[0134] Step S2: Input the temperature distribution data into the intelligent control unit and update the state vector of the digital twin thermal model through the Kalman filter;

[0135] Step S3: Based on the updated state vector, use the digital twin thermal model to predict the temperature evolution trajectory of each region in the future prediction time domain;

[0136] Step S4: With the optimization objectives of minimizing temperature tracking deviation, minimizing the rate of change of control variable, and minimizing total coolant flow rate, and under the premise of satisfying the control variable amplitude constraint, control variable rate of change constraint, and temperature upper limit constraint, solve the optimal coolant flow rate distribution sequence in the control time domain.

[0137] Step S5: Output the control command for the current moment in the optimal coolant flow distribution sequence to the flow regulating actuator array to adjust the coolant flow of each glass capillary branch.

[0138] Through this technical solution, firstly, the control method of the present invention can achieve accurate state estimation, laying a data foundation for predictive control. Specifically, in existing solutions, temperature sensor measurements inevitably contain noise interference. Directly using noisy measurements for control decisions leads to frequent actuator movements and decreased system stability. In the method of the present invention, step S2 fuses the current temperature measurement with the predicted value from the digital twin model using a Kalman filter to estimate the system state vector in a statistically optimal sense. This effectively filters out the influence of measurement noise on state estimation, providing accurate and reliable initial state values ​​for subsequent temperature prediction and optimization decisions, ensuring the robustness and decision accuracy of the control system.

[0139] Second, the control method of this invention can achieve forward-looking thermal evolution prediction and eliminate control lag. Specifically, existing control schemes only passively increase cooling flow after the temperature exceeds the limit, resulting in severe control lag in scenarios with millisecond-level power fluctuations in AI chips. In the method of this invention, step S3 uses a digital twin thermal model to forward-calculate the temperature evolution trajectory of each region in the future prediction time domain based on the updated state vector, enabling the system to "predict" the temperature change trend in the next few seconds. Simultaneously, step S4 solves for the optimal control sequence based on this, allowing the system to deploy cooling resources in advance based on the prediction results before the temperature exceeds the limit. This achieves a fundamental shift from "passive response" to "forward-looking thermal management," fundamentally eliminating the inherent control lag and temperature overshoot defects in traditional feedback control.

[0140] Third, the control method of this invention can achieve multi-objective collaborative optimization, taking into account temperature control accuracy, actuator lifespan, and system energy efficiency. Specifically, existing single-loop control often sacrifices one aspect for another when adjusting flow rate, making it difficult to simultaneously meet multiple performance requirements. In the method of this invention, step S4 minimizes the temperature tracking deviation to drive the chip temperature to approach the target value to ensure temperature control accuracy, minimizes the rate of change of the control quantity to suppress drastic fluctuations in flow rate to extend the lifespan of actuators such as piezoelectric valves, and minimizes the total coolant flow to avoid unnecessary pump power consumption and improve system energy efficiency. The three optimization objectives are synergistically balanced within a unified optimization framework, enabling the system of this invention to automatically find the overall optimal operating point while meeting heat dissipation requirements, avoiding the problem of sacrificing one aspect for another in single-objective optimization.

[0141] Fourth, the control method of this invention can ensure the safe operation of the system and the physical feasibility of the actuator. Specifically, existing solutions lack systematic constraint verification of control commands, which may output commands that exceed the physical capabilities of the actuator or cause the chip to overheat. In the method of this invention, step S4 simultaneously applies control amplitude constraints to ensure that the output flow command does not exceed the physical adjustment range of each branch regulating valve; applies control rate of change constraints to ensure that the flow adjustment rate does not exceed the response capability of the piezoelectric valve, avoiding sudden command changes that could cause the actuator to lose synchronization or be damaged; and applies a hard constraint on the upper limit of temperature to forcibly limit the chip temperature to the rated safe range at all times. In this way, even under transient conditions, the temperature is not allowed to exceed the limit, providing a hard guarantee for the operational safety of the chip and the system from the control law level.

[0142] Fifth, the control method of this invention can achieve rolling time-domain control, continuously adapting to dynamically changing operating conditions. Specifically, if an optimization solution is not updated after a single iteration (i.e., open-loop execution), it will deviate from the optimal trajectory due to model mismatch and external disturbances under dynamically changing load conditions. In the method of this invention, step S5 only outputs the current time-domain instruction of the optimal control sequence to the actuator, and step S6 re-executes all steps in the next control cycle. Thus, through this "short-sighted but continuously updated" rolling time-domain strategy, the system can re-evaluate the state, predict the temperature, and optimize the solution based on the latest temperature measurement value in each control cycle. This allows for continuous tracking of real-time changes in chip load and online correction of control instructions, ensuring that the control system maintains its adaptive capability to dynamic operating conditions throughout its entire lifecycle.

[0143] In one embodiment of the present invention, the method for solving the optimal coolant flow distribution sequence in step S4 can be as follows: a hybrid calculation method combining offline pre-calculation and online table lookup is adopted. In the offline stage, the solution of the optimization problem is pre-calculated as a piecewise affine control law on the state space partition through multi-parameter quadratic programming and stored as a lookup table. In the online stage, the optimal control quantity is obtained by looking up the table according to the current state vector and the target temperature. When the system state falls in the boundary region of the lookup table entry or encounters a working condition not covered in the offline stage, the online quadratic programming solver is started to solve the problem.

[0144] In this implementation, during the offline phase, the solution to the optimization problem is pre-calculated as a piecewise affine control law on the state space partition through multi-parameter quadratic programming and stored as a lookup table. This transforms the high-dimensional optimization problem that originally required online solution into pre-calculated data in read-only memory, reducing the computational load of a single decision from nonlinear optimization to the level of a lookup table operation. This enables the real-time deployment of the MPC algorithm on a low-cost embedded microcontroller. Simultaneously, when the system state falls within the boundary region of a lookup table entry or encounters a condition not covered in the offline phase, an online quadratic programming solver is activated to solve the problem. This ensures that the control system can obtain effective control quantities under various non-standard operating conditions, avoiding the control failure problems that may occur with pure lookup table schemes, and improving the robustness and applicability of the system.

[0145] In one embodiment of the present invention, the control method of the present invention may further include:

[0146] During chip operation, control input data and temperature sensor output data are continuously collected. A recursive least squares algorithm with a forgetting factor is used to update the state transition matrix and control input matrix of the digital twin thermal model online.

[0147]

[0148]

[0149]

[0150] In the formula, for The estimated value of the parameter vector to be identified at time step. for The estimated value of the parameter vector to be identified at time step. for Time-regression vector, for transpose, for Kalman gain vector at time step for Output vector at each time step for Time-varying covariance matrix for Time-varying covariance matrix Forgetting factor and This is used to control the rate at which historical data decays.

[0151] Calculate the relative change between the current model parameters and the calibration model parameters. :

[0152]

[0153] In the formula, This is the current online identification state transition matrix. This is the initially calibrated state transition matrix. It is the Frobenius norm;

[0154] when When the preset threshold is exceeded, the model re-identification process is triggered to generate an updated digital twin thermal model, and the piecewise affine control law is adjusted according to the updated model parameters.

[0155] In this implementation, the online adaptive update mechanism of the digital twin model employs a recursive least squares algorithm with a forgetting factor to continuously update the state transition matrix and control input matrix during chip operation. This algorithm only needs to save the estimation results and covariance matrix from the previous time step, eliminating the need to store all historical data. It has low computational complexity and low storage resource consumption, enabling real-time operation even with idle computing power in the ICU. The forgetting factor... By applying exponential weighted decay to historical data, the algorithm can adaptively "forget" outdated old data and quickly track changes in the thermal characteristics of the chip caused by slow time-varying factors such as electromigration aging and coolant property drift, thereby maintaining the prediction accuracy of the digital twin model throughout its entire life cycle.

[0156] At the same time, the relative changes between the current model parameters and the calibration model parameters are calculated. ,when When the preset threshold is exceeded, the model re-identification process is triggered and the piecewise affine control law is adjusted accordingly. This forms a two-layer adaptive architecture that combines "online continuous correction" and "significant change re-identification" to ensure the matching of the control system and the actual thermal characteristics of the chip throughout the entire life cycle.

[0157] The technical solution of the present invention will be further described below with reference to an exemplary embodiment.

[0158] First, it should be noted that the mathematical symbols, physical meanings, and units used in this exemplary embodiment are as follows:

[0159] express The system state vector at any given time includes the temperature and rate of temperature change of each region, in units of... and ;

[0160] express The constant-time control input vector, i.e., the coolant flow rate of each TGC branch, is expressed in units of... ; express The output vector at any given time is the temperature measured by the sensor, in °C.

[0161] This represents the state transition matrix, which describes the inherent dynamic characteristics of the thermal system. This represents the control input matrix, describing the effect of coolant flow rate on temperature. The output matrix is ​​dimensionless.

[0162] and They represent The time-process noise vector and the measurement noise vector have different units for their components;

[0163] This indicates the dimension of the state vector; This indicates the dimension of the control input vector (number of TGC branches). This indicates that the dimension of the output vector (the number of temperature sensors) is dimensionless;

[0164] This represents the value of the objective function for MPC optimization. Indicates the number of prediction steps in the time domain; Indicates the number of control time-domain steps; all are dimensionless.

[0165] This represents the target temperature vector, in °C.

[0166] express Rate of change of control quantity at any time: The unit is ,

[0167] This represents the temperature tracking deviation weighting matrix (positive definite matrix). This represents the weighted matrix for the penalty of the rate of change of control (positive definite matrix). This represents the control amplitude penalty weighting matrix (positive definite matrix), used for energy-saving optimization; all are dimensionless.

[0168] This represents the lower limit vector of traffic for each TGC branch; This represents the upper limit vector of traffic for each TGC branch; This represents the upper limit vector of the rate of change of traffic for each TGC branch; the unit is... ;

[0169] This vector represents the upper limit of the chip's rated operating temperature, in °C.

[0170] express The optimal control sequence at time, including One future control command; Indicates in Time prediction Optimal control quantity at any given time; all units are... ;

[0171] This represents a piecewise affine function that maps the state to the optimal control variable;

[0172] Representing the augmented state vector: The units of each component are different; This represents the number of entries in the lookup table; it is dimensionless.

[0173] Indicates the first The affine gain matrix of each partition, its reason The quantity varies depending on the unit of measurement; Indicates the first The affine bias vector for each partition, in units of ; Indicates the first Each state space partition (convex polyhedron); Indicates the total number of state space partitions;

[0174] Indicates the first The number of adaptive prediction time-domain steps for each temperature monitoring area, dimensionless;

[0175] Indicates the first The thermal time constant of each region, in units of ; Indicates the first The deviation between the current temperature and the target temperature in each area: The unit is ℃; Indicates the first The current rate of change norm of each region's control input, in units of ;

[0176] Represents the adaptive prediction time-domain mapping function;

[0177] express The estimated value of the parameter vector to be identified at any given time, the unit of each component depends on the identification structure; express The time-regression vector is composed of historical input and output data, and the unit of each component depends on the identification structure. express The Kalman gain vector at each time step, with the unit of each component determined by the identification structure; express The time-varying covariance matrix, the units of each component are determined according to the identification structure; Represents the forgetting factor, which controls the rate of decay of historical data and ;

[0178] This represents the relative change in model parameters; Represents the Forbenius norm (the square root and norm of a matrix). This represents the current online identification state transition matrix; This represents the initial calibration state transition matrix;

[0179] Indicates the first The current temperature of each temperature monitoring zone; Indicates the first The target temperature of each temperature monitoring area; all units are in °C. This represents the rate of temperature change, expressed in °C / s.

[0180] Represents the Hessian matrix in quadratic programming; Represents the vector of linear terms in a quadratic programming problem; Represents the partition constraint inequality matrix; This represents the vector of partition constraint inequalities.

[0181] I. System Composition.

[0182] Please see Figure 2 It includes:

[0183] The glass substrate 100 has an extremely low coefficient of thermal expansion and excellent high-frequency insulation.

[0184] The logic die 101, the core high-heat computing power source (such as the GPU core), is directly flip-chip soldered to the top of the glass substrate 100 via microbumps.

[0185] The high-bandwidth memory stack 102 is deployed adjacent to the logic die 101 and is also flip-chip mounted on the glass substrate 100.

[0186] The server motherboard 103 carries the lowest level of power supply and external communication, and is connected to the glass substrate 100 through BGA solder balls.

[0187] Solid glass metal via (TGV) 201, filled with electroplated copper, is responsible for delivering the current and signals of the server motherboard 103 vertically upward through the glass substrate 100 to the logic die 101 and the high-bandwidth memory stack 102.

[0188] The hollow microfluidic channel (TGC) 202, the core heat dissipation component of this invention, consists of numerous vertically or laterally interconnected hollow glass capillary tubes etched by a femtosecond laser within the gaps of the solid glass metal via array 201. High-pressure single-phase coolant flows through its interior.

[0189] Laser-fused micromanifolds 300 are tightly bonded to the bottom or side edge of glass substrate 100, serving as a pressure reduction and flow equalization hub for external CDU coolant to flow into hollow microfluidic channels 202.

[0190] A distributed temperature sensing array embeds miniature thin-film thermocouple temperature sensors in different areas within the glass substrate 100, with a sensor spacing of no more than 2 mm, forming a distributed temperature monitoring network covering the entire chip projection area. Sensor signals are transmitted to the bottom control unit via tiny metal leads inside the glass substrate.

[0191] Intelligent Control Unit (ICU). Integrated on the server motherboard 103 or deployed independently as a dedicated control chip, it includes a data acquisition module, a digital twin model module, an MPC decision module, and an execution drive module. The ICU is connected to the distributed temperature sensor array 401, the CDU water pump, and the flow regulating valve in the laser-welded micromanifold 300, forming a complete perception-prediction-decision-execution closed-loop control link.

[0192] The flow regulating actuator array, integrated at the outlet of each branch channel of the laser-welded micromanifold 300, includes miniature piezoelectric valves that receive control commands from the ICU and independently regulate the flow of coolant entering each TGC branch.

[0193] II. System Description.

[0194] This embodiment uses a special glass as the substrate, which possesses high dimensional stability, low thermal expansion coefficient matching, and high electrical insulation. A mesh-like cavity is created within the glass, which is tens to hundreds of micrometers thick, through femtosecond laser modification and wet etching. Please refer to [link to relevant documentation]. Figure 3 Inside the glass substrate, there are numerous solid metal pillars (electroplated copper) running vertically from top to bottom. The TGV is responsible for vertically transmitting the power and high-speed data signals from the bottom PCB motherboard to the top logic die and HBM stack. The diameter of the TGV is set between 10μm and 20μm to ensure low RF signal insertion loss.

[0195] In the gaps between the TGV metal array, this implementation alternately etches numerous vertically connected, hollow channel clusters (TGCs). High-specific-heat-capacity single-phase coolant (special deionized water or electronic fluorinated liquid) flows through the TGCs. These TG channels are directly adjacent to or exposed near the microbumps on the bottom of the upper silicon wafer, forming a heat dissipation network in direct contact with the heat source. The TGVs and TGCs are arranged asymmetrically and interwoven. Solid black dots represent TGVs carrying electrical signals, while hollow blue circles represent TGCs carrying coolant. They are physically isolated by glass walls tens of micrometers thick, achieving effective encapsulation and fusion. Furthermore, laser-welded bottom micromanifolds are located directly below or on the side edge of the glass substrate. To safely and leak-free inject high-pressure coolant generated by an external water pump into the TGC array, this system uses a laser-welded process to seamlessly bond a 3D-printed distribution manifold to the surface of the glass substrate, acting as a hub for fluid inflow and outflow.

[0196] Based on the aforementioned hardware architecture, this implementation further adds the following intelligent control layer components to the system:

[0197] 1. Distributed temperature sensing module.

[0198] Inside the glass substrate 100, directly beneath each functional block of the logic die 101 and each layer of the high-bandwidth memory stack 102, a micro-thin-film thermocouple temperature sensor is embedded, forming... Two-dimensional temperature sensing array ( The sensor is fabricated using a metal thin-film deposition process compatible with TGV technology, with a thickness of no more than 2μm, which does not affect the mechanical strength and electrical insulation properties of the glass substrate. The sensor has a spatial resolution of no more than 2mm and a temporal resolution of no more than 1ms, enabling it to capture minute changes in the temperature field on the chip surface in real time.

[0199] 2. Digital Twin Thermal Model Module.

[0200] Inside the ICU, a reduced-order digital twin model of the chip's thermal system is pre-built. This model, based on heat transfer principles and system identification technology, can accurately describe the temperature evolution dynamics of various regions of the chip with extremely low computational cost.

[0201] Digital twin models are described using discrete-time state-space equations:

[0202]

[0203]

[0204] State transition matrix and control input matrix The initial values ​​are obtained through a system identification method. During the chip calibration phase, a series of known power and flow perturbations are applied, temperature response data is collected, and the values ​​are estimated using a subspace identification algorithm or prediction error method. and The value. Subsequently, during the actual operation of the chip, and The system continuously corrects itself through an online adaptive update mechanism to cope with slow time-varying factors such as chip aging and changes in coolant properties.

[0205] 3. Lightweight Model Predictive Control (MPC) Decision Module.

[0206] Based on a digital twin model, the MPC decision module performs the following optimization calculations in each control cycle (typically 10ms to 50ms):

[0207] Step 1: Read the temperature measurement values ​​of each area at the current time. The state vector of the digital twin model is updated using a Kalman filter. .

[0208] Step 2: Based on the current state Using digital twin models to predict the time domain (Typical values: 5 to 20 steps, corresponding to 50 ms to 1000 ms) Temperature evolution trajectory of each region.

[0209] Step 3: In the control time domain ( Within the given constraints, solve the following optimization problem:

[0210]

[0211] In the formula, This represents the optimal control quantity calculated at time k and applied at time k. This represents the result calculated at the current time k, which affects the future k-th time. The optimal control quantity at any given time;

[0212] The constraints are as follows:

[0213] Control amplitude constraints:

[0214]

[0215] Control quantity change rate constraint:

[0216]

[0217] Hard constraint on upper temperature limit:

[0218]

[0219] Step 4: Solve the above optimization problem to obtain the optimal control sequence in the control time domain:

[0220] In the formula, This means that the constraint must apply to every moment in the set;

[0221] Step 5: Only the first element of the optimal control sequence The output to the actuator (i.e., the flow control valve of each TGC branch) is recalculated in the next control cycle, which is the rolling time domain control strategy.

[0222] III. Efficient MPC solution method.

[0223] The core challenge of implementing traditional MPC algorithms on embedded platforms is that each step requires solving a constrained quadratic programming (QP) problem online, resulting in high computational complexity. The computation time is on a scale of magnitude (i.e., the solution time increases cubically with the number of decision variables), making it difficult to complete within a millisecond-level control cycle. This implementation achieves a lightweight approach to the MPC algorithm through the following three technological innovations, enabling it to run in real time on low-cost ICU chips:

[0224] 1. Offline-online hybrid computing architecture.

[0225] The solution to the MPC optimization problem is decomposed into two stages: offline pre-computation and online table lookup.

[0226] Offline Phase: During the chip design / calibration phase, for all possible operating states (different power levels, different ambient temperatures, different flow boundaries), the corresponding optimal control law is pre-calculated and encoded into a piecewise affine function (PWA).

[0227]

[0228] In the formula, The optimal control variable represents the current state. and target parameters The optimal control action is obtained directly by looking up a table.

[0229] The piecewise affine function It is stored in the ICU's read-only memory in the form of a lookup table (LUT) or a decision tree.

[0230] In the online phase: During each control cycle, the ICU only needs to look up the optimal control value in a table based on the current sensor readings, reducing computational complexity from... Down to ( (Number of entries in the lookup table) This means that the query time increases logarithmically with the number of entries in the lookup table, achieving a microsecond-level decision response.

[0231] 2. Combining explicit MPC with online correction.

[0232] This implementation employs an explicit MPC method, that is, by pre-computing the solution to the optimization problem as a PWA function on the state-space partition through multi-parameter quadratic programming (mp-QP):

[0233]

[0234] Unlike pure lookup table schemes, this invention adds an online correction step to the explicit MPC: when the system state falls within the boundary region of a lookup table entry or encounters a condition not covered in the offline stage, a lightweight online QP solver (using gradient projection or alternating direction multiplier method ADMM) is activated, requiring only a few iterations to obtain a feasible solution. This hybrid strategy of "explicit MPC as the main method and online solution as a supplement" balances real-time performance and robustness.

[0235] 3. Adaptive prediction time domain reduction method.

[0236] Traditional MPC uses a fixed prediction time domain. However, in chip thermal management, the thermal time constants of different regions differ significantly. The thermal time constant of the logic die core region is short (milliseconds), while the thermal time constant of the HBM stack region is long (hundreds of milliseconds). The adaptive predictive time-domain reduction method proposed in this implementation is as follows:

[0237]

[0238] Regions with short thermal time constants use shorter prediction time domains (fast response), while regions with long thermal time constants use longer prediction time domains (advanced planning), minimizing the total computational load while ensuring control performance.

[0239] IV. Explanation of Working Principle (Please refer to the following) Figure 4 ).

[0240] 1. Thermal resistance analysis of traditional top heat dissipation.

[0241] In traditional top-mounted cooling systems, the total thermal resistance of the cooling system is... The calculation method is shown in the following formula. Heat must be transferred through layers to reach the fluid heat exchange layer, resulting in high thermal resistance:

[0242]

[0243] In the formula, The total thermal resistance of the system is expressed in K / W. The sum of the series thermal resistances of a traditional multilayer stacked structure; The thickness of each layer (e.g., silicone grease thickness, copper cover plate thickness). The inherent thermal conductivity of each layer of material; It is the cross-sectional area for heat conduction; The convective heat transfer thermal resistance represents the fluid contact surface. The convective heat transfer coefficient; This represents the wetting area of ​​the inner wall of the microchannel.

[0244] In this embodiment, the heat dissipation network at the bottom of the glass substrate is such that the heat-generating logic chip is only a few micrometers away from the coolant inside the TGC, with a series heat conduction path. With a heat exchange rate of almost zero, rapid and efficient heat exchange between the heat source and the fluid can be achieved.

[0245] 2. Microchannel convection heat transfer enhancement principle.

[0246] According to the definition of the Nusselt coefficient, the convective heat transfer coefficient increases as the fluid flows through a narrower TGC pipe. The larger:

[0247]

[0248] In the formula, For Nusselt numbers; The convective heat transfer coefficient; The hydraulic diameter of the microchannel; This is the inherent thermal conductivity of the cooling liquid.

[0249] While a smaller hydraulic diameter results in greater heat transfer intensity, it also leads to increased drag. Therefore, parallel TGC arrays must be used to distribute the flow rate, and their flow control dynamics are constrained by Poiseuille's equations.

[0250]

[0251] In the formula, The absolute pressure drop at the inlet and outlet is generated when the fluid passes through the glass substrate channel; The kinematic viscosity of the cooling fluid; The longitudinal depth length of the TGC microchannel; The volumetric flow rate that a single microchannel can handle; It is the hydraulic diameter.

[0252] Since the pressure drop is inversely proportional to the fourth power of the diameter, by increasing the number of parallel connections, the flow rate Q of a single channel can be reduced, thereby keeping the overall system pressure drop within the range that a standard CDU pump can withstand.

[0253] 3. Working principle of intelligent predictive control.

[0254] This implementation method, based on traditional hardware heat dissipation, further introduces intelligent predictive control, and its working principle is as follows:

[0255] In each control cycle, the ICU first reads the temperature distribution of each region of the chip using a distributed temperature sensor array. Subsequently, the digital twin model is based on the current state. and control input Foreseeing the future The temperature evolution trajectory is determined step by step. The MPC decision module minimizes the objective function shown in the formula, and solves for the optimal flow allocation sequence while satisfying constraints such as upper and lower flow limits, flow rate change limits, and upper temperature limits. Finally, only the optimal flow command at the current moment is processed. The flow regulation actuators distributed to each TGC branch enable on-demand distribution of coolant.

[0256] Compared to traditional feedback control (such as PID), the MPC controller of this invention is forward-looking, meaning it increases the cooling flow to hot spots in advance based on predictions before the temperature exceeds the limit, rather than responding passively after the temperature exceeds the limit. Simultaneously, its Multiple-Input Multiple-Output (MIMO) architecture can coordinate the flow distribution of all TGC branches, uniformly handling the thermal coupling effects between multiple heat sources such as logic dies and HBM stacks.

[0257] V. System Design Methodology (Controller Parameter Tuning and Digital Twin Modeling Process).

[0258] Please see Figure 5 This implementation provides a complete method for offline calibration and online adaptive update of the controller, ensuring that the MPC controller obtains optimal parameters before actual deployment.

[0259] 1. Offline calibration stage.

[0260] Step 1: Stimulation Experiment and Data Acquisition (see [link to relevant documentation]) Figure 6 ).

[0261] On the chip calibration test platform, a series of standardized power perturbation signals (such as step signals, sinusoidal sweep signals, and pseudo-random binary sequences PRBS) are applied, while each TGC branch is driven with a preset flow sequence, and the response data of the distributed temperature sensing array is recorded. The sampling frequency is no less than 1kHz, and the recording time covers the entire process of the system from the initial steady state to the final steady state.

[0262] Step 2, System Identification and Digital Twin Modeling (see [link]) Figure 7 ).

[0263] Using the collected input-output data, the state transition matrix A and control input matrix B of the state-space model are estimated using subspace identification methods (such as the N4SID algorithm) or prediction error methods (PEM). The model order n is determined by the Akaike Information Criterion (AIC) or cross-validation, typically between order 5 and 15. After identification, the model's prediction accuracy on the validation dataset is verified through simulation—the root mean square prediction error (RMSE) is required to be no greater than 1%.

[0264] Step 3: Explicit MPC offline solution (see [link]) Figure 8 ).

[0265] The identified digital twin model is then compared with the preset control weight matrices Q, R, S and constraint boundaries. , , , The solution is input into a multi-parameter quadratic programming (mp-QP) solver to calculate the state-space partitions and the corresponding PWA optimal control laws. The solution results are encoded into a lookup table and burned into the ICU's read-only memory.

[0266] Step 4, Hardware-in-the-Loop Verification (see [link]) Figure 9 ).

[0267] Connect the ICU with the programmed control law to a chip thermal simulator or actual test platform to perform full-condition traversal testing to verify the stability, robustness, and real-time performance of the control system. If the verification fails, adjust the weight matrices Q, R, and S and repeat steps 3 and 4.

[0268] 2. Online adaptive update phase.

[0269] Step 5: Online Recursive Least Squares Identification (see [link]) Figure 10 ).

[0270] During actual chip operation, the ICU continuously collects input (flow commands) and output (temperature sensor readings) data, and uses a recursive least squares algorithm with a forgetting factor (FF-RLS) to update the parameters A and B of the digital twin model online.

[0271]

[0272]

[0273]

[0274] Step 6: Model parameter variable detection (see [link]) Figure 11 ).

[0275] Calculate the relative change between the current model parameters and the calibration model parameters:

[0276]

[0277] when When the threshold is exceeded (e.g., 10%), it is determined that the system characteristics have changed significantly, triggering the model re-identification process—starting a more refined system identification algorithm in the background to generate an updated digital twin model.

[0278] Step 7: Online adjustment of the control law (see [link]). Figure 12 ).

[0279] When the digital twin model is significantly updated, the ICU recalculates the PWA control law of the explicit MPC in the background (or adjusts the lookup table entries) based on the updated model parameters, so as to realize the adaptive update of the controller and ensure that the control system always matches the actual thermal characteristics of the chip.

[0280] VI. Operating characteristics and performance.

[0281] This implementation exhibits the following significantly superior operational characteristics compared to the prior art within a heterogeneous package:

[0282] 1. Based on the breakdown resistance properties of insulating glass.

[0283] If traditional methods are used to etch water channels within the silicon interposer, silicon, being a semiconductor, is highly susceptible to electrical dendrite formation at the water-silicon interface under high-voltage DC power and high-frequency signals, leading to leakage and breakdown of the entire wafer. This invention, however, utilizes amorphous special glass with extremely high dielectric constant and breakdown voltage. Even with prolonged water flow within the TGC, the glass walls, only tens of micrometers in size, can withstand transient voltage surges, achieving zero leakage through water-electricity isolation.

[0284] 2. The risk of thermal stress warping is reduced by 3D heterogeneous stacking.

[0285] Traditional 1000W chips experience cooling at the top and heat buildup at the bottom, resulting in a temperature difference exceeding 60 degrees Celsius between the upper and lower surfaces. This vertical temperature difference causes CTE mismatch, leading to severe silicon wafer warping and potentially breaking the BGA solder balls at the bottom. This invention addresses this by injecting a constant-temperature coolant into the bottom layer of a glass substrate, allowing the heat-generating base to be in near-direct contact with the coolant, minimizing the vertical temperature gradient within the chip. Therefore, regardless of varying computing power conditions, the wafer package maintains its flatness, significantly improving the physical yield and mechanical durability limits of heterogeneous packaging.

[0286] 3. Proactive thermal management to eliminate control lag.

[0287] Traditional feedback control (such as PID) only increases cooling flow after the temperature exceeds the limit, resulting in inherent control lag. The MPC controller of this invention predicts future temperature evolution based on a digital twin model, increasing cooling flow to hotspot areas before the temperature exceeds the limit. During the transient process of a step change in AI chip power, this invention can reduce the peak temperature by 8°C to 12°C compared to PID control, and reduce temperature overshoot by more than 60%.

[0288] 4. Coordinated control of multiple heat sources to achieve uniform temperature across the entire area.

[0289] In 3D heterogeneous packaging, the thermal characteristics of logic dies and HBM stacks differ significantly. The MPC controller of this invention, through a multiple-input multiple-output (MIMO) architecture, simultaneously regulates the flow distribution of all TGC branches, coordinating the thermal coupling effects between heat sources. Under full load conditions, the chip surface temperature non-uniformity can be reduced from ±15℃ in conventional solutions to within ±3℃.

[0290] 5. Pump power consumption is significantly reduced, and energy efficiency ratio is improved.

[0291] This invention introduces a control magnitude weighting term into the objective function. This invention automatically minimizes the total coolant flow rate while meeting temperature constraints, achieving on-demand cooling. Compared to a constant maximum flow rate scheme, this invention can reduce CDU pump power consumption by 30% to 50%. Compared to traditional PID control, the forward-looking nature of MPC can reduce the average flow rate by 15% to 25%.

[0292] 6. Breakthrough in computational efficiency, enabling real-time operation on low-cost ICUs.

[0293] By employing an offline-online hybrid computing architecture and adaptive prediction time-domain reduction, the MPC algorithm of this invention can complete a full decision cycle within 10ms on an ARM Cortex-M4 microcontroller, consuming less than 20 KB of RAM and less than 100 KB of ROM, without relying on high-performance DSPs or FPGAs.

[0294] 7. Adaptive modeling to address chip aging and operating condition drift.

[0295] The online adaptive update mechanism of the digital twin model enables the control system to maintain modeling accuracy throughout the entire life cycle of the chip—even if the chip's thermal characteristics change due to electromigration aging, or the coolant's physical properties drift due to long-term use, the system can still automatically correct the model parameters and maintain optimal control performance.

[0296] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A microfluidic network heat dissipation system based on intelligent predictive control, characterized in that, include: A glass substrate, which has a solid glass-metal via array and a glass capillary heat dissipation network embedded inside, the glass capillary heat dissipation network including multiple hollow microfluidic channels; At least one logic die and / or a high-bandwidth memory stack is flip-chipped over the glass substrate; Laser-fused micromanifolds are bonded to the bottom or side edge of the glass substrate and communicate with the glass capillary for distributing coolant. A distributed temperature sensing array is embedded in different areas inside the glass substrate to monitor the temperature distribution of each area of ​​the chip in real time. An array of flow regulating actuators is integrated at the outlet of each branch channel of the micromanifold to independently regulate the flow rate of coolant entering each glass capillary branch. An intelligent control unit is signal-connected to the distributed temperature sensor array and the flow regulation actuator array. The intelligent control unit includes a digital twin thermal model module and a model predictive control decision module. The intelligent control unit is configured to: predict future temperature evolution trends based on real-time temperature measurements from the distributed temperature sensor array using the digital twin thermal model module, and solve for the optimal coolant flow distribution command using the model predictive control decision module, and output the command to the flow regulation actuator array.

2. The microfluidic network heat dissipation system based on intelligent predictive control according to claim 1, characterized in that, The distributed temperature sensing array is a two-dimensional temperature sensing array of M×N, where M≥5, N≥5, the sensor spacing is no greater than 2mm, and the time resolution is no greater than 1ms.

3. The microfluidic network heat dissipation system based on intelligent predictive control according to claim 1, characterized in that, The digital twin thermal model module uses discrete-time state-space equations to describe the temperature evolution dynamics of the chip's thermal system. , In the formula, for The system state vector at any given time includes the temperature and rate of temperature change of each region. for The system state vector at any given time. for The input vector, i.e., the coolant flow rate of each glass capillary branch, is controlled at all times. for Time-process noise vector, Here is the state transition matrix. To control the input matrix, for The output vector at any given time is the temperature measured by the sensor. For the output matrix, for Measure the noise vector at any time.

4. The microfluidic network heat dissipation system based on intelligent predictive control according to claim 3, characterized in that, The model predictive control decision module performs the following optimization in each control cycle: The objective function is minimized based on the following criteria: In the formula, To predict the number of time-domain steps, To control the number of time-domain steps, In order to be in Time prediction The output vector at any given time is the predicted temperature value. For the target temperature vector, In order to be in Time prediction The rate of change of the control quantity at all times , In order to be in Time prediction The constant-time control input vector, i.e., the predicted coolant flow rate of each glass capillary branch, This is the weighted matrix for temperature tracking deviation. To control the rate of change penalty weighting matrix, The weighting matrix is ​​used to control the amplitude penalty. And it satisfies the following constraints: Control amplitude constraints: ; Control variable change rate constraint: ; Hard constraint on upper temperature limit: ; In the formula, Let this be the lower limit vector of the flow rate for each glass capillary branch. Let the upper limit vector of the flow rate of each glass capillary branch be denoted as . This represents the upper limit vector of the flow rate change rate for each glass capillary branch. This is the upper limit vector of the chip's rated operating temperature.

5. The microfluidic network heat dissipation system based on intelligent predictive control according to claim 4, characterized in that, The model prediction control decision module adopts an offline-online hybrid computing architecture: In the offline phase, the solution to the optimization problem is pre-computed as a piecewise affine function: In the formula, This is the optimal control quantity. For piecewise affine functions, the state vector and target temperature The mapping is to the optimal control quantity; the piecewise affine function is stored in the read-only memory of the intelligent control unit in the form of a lookup table or decision tree; During the online phase, the intelligent control unit queries the piecewise affine function stored in the read-only controller based on the current sensor readings to obtain the optimal control quantity.

6. The microfluidic network heat dissipation system based on intelligent predictive control according to claim 5, characterized in that, The piecewise affine function is specifically: In the formula, This is the optimal control quantity. For augmented state vectors and The units of its components are based on and The components are determined; This represents the total number of state space partitions. This is the affine gain matrix for the first partition. This is the affine gain matrix for the second partition. For the first Affine gain matrix for each partition; Let be the affine bias vector for the first partition. This is the affine bias vector for the second partition. For the first Affine bias vectors for each partition; , , ..., Partition the state space.

7. The microfluidic network heat dissipation system based on intelligent predictive control according to claim 4, characterized in that, The model predictive control decision module employs an adaptive predictive time-domain reduction method. Prediction time steps for each temperature monitoring area Determined by the following function: In the formula, For adaptive prediction time-domain mapping function, For the first The thermal time constant of each temperature monitoring area For the first The deviation between the current temperature and the target temperature in each temperature monitoring area For the first The current temperature of each temperature monitoring area For the first The target temperature of each temperature monitoring area For the first The current rate of change norm of the temperature monitoring zone control input.

8. A control method for a microfluidic network heat dissipation system based on intelligent predictive control, characterized in that, The method applied to the microfluidic network heat dissipation system based on intelligent predictive control according to any one of claims 1-7 includes the following steps: Step S1: Real-time acquisition of current temperature distribution data in various regions of the chip through a distributed temperature sensing array embedded inside the glass substrate; Step S2: Input the temperature distribution data into the intelligent control unit and update the state vector of the digital twin thermal model through a Kalman filter; Step S3: Based on the updated state vector, use the digital twin thermal model to predict the temperature evolution trajectory of each region in the future prediction time domain; Step S4: With the optimization objectives of minimizing temperature tracking deviation, minimizing the rate of change of control variable, and minimizing total coolant flow rate, and under the premise of satisfying the control variable amplitude constraint, control variable rate of change constraint, and temperature upper limit constraint, solve the optimal coolant flow rate distribution sequence in the control time domain. Step S5: Output the control command for the current moment in the optimal coolant flow distribution sequence to the flow regulating actuator array to adjust the coolant flow of each glass capillary branch.

9. The control method for a microfluidic network heat dissipation system based on intelligent predictive control according to claim 8, characterized in that, The method for solving the optimal coolant flow distribution sequence in step S4 is as follows: a hybrid calculation method combining offline pre-calculation and online table lookup is adopted. In the offline stage, the solution of the optimization problem is pre-calculated as a piecewise affine control law on the state space partition through multi-parameter quadratic programming and stored as a lookup table. In the online stage, the optimal control quantity is obtained by looking up the table according to the current state vector and the target temperature. When the system state falls within the boundary region of the lookup table entry or encounters a working condition not covered in the offline stage, the online quadratic programming solver is started to solve the problem.

10. The control method for a microfluidic network heat dissipation system based on intelligent predictive control according to claim 8, characterized in that, The control method further includes: During chip operation, control input data and temperature sensor output data are continuously collected. A recursive least squares algorithm with a forgetting factor is used to update the state transition matrix and control input matrix of the digital twin thermal model online. , , In the formula, for The estimated value of the parameter vector to be identified at time step. for The estimated value of the parameter vector to be identified at time step. for Time-regression vector, for transpose, for Kalman gain vector at time step for Output vector at each time step for Time-varying covariance matrix for Time-varying covariance matrix Forgetting factor and This is used to control the rate at which historical data decays. Calculate the relative change between the current model parameters and the calibration model parameters. : In the formula, This is the current online identification state transition matrix. This is the initially calibrated state transition matrix. It is the Frobenius norm; when When the preset threshold is exceeded, the model re-identification process is triggered to generate an updated digital twin thermal model, and the piecewise affine control law is adjusted according to the updated model parameters.