A heat dissipation method and system for a microfluidic chip based on a reinforcement learning algorithm
By applying reinforcement learning algorithms in the microfluidic chip heat dissipation system, the water flow velocity and microfluidic pump power consumption are accurately controlled, and the problems of water flow velocity adjustment in the prior art are solved, thereby achieving efficient chip heat dissipation and low energy consumption heat dissipation system.
Patent Information
- Application Number
- CN202310568852.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-19
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2043-05-19
AI Technical Summary
The existing interlayer microchannel liquid cooling technology cannot finely adjust the water flow velocity, and will bring greater energy consumption while meeting the heat dissipation needs.
The microfluidic chip heat dissipation method based on reinforcement learning algorithm is adopted. Through the reinforcement learning algorithm interacting with the environment, iterative training is used to optimize the water flow speed, and the precise temperature control of each layer of chip is achieved, and the power consumption of the microfluidic pump is intelligently adjusted on this basis.
The maximum temperature of each layer of chip is close to the ideal temperature, and while meeting the temperature constraints, it intelligently saves the power consumption of the microfluidic pump and reduces the energy consumption cost of the microfluidic control system.
Smart Images

Figure CN116595879B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of integrated circuit design, and particularly relates to a microfluidic chip heat dissipation method and system based on a reinforcement learning algorithm. Background Art
[0002] The progress of semiconductor processes and technologies has promoted the development of modern electronic systems towards high performance and miniaturization. With the continuous improvement of the integration density of integrated circuits, three-dimensional integrated circuits have gradually developed into the main form of current integrated circuits. However, the power density of the chip has increased sharply due to three-dimensional stacking, and chip heat dissipation has become an important bottleneck restricting the release of chip performance. The sharp increase in the power consumption of large-scale integrated circuits leads to high operating temperatures and large thermal gradients, resulting in serious reliability problems. In some cases, there are local thermal gradients, which even lead to logic failures. Therefore, how to effectively dissipate heat from the chip has always been an important issue in order to maintain the normal operation of the chip without exceeding its temperature limit.
[0003] Common chip heat dissipation methods can be divided into three forms: natural heat dissipation, air cooling, and liquid cooling. Among the three forms of natural heat dissipation, air cooling, and liquid cooling, liquid cooling is the most efficient heat dissipation method. Currently, the widely used chip liquid cooling technologies mainly include liquid spray cooling technology, indirect water cooling plate technology, and microchannel water cooling technology, etc. Obviously, methods such as liquid spray cooling and indirect water cooling plates cannot effectively dissipate the heat generated inside the chip after multi-layer stacking. Therefore, the microchannel water cooling technology is considered to have a stronger heat dissipation ability.
[0004] Since the heat dissipation effect inside the package is stronger than that outside the package, in the application scenario of three-dimensional integrated circuits, interlayer integrated microchannel liquid cooling is a promising and scalable solution for enhancing heat dissipation. Existing interlayer microchannel liquid cooling technologies such as: compact silicon integrated circuit water-cooled integrated heat sink, 3D-IC technology with integrated microchannel cooling, microchannel technology on the back of the chip combining chip inlet and outlet through holes and compatible heat pipes, circulating embedded microfluidic channel chip cooling technology, and microfluidic interconnection integration technology, etc. Generally speaking, the above interlayer microchannel liquid cooling technologies can effectively eliminate the hot spots of multi-layer integrated circuits and ensure that their temperature requirements meet the standards. However, the above interlayer microchannel liquid cooling technologies cannot finely adjust the water flow rate, and will bring relatively large energy consumption while meeting the heat dissipation requirements. Summary of the Invention
[0005] The purpose of the present invention is to provide a microfluidic chip heat dissipation method and system based on a reinforcement learning algorithm, aiming to solve the problems that the existing interlayer microchannel liquid cooling technology cannot finely adjust the water flow rate and will bring relatively large energy consumption while meeting the heat dissipation requirements.
[0006] The present invention is implemented as follows. A microfluidic chip heat dissipation method based on a reinforcement learning algorithm, the method comprising the following steps:
[0007] S1. Obtain the parameters related to the on-chip network and initialize the cooling rate in the microfluidic channels;
[0008] S2. Through the reinforcement learning algorithm, perform the interaction between the algorithm and the reinforcement learning environment and the iterative training of the algorithm;
[0009] S3. Taking the minimization of the power consumption within the normal operating temperature range of the chip as the optimization goal, test the actual effect of the reinforcement learning algorithm, and update the parameters according to the test results.
[0010] Preferably, in step S1, combining the parameters involved in the thermal model part with the empirical values to determine a set of process parameter values of the thermal model.
[0011] Preferably, in step S2, the reinforcement learning environment includes a state space, an action space, a reward function, and a thermal model; wherein,
[0012] The state space includes the module number of the chip, the heat generation power of the chip module, the current speeds of the 1st to nth water pipes, and the average temperature of the module after passing through the coolant;
[0013] The action space is the water flow speed transformation of n water channels, and the change of each water channel includes an increase of 1 and a decrease of 1 in speed;
[0014] The formula of the reward function is: R = -V + γ(T t-1 -T t ), γ = β / |α| > 0, where α and β are coefficients and α < 0, β > 0, and the value of γ can be defined according to the specific requirements, T t-1 -T t is the difference between the average temperature of the chip in the previous step and the average temperature of the current chip. When the water valve pressure is constant, the final power P ∝ V, and V is simplified to a known value as a dynamic function.
[0015] Preferably, in step S2, the reinforcement learning algorithm includes the following steps:
[0016] M1. Invoke the reinforcement learning environment to obtain the environmental state information, including the state space and the action space;
[0017] M2. Determine the policy of the action according to the state, which are the greedy policy for updating the Q-table and the greedy policy for actually promoting the Q-Learning algorithm;
[0018] M3. Q-Learning algorithm training, used to continuously update the Q-table, and in this module, the previous two steps M1 and M2 are invoked.
[0019] Preferably, step S3 includes the following specific steps:
[0020] S31. Call the code of the reinforcement learning algorithm to train a better Q-table;
[0021] S32. Test the actual effect of the reinforcement learning algorithm. First, set an ideal temperature that the on-chip network needs to reach, and then randomly generate the coolant speed driven by the microfluidic pump to meet the temperature requirement, and compare the total power consumption with the result of the reinforcement learning;
[0022] S33. If the effect is not good, improve the algorithm code and retrain the Q-table. If the effect is good, it proves that the algorithm is relatively effective and can be continued to be used.
[0023] The present invention further discloses a microfluidic chip heat dissipation system based on a reinforcement learning algorithm, and the system includes:
[0024] An initialization module, configured to obtain on-chip network related parameters and initialize the cooling speed in the microfluidic channel;
[0025] An iterative training module, configured to perform the interaction between the algorithm and the reinforcement learning environment and the iterative training of the algorithm through the reinforcement learning algorithm;
[0026] A testing module, configured to test the actual effect of the reinforcement learning algorithm with the goal of minimizing the power consumption within the normal operating temperature range of the chip, and update the parameters according to the test results.
[0027] Preferably, in the initialization module, a set of process parameter values of the thermal model is determined by combining the parameters and empirical values involved in the thermal model part.
[0028] Preferably, in the iterative training module, the reinforcement learning environment includes a state space, an action space, a reward function, and a thermal model; wherein,
[0029] The state space includes the module number of the chip, the heat generation power of the chip module, the current speeds of the 1st to nth water pipes, and the average temperature of the module after passing through the coolant;
[0030] The action space is the water flow speed transformation of n water channels, and the change of each water channel includes an increase of 1 and a decrease of 1 in speed;
[0031] The formula of the reward function is: R = -V + γ(T t-1 -T t ), γ = β / |α| > 0, where α and β are coefficients and α < 0, β > 0, and the value of γ can be defined according to the specific requirements, T t-1 -T tis the difference between the average temperature of the chip in the previous step and the average temperature of the current chip. When the water valve pressure is constant, the final power P ∝ V, and V is a dynamic function and simplified to a known value.
[0032] Preferably, in the iterative training module, the reinforcement learning algorithm includes the following steps:
[0033] M1. Invoke the reinforcement learning environment to obtain environmental state information, including the state space and the action space;
[0034] M2. The strategies for determining actions according to the state are the greedy strategy for updating the Q-table and the greedy strategy for actually promoting the Q-Learning algorithm;
[0035] M3. Train with the Q-Learning algorithm to continuously update the Q-table. In this module, the previous two steps M1 and M2 are invoked.
[0036] Preferably, the test module includes:
[0037] An invocation module for invoking the reinforcement learning algorithm code to train a better Q-table;
[0038] A comparison module for testing the actual effect of the reinforcement learning algorithm. First, set an ideal temperature that the on-chip network needs to reach, and then randomly generate the coolant velocity driven by the microfluidic pump to meet the temperature requirement, and compare the total power consumption with the result of the reinforcement learning;
[0039] An improvement module for improving the algorithm code and retraining the Q-table if the effect is not good, and if the effect is good, it proves that the algorithm is relatively effective and can continue to be used.
[0040] The present invention overcomes the deficiencies of the prior art and provides a microfluidic chip heat dissipation method and system based on a reinforcement learning algorithm. The essence of the present invention is to introduce a reinforcement learning algorithm and a layer-by-layer microchannel liquid cooling technology for precisely controlling the liquid flow rate of each layer of a three-dimensional chip. This layer-by-layer microchannel liquid cooling technology can make the highest temperature of each layer of the chip as close as possible to its respective ideal temperature, and while meeting the temperature constraints, it can intelligently save the power consumption of the microfluidic pump and reduce the energy consumption cost of the microfluidic system.
[0041] Regarding the construction of a microfluidic heat dissipation system model, in the present invention, a common silicon-based chip is taken as an example for research, that is, it is considered that the same structures on the plane are stacked vertically in a complete chip structure through a certain packaging method. In the present invention, the heat dissipation system of the microfluidic chip is constructed by etching the microchannels and TSVs on the back of a single chip, aligning them, and stacking them through a bonding process; in the present invention, the chip has a cavity geometry with a double-port straight channel. From the plane view, it shows that the coolant flows through the microchannels etched under the chip for heat dissipation. Each microfluid has a fixed inlet and a corresponding unique outlet, and the path of the coolant flow is a straight line segment determined at both ends.
[0042] It should be noted that in order to precisely control the flow rate of the water and optimize the heat dissipation efficiency, in the present invention, each microchannel has a unique corresponding water pump to control the coolant, that is, the flow rate of the coolant in each microchannel can be individually controlled. Further, each layer of the microfluidic heat dissipation structure has an independent cooling system and a cluster of microfluid pumps to form an efficient, stable and easy-to-precisely-regulate coolant circulation structure.
[0043] In the present invention, a reinforcement learning algorithm is used to control the water pumps in the heat dissipation system in detail, forming an intelligent control system to regulate the parameters of the microfluid pumps to optimize the heat dissipation capacity. The parameters that can be adjusted in the art include the heat transfer coefficient of the coolant, the flow rate of the coolant generated by the water pump, the size of the heat units in the chip, the construction method of the basic heat model, etc. Among them, the most operable and influential parameter is the flow rate of the coolant. Therefore, the present invention mainly controls the heat dissipation capacity of the microfluidic system according to the change of the flow rate of the coolant. In addition, the present invention aims to minimize the power consumption of the microfluidic heat dissipation system, ensure that the chip is within the normal operating temperature range while the coolant circulation structure consumes the least power, and uses a reinforcement learning algorithm to simulate and optimize this process.
[0044] In the present invention, by separating the heat units, a two-resistor heat model is constructed and the heat transfer problem is simplified using classical electrical formulas to obtain an intuitive and adjustable parameter matrix to characterize the thermal state of the chip. Combining with the classical electrical theory Ohm's law, the heat transfer problem in the heat unit can be simplified to analyze the functional relationship between thermal resistance, temperature and heat flow. Therefore, in the present invention, based on the premise of hot spot discretization, the temperature of the chip is discretized at different hot spots as the observation parameter; in the present invention, based on the premise that the microfluid pumps and the microchannels are in a one-to-one correspondence, the coolant velocities in all microchannels are used as the controllable parameters; in the present invention, based on the premise that the microfluid pumps and the microchannels are in a one-to-one correspondence, the sum of the powers consumed by the coolant circulation system is used as an additional observation parameter.
[0045] On this basis, the present invention proposes a heat dissipation method and system for a microfluidic chip based on a reinforcement learning algorithm, including: obtaining parameters related to the on-chip network and initializing the cooling rate in the microfluidic channels; writing reinforcement learning environment and algorithm codes to train a relatively accurate Q-table; testing the actual effect of this reinforcement learning algorithm and updating parameters according to the results to achieve the optimization goal.
[0046] The present invention combines training and testing through the reinforcement learning method (Q-learning). Given an input file based on the chip thermal model parameters, this file is used to describe the topology in the basic thermal model, the externally randomly determined chip thermal power, the optimization constraints and target conditions, and the design parameters. First, an effective parameter characterization and transformation matrix are obtained in combination with the thermal model, and then a reward function is given through reinforcement learning to obtain the optimal path that meets the design requirements.
[0047] In the present invention, the reinforcement learning algorithm mainly consists of the following three modules:
[0048] M1: Invoking the reinforcement learning environment to obtain environmental state information therefrom, including the state space and the action space;
[0049] M2: The policies for determining actions according to the state are the greedy policy for updating the Q-table and the ε-greedy policy for actually promoting the Q-Learning algorithm;
[0050] M3: Q-Learning algorithm training, used to continuously update the Q-table, and the above two modules are invoked in this module.
[0051] Furthermore, the Q-Learning algorithm is a reinforcement learning method for the state-action value function, belonging to model-free learning, that is, there is no need to model the external environment in detail. Only by providing sufficient training samples, the optimal policy can be obtained through the interaction between the agent and the environment. Using the Q-Learning algorithm for on-chip network microfluidic control can not only effectively eliminate hot spots, make the on-chip temperature as balanced as possible, ensure the chip working efficiency, but also reduce the total power consumption of the microfluidic heat dissipation system.
[0052] Compared with the disadvantages and deficiencies of the prior art, the present invention has the following beneficial effects: The present invention can make the highest temperature of each layer of the chip as close as possible to their respective ideal temperatures, and intelligently save the power consumption of the pump while meeting the temperature constraints, reducing the energy consumption cost of the microfluidic system; in addition, the present invention can prevent the negative impacts on its performance, life, etc. caused by the uneven heat distribution and excessive temperature of the chip, so as to achieve the purpose of fully releasing the working performance of the CPU. Description of the Drawings
[0053] Figure 1It is a two-resistor model diagram based on thermal model abstraction processing in an embodiment of the present invention;
[0054] Figure 2 It is a three-dimensional model diagram of a basic thermal unit in an embodiment of the present invention;
[0055] Figure 3 It is a cross-sectional view of a complete chip structure including a liquid cooling system and an intelligent control system in an embodiment of the present invention;
[0056] Figure 4 It is a three-dimensional model diagram of an intelligent controllable single-layer microfluidic heat dissipation structure in an embodiment of the present invention;
[0057] Figure 5 It is a three-dimensional schematic diagram of a complete chip structure with packaging levels, a liquid cooling system, and an intelligent control system added in an embodiment of the present invention (note: the pump marked in the figure is a microfluidic pump for controlling the coolant);
[0058] Figure 6 It is a schematic diagram of the stacked materials and specifications of each layer in a complete chip structure in an embodiment of the present invention;
[0059] Figure 7 It is a step flow chart of an embodiment of the method of the present invention;
[0060] Figure 8 It is a schematic structural diagram of an embodiment of the system of the present invention. Specific embodiments
[0061] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0062] An embodiment of the present invention takes a common silicon-based chip as an example. For the construction of a microfluidic heat dissipation system model, first, a two-resistor thermal model is constructed, as shown in the accompanying Figure 1 figures. In the embodiment of the present invention, it is assumed that the topological structures of the hot spots on the chip are all the same, and the chip can be abstracted as a stack of repeated thermal unit combinations; combining with the classical electrical theory Ohm's law, the heat transfer problem in the thermal unit can be simplified to the formula GT(t) = U(t), where the thermal resistance G, the function of temperature with respect to time T(t), and the function of heat flow with respect to time U(t) are all in matrix form corresponding to the size structure of the thermal model unit. Combining with the accompanying Figure 1It can be seen that when establishing the thermal model, it is considered that the heat flow in the thermal unit occurs on the equivalently derived thermal resistance, that is, the heat transfer process occurs on the thermal resistance shown in the figure. It is worth mentioning that based on the two-resistance thermal model, the problem of the order of magnitude of the microchannel size not matching the IC can be effectively overcome, which is achieved by homogenizing the entire microchannel cavity layer in the 3D IC into a single porous medium. As shown in the appendix Figure 2 it can be seen that the cross-sectional size of the microchannel matches the IC. Combining the appendix Figure 1 , by projecting the heat entering from the side wall onto the top wall and the bottom wall to calculate the effective heat transfer coefficient, the parameters in the heat transfer matrix can be obtained more conveniently.
[0063] Regarding the important heat transfer formula in the thermal model, in the embodiment of the present invention, based on the premise of hot spot discretization, the thermal power of the chip is discretized on different hot spots as the input parameters of the heat flow matrix. Combining with the reinforcement learning algorithm, the random power of the chip operation can be set as the state parameter. By continuously learning and iteratively updating the Q table, the optimal strategy can be obtained through interaction with the environment; it should be noted that although both the temperature and the heat flow in the heat transfer formula are functions related to time, in the state environment of reinforcement learning, the temperature and heat flow parameters in the steady state are selected, that is, the parameters at a certain moment when the thermal state of the system is relatively stable after a period of time are used as the state result of one action.
[0064] The above is the elaboration on the establishment of the basic thermal unit model. Further, the overall chip model in the embodiment of the present invention is as shown in the appendix Figure 5 and the appendix Figure 6 . It can be seen that in the chip structure of the conventional package, for the silicon layer that needs to dissipate heat, in the embodiment of the present invention, multiple layers of microchannels with mutually perpendicular directions are used for heat dissipation, and the parameters of the microfluidic pump clusters on each layer are precisely controlled through an intelligent control system to optimize the best state. The microfluidic heat dissipation structure of each layer is as shown in the appendix Figure 4 . In Figure 4 , the encapsulation layer and the BEOL layer are omitted, and only the single-layer microfluidic heat dissipation structure and the corresponding coolant circulation system are analyzed and described. For the single-layer microfluidic heat dissipation structure formed by the combination of n*n thermal units, the coolant driven by the microfluidic pump cluster flows through the microchannels penetrating the silicon substrate layer to play the role of heat dissipation, and after flowing out of the IC, it will enter the liquid cooling system to be cooled and is ready for the next heat dissipation cycle under the regulation of the intelligent control system. It is worth mentioning that as shown in the appendix Figure 4 , the microchannel does not affect the TSV in the chip package and mainly plays the role of heat dissipation in the chip structure without interfering with the normal operation of the chip. Appendix Figure 3The cross-sectional structure of the complete chip model in the embodiment is shown, and it can be seen that there is a liquid cooling system around the chip to ensure a certain temperature of the coolant flowing into the microchannels; relying on the microfluidic pump clusters of each layer of the microfluidic heat dissipation structure to drive the coolant to flow through the microchannels etched under the silicon layer to perform the heat dissipation function. It should be noted that the coolant completes a cyclic flow in this process, and after completing the heat exchange with the hot spots on the chip, it re-enters the liquid cooling system to be cooled down, and this cycle is intelligently learned and controlled by the reinforcement learning algorithm.
[0065] The following describes a microfluidic chip heat dissipation method based on a reinforcement learning algorithm according to an embodiment of the present invention with reference to the accompanying drawings. As shown in the accompanying Figure 7 drawings, the method includes the following steps:
[0066] S1. Obtain the relevant parameters of the on-chip network and initialize the cooling rate in the microfluidic channels;
[0067] In step S1, obtain the relevant parameters of the on-chip network and initialize the cooling rate in the microfluidic channels. Combine the parameters and empirical values involved in the thermal model part to determine a set of process parameter values of the thermal model. The process parameters are as follows:
[0068] AW: Actual wetted surface area;
[0069] AP: Projected area for heat transfer from the wetted surface;
[0070] K: Thermal conductivity of the coolant;
[0071] Nu: Nusselt number of the microfluid;
[0072] AR: Aspect ratio of the channel;
[0073] Re: Reynolds number of the microfluid;
[0074] Pr: Prandtl number of the microfluid;
[0075] Hconv: Effective heat transfer coefficient;
[0076] Heff: Surface heat transfer coefficient.
[0077] S2. Perform the interaction between the algorithm and the reinforcement learning environment and the iterative training of the algorithm through the reinforcement learning algorithm (Q-learning algorithm)
[0078] In step S2, write the reinforcement learning environment code, including the state space S t , action space A t , reward function and the implementation of the thermal model; specifically, the state space includes the module number Num{j} of the chip, the heat generation power Pd{j} of the chip module, and the current speeds V of the 1st to nth water pipes tThe average temperature T of the {i} - th module after passing through the coolant t , where j is the number of the thermal unit module and i is the number of the water channel.
[0079] S t ={Num{j}, Pd{j}, V t {i}, T t}, j = 1, 2…N, i = 1, 2…n;
[0080] The action space is the transformation of the water flow velocity in n water channels, and the change in each water channel includes an increase of 1 and a decrease of 1 in velocity.
[0081] A t ={u{i}, d{i}|i ∈ {1, 2…n}}
[0082] where u{i} represents increasing the water flow velocity of the i - th water pipe, and d{i} represents decreasing the water flow velocity of the i - th water pipe.
[0083] To make the reward function have a better association with the action, the embodiment of the present invention sets V t {i} rather than the final power P as an important parameter of the reward function;
[0084] Since when the water valve pressure is constant, P ∝ V, and V, as a dynamic function, can better reflect the system change than P as a static function, using V as a parameter can better reflect the relationship between the reward function required by the present invention and the final power.
[0085] To also consider the optimization of temperature into the reward function, the embodiment of the present invention takes the difference T t-1 -T t between the average temperature of the chip in the previous step and the current average temperature as another parameter consideration item of the reward function.
[0086] Therefore, the embodiment of the present invention can list the reward function formula: R = αV+β(T t-1 -T t ), where α and β are coefficients. The present invention expects the total power to be small and the final temperature to be lower, so α < 0, β > 0 and there is a relationship between the two.
[0087] From the perspective of the simplicity of the formula, the embodiment of the present invention can make the magnitude of each change in the water flow also be 1 unit. This unit can be made small enough to ensure that the present invention will not miss a large enough reward value due to excessive changes in the water flow velocity. In this way, V in the reward function of the present invention can participate in the formula as a known number, so the reward formula can be further simplified to:
[0088] R=-V + γ(T t-1 -Tt )
[0089] where γ = β / |α| > 0.
[0090] It should be noted that the value of γ can be defined according to specific requirements, such as whether the temperature requirement is clear and whether the energy consumption requirement is strict.
[0091] In step S2, formulate the reinforcement learning algorithm process and implement the interaction between the algorithm and the reinforcement learning environment and the iterative training of the algorithm.
[0092] Specifically, the reinforcement learning algorithm includes the following steps:
[0093] M1. Call the reinforcement learning environment to obtain environmental state information, including the state space and the action space.
[0094] M2. The policies for determining actions according to the state are the greedy policy for updating the Q-table and the greedy policy for actually promoting the Q-Learning algorithm, respectively.
[0095] M3. Train the Q-Learning algorithm to continuously update the Q-table, and call the previous two steps M1 and M2 in this module.
[0096] The Q-table training process of this reinforcement learning algorithm is as follows:
[0097] (1) Initialize the Q-table and assign all Q values to 0.
[0098] (2) for i in range(num) (train for a total of num rounds);
[0099] (3) S t ← Reset the reinforcement learning environment to obtain the initial state;
[0100] (4) A t ← A random action selected according to the initial state;
[0101] (5) Termination information ← False;
[0102] (6) Call the step function in the reinforcement learning environment to obtain the next state S t+1 , the reward value R, the termination information, and the debug information;
[0103] (7) while termination information == False and the number of algorithm steps < 20
[0104] (8) Q(S t , A t ) ← Q(S t , A t) + α[R + γmax a (S t+1 ,a) - Q(S t ,A t ) (Update the Q - table according to the greedy strategy);
[0105] (9) A t+1 ← Select the action of S according to the ε - greedy strategy t+1 ;
[0106] (10) Call the step function in the reinforcement learning environment;
[0107] (11) Increment the algorithm step count by 1;
[0108] (12) Return the trained Q - table.
[0109] Among them, the above - mentioned step M3 includes the following specific processes:
[0110] Algorithm: Q - Learning;
[0111] Input: The number of algorithm training rounds num, the parameter α required for Q - value update, the parameter ε in the ε - greedy strategy;
[0112] Output: The trained Q - table (that is, the set of all trained Q - values. The Q - value represents the value corresponding to the state - action, that is, the value that can be brought by taking an action in the current state. In the Q - table, the action value At and the reward value R for each state St are obtained by iterating through the states, and the values obtained through update and iteration during the entire optimization process are the set of all Q - values).
[0113] S3. Taking the minimization of the power consumption of the chip within the normal operating temperature range as the optimization goal, test the actual effect of this reinforcement learning algorithm, and update the parameters according to the test results
[0114] In step S3, the following specific steps are included:
[0115] S31. Call the reinforcement learning algorithm code to train a better Q - table;
[0116] S32. Test the actual effect of this reinforcement learning algorithm. First, set an ideal temperature that the on - chip network needs to reach, then randomly generate the coolant speed driven by the micro - pump to meet the temperature requirement, and compare the total power consumption with the result of the reinforcement learning;
[0117] S33. If the effect is not good, improve the algorithm code and retrain the Q - table. If the effect is good, it proves that the algorithm is relatively effective and can be continued to be used.
[0118] The present invention further provides a microfluidic chip heat dissipation system based on a reinforcement learning algorithm, asFigure 8 As shown in the figure, the system includes:
[0119] An initialization module 1, configured to obtain on-chip network related parameters and initialize the cooling rate in the microfluidic channel;
[0120] An iterative training module 2, configured to perform the interaction between the algorithm and the reinforcement learning environment and the iterative training of the algorithm through a reinforcement learning algorithm;
[0121] A testing module 3, configured to test the actual effect of the reinforcement learning algorithm with the goal of minimizing the power consumption when the chip is within the normal operating temperature range, and update the parameters according to the test results.
[0122] In an embodiment of the present invention, the testing module specifically includes:
[0123] A calling module, configured to call the reinforcement learning algorithm code to train a better Q-table;
[0124] A comparison module, configured to test the actual effect of the reinforcement learning algorithm. First, set an ideal temperature that the on-chip network needs to reach, then randomly generate the coolant speed driven by the microfluidic pump to meet the temperature requirement, and compare the total power consumption with the result of the reinforcement learning;
[0125] An improvement module, configured to improve the algorithm code and retrain the Q-table if the effect is not good, and prove that the algorithm is relatively effective and can continue to be used if the effect is good.
[0126] The implementation process and effect of the above-mentioned microfluidic chip heat dissipation are used to explain the implementation process and effect of the microfluidic chip heat dissipation system of the present invention, and will not be elaborated here.
[0127] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A heat dissipation method for a microfluidic chip based on a reinforcement learning algorithm, characterized in that, the method comprises the following steps: S1. Obtain the parameters related to the on-chip network and initialize the cooling rate in the microfluidic channel; S2. Through the reinforcement learning algorithm, perform the interaction between the algorithm and the reinforcement learning environment and the iterative training of the algorithm; S3. Taking the minimization of the power consumption within the normal operating temperature range of the chip as the optimization goal, test the actual effect of the reinforcement learning algorithm, and update the parameters according to the test results; In step S2, the reinforcement learning environment includes a state space, an action space, a reward function, and a thermal model; wherein, the state space includes the module number of the chip, the heat generation power of the chip module, the current speeds of the 1st to nth water pipes, and the average temperature of the module after passing through the coolant; the action space is the water flow speed transformation of n water channels, and the change of each water channel includes an increase of 1 and a decrease of 1 in speed; The reward function formula is: R = -V + γ(T t-1 -T t ), γ = β / |α| > 0, where α and β are coefficients and α < 0, β > 0, and the value of γ can be defined according to specific requirements. T t-1 -T t is the difference between the average temperature of the chip in the previous step and the average temperature of the current chip. When the water valve pressure is constant, the final power P ∝ V, and V is a dynamic function and simplified to a known value.
2. The heat dissipation method for a microfluidic chip according to claim 1, characterized in that, in step S1, combining the parameters involved in the thermal model part with the empirical values to determine a set of process parameter values of the thermal model.
3. The heat dissipation method for a microfluidic chip according to claim 1, characterized in that, in step S2, the reinforcement learning algorithm comprises the following steps: M1. Invoke the reinforcement learning environment to obtain the environmental state information therefrom, including the state space and the action space; M2. The policies for determining actions according to the state are respectively the greedy policy for updating the Q-table and the greedy policy for actually promoting the Q-Learning algorithm; M3. Q-Learning algorithm training, which is used to continuously update the Q-table, and in this module, the above two steps M1 and M2 are invoked.
4. The heat dissipation method for a microfluidic chip according to claim 1, characterized in that, step S3 comprises the following specific steps: S31. Invoke the reinforcement learning algorithm code to train a better Q-table; S32. Test the actual effect of the reinforcement learning algorithm. First, set an ideal temperature that the on-chip network needs to reach, and then randomly generate the coolant speed driven by the microfluidic pump to meet the temperature requirement, and compare the total power consumption with the result of the reinforcement learning; S33. If the effect is not good, improve the algorithm code and retrain the Q-table. If the effect is good, it proves that the algorithm is relatively effective and can continue to be used.
5. A heat dissipation system for a microfluidic chip based on a reinforcement learning algorithm, characterized in that, the system comprises: an initialization module, which is used to obtain the parameters related to the on-chip network and initialize the cooling rate in the microfluidic channel; an iterative training module, which is used to perform the interaction between the algorithm and the reinforcement learning environment and the iterative training of the algorithm through the reinforcement learning algorithm; a test module, which is used to take the minimization of the power consumption within the normal operating temperature range of the chip as the optimization goal, test the actual effect of the reinforcement learning algorithm, and update the parameters according to the test results; In the iterative training module, the reinforcement learning environment includes a state space, an action space, a reward function, and a thermal model; wherein, The state space includes the module number of the chip, the heating power of the chip module, the current speeds of the water pipes numbered from 1 to n, and the average temperature of the module after passing through the coolant; The action space is the water flow speed transformation of n water channels, and the change of each water channel includes an increase of 1 and a decrease of 1 in speed; The formula for the reward function is: R = -V + γ(T t-1 -T t ), γ = β / |α| > 0, where α and β are coefficients and α < 0, β > 0, and the value of γ can be defined according to specific requirements. T t-1 -T t is the difference between the average temperature of the chip in the previous step and the average temperature of the current chip. When the water valve pressure is constant, the final power P ∝ V, and V is a dynamic function and simplified to a known value.
6. The system according to claim 5, characterized in that, In the initialization module, a set of process parameter values of the thermal model is determined by combining the parameters involved in the thermal model part with empirical values.
7. The system according to claim 6, characterized in that, In the iterative training module, the reinforcement learning algorithm includes the following steps: M1. Call the reinforcement learning environment to obtain environmental state information therefrom, including the state space and the action space; M2. The policies for determining actions according to the state are respectively the greedy policy for updating the Q-table and the greedy policy for actually promoting the Q-Learning algorithm; M3. Q-Learning algorithm training is used to continuously update the Q-table, and the above two steps M1 and M2 are called in this module.
8. The system according to claim 5, characterized in that, The test module includes: A calling module for calling the reinforcement learning algorithm code to train a better Q-table; A comparison module for testing the actual effect of the reinforcement learning algorithm. First, set an ideal temperature that the on-chip network needs to reach, and then randomly generate the coolant speed driven by the microfluidic pump to meet the temperature requirement, and compare the total power consumption with the result of the reinforcement learning; An improvement module for improving the algorithm code and retraining the Q-table if the effect is not good, and if the effect is good, it proves that the algorithm is relatively effective and can continue to be used.
Citation Information
Patent Citations
Online fault detection method for digital microfluidic biochip on basis of reinforcement learning
CN111141920A
Multi-objective optimization design method for hybrid thermal management system based on regression model algorithm
CN113435016A