Display card liquid cooling module with intelligent temperature control function
Through the intelligent temperature-controlled graphics card liquid cooling module, the data detection, analysis and prediction and cooling execution unit is used to achieve coordinated cooling of multiple different models of graphics cards, solving the problem of insufficient cooling coordination capabilities in traditional technology, and improving the accuracy of temperature prediction and cooling efficiency.
Patent Information
- Application Number
- CN202510360992.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing graphics card liquid cooling technology is mainly aimed at a single graphics card or multiple graphics cards of the same model, and lacks the ability to coordinate cooling for multiple graphics cards of different models.
An intelligent temperature-controlled graphics card liquid cooling module is designed, including a data detection unit, an analysis and prediction unit, a cooling execution unit and a terminal interaction unit. By obtaining the hardware and software data of the target graphics card, feature extraction and temperature prediction are performed, the coolant flow rate is adjusted in real time, and the coolant is coordinated distributing is performed through the ADMM consistency algorithm.
Coordinated cooling of multiple graphics cards of different models is achieved, the accuracy and cooling efficiency of temperature prediction are improved, and the problem of incomplete or conflicts in data acquisition caused by differences in graphics card drivers in traditional technology is solved.
Smart Images

Figure CN120215658A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of graphics card cooling, and particularly to a graphics card liquid cooling module with intelligent temperature control. Background Art
[0002] In recent years, the fields of computer technology and artificial intelligence have continued to develop at a high speed. The application popularity of high-performance graphics cards has been continuously improved, and arrays composed of multiple graphics cards have become more and more commonly used. Limited by the current graphics chip technology, graphics card heating is still a problem that is difficult to avoid. In actual applications, graphics card heating will lead to performance degradation and hardware loss, and it is difficult to release all the computing power potential. Graphics card liquid cooling can break through the technical bottleneck of traditional air cooling, and improve the cooling efficiency and stability by efficiently conducting heat through a liquid medium. Especially when used under high load and overclocking, it can maintain the graphics card core temperature within a reasonable range, maintain high-frequency operation and reduce hardware loss. At the same time, it reduces the dependence on fans, reduces noise pollution and improves energy efficiency, and increases the use stability and lifespan of the graphics card when performing high-load tasks such as high-performance computing, artificial intelligence training, and graphics rendering. Therefore, a lightweight multi-graphics card liquid cooling module that can achieve intelligent temperature control is a promising research direction.
[0003] The significance of graphics card liquid cooling lies in breaking through the technical bottleneck of traditional air cooling, efficiently conducting heat through a liquid medium, significantly improving the cooling efficiency and stability. Especially in high-load or overclocking scenarios, it can maintain the graphics card core temperature within a reasonable range, avoid performance degradation or hardware loss caused by overheating, and thus fully release the computing power potential. At the same time, the liquid cooling system reduces noise pollution by reducing the dependence on fans, optimizes the device operation environment, and supports higher-density hardware deployment through a compact design, providing a sustainable cooling solution for heavy-load fields such as high-performance computing, artificial intelligence training, and graphics rendering. Its energy-saving characteristics also conform to the development trend of green data centers, promoting the iterative upgrade of the computing power infrastructure towards high efficiency, low power consumption, and long lifespan.
[0004] At present, a Chinese invention patent application with the application number CN109308109A discloses a water-cooled radiator and method for a computer graphics card. This application includes: a water block, a fixing plate is provided outside the water block, an inlet metal pipe and an outlet metal pipe are provided on the side wall of the water block, transmission pipes are provided at the ends of the inlet metal pipe and the outlet metal pipe, a water-cooling radiator is provided at the end of the transmission pipe, the water block includes a housing, a motor partition is provided at the bottom of the inner cavity of the housing, a driving motor is provided inside the motor partition, an inner housing is provided on the top of the driving motor, a stirring blade is provided inside the inner housing, a rotating shaft is provided at the bottom of the stirring blade, a mechanical seal is sleeved outside the rotating shaft, and through holes are provided on the side wall of the inner housing. The present invention enables the unevenly heated part of the cooling water in the water block to be mixed, reduces the overall temperature of the cooling water, makes the temperature inside the entire water block more uniform, and improves the heat dissipation effect on the graphics card part. However, when this application cools the graphics card with water, there are only two states: startup and standby. Summary of the Invention
[0005] The technical problem solved by the present invention is that the application objects of the existing graphics card liquid cooling technology mainly target single graphics cards or multiple graphics cards of the same model, lacking coordinated cooling for multiple different models of graphics cards.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: A graphics card liquid cooling module with intelligent temperature control, including: A data detection unit, an analysis and prediction unit, a cooling execution unit, and a terminal interaction unit; The data detection unit is used to obtain the hardware layer data set, software data set, and condensation data set of the target graphics card, and add a unified time stamp to all the collected data; The analysis and prediction unit is used to perform preliminary feature extraction on the output data of the data detection unit to obtain the first feature vector set, the second feature vector set, and the target modal component, and calculate the predicted temperature of the target graphics card in combination with the temperature rise prediction model; The cooling execution unit is used to adjust the coolant flow rate in real time according to the predicted temperature value, and perform coolant collaborative distribution on all target graphics cards in combination with the ADMM consistency algorithm; The terminal interaction unit is used to provide a display interface for the user and receive user operation instructions, and receive all the output data of the data detection unit, the analysis and prediction unit, and the cooling execution unit for storage.
[0007] Preferably, the data detection unit includes a hardware data acquisition sub-unit, a software data acquisition sub-unit, and a time synchronization sub-unit; The hardware data acquisition sub-unit is used to respectively collect the physical layer data set for all target graphics cards, and collect the condensation data set for the coolant pipeline; The physical layer data set includes a coolant flow rate data set and a voltage data set, and the processing logic includes: A coolant flow rate data set of the target graphics card is acquired through a Hall flow meter, and a voltage data set of the target graphics card is acquired through a voltage sensor; The condensation data set includes coolant inlet temperature, coolant pressure difference, and speed regulating liquid pump duty cycle.
[0008] Preferably, the software data acquisition subunit is used to acquire system layer data sets from all target graphics cards respectively, wherein the system layer data sets include core temperature data sets, operating frequency data sets, video memory occupancy ratio data sets and video memory reading amount data sets, and the processing logic includes: The software data acquisition subunit is used to establish an independent acquisition thread for each target graphics card, obtain the manufacturer ID data and device ID data of all target graphics cards through PCI bus scanning, initialize the API interface based on the manufacturer ID data and device ID data of the target graphics card, and obtain the core temperature data set, operating frequency data set, video memory occupancy ratio data set and video memory reading amount data set through the corresponding API interface of each target graphics card.
[0009] Preferably, the time synchronization subunit is used to provide a unified time source for the hardware data acquisition subunit and the software data acquisition subunit through the hardware clock synchronization protocol PTS, and add a unified time stamp to all data acquired by the hardware data acquisition subunit and the software data acquisition subunit.
[0010] Preferably, the analysis and prediction unit includes a feature extraction subunit and a temperature prediction subunit; The feature extraction subunit is used to perform preliminary feature extraction on the output data of the data detection unit to obtain a preliminary feature data set, wherein the preliminary feature data set includes a first feature vector set, a second feature vector set and a target modal component. The processing logic of the preliminary feature extraction includes: A reference time slot is established based on a preset reference time window step. The physical layer data set and the system layer data set are divided by the reference time slot. Time phase alignment is performed by linear interpolation. The maximum tolerance delay is set. The data that exceeds the maximum tolerance delay is directly marked as delayed data. The delayed data is removed to obtain a synchronous time series data set. The synchronization timing data set includes coolant flow rate synchronization timing, voltage synchronization timing, core temperature synchronization timing, operating frequency synchronization timing, video memory occupancy ratio synchronization timing and video memory reading amount synchronization timing.
[0011] Preferably, the temperature change rate corresponding to each target graphics card is calculated based on the preset first time window step and the core temperature synchronization timing, and the memory bandwidth utilization rate corresponding to each target graphics card is calculated based on the preset second time window step and the memory reading amount synchronization timing; Establish a first set of eigenvectors based on the coolant flow rate synchronization timing sequence, voltage synchronization timing sequence, core temperature synchronization timing sequence, and temperature change rate, and establish a second set of eigenvectors based on the operating frequency synchronization timing sequence, video memory occupancy ratio synchronization timing sequence, and video memory bandwidth utilization rate; Perform variational mode decomposition (VMD) processing on the core temperature synchronization timing sequence to obtain a set of temperature intrinsic mode components. Calculate the energy entropy values of each mode component based on the set of temperature intrinsic mode components, and perform modal component screening processing to retain the modal components whose energy entropy values are less than or equal to a preset energy entropy threshold to obtain target modal components.
[0012] Preferably, the temperature prediction subunit is used to calculate the predicted temperature value of the target graphics card according to the preliminary feature data set and the temperature rise prediction model. The processing logic of the temperature rise prediction model includes: Subsequently, process the timing data through a two-stage gated LSTM architecture. Calculate the variance sensitivity of each data in the first set of eigenvectors through a spatial attention mechanism to obtain a spatial attention weight matrix, and perform a Hadamard product operation on the weight matrix and the original eigenvector to obtain a hardware eigenvector; Perform causal convolution calculation with a dilation coefficient of K on the second set of eigenvectors through temporal dilated convolution to obtain a multi-scale convolution result, and perform non-linear fusion processing on the multi-scale convolution result using a gated linear unit (GLU) to obtain a software eigenvector; Perform second sliding window feature extraction calculation on the hardware eigenvector through a short-term memory unit, and the short-term memory unit is a window LSTM network with a duration of seconds; Perform second periodic cumulative modeling on the software eigenvector through a long-term memory unit, and perform downsampling calculation to obtain a long-term trend hidden state. The long-term memory unit is a window LSTM network with a duration of seconds; Use the long-term dynamic hidden state as the query vector Query, use the short-term dynamic hidden state as the key-value pair Key-Value, calculate the long-term-short-term attention distribution weight through a cross-attention mechanism, and perform weighted summation on the short-term dynamic hidden state based on the long-term-short-term attention distribution weight to obtain a scale fusion eigenvector; Concatenate the target modal component and the multi-scale feature vector according to the reference time slot to obtain a multi-dimensional joint feature vector. Perform dimensionality reduction processing on the multi-dimensional joint feature vector through one-dimensional average pooling in the time dimension, and perform regression and residual correction processing through a fully connected layer and inverse normalization reduction to obtain the predicted temperature value.
[0013] Perform one-dimensional average pooling dimensionality reduction on the multi-dimensional joint feature vector through a fully connected layer, complete it by superimposing the core temperature synchronous time series data through residual connection, and obtain the predicted temperature value through the Tanh activation function and inverse normalization reduction processing.
[0014] Preferably, the cooling execution unit includes a flow rate control subunit and a collaborative distribution subunit; The flow rate control subunit is used to adjust the coolant flow rate in real time according to the predicted temperature value, and the processing logic includes: Calculate the predicted temperature deviation by taking the difference between the predicted temperature value corresponding to each target graphics card and the core temperature data. When the predicted temperature deviation in consecutive J reference time windows is less than or equal to the preset first regulation temperature threshold, keep the coolant flow rate unchanged; When the predicted temperature deviation in consecutive J reference time windows is greater than the preset first regulation temperature threshold, trigger the coolant flow rate regulation, and the processing logic of the flow rate regulation includes: Calculate the coolant flow rate demand value based on the predicted temperature value and the thermal model, and obtain the feedforward PWM duty ratio data based on the duty ratio-flow characteristic curve corresponding to the speed control liquid pump; Deploy a PID controller for each target graphics card node, establish a PID control equation with the predicted temperature deviation as the real-time error e(t), calculate the feedback control signal u(t) through the PID control equation, perform linear mapping processing on the feedback control signal u(t) to obtain the feedback PWM duty ratio data, perform superposition calculation on the feedforward control signal and the feedback control signal to obtain the PWM execution data, and adjust the speed of the speed control liquid pump based on the PWM execution data to achieve coolant flow rate control; Its calculation expression is: ; ; ; Among them, represents the coolant flow rate demand value, represents the heat capacity of the liquid cooling system, represents the predicted temperature value, represents the coolant inlet temperature, represents the preset reference time window step size, represents the thermal resistance of the liquid cooling system, represents the coolant density, represents the coolant heat capacity, represents the allowable temperature rise of the coolant, represents the feedback control signal, represents the predicted temperature deviation, represents the PID proportional coefficient, Represents the PID integral coefficient, Represents the PID differential coefficient, Represents, Represents the minimum duty cycle, Represents the maximum feedback value of the PID theoretical output, Represents the minimum feedback value of the PID theoretical output, Represents the feedback control signal.
[0015] Preferably, the collaborative allocation subunit is used for collaborative coolant allocation for the target graphics card, and the processing logic includes: When the predicted temperature deviation corresponding to a graphics card in J consecutive reference time windows is greater than the preset second regulation temperature threshold, mark the graphics card as a high-priority graphics card and trigger temporary collaborative allocation of coolant flow. The processing logic includes: Increase the flow quota by H% for the high-priority graphics card, count the coolant flow demand values and coolant pressure differences of each target graphics card, perform cyclic iterative calculations with the preset constraint pressure difference P% through the ADMM consensus algorithm to obtain the collaborative flow allocation weight, and obtain the collaborative flow allocation data set based on the collaborative flow allocation weight.
[0016] Preferably, the terminal interaction unit includes a display interaction subunit and a data storage subunit; The display interaction subunit is used to provide a display interface for the user and receive user operation instructions, and receive instructions issued by the user for parameter adjustment settings of the data detection unit, analysis and prediction unit, and cooling execution unit; The data storage subunit is used to receive all output data of the data detection unit, analysis and prediction unit, and cooling execution unit for storage, and generate a historical log of the operation of the liquid cooling module Advantages of the present invention: This application focuses on coordinating the cooling of multiple different models of graphics cards. By independently establishing a collection thread for each graphics card and scanning the manufacturer ID and device ID of the graphics card to initialize a dedicated API interface, it is beneficial for accurate adaptation and collection of software layer data of multiple different manufacturers' graphics cards, solving the problem of incomplete or conflicting data collection caused by differences in graphics card drivers in the prior art. Although the physical layer data set and the system layer data set have completed time synchronization during the collection process, the sampling frequencies of software data collection and hardware data collection are different. The feature extraction subunit first performs time axis normalization, divides data slots through the benchmark time window step size, and combines linear interpolation and the maximum tolerance delay mechanism to achieve dynamic time axis normalization of hardware and software data, which is beneficial for avoiding timing misalignment problems caused by sampling frequency differences during multi-source heterogeneous data collection and providing a strictly aligned timing data set for subsequent processing. Among them, the traditional LSTM only relies on a single time window. Through hierarchical memory and modal fusion, the two-stage LSTM and spatio-temporal attention mechanism can effectively capture the laws of hardware state mutations such as voltage fluctuations of the target graphics card and long-term software loads such as rendering tasks. Through spatial attention weighting of hardware features and multi-scale causal convolution of software features, and then cross-attention fusion through long short-term memory units, the accuracy of temperature prediction is improved. In addition, the two-way superposition processing of the feedforward of the thermal model prediction and the PID control feedback, dynamically superimposing the feedforward duty cycle and the PID feedback signal based on the predicted temperature deviation, breaks through the limitation of the traditional PID control relying on lagging temperature feedback, which is beneficial for predictive adjustment of the coolant flow rate and improving the liquid cooling efficiency of the graphics card. Brief Description of the Drawings
[0017] Figure 1 FIG. is a schematic basic flow diagram of a graphics card liquid cooling module with intelligent temperature control provided by an embodiment of the present invention. Detailed Embodiments
[0018] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be given in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments.
[0019] Referring to Figure 1 , an embodiment of the present invention provides a graphics card liquid cooling module with intelligent temperature control, including: a data detection unit, an analysis and prediction unit, a cooling execution unit, and a terminal interaction unit; The data detection unit is used to obtain the hardware layer data set, software data set, and condensation data set of the target graphics card, and add a unified timestamp to all the collected data; The analysis and prediction unit is used to perform preliminary feature extraction on the output data of the data detection unit to obtain the first feature vector set, the second feature vector set, and the target modal component, and calculate the predicted temperature of the target graphics card in combination with the temperature rise prediction model; The cooling execution unit is used to adjust the coolant flow rate in real time according to the predicted temperature value, and perform coolant collaborative allocation on all target graphics cards in combination with the ADMM consistency algorithm; The terminal interaction unit is used to provide a display interface to the user and receive user operation instructions, and store all the output data of the data detection unit, the analysis and prediction unit, and the cooling execution unit.
[0020] In this embodiment, the data detection unit includes a hardware data acquisition sub-unit, a software data acquisition sub-unit, and a time synchronization sub-unit; The hardware data acquisition sub-unit is used to respectively collect all target graphics cards to obtain a physical layer data set, and collect data from the coolant pipeline to obtain a condensation data set; The physical layer data set includes a coolant flow rate data set and a voltage data set, and the processing logic includes: The coolant flow rate data set of the target graphics card is collected by a Hall flowmeter, and the voltage data set of the target graphics card is collected by a voltage sensor; The condensation data set includes the coolant inlet temperature, the coolant pressure difference, and the duty ratio of the speed control liquid pump.
[0021] In this embodiment, the software data acquisition sub-unit is used to respectively collect all target graphics cards to obtain a system layer data set. The system layer data set includes a core temperature data set, an operating frequency data set, a video memory occupancy ratio data set, and a video memory read volume data set. The processing logic includes: The software data acquisition sub-unit is used to establish independent acquisition threads for each target graphics card, scan the manufacturer ID data and device ID data of all target graphics cards through the PCI bus, initialize the API interface based on the manufacturer ID data and device ID data of the target graphics card, and collect the core temperature data set, the operating frequency data set, the video memory occupancy ratio data set, and the video memory read volume data set through the corresponding API interfaces of each target graphics card.
[0022] Among them, by independently establishing an acquisition thread for each graphics card and scanning the manufacturer ID and device ID of the graphics card to initialize the exclusive API interface, it is beneficial to accurately adapt and collect the software layer data of multiple graphics cards from different manufacturers, and solves the problem of incomplete or conflicting data collection caused by graphics card driver differences in the prior art.
[0023] In this embodiment, the time synchronization subunit is used to provide a unified time source for the hardware data acquisition subunit and the software data acquisition subunit through the hardware clock synchronization protocol PTS, and add a unified timestamp to all data acquired by the hardware data acquisition subunit and the software data acquisition subunit.
[0024] In this embodiment, the analysis and prediction unit includes a feature extraction subunit and a temperature prediction subunit; The feature extraction subunit is used to perform preliminary feature extraction on the output data of the data detection unit to obtain a preliminary feature data set, wherein the preliminary feature data set includes a first feature vector set, a second feature vector set and a target modal component. The processing logic of the preliminary feature extraction includes: A reference time slot is established based on a preset reference time window step. The physical layer data set and the system layer data set are divided by the reference time slot. Time phase alignment is performed by linear interpolation. The maximum tolerance delay is set. The data that exceeds the maximum tolerance delay is directly marked as delayed data. The delayed data is removed to obtain a synchronous time series data set. The synchronization timing data set includes coolant flow rate synchronization timing, voltage synchronization timing, core temperature synchronization timing, operating frequency synchronization timing, video memory occupancy ratio synchronization timing and video memory reading amount synchronization timing.
[0025] Among them, although the physical layer dataset and the system layer dataset have completed time synchronization during the collection process, the sampling frequencies of software data collection and hardware data collection are different. The feature extraction subunit first performs time axis normalization, divides the data slots by the reference time window step, and combines linear interpolation and maximum tolerance delay mechanism to achieve dynamic time axis normalization of hardware and software data. This is conducive to avoiding the timing misalignment problem caused by sampling frequency differences when collecting multi-source heterogeneous data, and providing a strictly aligned time series dataset for subsequent processing.
[0026] In this embodiment, the temperature change rate corresponding to each target graphics card is calculated based on the preset first time window step and the core temperature synchronization timing, and the memory bandwidth utilization rate corresponding to each target graphics card is calculated based on the preset second time window step and the memory reading amount synchronization timing; A first feature vector set is established based on the coolant flow rate synchronization timing, voltage synchronization timing, core temperature synchronization timing and temperature change rate, and a second feature vector set is established based on the operating frequency synchronization timing, video memory occupancy ratio synchronization timing and video memory bandwidth utilization; The core temperature synchronization time series is processed by variational mode decomposition (VMD) to obtain the temperature intrinsic mode component set. The energy entropy value of each mode component is calculated based on the temperature intrinsic mode component set, and the modal component screening process is performed to retain the modal components that are less than or equal to the preset energy entropy threshold to obtain the target modal components.
[0027] Among them, variational mode decomposition (VMD) and energy entropy screening are innovatively introduced into the analysis of the graphics card temperature to separate the noise and change trends in the temperature fluctuations of the temperature signal, and the modal components with prominent physical meanings are screened out through the energy entropy threshold, which is beneficial to solving the problems of temperature time-series noise interference and multi-scale feature coupling.
[0028] In this embodiment, the temperature prediction subunit is used to calculate the predicted temperature value of the target graphics card according to the preliminary feature dataset and the temperature rise prediction model. The processing logic of the temperature rise prediction model includes: Subsequently, the time-series data is processed through a two-stage gated LSTM architecture. The variance sensitivity of each data in the first feature vector set is calculated through a spatial attention mechanism to obtain a spatial attention weight matrix, and the weight matrix and the original feature vector are subjected to a Hadamard product operation to obtain a hardware feature vector; Causal convolution with a dilation coefficient of K is performed on the second feature vector set through temporal dilated convolution to obtain a multi-scale convolution result, and a gated linear unit (GLU) is used to perform non-linear fusion processing on the multi-scale convolution result to obtain a software feature vector; The hardware feature vector is processed through a short-term memory unit to perform sliding window feature extraction calculation for seconds to obtain a short-term dynamic hidden state. The short-term memory unit is a window LSTM network with a duration of seconds; The software feature vector is subjected to second periodic cumulative modeling through a long-term memory unit, and downsampling calculation is performed to obtain a long-term trend hidden state. The long-term memory unit is a window LSTM network with a duration of seconds; The long-term dynamic hidden state is used as the query vector Query, and the short-term dynamic hidden state is used as the key-value pair Key-Value. The long-term-short-term attention distribution weight is calculated through a cross-attention mechanism, and the short-term dynamic hidden state is weighted and summed based on the long-term-short-term attention distribution weight to obtain a scale fusion feature vector;
[0029] The target modal component and the multi-scale feature vector are vector-concatenated according to the reference time slot to obtain a multi-dimensional joint feature vector. One-dimensional average pooling in the time dimension is used to perform dimensionality reduction processing on the multi-dimensional joint feature vector, and regression and residual correction processing are performed through a fully connected layer, and then inverse normalization reduction is performed to obtain the predicted temperature value.
[0030] Among them, the traditional LSTM only relies on a single time window. In this solution, through hierarchical memory and modal fusion, the dual-stage LSTM and spatio-temporal attention mechanism can effectively capture the sudden changes in hardware states such as voltage fluctuations of the target graphics card and the laws of long-term software loads such as rendering tasks. Through spatial attention weighting of hardware features and multi-scale causal convolution of software features, and then through cross-attention fusion of long short-term memory units, the accuracy of temperature prediction is improved.
[0031] In this embodiment, the cooling execution unit includes a flow rate control subunit and a collaborative allocation subunit; The flow rate control subunit is used to adjust the coolant flow rate in real time according to the predicted temperature value, and the processing logic includes: Calculate the difference between the predicted temperature value corresponding to each target graphics card and the core temperature data to obtain the predicted temperature deviation. When the predicted temperature deviation is less than or equal to the preset first regulation temperature threshold in consecutive J reference time windows, keep the coolant flow rate unchanged; When the predicted temperature deviation is greater than the preset first regulation temperature threshold in consecutive J reference time windows, trigger the regulation of the coolant flow rate. The processing logic of the flow rate regulation includes: Calculate the coolant flow rate demand value based on the predicted temperature value and the thermal model, and obtain the feedforward PWM duty ratio data based on the duty ratio-flow characteristic curve corresponding to the speed control liquid pump; Deploy a PID controller for each target graphics card node, establish a PID control equation with the predicted temperature deviation as the real-time error e(t), calculate the feedback control signal u(t) through the PID control equation, perform linear mapping processing on the feedback control signal u(t) to obtain the feedback PWM duty ratio data, perform superposition calculation on the feedforward control signal and the feedback control signal to obtain the PWM execution data, and adjust the rotation speed of the speed control liquid pump based on the PWM execution data to achieve coolant flow rate control; Its calculation expression is: ; ; ; Among them, represents the coolant flow rate demand value, represents the heat capacity of the liquid cooling system, represents the predicted temperature value, represents the coolant inlet temperature, represents the preset reference time window step size, represents the thermal resistance of the liquid cooling system, represents the coolant density, represents the coolant heat capacity, represents the allowable temperature rise of the coolant, represents the feedback control signal, represents the predicted temperature deviation, represents the PID proportional coefficient, represents the PID integral coefficient, represents the PID derivative coefficient, represents, represents the minimum duty cycle, represents the maximum feedback value of the PID theoretical output, represents the minimum feedback value of the PID theoretical output, represents the feedback control signal.
[0032] Among them, the two-way superposition processing of the feedforward predicted by the thermal model and the PID control feedback, dynamically superimposing the feedforward duty cycle and the PID feedback signal based on the predicted temperature deviation, breaks through the limitation of the traditional PID control relying on the lagging temperature feedback, is conducive to the predictive adjustment of the coolant flow rate, and improves the liquid cooling efficiency of the graphics card.
[0033] In this embodiment, the collaborative allocation subunit is used for the collaborative allocation of the coolant for the target graphics card, and the processing logic includes: When the predicted temperature deviation corresponding to a graphics card in J consecutive reference time windows is greater than the preset second regulation temperature threshold, mark this graphics card as a high-priority graphics card and trigger the temporary collaborative allocation of the coolant flow rate. The processing logic includes: Increase the flow quota by H% for the high-priority graphics card, count the coolant flow rate demand value and the coolant pressure difference of each target graphics card, perform cyclic iterative calculation with the preset constraint pressure difference P% through the ADMM consensus algorithm to obtain the collaborative flow allocation weight, and obtain the collaborative flow allocation data set based on the collaborative flow allocation weight.
[0034] Among them, the processing logic of the ADMM consensus algorithm is: use the preset constraint pressure difference P% as the global constraint condition of the ADMM algorithm, use the coolant pressure differences of all target graphics cards as the optimization target, perform alternating direction iterative calculation on the global constraint condition and the local optimization target through the ADMM algorithm to obtain the collaborative flow allocation weight, and then calculate the collaborative flow allocation reference value. The node corrects the local flow execution data based on the reference value and sends back the actual pressure difference feedback. The coordinator updates the Lagrange multiplier parameter based on the feedback data for secondary optimization, and performs cyclic iteration until the flow data of each node and the global pressure difference converge within the range of P%. It is conducive to realizing the collaborative control of the rapid cooling of the locally overheated graphics card and the stability of the global liquid cooling system through distributed convergence.
[0035] In this embodiment, the terminal interaction unit includes a display interaction subunit and a data storage subunit; The display interaction subunit is used to provide a display interface for the user and receive user operation instructions, and receive instructions issued by the user for parameter adjustment settings of the data detection unit, the analysis and prediction unit, and the cooling execution unit; The data storage subunit is used to receive all the output data of the data detection unit, the analysis and prediction unit, and the cooling execution unit for storage, and generate a historical log of the operation of the liquid cooling module.
[0036] Those skilled in the art should understand that the embodiments of the present invention can provide a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. Among them, the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, abbreviated as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read Only Memory, abbreviated as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, abbreviated as EPROM), programmable read-only memory (Programmable Red Only Memory, abbreviated as PROM), read-only memory (Read Only Memory, abbreviated as ROM), magnetic memory, flash memory, magnetic disk or optical disk. These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, and the instruction device implements the functions specified in one process Figure 1 one process or multiple processes and / or boxes Figure 1 functions specified in one box or multiple boxes.
[0037] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A graphics card liquid cooling module with intelligent temperature control, characterized in that: include: Data detection unit, analysis and prediction unit, cooling execution unit and terminal interaction unit; The data detection unit is used to obtain the hardware layer data set, software data set and condensed data set of the target graphics card, and add a unified timestamp to all the collected data; The analysis and prediction unit is used to perform preliminary feature extraction on the output data of the data detection unit to obtain a first feature vector set, a second feature vector set and a target modal component, and calculate the predicted temperature of the target graphics card in combination with the temperature rise prediction model; The cooling execution unit is used to adjust the coolant flow rate in real time according to the predicted temperature value, and to coordinate the distribution of coolant to all target graphics cards in combination with the ADMM consistency algorithm; The terminal interaction unit is used to provide a display interface to the user and receive user operation instructions, receive all output data of the data detection unit, the analysis and prediction unit and the cooling execution unit and store them.
2. A graphics card liquid cooling module with intelligent temperature control as claimed in claim 1, characterized in that: The data detection unit includes a hardware data acquisition subunit, a software data acquisition subunit and a time synchronization subunit; The hardware data acquisition subunit is used to acquire physical layer data sets from all target graphics cards respectively, and acquire condensation data sets from the coolant pipelines; The physical layer data set includes a coolant flow rate data set and a voltage data set, and the processing logic includes: A coolant flow rate data set of the target graphics card is acquired through a Hall flow meter, and a voltage data set of the target graphics card is acquired through a voltage sensor; The condensation data set includes coolant inlet temperature, coolant pressure difference, and speed regulating liquid pump duty cycle.
3. A graphics card liquid cooling module with intelligent temperature control as claimed in claim 2, characterized in that: The software data acquisition subunit is used to acquire system layer data sets from all target graphics cards respectively. The system layer data sets include core temperature data sets, operating frequency data sets, video memory occupancy ratio data sets and video memory reading amount data sets. The processing logic includes: The software data acquisition subunit is used to establish an independent acquisition thread for each target graphics card, obtain the manufacturer ID data and device ID data of all target graphics cards through PCI bus scanning, initialize the API interface based on the manufacturer ID data and device ID data of the target graphics card, and obtain the core temperature data set, operating frequency data set, video memory occupancy ratio data set and video memory reading amount data set through the corresponding API interface of each target graphics card.
4. A graphics card liquid cooling module with intelligent temperature control as claimed in claim 3, characterized in that: The time synchronization subunit is used to provide a unified time source for the hardware data acquisition subunit and the software data acquisition subunit through the hardware clock synchronization protocol PTS, and to add a unified time stamp to all data acquired by the hardware data acquisition subunit and the software data acquisition subunit.
5. A graphics card liquid cooling module with intelligent temperature control as claimed in claim 1, characterized in that: The analysis and prediction unit includes a feature extraction subunit and a temperature prediction subunit; The feature extraction subunit is used to perform preliminary feature extraction on the output data of the data detection unit to obtain a preliminary feature data set, wherein the preliminary feature data set includes a first feature vector set, a second feature vector set and a target modal component. The processing logic of the preliminary feature extraction includes: A reference time slot is established based on a preset reference time window step. The physical layer data set and the system layer data set are divided by the reference time slot. Time phase alignment is performed by linear interpolation. The maximum tolerance delay is set. The data that exceeds the maximum tolerance delay is directly marked as delayed data. The delayed data is removed to obtain a synchronous time series data set. The synchronization timing data set includes coolant flow rate synchronization timing, voltage synchronization timing, core temperature synchronization timing, operating frequency synchronization timing, video memory occupancy ratio synchronization timing and video memory reading amount synchronization timing.
6. A graphics card liquid cooling module with intelligent temperature control as claimed in claim 5, characterized in that: The temperature change rate corresponding to each target graphics card is calculated based on the preset first time window step and the core temperature synchronization timing, and the memory bandwidth utilization rate corresponding to each target graphics card is calculated based on the preset second time window step and the memory reading amount synchronization timing; A first feature vector set is established based on the coolant flow rate synchronization timing, voltage synchronization timing, core temperature synchronization timing and temperature change rate, and a second feature vector set is established based on the operating frequency synchronization timing, video memory occupancy ratio synchronization timing and video memory bandwidth utilization; The core temperature synchronization time series is processed by variational mode decomposition (VMD) to obtain the temperature intrinsic mode component set. The energy entropy value of each mode component is calculated based on the temperature intrinsic mode component set, and the modal component screening process is performed to retain the modal components that are less than or equal to the preset energy entropy threshold to obtain the target modal components.
7. A graphics card liquid cooling module with intelligent temperature control as claimed in claim 5, characterized in that: The temperature prediction subunit is used to calculate the predicted temperature value of the target graphics card according to the preliminary feature data set and the temperature rise prediction model. The processing logic of the temperature rise prediction model includes: Then, the time series data is processed through a two-stage gated LSTM architecture, and the variance sensitivity of each data in the first eigenvector set is calculated through the spatial attention mechanism to obtain the spatial attention weight matrix. The weight matrix is then Hadamard-producted with the original eigenvector to obtain the hardware eigenvector. The second feature vector set is subjected to causal convolution calculation with a dilation coefficient of K through time-dwelling convolution to obtain a multi-scale convolution result, and the multi-scale convolution result is subjected to nonlinear fusion processing using a gated linear unit GLU to obtain a software feature vector; The hardware feature vector is processed through the short-term memory unit The short-term dynamic hidden state is calculated by sliding window feature extraction of seconds, and the short-term memory unit is of duration Window LSTM network of seconds; The software feature vector is stored in the long-term memory unit The long-term trend hidden state is obtained by downsampling calculation. The long-term memory unit is the duration Window LSTM network of seconds; The long-term dynamic hidden state is used as the query vector Query, and the short-term dynamic hidden state is used as the key-value pair Key-Value. The long-term-short-term attention distribution weight is calculated through the cross-attention mechanism. The short-term dynamic hidden state is weighted summed based on the long-term-short-term attention distribution weight to obtain the scale fusion feature vector. The target modal component and the multi-scale feature vector are concatenated according to the reference time slot to obtain a multi-dimensional joint feature vector. The multi-dimensional joint feature vector is reduced in dimension by one-dimensional average pooling in the time dimension. The predicted temperature value is obtained by regression and residual correction through the fully connected layer and then denormalization. The multi-dimensional joint feature vector is subjected to one-dimensional average pooling dimensionality reduction processing through the fully connected layer, and the core temperature synchronization time series data is superimposed through the residual connection. The predicted temperature value is obtained through the Tanh activation function and anti-normalization restoration processing.
8. A graphics card liquid cooling module with intelligent temperature control as claimed in claim 1, characterized in that: The cooling execution unit includes a flow rate control subunit and a coordination distribution subunit; The flow rate control subunit is used to adjust the coolant flow rate in real time according to the predicted temperature value. The processing logic includes: The predicted temperature value corresponding to each target graphics card is calculated by subtracting the core temperature data from the predicted temperature value, and when the predicted temperature deviation in J consecutive benchmark time windows is less than or equal to a preset first control temperature threshold, the coolant flow rate is kept unchanged; When the predicted temperature deviation in J consecutive reference time windows is greater than the preset first control temperature threshold, the coolant flow rate control is triggered, and the processing logic of the flow rate control includes: The coolant flow demand value is calculated based on the predicted temperature value and the thermal model, and the feedforward PWM duty cycle data is obtained based on the coolant flow demand value and the duty cycle-flow characteristic curve corresponding to the speed regulating liquid pump; Deploy a PID controller for each target graphics card node, use the predicted temperature deviation as the real-time error e(t) to establish a PID control equation, calculate the feedback control signal u(t) through the PID control equation, perform linear mapping on the feedback control signal u(t) to obtain feedback PWM duty cycle data, superimpose the feedforward control signal and the feedback control signal to obtain PWM execution data, and adjust the speed of the speed regulating liquid pump based on the PWM execution data to achieve coolant flow rate control; Its calculation expression is: ; ; ; in, Indicates the coolant flow demand value, represents the heat capacity of the liquid cooling system, represents the predicted temperature value, Indicates the coolant inlet temperature, Represents the preset reference time window step, represents the thermal resistance of the liquid cooling system, Indicates the coolant density, represents the heat capacity of the coolant, Indicates that the coolant is allowed to heat up. represents the feedback control signal, represents the predicted temperature deviation, Indicates the PID proportional coefficient, Indicates PID integral coefficient, represents the PID differential coefficient, express, Indicates the minimum duty cycle, Indicates the maximum feedback value of PID theoretical output, Indicates the minimum feedback value of PID theoretical output, Represents the feedback control signal.
9. A graphics card liquid cooling module with intelligent temperature control as claimed in claim 8, characterized in that: The co-allocation subunit is used to co-allocate coolant to the target graphics card. The processing logic includes: When there is a graphics card whose corresponding predicted temperature deviation in J consecutive benchmark time windows is greater than the preset second control temperature threshold, the graphics card is marked as a high-priority graphics card, triggering temporary coordinated allocation of coolant flow, and the processing logic includes: Increase the flow quota H% for high-priority graphics cards, count the coolant flow demand value and coolant pressure difference of each target graphics card, and use the ADMM consistency algorithm to perform cyclic iterative calculations with the preset constraint pressure difference P% to obtain the collaborative flow allocation weight, and obtain the collaborative flow allocation data set based on the collaborative flow allocation weight.
10. A graphics card liquid cooling module with intelligent temperature control as claimed in claim 1, characterized in that: The terminal interaction unit includes a display interaction subunit and a data storage subunit; The display interaction subunit is used to provide a display interface to the user and receive user operation instructions, and receive instructions issued by the user to adjust parameters of the data detection unit, the analysis and prediction unit, and the cooling execution unit; The data storage subunit is used to receive all output data from the data detection unit, the analysis and prediction unit, and the cooling execution unit for storage, and to generate a historical log of the operation of the liquid cooling module.
Citation Information
Patent Citations
A water-cooled heat sink and a method of a computer graphics card
CN109308109A