Hardware chip temperature monitoring and automatic protection system based on artificial intelligence
Through an artificial intelligence-based hardware chip temperature monitoring system, combined with multi-dimensional data analysis and dynamic protection strategies, the lack of flexibility in the existing technology and the limitations of traditional monitoring methods are solved, and more efficient and stable temperature management is achieved to ensure the safety and reliability of the chip.
Patent Information
- Application Number
- CN202510215146.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art lacks flexibility in hardware chip temperature management, and cannot adapt to the thermal change characteristics under different workloads and environmental conditions, resulting in too late or unnecessary performance losses for protection triggers, and traditional monitoring methods cannot capture local hot spots in time, increasing system operation uncertainty.
Using a hardware chip temperature monitoring system based on artificial intelligence, through data acquisition, preprocessing, feature extraction, temperature prediction, abnormality detection and protection modules, combined with physically guided neural networks and deep learning, a comprehensive analysis of the temperature distribution and operating state of multiple areas of the chip is realized, and protection strategies are dynamically adjusted to adapt to complex environments.
It improves the accuracy and real-time nature of temperature management, reduces the risk of misjudgment, optimizes the adaptability of temperature control strategies, ensures the operational safety and reliability of the chip, and extends the service life of the chip.
Smart Images

Figure CN120336111A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and more particularly, to a temperature monitoring and automatic protection system for hardware chips based on artificial intelligence. Background Art
[0002] In modern electronic devices and high-performance computing systems, the working stability of hardware chips directly affects the reliability of the entire system. With the continuous evolution of chip manufacturing processes and the improvement of computing capabilities, the integration and power consumption of a single chip have been increasing continuously, resulting in more serious heat accumulation inside the chip. Under high load and long-term operation conditions, the temperature of the chip may rise rapidly, affecting its normal operation and even causing irreversible physical damage. Therefore, the monitoring and management of chip temperature have become an important link to ensure the stable operation of hardware.
[0003] In related technologies, common chip temperature management methods usually rely on temperature sensors arranged fixedly to collect temperature data and trigger an alarm or protection mechanism after the temperature reaches a set critical value. For example, some systems adopt a temperature control scheme based on simple threshold judgment. When it is detected that the chip temperature exceeds the safe range, operations such as frequency reduction, voltage reduction, or shutting down some computing units are performed to reduce power consumption and heat generation. However, such methods have obvious limitations. On the one hand, the fixed threshold setting lacks flexibility and cannot adapt to the thermal change characteristics of the chip under different workloads and environmental conditions, which may lead to too late protection triggering or unnecessary performance loss. On the other hand, relying solely on the data of temperature sensors may have hysteresis. Especially in the case where local hot spots are formed rapidly inside the chip, traditional monitoring methods often cannot capture sudden temperature anomalies in time, thus increasing the uncertainty of system operation.
[0004] In addition, related technologies usually only perform temperature management for the entire chip and do not fully consider the thermal coupling effect between different parts of the chip in the chip system, and thus cannot effectively optimize the system-level thermal distribution. Therefore, in a complex computing environment, the applicability and scalability of traditional temperature management schemes still have certain limitations.
[0005] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0006] The purpose of the embodiments of the present disclosure is to provide a temperature monitoring and automatic protection system for hardware chips based on artificial intelligence, so as to improve the accuracy and real-time performance of temperature control management and ensure the operation safety of the chip.
[0007] According to the first aspect of the embodiments of the present disclosure, there is provided an artificial intelligence-based hardware chip temperature monitoring and automatic protection system, including:
[0008] A data acquisition module, configured to obtain the current temperature distribution data of multiple regions of the chip and synchronously obtain the operating state data of the chip;
[0009] A data preprocessing module, configured to preprocess the current temperature distribution data and the operating state data to obtain a chip monitoring data set;
[0010] A feature extraction module, configured to extract features from the chip monitoring data set to obtain a monitoring data feature vector of at least one feature dimension, and construct a multi-dimensional feature vector according to the monitoring data feature vector;
[0011] A temperature prediction module, configured to input the multi-dimensional feature vector into a pre-trained temperature prediction model to infer the future temperature distribution data of the chip at the next moment;
[0012] An anomaly detection module, configured to determine the working state of the chip according to the current temperature distribution data and the future temperature distribution data, and calculate a chip temperature anomaly score when the chip is in an abnormal working state;
[0013] An anomaly analysis module, configured to determine the chip protection level through the chip temperature anomaly score, and determine the chip temperature protection operation based on the chip protection level;
[0014] A temperature protection module, configured to optimize and adjust the operating parameters of the chip according to the chip temperature protection operation to achieve temperature protection of the chip.
[0015] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0016] In the artificial intelligence-based hardware chip temperature monitoring and automatic protection system in the exemplary embodiments of the present disclosure, temperature data is collected in multiple thermal field regions of the chip, and comprehensive analysis is performed in combination with the operating state data, so that the temperature distribution information can be accurately perceived within a larger range, avoiding the situation where temperature anomalies are not detected in a timely manner due to local heat accumulation in the chip or limited layout of local sensors; through the feature fusion of multi-dimensional data, the temperature prediction can have stronger adaptability, reduce the prediction deviation caused by the fluctuation of a single variable, and improve the ability to judge the temperature change trend under complex operating conditions; by constructing a physics-guided neural network and optimizing the balance relationship between the data-driven loss and the physical constraint loss during the model training process, the temperature prediction can maintain an accurate simulation of the internal heat conduction characteristics of the chip and still provide stable prediction results under long-term operation or large changes in environmental conditions; further, after comparing the predicted temperature data with the current temperature data, comprehensively considering the temperature change rate, historical operating state and statistical distribution characteristics, calculate the temperature anomaly score, so that the anomaly judgment can be comprehensively evaluated from multiple dimensions, reduce the misjudgment risk brought by the single threshold judgment method, and improve the adaptability to temperature anomalies under different load modes; after detecting a temperature anomaly, determine the protection level of the chip based on a preset score interval, and combine the strategies corresponding to different protection levels to execute targeted temperature control measures, so that the temperature control is not limited to simply reducing the frequency or shutting down the computing unit, but can be dynamically adjusted between different protection strategies based on the current operating requirements; during the long-term operation of temperature management, by continuously monitoring the operating state of the chip and analyzing the execution effect of the temperature control strategy, the system can dynamically optimize the temperature management model, and through continuous recording and analysis of historical temperature data, anomaly detection records and the execution effect of protection strategies, the temperature control system can gradually learn and adapt to different chip architectures and operating environments, so that the temperature control strategy has stronger optimization ability during long-term operation, and through multi-level collaborative optimization of real-time monitoring, temperature prediction, anomaly detection, protection level determination and protection strategy execution of the chip, the temperature management process can maintain efficient, stable and intelligent control ability, thereby effectively improving the reliability and service life of the chip.
[0017] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 Schematically shows a structural diagram of an artificial intelligence-based hardware chip temperature monitoring and automatic protection system according to some embodiments of the present disclosure.
[0020] Figure 2 Schematically shows a structural diagram of a data acquisition module according to some embodiments of the present disclosure.
[0021] Figure 3 Schematically shows a structural diagram of a feature extraction module according to some embodiments of the present disclosure.
[0022] Figure 4 Schematically shows a structural diagram of a temperature prediction module according to some embodiments of the present disclosure.
[0023] Figure 5 Schematically shows a structural diagram of a physical guidance training unit according to some embodiments of the present disclosure.
[0024] Figure 6 Schematically shows a structural diagram of an anomaly detection module according to some embodiments of the present disclosure.
[0025] Figure 7 Schematically shows a structural diagram of an anomaly analysis module according to some embodiments of the present disclosure.
[0026] Figure 8 Schematically shows a structural diagram of an artificial intelligence-based hardware chip temperature monitoring and automatic protection system according to some other embodiments of the present disclosure.
[0027] Figure 9 Schematically shows a structural diagram of a temperature protection module according to some embodiments of the present disclosure.
[0028] Figure 10 Schematically shows a structural diagram of a real-time monitoring module according to some embodiments of the present disclosure.
[0029] Figure 11 Schematically shows a structural diagram of a data preprocessing module according to some embodiments of the present disclosure.
[0030] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts. Detailed Implementation Manner
[0031] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of this specification as detailed in the appended claims.
[0032] In addition, the drawings are only schematic diagrams and are not necessarily drawn to scale. The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0033] In the present exemplary embodiment, first, a hardware chip temperature monitoring and automatic protection system based on artificial intelligence is provided. The hardware chip temperature monitoring and automatic protection system based on artificial intelligence can be applied to terminal devices, such as electronic devices like mobile phones and computers, and can also be applied to servers, especially deep learning servers with large computing volumes. Figure 1 The structural schematic diagram of a hardware chip temperature monitoring and automatic protection system based on artificial intelligence according to some embodiments of the present disclosure is schematically shown. Refer to Figure 1 As shown, the hardware chip temperature monitoring and automatic protection system 100 based on artificial intelligence may include:
[0034] A data acquisition module 110, configured to obtain the current temperature distribution data of multiple regions of the chip 10, and synchronously obtain the operation state data of the chip 10;
[0035] A data preprocessing module 120, configured to preprocess the current temperature distribution data and the operation state data to obtain a chip monitoring data set;
[0036] A feature extraction module 130, configured to extract features from the chip monitoring data set to obtain a monitoring data feature vector of at least one feature dimension, and construct a multi-dimensional feature vector according to the monitoring data feature vector;
[0037] A temperature prediction module 140, configured to input the multi-dimensional feature vector into a pre-trained temperature prediction model to infer the future temperature distribution data of the chip 10 at the next moment;
[0038] Anomaly detection module 150 is configured to determine the operating state of the chip 10 based on the current temperature distribution data and the future temperature distribution data, and calculate a chip temperature anomaly score when the chip 10 is in an abnormal operating state;
[0039] Anomaly analysis module 160 is configured to determine a chip protection level based on the chip temperature anomaly score, and determine a chip temperature protection operation based on the chip protection level;
[0040] Temperature protection module 170 is configured to optimize and adjust the operating parameters of the chip 10 according to the chip temperature protection operation to achieve temperature protection for the chip 10.
[0041] Based on the artificial intelligence-based hardware chip temperature monitoring and automatic protection system 100 in this exemplary embodiment, temperature data is collected in multiple thermal field regions of the chip, and comprehensive analysis is performed in combination with the operating state data, so that the temperature distribution information can be accurately perceived within a larger range, avoiding the situation where temperature anomalies are not detected in time due to local heat accumulation of the chip 10 or limited local sensor layout; through the feature fusion of multi-dimensional data, the temperature prediction can have stronger adaptability, reduce the prediction deviation caused by the fluctuation of a single variable, and improve the ability to judge the temperature change trend under complex operating conditions; by constructing a physics-guided neural network and optimizing the balance relationship between the data-driven loss and the physical constraint loss during the model training process, the temperature prediction can maintain an accurate simulation of the internal heat conduction characteristics of the chip and still provide stable prediction results under long-term operation or large changes in environmental conditions; further, after comparing the predicted temperature data with the current temperature data, comprehensively considering the temperature change rate, historical operating state, and statistical distribution characteristics, calculate the temperature anomaly score, so that the anomaly judgment can be comprehensively evaluated from multiple dimensions, reduce the misjudgment risk brought by the single threshold judgment method, and improve the adaptability to temperature anomalies under different load modes; after detecting a temperature anomaly, determine the protection level of the chip based on a preset score interval, and combine the strategies corresponding to different protection levels to execute targeted temperature control measures, so that the temperature regulation is not limited to simply reducing the frequency or shutting down the computing unit, but can be dynamically adjusted between different protection strategies based on the current operating requirements; during the long-term operation of temperature management, by continuously monitoring the operating state of the chip and analyzing the execution effect of the temperature control strategy, the system can dynamically optimize the temperature management model. Through continuous recording and analysis of historical temperature data, anomaly detection records, and the execution effect of protection strategies, the temperature control system can gradually learn and adapt to different chip architectures and operating environments, so that the temperature control strategy has stronger optimization ability during long-term operation. Through multi-level collaborative optimization of real-time monitoring, temperature prediction, anomaly detection, protection level determination, and protection strategy execution of the chip, the temperature management process can maintain efficient, stable, and intelligent control ability, ensure the safe operation of the chip, and thus effectively improve the reliability and service life of the chip.
[0042] Next, the artificial intelligence-based hardware chip temperature monitoring and automatic protection system 100 in this exemplary embodiment will be further described.
[0043] In an exemplary embodiment of the present disclosure, referring to Figure 2 As shown, the data acquisition module 110 may include a temperature sensor array 111 and a dynamically adjustable sensor control unit 112. Specifically:
[0044] A temperature sensor array 111, including at least two temperature sensors, is used to be arranged in a plurality of thermal field regions corresponding to the chip 10 determined in advance to collect temperature distribution data;
[0045] A dynamically adjustable sensor control unit 112, electrically connected to the temperature sensor array 111, is used to adaptively adjust the sampling frequency of each temperature sensor according to the real-time thermal field distribution of the chip 10.
[0046] Among them, the temperature sensor array 111 can be used to collect temperature distribution data in a plurality of thermal field regions of the chip 10 to ensure the comprehensiveness and accuracy of temperature monitoring. The temperature sensor array 111 can include at least two temperature sensors, and each temperature sensor is respectively arranged at a predetermined monitoring point or a predetermined thermal field region of the chip 10 to detect the temperature information of this region. The arrangement method of the temperature sensors can be optimized according to the heat conduction characteristics, power consumption distribution and the positions of key functional modules of the chip 10, so that the temperature acquisition covers the high heat density regions and the regions where local hot spots may occur on the chip 10, thereby improving the timeliness and accuracy of temperature monitoring.
[0047] In practical applications, the temperature sensor can be a micro sensor integrated inside the chip 10, or an external patch temperature sensor or an infrared temperature sensor to meet the measurement requirements in different chip packaging forms and different heat dissipation environments. This embodiment does not make special limitations on the type of temperature sensor; for high-precision application scenarios, high-resolution temperature sensors can be used and combined with a multi-point measurement calibration algorithm to reduce the temperature deviation caused by measurement errors or environmental interference.
[0048] In the process of obtaining temperature data, each temperature sensor needs to have the characteristics of high precision, high sensitivity and low power consumption to meet the real-time requirements of in-chip thermal management. The measurement accuracy of the temperature sensor is usually limited by its own material characteristics, environmental noise and the stability of the measurement circuit. Therefore, it is necessary to combine a high-precision analog front-end (AFE) circuit to amplify and filter the temperature sensor signal to improve the stability of temperature measurement. The measured data can be converted into digital signals by an analog-to-digital converter (ADC) and then transmitted to the data acquisition module for subsequent processing. The transmission of temperature data can adopt a serial peripheral interface (SPI), a serial communication bus I 2C (Inter-Integrated Circuit), Bluetooth Low Energy (BLE), or other data bus protocols suitable for embedded systems to reduce latency during data transmission and improve the real-time response ability of the system.
[0049] The dynamically adjustable sensor control unit 112 is electrically connected to the temperature sensor array 111 and is used to adaptively adjust the sampling frequency of each temperature sensor according to the real-time thermal field distribution of the chip, so as to optimize the energy efficiency ratio and monitoring accuracy of data acquisition.
[0050] During the operation of the chip, the temperature distribution has dynamic change characteristics. Different load modes, heat dissipation conditions, and environmental temperatures may all cause changes in the chip's thermal field. Therefore, a temperature monitoring scheme with a fixed sampling frequency may not be able to adapt to the rapid changes of local hot spots inside the chip or may cause resource waste. The dynamically adjustable sensor control unit 112 can calculate the temperature change rate of the area where each temperature sensor is located by obtaining the chip operation state data and combining the current temperature change trend, and adaptively change the sampling frequency based on the set dynamic adjustment strategy. For example, when it detects that the temperature change rate exceeds a preset threshold, it increases the sampling frequency of the corresponding sensor to capture the rapidly changing temperature dynamics; in areas with low-load operation or small temperature changes, it reduces the sensor sampling frequency to reduce data redundancy and power consumption burden. The adjustment of the sampling frequency can be carried out in ways such as fixed interval adjustment, exponential adjustment, or adaptive feedback adjustment to meet the requirements of different chip architectures and application scenarios. This exemplary embodiment does not make special limitations on this.
[0051] The dynamically adjustable sensor control unit 112 can adopt a control strategy based on Digital Signal Processing (DSP) to calculate the temperature change rate in real time and dynamically adjust the sampling interval through a preset adaptive algorithm. In specific implementation, the dynamically adjustable sensor control unit 112 can evaluate the trend of chip temperature change based on statistical analysis methods such as moving average filtering or time window sliding calculation, and combine a prediction model to judge the future temperature fluctuation situation, so as to optimize the sampling strategy in advance. In addition, the dynamically adjustable sensor control unit 112 can also adopt a pattern recognition method based on neural networks to classify the thermal distribution patterns under different working states and dynamically adjust the temperature monitoring strategy to improve the sampling efficiency and monitoring accuracy. This exemplary embodiment does not make special limitations on the way of evaluating the trend of chip temperature change.
[0052] During the dynamic adjustment process, to avoid the system overhead caused by frequent changes in the sampling frequency, a minimum adjustment interval can be set to ensure that each sampling adjustment is completed within a reasonable range. In addition, to reduce the computational overhead, the sensor control unit can be implemented in hardware, such as an embedded system based on a System on Chip (SoC) or a Microcontroller Unit (MCU), to provide low-power and high-performance data acquisition and control capabilities. In a High-Performance Computing (HPC) or server environment, hardware acceleration can be achieved by combining a Field-Programmable Gate Array (FPGA) or an Application-Specific Integrated Circuit (ASIC) to meet the high-throughput and high-concurrency data acquisition requirements.
[0053] In a multi-chip system, the dynamically adjustable sensor control unit 112 can further perform global optimization based on the cross-chip temperature distribution information to adjust the sampling strategies of different chips and improve the system-level temperature management capabilities. For example, when local overheating of a certain chip 10 is detected, the temperature sampling density of adjacent chips 10 can be increased to evaluate the degree of thermal interference between chips and provide support for subsequent temperature protection strategies. In a mobile device or a low-power embedded system, the sampling frequency can be reduced during the execution of non-critical tasks in combination with the low-power mode adjustment strategy to reduce energy consumption and improve the overall battery life of the system.
[0054] Through the collaborative work of the temperature sensor array 111 and the dynamically adjustable sensor control unit 112, the temperature monitoring of the chip 10 can not only cover the temperature distribution in different regions but also adapt to dynamic thermal field changes, improving the monitoring accuracy and energy efficiency and providing high-quality data input for subsequent temperature prediction, anomaly detection, and optimization of protection strategies.
[0055] In an exemplary embodiment of the present disclosure, as shown in Figure 3 the feature extraction module 130 may include a frequency domain analysis unit 131, a causal inference unit 132, and a feature encoding unit 133, where:
[0056] The frequency domain analysis unit 131 is configured to perform a short-time Fourier transform on the data in the chip monitoring dataset to extract temperature change trend features and voltage and frequency fluctuation features;
[0057] The causal inference unit 132 is configured to analyze the correlation between the changes in the chip operating state and the temperature change by using a causal inference method to generate causal association features;
[0058] The feature encoding unit 133, electrically connected to the frequency domain analysis unit 131 and the causal inference unit 132, is configured to perform feature fusion on the causal association feature, the temperature change trend feature, and the voltage and frequency fluctuation feature, construct a multi-dimensional feature vector, and send the multi-dimensional feature vector to the temperature prediction module.
[0059] Among them, in the process of processing the chip monitoring data set, the frequency domain analysis unit 131 can perform Short-Time Fourier Transform (STFT) processing on the temperature distribution data and related operating state data to extract the frequency domain features of temperature changes. The heat conduction process of the chip is affected by various factors, such as transient changes in circuit load, adjustment of power consumption patterns, and fluctuations in heat dissipation conditions, making the temperature signal exhibit complex time-frequency characteristics. Traditional time-domain analysis methods cannot effectively separate high-frequency and low-frequency temperature change signals. Therefore, the short-time Fourier transform can convert the temperature data from the time domain to the frequency domain, enabling the contribution of different frequency components to temperature changes to be analyzed.
[0060] In practical applications, the frequency domain analysis unit 131 can select different window functions, such as the Hanning Window and the Kaiser Window, to perform frequency domain analysis on the data in the chip monitoring data set to optimize the time-frequency resolution and ensure the accurate capture of short-time temperature change signals. In addition, during the feature extraction process, empirical mode decomposition (EMD) or wavelet transform can be combined to perform multi-scale decomposition on the temperature signal to further improve the ability to identify non-linear temperature change patterns.
[0061] After obtaining the frequency-domain features of temperature changes, the features of the chip's operating state can be further extracted to analyze the impact of voltage and frequency fluctuations on temperature changes. The power consumption change of chip 10 directly affects its temperature distribution, and the power consumption is controlled by the dynamic adjustment of voltage and frequency. Therefore, establishing the mapping relationship between voltage, frequency, power consumption, and temperature plays an important role in the accuracy of temperature prediction. Through the power consumption modeling method, the instantaneous power consumption can be calculated based on the time-series data of voltage and frequency, and the theoretical temperature change trend of different regions of the chip can be deduced in combination with the heat conduction model. In the specific implementation process, a linear regression model or a neural network model, such as a Long Short-Term Memory (LSTM) network, can be used to model the relationship between power consumption fluctuations and temperature changes, and the feature extraction results can be optimized through time-series prediction methods. In addition, for different chip architectures, the dynamic voltage and frequency scaling (DVFS) strategy can be combined to model the power consumption adjustment mode of the chip, and features highly correlated with temperature changes can be extracted based on different operating states to improve the quality of the input data of the prediction model.
[0062] After obtaining the temperature change trend features and operating state features, the causal relationship between different variables can be further analyzed through the causal reasoning unit 132 to generate causal association features. Since the temperature change of the chip is not only directly affected by the current operating state, but may also be affected by the combined effects of historical operating states and external environmental factors, traditional statistical correlation analysis methods are difficult to accurately describe the causal mechanism of chip thermal evolution. Through causal reasoning methods, for example, a causal graph can be constructed based on time-series data to infer the contribution degree of factors such as voltage mutation and load switching to temperature changes, and the impact of each factor on the temperature evolution process can be quantified. In the specific implementation process, Granger causality analysis or a structural causal model (SCM) can be used to model the operating state and temperature change data of chip 10, and feature components with strong causal relevance can be extracted. In addition, to further enhance the robustness of causal reasoning, a heat conduction modeling method based on a graph neural network (GNN) can be combined to learn the heat interaction mode between different chip regions, and spatial causal relationships can be introduced in the feature extraction process, enabling the prediction model to adapt to the temperature evolution trend under different heat dissipation conditions.
[0063] After completing the above feature extraction steps, the extracted features can be fused to construct an efficient multi-dimensional feature vector. The feature encoding unit 133 can receive the feature data from the frequency domain analysis unit 131 and the causal reasoning unit 132, and normalize the features from different sources to ensure the consistency of the numerical scales of different features. In the specific implementation process, principal component analysis (PCA) or autoencoder (AE) can be used to reduce the dimensionality of the feature vector, so as to reduce redundant information and improve the compactness of feature representation. In addition, in order to enhance the adaptability of the model to the feature data, the attention mechanism can also be used to weight different feature dimensions, so that the prediction model can dynamically adjust the degree of attention to each feature under different operating conditions, thereby optimizing the prediction effect. In the feature fusion process, a multi-modal fusion strategy based on time series can be adopted, such as a feature alignment method based on recurrent neural network (RNN), so that features of different time scales can be integrated in a unified feature space, improving the consistency and stability of feature expression.
[0064] After constructing the high-dimensional feature vector, the feature encoding unit 133 can send the final feature data to the temperature prediction module 140 to support subsequent temperature prediction calculations. Through the above feature extraction scheme, the chip monitoring data not only contains traditional temperature trend information, but also can comprehensively consider voltage, frequency fluctuations and causal relationships, improving the accuracy and adaptability of temperature prediction, and providing more stable input data for subsequent anomaly detection and temperature protection.
[0065] In an exemplary embodiment of the present disclosure, referring to Figure 4 as shown, the temperature prediction module 140 may include a physics-guided training unit 141 and a deep learning prediction unit 142. Specifically:
[0066] The physics-guided training unit 141 is configured to train a pre-constructed temperature prediction model according to the training data set corresponding to the chip. The temperature prediction model includes a physics-guided neural network constructed based on the heat conduction principle, and obtain a trained temperature prediction model;
[0067] The deep learning prediction unit 142 is configured to infer the future temperature distribution data of the chip at the next moment according to the trained temperature prediction model and in combination with the multi-dimensional feature vector.
[0068] Among them, the temperature prediction module 140 can input the multi-dimensional feature vector after feature extraction and perform inference through a prediction model based on a Physics-Informed Neural Network (PINN) to obtain the chip temperature distribution data at future moments. This module mainly includes a physics-guided training unit 141 and a deep learning prediction unit 142. Among them, the physics-guided training unit 141 is used to construct and train a temperature prediction model based on the heat conduction principle, while the deep learning prediction unit 142 is used to perform real-time inference using the trained model to predict the future temperature change of the chip.
[0069] Through the construction of the physics-guided training unit 141 and the deep learning prediction unit 142, chip temperature prediction can not only use deep learning methods for data-driven modeling, but also combine physical constraints to improve the credibility and stability of the prediction. In practical applications, the temperature prediction module 140 can be widely applied to scenarios with high requirements for temperature monitoring accuracy, such as server chips, high-performance computing devices, autonomous driving computing units, and intelligent mobile devices, providing efficient and stable chip temperature prediction capabilities and providing high-quality decision-making support for subsequent anomaly detection and temperature protection.
[0070] Optionally, as shown in Figure 5 the physics-guided training unit 141 may include a physical constraint modeling subunit 1411, a data-driven prediction subunit 1412, a loss calculation subunit 1413, and a model training subunit 1414. Specifically:
[0071] The physical constraint modeling subunit 1411 is used to add heat conduction physical constraints to the pre-constructed initial neural network using the material thermal conductivity, material specific heat capacity, and chip power consumption density of the chip to obtain a physics-guided neural network;
[0072] The data-driven prediction subunit 1412 is used to obtain the sample multi-dimensional feature vector in the training dataset and calculate the temperature distribution prediction result of the chip at future moments based on the physics-guided neural network and the sample multi-dimensional feature vector;
[0073] The loss calculation subunit 1413 is used to determine the data-driven loss based on the temperature distribution prediction result and the sample temperature prediction result in the training dataset, construct a temperature prediction loss function by combining the physical constraint loss determined by the heat conduction physical constraints, balance the data-driven loss and the physical constraint loss using a weight factor, and optimize the neural network parameters through the gradient descent algorithm;
[0074] The model training subunit 1414 is used to iteratively train the physics-guided neural network based on the optimized temperature prediction loss function to obtain a trained temperature prediction model.
[0075] Among them, the physical guidance training unit 141 can physically model the temperature change by means of physical constraint modeling, data-driven prediction, loss calculation, and model training optimization, combined with the partial differential equation (PDE) of heat conduction, and impose physical constraints during the training process of the deep learning model to achieve high-precision and high-stability temperature prediction.
[0076] During the physical guidance training process, a physical-guided neural network can be constructed based on the heat conduction characteristics of the chip 10, enabling the model to follow thermodynamic constraints while learning the temperature distribution characteristics, so as to avoid prediction results that do not conform to physical laws in extreme cases for a pure data-driven model. During the physical modeling process, the physical constraint modeling subunit 1411 can construct a partial differential equation (PDE) of heat conduction based on the material properties, power consumption distribution, and structural information of the chip 10 to describe the heat diffusion behavior inside the chip 10. For example, the general form of the partial differential equation of heat conduction can be expressed as:
[0077]
[0078] Among them, ρ can represent the density of the chip material, c p can represent the specific heat capacity of the chip material, T can represent the temperature corresponding to time t, k can represent the thermal conductivity of the chip material, can represent the partial derivative of the temperature field with respect to time, can represent the Laplace operator of the temperature. In the actual implementation process, the heat conduction equation can be discretized according to the specific physical structure of the chip, and numerically solved in combination with the finite difference method (FDM) or the finite element method (FEM) to obtain the temperature distribution under steady-state and transient conditions and ensure the calculation efficiency. During the training process of the deep learning model, the physical constraint modeling subunit 1411 can calculate the time gradient and spatial gradient of the neural network predicted temperature based on the above equation, and use the automatic differentiation method to calculate the partial derivative of the heat conduction equation to construct a loss function that conforms to physical laws, thereby guiding the parameter optimization direction of the neural network model.
[0079] After establishing the physical constraints, the data-driven prediction subunit 1412 can utilize the sample multi-dimensional feature vectors in the training dataset, combine with the physics-informed neural network for temperature prediction, and output the temperature distribution of the chip 10 at future moments. Since the change of chip temperature is affected by multiple factors such as power consumption, heat dissipation conditions, and workload, it is difficult to accurately predict the complex temperature evolution process of the chip relying solely on physical models. Therefore, it is necessary to combine deep learning methods for data-driven modeling. The data-driven prediction subunit 1412 can receive the sample multi-dimensional feature vectors generated by the feature extraction module 130 and input them into the physics-informed neural network for forward propagation calculation to predict the temperature distribution of the chip 10 at future moments.
[0080] The input layer of the physics-informed neural network can include temperature change trend features, voltage and frequency fluctuation features, and causal association features, and use a Deep Neural Network (DNN) or a Recurrent Neural Network (RNN) to model the temporal variation characteristics of chip temperature. In different implementation manners, a Long Short-Term Memory (LSTM) or a Variational Autoencoder (VAE) can also be used to model time series data, and a Convolutional Neural Network (CNN) is used to extract spatial features to improve the accuracy of temperature prediction. For high-performance computing scenarios, a spatial relationship modeling method based on a Graph Neural Network (GNN) can be combined to perform feature learning on the heat conduction modes in different regions of the chip and optimize the spatial resolution of the prediction results.
[0081] In the process of constructing the physics-informed neural network, in order to ensure that the data-driven loss and the physical constraint loss can be optimized simultaneously during the training process, the loss calculation subunit 1413 can evaluate the error of the prediction result of the physics-informed neural network and construct a temperature prediction loss function. Among them, the temperature prediction loss function can be jointly composed of the data-driven loss and the physical constraint loss. The data-driven loss can be used to minimize the error between the model prediction value and the historical temperature data and can be represented by the following relational expression:
[0082]
[0083] Among them, L data can represent the loss function corresponding to the data-driven loss, T pred (x i ,y i ,t i) can represent the temperature value at the spatial coordinates (x i at the chip at time t during the i-th iteration training process predicted by the physics-informed neural network i , y i ). T true (x i , y i , t i ) can represent the actual temperature value (i.e., the sample label) of the chip corresponding to the spatial coordinates and the time.
[0084] The physical constraint loss can be used to ensure that the prediction result conforms to the partial differential equation of heat conduction, and the temperature gradient of the neural network output is calculated by automatic differentiation to satisfy:
[0085]
[0086] where L pde can represent the loss function corresponding to the physical constraint loss, can represent the partial derivative of the predicted temperature field with respect to time, can represent the second-order spatial partial derivative of the predicted temperature field, i.e., the Laplace operator.
[0087] Furthermore, the total loss function of the physics-informed neural network can be defined as:
[0088] L total = L data + λL pde ;
[0089] where L total can represent the total loss function of the physics-informed neural network, L data can represent the loss function corresponding to the data-driven loss, L pde can represent the loss function corresponding to the physical constraint loss, and λ can represent an adjustable weight parameter used to achieve an optimal balance between the data-driven loss and the physical constraint loss. The loss calculation subunit 1413 continuously adjusts the value of λ during the training process to ensure that the model follows physical laws while learning the historical data pattern, thereby improving the stability and reliability of the prediction result.
[0090] After constructing the temperature prediction loss function, the model training subunit 1414 can perform iterative training on the physics-informed neural network based on the optimized temperature prediction loss function and optimize the model parameters using the gradient descent algorithm. In the specific training process, an adaptive gradient optimization algorithm can be adopted. For example, the Adam algorithm or the RMSprop algorithm can be used for weight update, and the learning rate decay strategy can be combined to improve the training convergence speed. This embodiment does not make special limitations on the adopted adaptive gradient optimization algorithm.
[0091] The model training subunit 1414 can also combine transfer learning methods, enabling the trained neural network to adaptively adjust model parameters under different chip architectures or heat dissipation environments and improve generalization ability. In addition, to improve training efficiency, the model training process can also use a Graphics Processing Unit (GPU) or a Tensor Processing Unit (TPU) for computing acceleration to reduce the computational overhead under large-scale training data.
[0092] The training process is implemented through the physical constraint modeling subunit 1411, the data-driven prediction subunit 1412, the loss calculation subunit 1413, and the model training subunit 1414, enabling the physical-guided training unit 141 to not only perform data-driven modeling by combining deep learning methods but also introduce physical constraints to optimize the prediction ability of the model and improve the physical consistency and stability of chip temperature prediction.
[0093] In an exemplary embodiment of the present disclosure, referring to Figure 6 as shown, the anomaly detection module 150 may include an error calculation unit 151 and an anomaly scoring unit 152. Specifically:
[0094] The error calculation unit 151 is configured to calculate the temperature deviation between the future temperature distribution data and the current temperature distribution data. If it is determined that the temperature deviation is greater than or equal to a preset anomaly temperature threshold, it is determined that the operating state of the chip is an abnormal operating state.
[0095] The anomaly scoring unit 152 is configured to perform weighted fusion on the temperature deviation, the temperature change gradient determined by the future temperature distribution data and the current temperature distribution data, the historical operating state consistency, and the statistical distribution deviation scoring to determine the chip temperature anomaly score.
[0096] Among them, the anomaly detection module 150 can utilize the future temperature data provided by the temperature prediction module 140 and combine the current chip temperature measurement value to perform quantitative analysis on the temperature change trend to form a complete temperature anomaly evaluation system and provide a basis for the protection strategy of the anomaly analysis module 160.
[0097] During the process of temperature anomaly assessment, the error calculation unit 151 can be used to calculate the temperature deviation between the future temperature distribution data and the current temperature distribution data, and determine whether the chip 10 is in an abnormal state based on this deviation value. Since the temperature of the chip 10 is usually affected by dynamic workload, ambient temperature, and power consumption changes, relying solely on a fixed temperature threshold for anomaly judgment may lead to false alarms or missed alarms. The error calculation unit 151 can calculate the temperature deviation between the predicted future temperature distribution data and the measured previous temperature distribution data, and compare the temperature deviation with a preset abnormal temperature threshold. When the temperature deviation is greater than or equal to the preset abnormal temperature threshold, it can be considered that the chip is currently in an abnormal working state; of course, when the temperature deviation is less than the preset abnormal temperature threshold, it can be considered that the chip is currently in a normal working state.
[0098] In actual implementation, the preset abnormal temperature threshold can be dynamically adjusted based on statistical analysis methods to adapt to the temperature change patterns in different working scenarios. The error calculation unit 151 can use the Weighted Moving Average (WMA) method to smooth the temperature data to reduce the interference of temperature fluctuations in a short period on anomaly detection. In specific application scenarios, an adaptive temperature threshold adjustment strategy can be combined to enable anomaly judgment to dynamically adapt to the working environment of the chip.
[0099] After calculating the temperature deviation, the dynamic trend of the temperature change can be further analyzed to evaluate the temperature change rate of the chip 10 at different time points. The anomaly scoring unit 152 can identify the rate of temperature increase or decrease by calculating the temperature change gradient and determine whether there is an abnormal temperature fluctuation. For example, the temperature change gradient can be calculated by the following relational expression:
[0100]
[0101] where S gradient can represent the temperature change gradient, T t+1 and T t can represent the temperature values at adjacent time points, and Δt can represent the time interval. When S gradient is greater than or equal to the preset threshold, it indicates that a sudden anomaly may occur in the chip temperature. In terms of implementation, the temperature change rate can be calculated using a sliding window method to reduce the error of individual data points and improve the robustness of anomaly detection. For high-precision scenarios, the frequency characteristics of the temperature signal can be analyzed by combining the Fourier Transform or Wavelet Transform to identify periodic temperature fluctuations and optimize the sensitivity of anomaly judgment.
[0102] Based on the analysis of the temperature change rate, the matching degree between the historical operating state of the chip and the current temperature data can be evaluated to detect whether there is an abnormal operating state. The anomaly scoring unit 152 can use a causal inference model to analyze the correlation between the voltage, frequency, and power consumption parameters of the chip and the temperature, and calculate the consistency score of the chip operating state based on historical data. The historical operating state consistency can be calculated by the following formula:
[0103]
[0104] where S history can represent the historical operating state consistency score, T expected can represent the theoretical temperature calculated based on historical data, and T current can represent the currently measured temperature. When the historical operating state consistency score exceeds the preset threshold, it indicates that there may be an abnormality in the current operating state. In the specific implementation process, the theoretical temperature can be calculated by a regression model or a deep learning prediction model and adjusted in combination with long-term data. For complex chip architectures, a method based on a Graph Neural Network (GNN) can be used to establish the heat conduction relationship between different chip regions and optimize the accuracy of the operating state consistency analysis.
[0105] After completing the above analysis, statistical methods can also be used to detect whether the temperature data deviates from the normal distribution to identify potential temperature anomaly patterns. The anomaly scoring unit 152 can evaluate the degree of deviation from the historical data distribution by calculating the Z-score normalization value of the current temperature data, that is, the statistical distribution deviation score. For example, the statistical distribution deviation score can be determined by the following relationship:
[0106]
[0107] where S stat can represent the statistical distribution deviation score, μ can represent the mean of the historical temperature data, and σ can represent the standard deviation of the historical temperature data. When the statistical distribution deviation score exceeds the threshold, it indicates that the current temperature significantly deviates from the normal distribution and there may be an abnormality. In practical applications, the calculation of the statistical distribution deviation can be extended based on the Bayesian Method or the Gaussian Mixture Model (GMM) to improve the accuracy of anomaly detection. In addition, to avoid the influence of environmental temperature changes on anomaly detection, external sensor data can be combined for temperature compensation to reduce the interference of external factors on statistical analysis.
[0108] After comprehensively calculating various scores, the anomaly scoring unit 152 can perform weighted fusion on the temperature deviation, temperature change gradient, historical operation state consistency, and statistical distribution deviation scores to determine the final chip temperature anomaly score. In practical applications, the anomaly scoring unit 152 can combine an adaptive weight adjustment strategy to optimize the sensitivity of anomaly score calculation, making it highly adaptable in different operating environments.
[0109] Through the above multi-dimensional analysis, the anomaly detection module 150 can not only accurately determine whether the chip is in an abnormal state, but also identify the main influencing factors of temperature change and quantitatively evaluate the temperature anomaly situation, thereby providing high-precision anomaly detection capabilities for chip temperature management and providing data support for subsequent temperature protection strategies.
[0110] In an exemplary embodiment of the present disclosure, referring to Figure 7 as shown, the anomaly analysis module 160 may include a protection level determination unit 161, a protection strategy generation unit 162, and an anomaly log storage unit 163. Specifically:
[0111] The protection level determination unit 161 is configured to determine the chip protection level of the chip based on the anomaly score range in which the chip temperature anomaly score is located, and there is a mapping relationship between the chip protection level and the anomaly score range;
[0112] The protection strategy generation unit 162 is configured to call the preset chip temperature protection strategy corresponding to the chip protection level according to the chip protection level, and determine the chip temperature protection operation based on the chip temperature protection strategy;
[0113] The anomaly log storage unit 163 is configured to record the anomaly detection results, chip protection levels, and chip temperature protection operations of the chip.
[0114] Among them, the anomaly analysis module 160 can comprehensively analyze the temperature anomaly score through the protection level determination unit 161, the protection strategy generation unit 162, and the anomaly log storage unit 163, and match appropriate temperature management strategies based on different anomaly levels to optimize the operation safety and stability of the chip.
[0115] During the determination process of the chip protection level, the protection level determination unit 161 can determine the abnormal score range in which the chip is currently located based on the chip temperature abnormal score calculated by the abnormal detection module 150, and divide different protection levels accordingly. The chip temperature abnormal score is a comprehensive evaluation value of temperature abnormality, reflecting the suddenness, persistence of temperature change, and the degree of deviation from the normal state. The protection level determination unit 161 can match the temperature abnormal score with a preset abnormal score range to determine the protection level in which the chip is currently located. For example, when the chip temperature abnormal score is low, it indicates that the temperature change is within the normal range and the chip does not require additional protection; when the score reaches the set threshold, the corresponding temperature protection mechanism needs to be triggered to prevent the temperature from rising further and affecting the stable operation of the chip. The score thresholds of different chip systems can be adjusted according to the chip manufacturing process, maximum allowable operating temperature, and heat dissipation characteristics to optimize the accuracy of protection level division. In the specific implementation process, the protection level determination unit can adopt a method based on statistical analysis, optimize the score range by analyzing historical abnormal data, and perform adaptive threshold adjustment in combination with machine learning algorithms to improve the intelligence level of protection level determination.
[0116] After completing the determination of the protection level, the protection strategy generation unit 162 can select a preset chip temperature protection strategy that matches the determined chip protection level, and determine specific chip temperature protection operations accordingly. The chip temperature protection operation can be understood as a chip temperature regulation instruction generated in combination with the preset chip temperature protection strategy, used to perform corresponding operations on the chip system. Since different temperature abnormal levels may have different degrees of impact on the operation stability of the chip, the protection strategy generation unit 162 may need to call the corresponding protection strategy according to the severity of the protection level and dynamically adjust the parameters of the protection operation to ensure the effectiveness of the temperature management measures. In the case of temperature abnormalities at different levels, the protection strategy may include measures such as reducing chip power consumption, adjusting fan speed, dynamically adjusting the working frequency, shutting down some non-critical computing units, or triggering hardware fusing. For example, the mapping relationship between the chip protection level, abnormal score range, and chip temperature protection strategy can be represented by Table 1:
[0117] Table 1 Mapping relationship between chip protection level, abnormal score range, and chip temperature protection strategy
[0118]
[0119] When the chip is at a low risk level, temperature control can be achieved by adjusting the fan speed or optimizing the power management strategy; when the protection level is increased, stronger protection measures are required, such as reducing the chip's operating voltage and adjusting the Dynamic Voltage and Frequency Scaling (DVFS) strategy; when the temperature anomaly score reaches the high risk level, some computing units can be turned off to reduce the chip's power consumption and heat accumulation; in extreme cases, such as when the temperature anomaly score exceeds the safety threshold, it may be necessary to trigger the hardware fuse mechanism and suspend the chip's computing tasks to prevent the chip from being damaged due to overheating. In practical applications, the protection strategy generation unit 162 can optimize and adjust the protection strategy in combination with the chip's current task load status to minimize the impact on system performance while ensuring the chip's safety. For example, in a server chip or a high-performance computing environment, the high-load computing tasks can be migrated to a chip with a lower temperature in combination with the task scheduling strategy to optimize the system's heat dissipation balance ability.
[0120] After the temperature protection strategy is executed, in order to ensure that the system can continuously monitor the execution effect of the protection strategy and make adjustments when necessary, the anomaly log storage unit 163 can record the chip's anomaly detection results, protection level, and corresponding temperature protection operations, and store them in the system log database. The main function of the anomaly log storage unit 163 is to archive historical data of temperature anomaly events for subsequent anomaly analysis, fault troubleshooting, and temperature protection strategy optimization. The recorded log information can include data such as the chip's current temperature data, temperature anomaly score, protection level, executed protection strategy, execution time, and temperature change after execution. The anomaly log can be stored locally or in the cloud to ensure data traceability. In practical applications, the anomaly log storage unit 163 can perform statistical analysis on the historical log data in combination with the anomaly analysis model and optimize the temperature protection strategy based on data mining techniques. For example, a time series model of the chip's operating state can be constructed based on the log data to analyze the execution effects of different protection strategies, and the parameters of the protection strategy can be adjusted accordingly to improve the intelligence level of the temperature management system. In addition, in a multi-chip system or a distributed computing environment, the anomaly log storage unit 163 can also interact with the remote monitoring system to share temperature anomaly information among multiple chips to improve the overall system's temperature management ability.
[0121] Through the combination of the protection level determination unit 161, the protection strategy generation unit 162, and the exception log storage unit 163, the exception analysis module 160 can not only accurately determine the protection level of the chip, but also select the optimal protection strategy according to different temperature exception situations, and record and analyze the execution of the protection strategy, thereby optimizing the temperature management ability of the chip, improving the operation stability of the system, and effectively preventing the chip from malfunctioning or being damaged due to temperature anomalies in extreme cases.
[0122] In an exemplary embodiment of the present disclosure, as shown in Figure 8 the system may further include a real-time monitoring module 180. As shown in Figure 9 the temperature protection module 170 may include a protection operation execution unit 171 and an operating state feedback unit 172. Specifically:
[0123] The protection operation execution unit 171 is configured to adjust the operating parameters of the chip based on the chip temperature protection operation control. Wherein, the chip temperature protection operation includes: in response to a first-level protection operation, reducing the power consumption of the chip and adjusting the fan speed; or in response to a second-level protection operation, adjusting the voltage and operating frequency of the chip; or in response to a third-level protection operation, turning off some non-critical computing units; or in response to a fourth-level protection operation, triggering a hardware fusing mechanism to pause the current computing task.
[0124] The operating state feedback unit 172 is configured to collect the latest operating state data of the chip after executing the temperature protection operation and send the latest operating state data to the real-time monitoring module.
[0125] Among them, the operating parameters of the chip can be dynamically adjusted through the protection operation execution unit 171 to respond to the temperature protection operations corresponding to different temperature exception levels. At the same time, after executing the temperature protection measures, the latest operating state data of the chip can be continuously collected through the operating state feedback unit 172 and data interaction with the real-time monitoring module 180 to optimize the execution effect of the temperature protection strategy. The temperature protection module 170 ensures that the chip can be effectively protected after a temperature anomaly occurs and reasonably adjusts the operating state after the temperature returns to normal by executing the chip temperature protection strategy, dynamically adjusting the operating parameters, and real-time monitoring the chip state, so as to improve the temperature management efficiency of the system and the operating stability of the chip.
[0126] When a chip temperature anomaly occurs, the protection operation execution unit 171 can optimize and adjust the operating parameters of the chip based on the chip temperature protection operation determined by the anomaly analysis module 160, so as to reduce the chip temperature level and reduce the impact of the temperature anomaly on the chip performance. During the execution of the protection operation, different levels of temperature anomalies correspond to different protection measures, including operations such as adjusting the fan speed, reducing the power consumption, adjusting the operating frequency, shutting down some computing units, and triggering hardware fusing.
[0127] After the chip executes the temperature protection measures, in order to ensure the execution effect of the protection strategy and optimize and adjust according to the actual operating state, the operating state feedback unit 172 can collect the latest operating state data of the chip and send the data to the real-time monitoring module to form a closed-loop control system. The operating state feedback unit 172 can evaluate the actual effect of the protection operation by obtaining the latest temperature data, power consumption data, and load state of the chip, and adjust the temperature control strategy based on the data feedback. In the specific implementation process, the operating state feedback unit 172 can adopt an adaptive control algorithm based on machine learning, and dynamically adjust the protection parameters by analyzing the execution effects of different protection strategies to optimize the temperature management efficiency of the chip. In addition, in a multi-chip system, the temperature data of multiple chips can be collected in real time based on a distributed monitoring architecture, and the temperature management scheme between different chips can be optimized based on a global temperature control strategy to improve the overall thermal management ability of the system.
[0128] Through the combination of the protection operation execution unit 171 and the operating state feedback unit 172, the temperature protection module 170 can not only take corresponding protection measures for different levels of temperature anomalies, but also dynamically adjust the protection strategy based on the operating state feedback of the chip, so as to improve the accuracy of temperature control and optimize the long-term stability of the chip. In addition, in a multi-chip system or a high-performance computing platform, the temperature protection module 170 can combine task scheduling and thermal balance strategies to achieve temperature management optimization at the whole system level, so as to improve the reliability of the computing system and reduce the safety risks brought by chip overheating.
[0129] In an exemplary embodiment of the present disclosure, referring to Figure 10 as shown, the real-time monitoring module 180 can include a temperature change calculation unit 181, a dynamic adjustment unit 182, and a monitoring log storage unit 183. Specifically:
[0130] The temperature change calculation unit 181 is configured to determine the chip temperature change trend based on the temperature data before and after the execution of the temperature protection operation.
[0131] A dynamic adjustment unit 182, configured to, based on the chip temperature change trend, after the temperature protection operation is executed, perform real-time adjustment on the chip temperature protection policy corresponding to the chip temperature protection operation according to a preset adjustment step size and based on the latest operating state data;
[0132] A monitoring log storage unit 183, configured to store the temperature protection execution result and update the chip temperature protection policy of the chip.
[0133] Among them, in the chip temperature protection system, in order to ensure that the execution effect of the temperature management policy meets the real-time operation requirements and can adapt to the heat dissipation changes of the chip under different workloads, it is necessary to combine the real-time monitoring module 180 to continuously track the temperature change trend of the chip 10, and optimize the temperature control parameters based on the execution effect of the chip temperature protection policy to achieve high-precision and dynamic temperature management. The real-time monitoring module 180 can perform real-time analysis on the chip temperature change through the collaborative work of the temperature change calculation unit 181, the dynamic adjustment unit 182, and the monitoring log storage unit 183, and optimize the temperature protection policy according to the monitoring results to ensure the operation stability of the chip.
[0134] During the temperature management process, the temperature change calculation unit 181 can calculate the chip temperature change trend based on the temperature data before and after the temperature protection execution, and analyze the actual effect of the temperature control measures. During the temperature change analysis process, the temperature change calculation unit 181 can first receive the chip temperature data before the protection operation is executed by the temperature protection module, and combine the latest temperature data obtained by the temperature sensor to calculate the temperature change amplitude of the chip at different time points. In the specific implementation process, the temperature change calculation unit 181 can combine the sliding window method to smooth the temperature change rate to reduce the interference of short-term temperature fluctuations on the analysis result. In addition, the exponentially weighted moving average (EWMA) method can be combined to assign different weights to the historical temperature data to enhance the analysis ability of the long-term temperature change trend.
[0135] After calculating the temperature change trend of the computing chip, the dynamic adjustment unit 182 can, based on the result of the temperature change trend analysis, after the temperature protection operation is executed, optimize the temperature protection strategy in real time according to a preset adjustment step, and dynamically adjust the temperature protection parameters according to the latest operating state data. Since the heat dissipation efficiency of the chip may be affected by environmental factors, workload, and the state of heat dissipation devices, a fixed temperature protection strategy may not be able to maintain the optimal control effect in all operating environments. The dynamic adjustment unit 182 can perform iterative optimization of the chip temperature protection operation based on a feedback control mechanism, enabling the protection strategy to be adaptively adjusted according to the actual temperature change. In the specific implementation process, the temperature control parameters can be adjusted in real time based on the Proportional-Integral-Derivative (PID) control algorithm, and the protection measures of the chip can be optimized based on the temperature recovery curve. Of course, the Reinforcement Learning (RL) method can also be combined to perform adaptive learning on the protection strategies for different temperature anomalies and optimize the adjustment accuracy of the temperature control parameters to improve the intelligent level of chip temperature management. This embodiment is not limited thereto.
[0136] During the dynamic adjustment of the temperature protection strategy, in order to ensure that the chip can maintain the optimality of the temperature management strategy under different operating states, the monitoring log storage unit 183 can be used to store the temperature protection execution results and update the chip's temperature protection strategy to optimize the system's long-term temperature management ability. The monitoring log storage unit 183 can record the chip's temperature data, temperature change trend, protection strategy execution results, and adjustment parameters, and store them in the system database for subsequent analysis and optimization. In the specific implementation process, the monitoring log storage unit 183 can adopt a structured log storage method to ensure efficient retrieval and access of data, and combine a data compression algorithm to optimize the storage space occupancy. In addition, in order to improve the long-term optimization ability of the temperature management strategy, the monitoring log storage unit 183 can combine a historical temperature data analysis model to statistically analyze the temperature management effects in different operating environments, and optimize the chip's temperature control algorithm based on data mining techniques. For example, the Clustering Analysis method can be used to classify the chip temperature change patterns and optimize the temperature protection parameters under different chip architectures to enhance the system's adaptive adjustment ability.
[0137] Through the combination of the temperature change calculation unit 181, the dynamic adjustment unit 182, and the monitoring log storage unit 183, the real-time monitoring module 180 can not only dynamically track the execution effect of the temperature protection measures, but also optimize the temperature protection strategy based on the temperature change trend of the chip. By accumulating and analyzing long-term monitoring data, the accuracy and self-adaptability of chip temperature management can be improved. In addition, in server-level chips, multi-chip systems, and high-performance computing platforms, the real-time monitoring module 180 can combine the system-level thermal management strategy to achieve collaborative temperature management of multiple chips, improve the system-level thermal balance ability, and reduce the impact of local overheating on chip life and computing performance.
[0138] In an exemplary embodiment of the present disclosure, referring to Figure 11 as shown, the data preprocessing module 120 may include a data cleaning unit 121, a time alignment unit 122, and a normalization unit 123. Specifically:
[0139] The data cleaning unit 121 is configured to filter redundant information in the current temperature distribution data and the operating state data, and perform noise removal processing on the current temperature distribution data by using a filtering algorithm;
[0140] The time alignment unit 122 is configured to perform time alignment on the current temperature distribution data and the operating state data after cleaning based on a timing analysis method;
[0141] The normalization unit 123 is configured to convert the current temperature distribution data and the operating state data after time alignment into a unified standardized data format to form a chip monitoring data set.
[0142] Among them, during the chip temperature monitoring process, in order to ensure that the collected temperature data can accurately reflect the thermal distribution state of the chip and provide high-quality data input for subsequent temperature prediction, anomaly detection, and temperature management, it is necessary to process the collected temperature data and operating state data through the data preprocessing module 120 to eliminate data noise, correct timing deviation, and standardize the data format, so that data from different sources can be analyzed and processed under a unified data framework. The data preprocessing module 120 may include a data cleaning unit 121, a time alignment unit 122, and a normalization unit 123. Each unit ensures the consistency, accuracy, and availability of the data through multi-level preprocessing of the data.
[0143] In the process of collecting temperature data and operating status data, due to the influence of sensor measurement error, signal interference and environmental noise, the original data may contain abnormal values, missing values and repeated data. These data problems may affect the accuracy of temperature prediction. The collected data can be cleaned by the data cleaning unit 121 to remove invalid data and optimize data quality. The data cleaning unit 121 can first perform an integrity check on the temperature data and operating status data of the chip, detect whether there are missing values, and use interpolation methods to fill the missing data. In the specific implementation process, linear interpolation (Linear Interpolation), spline interpolation (Spline Interpolation) or Gaussian Process Regression (Gaussian Process Regression, GPR) methods can be used to estimate missing values based on temperature data at adjacent moments to reduce the impact of data loss on subsequent analysis. In addition, in order to remove abnormal values in the data, the data cleaning unit 121 can use an abnormality detection method based on statistical analysis, such as based on Z-score standardization to detect whether the temperature data exceeds the normal distribution range, or use a box plot (Box Plot) method to detect outliers in the temperature data, and remove or correct the abnormal data. In a chip environment with high dynamic load, deep learning-based anomaly detection algorithms, such as autoencoder (AE) or isolation forest (Isolation Forest), can also be combined to identify data anomaly patterns and optimize data cleaning strategies. This example embodiment does not specifically limit this.
[0144] After data cleaning, since the data acquisition frequency of the temperature sensor may not match the sampling period of the operating state data, resulting in a timing deviation in the data, the time alignment unit 122 can perform time alignment processing on the cleaned data to make the temperature data and the operating state data consistent in the time dimension. The time alignment unit 122 can adjust the timestamps of different data sources based on the clock synchronization mechanism of the chip and map the data to a unified time axis using the time series interpolation method. In the specific implementation process, the time window method can be used to block-process data of different time scales, and weighted average or time series interpolation methods can be used to align the data to ensure that data from different sources can be processed synchronously. Of course, in a multi-chip system or a distributed computing environment, the time alignment unit 122 can also combine the Network Time Protocol (NTP) or the Precision Time Protocol (PTP) to ensure the timestamp accuracy of multiple data sources and optimize the temperature data alignment ability across chips. In some high-precision scenarios, the Dynamic Time Warping (DTW) method can be used to match the timing patterns of the temperature data and the operating state data and optimize the time alignment effect.
[0145] After data cleaning and time alignment are completed, to ensure the consistency of data on the numerical scale and improve the computational efficiency of the data, the normalization unit 123 can perform normalization processing on the temperature data and the operating state data after time alignment, so that data of different physical quantities can be operated within the same numerical range and the training efficiency of the deep learning model can be improved. The normalization unit 123 can select an appropriate normalization method to transform the data based on the distribution of different data features. In the specific implementation process, the Min-Max Normalization method can be used to map the data to the [0, 1] interval to maintain the relative proportional relationship of the data, or the Z-score standardization method can be used to make the data conform to the normal distribution with zero mean and unit variance, thereby reducing the impact of data skewness on model training. In addition, in a dataset containing multiple feature dimensions, the normalization unit 123 can use the Principal Component Analysis (PCA) or Feature Scaling method to perform dimensionality reduction processing on the data to reduce the computational complexity and improve the separability of the data.
[0146] After the normalization process is completed, the normalization unit 123 can store the normalized data in the chip monitoring data set and provide it to the subsequent feature extraction module 130 for further analysis. Through the above data preprocessing process, the chip temperature data can not only effectively remove noise and outliers, but also maintain synchronization in the time dimension and consistency in the numerical scale, providing high-quality data input for subsequent temperature prediction, anomaly detection, and temperature management. In addition, in a multi-chip system or a high-performance computing environment, the data preprocessing module 120 can combine cross-platform data alignment and distributed data storage optimization strategies to achieve efficient management of large-scale temperature data and improve the real-time performance and stability of data processing.
[0147] It should be noted that although several modules or units of the artificial intelligence-based hardware chip temperature monitoring and automatic protection system are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0148] Those skilled in the art will readily think of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0149] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. An artificial intelligence-based hardware chip temperature monitoring and automatic protection system, characterized in that, The system includes: A data acquisition module, configured to obtain the current temperature distribution data of multiple regions of the chip and synchronously obtain the operating status data of the chip; A data preprocessing module, configured to preprocess the current temperature distribution data and the operating status data to obtain a chip monitoring data set; A feature extraction module, configured to extract features from the chip monitoring data set to obtain a monitoring data feature vector of at least one feature dimension, and construct a multi-dimensional feature vector according to the monitoring data feature vector; A temperature prediction module, configured to input the multi-dimensional feature vector into a pre-trained temperature prediction model to infer the future temperature distribution data of the chip at the next moment; An anomaly detection module, configured to determine the working status of the chip according to the current temperature distribution data and the future temperature distribution data, and calculate the chip temperature anomaly score when the chip is in an abnormal working state; An anomaly analysis module, configured to determine the chip protection level through the chip temperature anomaly score, and determine the chip temperature protection operation based on the chip protection level; A temperature protection module, configured to optimize and adjust the operating parameters of the chip according to the chip temperature protection operation to achieve temperature protection of the chip.
2. The temperature monitoring and automatic protection system for hardware chips based on artificial intelligence according to claim 1, characterized in that, The data acquisition module includes: A temperature sensor array, including at least two temperature sensors, configured to collect temperature distribution data by being arranged in multiple pre-determined thermal field regions corresponding to the chip; A dynamically adjustable sensor control unit, electrically connected to the temperature sensor array, configured to adaptively adjust the sampling frequency of each temperature sensor according to the real-time thermal field distribution of the chip.
3. The temperature monitoring and automatic protection system for hardware chips based on artificial intelligence according to claim 1, characterized in that, The feature extraction module includes: A frequency domain analysis unit, configured to perform short-time Fourier transform on the data in the chip monitoring data set to extract temperature change trend features and voltage and frequency fluctuation features; A causal inference unit, configured to analyze the correlation between the change of the chip operating status and the temperature change by using a causal inference method to generate causal association features; A feature encoding unit, electrically connected to the frequency domain analysis unit and the causal inference unit, configured to fuse the causal association features, the temperature change trend features, and the voltage and frequency fluctuation features to construct a multi-dimensional feature vector, and send the multi-dimensional feature vector to the temperature prediction module.
4. The temperature monitoring and automatic protection system for a hardware chip based on artificial intelligence according to claim 1, wherein The temperature prediction module includes: A physics-guided training unit, configured to train a pre-constructed temperature prediction model according to the training data set corresponding to the chip, where the temperature prediction model includes a physics-guided neural network constructed based on the heat conduction principle, to obtain a trained temperature prediction model; A deep learning prediction unit, configured to infer the future temperature distribution data of the chip at the next moment according to the trained temperature prediction model and in combination with the multi-dimensional feature vector.
5. The temperature monitoring and automatic protection system for hardware chips based on artificial intelligence according to claim 4, characterized in that The physics-guided training unit includes: A physical constraint modeling sub-unit, configured to add heat conduction physical constraints to a pre-constructed initial neural network by using the material thermal conductivity, material specific heat capacity, and chip power consumption density of the chip to obtain a physics-guided neural network; A data-driven prediction sub-unit, configured to obtain the sample multi-dimensional feature vectors in the training data set, and calculate the predicted result of the temperature distribution of the chip at a future moment based on the physics-informed neural network and the sample multi-dimensional feature vectors; A loss calculation sub-unit, configured to determine a data-driven loss based on the predicted result of the temperature distribution and the predicted result of the sample temperature in the training data set, and construct a temperature prediction loss function by combining the physical constraint loss determined by the heat conduction physical constraint, balance the data-driven loss and the physical constraint loss by using a weight factor, and optimize the neural network parameters by using a gradient descent algorithm; A model training sub-unit, configured to iteratively train the physics-informed neural network based on the optimized temperature prediction loss function to obtain a trained temperature prediction model.
6. The temperature monitoring and automatic protection system for hardware chips based on artificial intelligence according to claim 1, characterized in that, The anomaly detection module includes: An error calculation unit, configured to calculate the temperature deviation between the future temperature distribution data and the current temperature distribution data, and if it is determined that the temperature deviation is greater than or equal to a preset anomaly temperature threshold, determine that the working state of the chip is an abnormal working state; An anomaly scoring unit, configured to perform weighted fusion on the temperature deviation, the temperature change gradient determined by the future temperature distribution data and the current temperature distribution data, the historical operation state consistency, and the statistical distribution deviation scoring to determine the chip temperature anomaly score.
7. The temperature monitoring and automatic protection system for hardware chips based on artificial intelligence according to claim 1, characterized in that, The anomaly analysis module includes: A protection level determination unit, configured to determine the chip protection level of the chip based on the anomaly score interval where the chip temperature anomaly score is located, and there is a mapping relationship between the chip protection level and the anomaly score interval; A protection strategy generation unit, configured to call the preset chip temperature protection strategy corresponding to the chip protection level according to the chip protection level, and determine the chip temperature protection operation based on the chip temperature protection strategy; An anomaly log storage unit, configured to record the anomaly detection result, the chip protection level, and the chip temperature protection operation of the chip.
8. The artificial intelligence-based hardware chip temperature monitoring and automatic protection system according to claim 7, characterized in that The system further includes a real-time monitoring module, and the temperature protection module includes: A protection operation execution unit, configured to control and adjust the operating parameters of the chip based on the chip temperature protection operation; Wherein, the chip temperature protection operation includes: Responding to a first-level protection operation, reducing the power consumption of the chip and adjusting the fan speed; or Responding to a second-level protection operation, adjusting the voltage and operating frequency of the chip; or Responding to a third-level protection operation, turning off some non-critical computing units; or Responding to a fourth-level protection operation, triggering a hardware fusing mechanism to pause the current computing task; An operation state feedback unit, configured to collect the latest operation state data of the chip after executing the temperature protection operation, and send the latest operation state data to the real-time monitoring module.
9. The temperature monitoring and automatic protection system for hardware chips based on artificial intelligence according to claim 8, characterized in that, The real-time monitoring module includes: A temperature change calculation unit, configured to determine the chip temperature change trend based on the temperature data before and after executing the temperature protection operation; A dynamic adjustment unit, configured to, based on the chip temperature change trend, after the temperature protection operation is executed, perform real-time adjustment on the chip temperature protection strategy corresponding to the chip temperature protection operation at a preset adjustment step and based on the latest operation status data; A monitoring log storage unit, configured to store the temperature protection execution result and update the chip temperature protection strategy of the chip.
10. The temperature monitoring and automatic protection system for hardware chips based on artificial intelligence according to claim 1, characterized in that, The data preprocessing module includes: A data cleaning unit, configured to filter redundant information in the current temperature distribution data and the operation status data, and perform noise removal processing on the current temperature distribution data by using a filtering algorithm; A time alignment unit, configured to perform time alignment on the current temperature distribution data and the operation status data after cleaning based on a time series analysis method; A normalization unit, configured to convert the current temperature distribution data and the operation status data after time alignment into a unified standardized data format to form a chip monitoring data set.
Citation Information
Cited By
Intelligent monitoring system for energy consumption of server cluster radiator
CN120891748A
Power module multi-chip parallel real-time junction temperature monitoring method
CN120928092A
A method for monitoring real-time junction temperature of a power module multi-chip parallel connection
CN120928092B
Engineering test method and system for storage chip and medium
CN120998286A
Server mainboard working environment automatic adaptation system and method
CN121008673A