Edge ai hot parameter online identification and stability prediction control method and control system

By using a controllable inference load to generate an excitation sequence on an edge AI device, online identification and stability prediction control of thermal parameters are performed. This solves the stability and reliability problems of thermal management in existing technologies, achieves high-precision thermal parameter identification and prediction, reduces the difficulty of equipment maintenance, and improves the stability and performance of equipment operation.

CN122431187APending Publication Date: 2026-07-21KEYU INTELLIGENT ENVIRONMENTAL TECH SERVICES
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
KEYU INTELLIGENT ENVIRONMENTAL TECH SERVICES
Filing Date
2026-04-30
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing edge AI devices' thermal management technology cannot achieve on-site reproducible active stimulation without additional hardware, rapid and high-precision online identification of thermal parameters, confidence gating of identification results, online detection and local updating of parameter drift, early prediction of thermal risks, closed-loop active control with minimal performance cost, and full-process traceable evidence chain recording, leading to problems with device operational stability and performance fluctuations.

Method used

Using the controllable inference load of edge AI devices as the active excitation source, a reproducible excitation sequence is generated through a multi-condition excitation matrix. Multi-source telemetry data is collected, and time-series alignment and validity screening are performed. Online parameter identification is carried out based on the thermal dynamics model, the real-time confidence of the identification results is calculated, and a gating threshold is set. Parameter drift is detected and local updates are triggered. Thermal failure time is predicted through the thermodynamic model, and prediction results with confidence intervals are output. A hierarchical closed-loop control strategy with minimum performance cost is executed, and a fully traceable evidence chain report is automatically generated.

Benefits of technology

It enables early warning 15-30 seconds before thermal throttling occurs, avoiding passive frequency reduction, reducing troubleshooting time and maintenance costs, improving equipment operation stability and real-time performance, ensuring identification accuracy and reliability, providing a traceable evidence chain record throughout the process, and optimizing the performance and reliability of the thermal management system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122431187A_ABST
    Figure CN122431187A_ABST
Patent Text Reader

Abstract

The application discloses an edge AI hot parameter online identification and stability prediction control method and a control system, which comprises the following steps: taking the controllable inference load of an edge AI device as a main excitation source, generating a reproducible excitation sequence according to a preset multi-working-condition excitation matrix, and collecting multi-source operation telemetry data of the device; performing time sequence alignment and effectiveness screening on the collected telemetry data to generate a standardized telemetry data set; performing online parameter identification on the standardized data set based on a discrete thermal dynamics model, calculating the real-time confidence of the identification result, and setting a gating threshold; performing parameter drift detection based on the identification prediction residual, triggering local parameter update when drift is detected, and blocking the risk prediction output through the confidence gating during the update; predicting the remaining time for the device to reach the thermal failure threshold through the thermal dynamics model, and outputting the prediction result with a confidence interval; when the predicted remaining time is less than a preset warning window and the confidence meets the gating requirement, performing a hierarchical closed-loop control strategy according to the principle of minimum performance cost; and automatically generating a full-process traceable structured evidence chain report.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of edge computing and embedded AI, specifically involving a method and control system for online identification and stability prediction control of edge AI thermal parameters. Background Technology

[0002] With the rapid development of edge computing and embedded artificial intelligence technologies, edge AI platforms, represented by the NVIDIA Jetson series, are widely used in scenarios with extremely high requirements for real-time performance and reliability, such as autonomous driving, smart healthcare, industrial vision, and security monitoring. In these scenarios, edge AI devices typically need to run high-load AI inference tasks continuously for extended periods. The high load operation of core computing units such as GPUs generates a large amount of heat, causing the chip temperature to rise rapidly. When the temperature exceeds a preset threshold, the device triggers a dynamic voltage-frequency scaling (DVFS) mechanism to passively reduce the frequency, i.e., thermal throttling. This directly leads to a sudden increase in AI inference latency, a sharp drop in frame rate, fluctuations in inference accuracy, and may even cause business interruptions and system crashes, severely impacting edge AI services with high reliability requirements.

[0003] Currently, thermal management and stability assurance technologies for edge AI devices are mainly divided into two categories: one is passive thermal control technology based on DVFS, and the other is parameter identification and prediction technology based on thermal models.

[0004] The first type of passive thermal control technology is the mainstream thermal management solution currently used in edge AI devices. Its core principle is to set a fixed threshold for chip temperature. When the temperature exceeds the threshold, power consumption is reduced by lowering the operating frequency and supply voltage of core units such as GPU and CPU, thereby suppressing temperature rise. This type of technology has fundamental defects: First, the response time is severely delayed. It can only passively intervene after the temperature has exceeded the limit, and cannot actively prevent thermal throttling before it occurs. This will inevitably lead to a temporary performance drop, which cannot meet the strict requirements of edge AI business for inference latency stability. Second, it has extremely poor adaptability. The fixed temperature threshold and frequency reduction strategy cannot adapt to the differences in thermal characteristics between different individual devices, nor can it adapt to changes in thermal behavior caused by changes in ambient temperature, peripheral configuration, and operating conditions. It is very easy to have problems such as excessive frequency reduction leading to performance waste, or insufficient frequency reduction leading to ineffective suppression of thermal throttling. Third, the control logic is a black box mode, without complete action triggering and execution records, making it impossible to perform post-event auditing and problem reproduction. When performance abnormalities occur, engineers cannot quickly distinguish whether it is a thermal problem or a problem with the model or hardware itself, making troubleshooting extremely difficult.

[0005] For the second type of thermal parameter identification and prediction technology, most existing solutions adopt offline laboratory testing for thermal model parameter calibration. Specifically, in a laboratory environment, high-precision temperature sensors, power analyzers, and other specialized equipment are used to conduct multi-condition thermal characteristic tests on the equipment. Thermal model parameters are obtained through offline data fitting, and then the calibrated parameters are solidified into the equipment's thermal management system. This type of technology has the following insurmountable drawbacks: First, offline testing is extremely costly, requiring specialized testing equipment and environments, and has a long testing cycle. It cannot be reproduced and recalibrated on-site, and for mass-deployed edge AI devices, the cost and time expenditure of offline calibration for each device is unacceptable. Second, parameter portability is extremely poor. When peripherals are replaced, hardware versions are upgraded, or the installation environment is changed, the thermal characteristics of the equipment will change, and the offline calibrated parameters will no longer be applicable. Engineers need to manually adjust and calibrate on-site, resulting in extremely high maintenance costs and troubleshooting cycles of several hours. Third, it cannot adapt to thermal characteristic drift throughout the entire lifecycle of the equipment. During long-term operation, thermal grease aging and wind... Factors such as dust accumulation and component aging can cause a slow drift in thermal characteristics, gradually invalidating offline calibrated parameters. Existing technologies cannot detect this drift online and update the parameters. Fourth, existing online identification schemes suffer from low signal-to-noise ratio and insufficient identification accuracy. Because they do not employ an active excitation mechanism and rely solely on random load fluctuations during normal equipment operation as identification input, the power consumption excitation amplitude is extremely small, resulting in a high condition number in the identification matrix and large errors in the identification results, making them unsuitable for reliable thermal risk prediction. Fifth, existing technologies do not include confidence assessment and gating mechanisms for identification results. When the identification results are unstable, they will still output incorrect prediction results, which can easily lead to miscontrol and reduce the stability of equipment operation.

[0006] In practical engineering scenarios of edge AI deployment, edge AI engineers face a fundamental industry dilemma when deploying and maintaining customized edge AI hardware: the thermal behavior of each edge AI development board varies individually, and currently no tools or methods can complete online identification of device thermal parameters and prediction of thermal risks within 60 seconds. Engineers can only rely on traditional trial-and-error methods to troubleshoot thermal problems, with troubleshooting cycles lasting several hours. Furthermore, calibrated parameters cannot be migrated across hardware versions or deployment environments. For example, when YOLO series object detection models run continuously on the Jetson platform, prolonged high loads can easily cause GPU temperature to rise, triggering DVFS thermal throttling, leading to a sudden increase in inference latency and a sharp drop in frame rate. Existing technologies cannot predict this risk in advance and intervene proactively; they can only handle it post-failure, severely impacting the continuous and stable operation of edge AI services.

[0007] Currently, no relevant technical solution can simultaneously solve all the defects of the above-mentioned existing technologies, that is, it is impossible to simultaneously achieve: field-reproducible active stimulation without additional hardware, fast and high-precision online identification of thermal parameters, confidence gating of identification results, online detection and local update of parameter drift, early prediction of thermal risks, closed-loop active control with minimal performance cost, and full-process traceable evidence chain recording. Therefore, it is urgent to develop an edge AI thermal parameter online identification and stability prediction control technology that can comprehensively solve the above problems. Summary of the Invention

[0008] The purpose of this invention is to provide a method and system for online identification and stability prediction control of edge AI thermal parameters, so as to solve the problems mentioned in the background art.

[0009] To solve the above-mentioned technical problems, the present invention provides the following technical solution: A method for online identification and stability prediction control of edge AI thermal parameters includes: S1. Using the controllable inference load of the edge AI device as the active excitation source, generate a reproducible excitation sequence according to the preset multi-condition excitation matrix, and collect multi-source telemetry data of the device operation. S2. Perform time-series alignment and validity screening on the collected telemetry data to generate a standardized telemetry dataset; S3. Based on the heat dissipation dynamics model, perform online parameter identification on the standardized dataset, calculate the real-time confidence of the identification results, and set the gating threshold; S4. Parameter drift detection is performed based on the identification prediction residual. When drift is detected, local parameter update is triggered. During the update, risk prediction output is blocked by confidence gating. S5. Based on the identified thermal parameters, predict the remaining time for the equipment to reach the thermal failure threshold through a thermodynamic model, and output the prediction results with confidence intervals; S6. When the predicted remaining time is less than the preset warning window and the confidence level meets the gating requirements, the hierarchical closed-loop control strategy is executed according to the principle of minimum performance cost. S7. Automatically generate a fully traceable, structured chain of evidence report.

[0010] Furthermore, in step S1, the controllable inference load is a YOLO series target detection inference task, and the preset multi-condition excitation matrix is ​​a combination of multiple conditions covering three dimensions: GPU frequency level, inference accuracy, and batch size. The excitation duration is not less than twice the device thermal time constant, and the excitation power consumption span is not less than 2W.

[0011] Furthermore, the multi-condition stimulus matrix includes 12 stimulus sequences formed by combining 3 GPU frequencies, 2 inference precisions, and 2 batch sizes. The GPU frequency levels include 305MHz, 713MHz, and 914MHz, the inference precisions include FP16 and FP32, and the batch sizes include bs1 and bs4. The duration of each stimulus sequence is 180s.

[0012] Furthermore, in step S1, the multi-source telemetry data is collected through the telemetry interface built into the edge AI device system, without the need for external hardware sensors. The collection interface includes tegrastats and jtop, and the collection fields include timestamp, GPU temperature, GPU power consumption, GPU frequency, GPU utilization, and end-to-end inference latency.

[0013] Furthermore, in step S2, the telemetry data is binned and time-aligned at a frequency of 1Hz. Standardized fields are defined, including timestamp, GPU temperature, GPU power consumption, GPU frequency, GPU utilization, and end-to-end inference latency. At the same time, a validity flag is set to retain only valid data with a temperature between 20°C and 90°C, power consumption greater than 0, and inference latency greater than 0.

[0014] Furthermore, in step S3, an autoregressive ARX thermal dynamics model with external input is used for online parameter identification. The ARX model expression is: T_t = a·T_{t-1} + b·P_t + c, where T_t is the current GPU temperature, T_{t-1} is the previous GPU temperature, P_t is the current GPU power consumption, and a, b, and c are the parameters to be identified. Simultaneously, the inherent thermal sensitivity coefficient k and the equivalent ambient temperature T_amb are calculated through cross-condition steady-state linear regression, with the regression formula: T_ss = k·P_ss +T_amb, where T_ss is the steady-state operating temperature and P_ss is the steady-state operating power consumption; in step S3, the real-time confidence is calculated as: confidence_t=max(0,1−RMSE_t / σ_T), where RMSE_t is the root mean square error between the identified predicted value and the measured temperature within the current sliding window, and σ_T is the standard deviation of the temperature signal within the current sliding window; the confidence threshold is set to 0.80, and thermal failure time prediction output and closed-loop control execution are only allowed when confidence_t≥0.80.

[0015] Furthermore, in step S4, the exponentially weighted moving variance (EWMV) is used to detect drift in the identified prediction residuals. When the EWMV exceeds three times the threshold of the benchmark steady-state variance, it is determined to be a parameter drift, triggering a local parameter update. The local parameter update strategy is to keep the structure of the heat map unchanged and only update the parameters affected by the drift, with the parameter reconvergence time not exceeding 40 seconds.

[0016] Furthermore, in step S5, a first-order thermodynamic model is used to predict temperature evolution, and the model expression is: Where T(t) is the predicted temperature after time t, T_ss is the steady-state temperature under current operating conditions, and T_now is the current measured temperature. Where is the thermal time constant; the formula for calculating the thermal failure time (TTF) is: , where T_thresh is the thermal failure threshold.

[0017] Furthermore, in step S6, the preset warning window is 30 seconds, and the hierarchical closed-loop control strategy is set according to the principle of minimum performance cost, with the execution priorities from low to high as follows: GPU frequency soft limit, reduce batch size, limit inference FPS, and switch NVPModel power consumption mode. Each control action records the execution timestamp, triggering condition, device status before and after execution, and execution result.

[0018] This application also discloses an electronic device, including: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the above-described edge AI thermal parameter online identification and stability prediction control method of the present invention.

[0019] This application also discloses an edge AI thermal parameter online identification and stability prediction control system, including: The active excitation module is used to generate a reproducible excitation sequence based on a preset multi-condition excitation matrix, using the controllable inference load of the edge AI device as the active excitation source, and to collect multi-source telemetry data of the device. The telemetry alignment module is used to perform time-series alignment and validity screening on the collected telemetry data to generate a standardized telemetry dataset. The online identification and confidence assessment module is used to identify parameters of a standardized dataset online based on a heat dissipation dynamics model, calculate the real-time confidence of the identification results, and set a gating threshold. The drift detection and parameter update module is used to detect parameter drift based on the identified prediction residual. When drift is detected, a local parameter update is triggered. During the update, the risk prediction output is blocked by confidence gating. The thermal failure time prediction module is used to predict the remaining time for the device to reach the thermal failure threshold based on the identified thermal parameters and a thermodynamic model, and output the prediction results with confidence intervals. The closed-loop control module is used to execute a hierarchical closed-loop control strategy according to the principle of minimum performance cost when the predicted remaining time is less than the preset warning window and the confidence level meets the gating requirements. The evidence chain generation module is used to automatically generate a structured evidence chain report that is traceable throughout the entire process.

[0020] Positive and beneficial effects: This invention predicts the thermal failure time (TTF) using a first-order thermodynamic model, allowing for a 15-30 second warning window before thermal throttling occurs. This enables proactive control strategies to be implemented in advance, avoiding passive frequency reduction after temperature limits are exceeded. Real-world testing shows that the inference latency jitter σ in BS1 mode can be controlled within 0.74ms, significantly improving the operational stability and real-time performance of edge AI services. It achieves proactive intervention in thermal risks, solving the performance jitter problem caused by the passive response of existing technologies.

[0021] This invention employs a purely software-based, controllable inference load as the active stimulus, requiring no external professional testing equipment. Data can be collected solely through the device's built-in telemetry interface, enabling full-process identification on-site. Full calibration of 12 stimulus sequences can be completed in just 36 minutes, and single-condition identification can converge within 38 seconds. This reduces the time for troubleshooting thermal issues from hours to minutes, significantly lowering testing costs and maintenance complexity. It achieves rapid on-site thermal parameter identification without additional hardware, solving the problems of existing technologies that rely on offline laboratory testing, are costly, time-consuming, and unable to adapt to individual device differences.

[0022] This invention, through experimental verification, demonstrates that the steady-state thermal sensitivity coefficient k is the inherent thermal resistance of the device, with a difference of less than 1.5% across different power consumption modes. It is unaffected by operating conditions, peripherals, or hardware versions, and the parameter can be directly transferred across devices. Simultaneously, through EWMV drift detection, it can detect drift in the device's thermal characteristics within 14 seconds, triggering local parameter updates, and completing reconvergence within 38 seconds. This adapts to changes in the device's thermal characteristics throughout its entire lifecycle, eliminating the need for manual recalibration and significantly reducing long-term maintenance costs. It achieves high transferability of thermal parameters and adaptive updates throughout the entire lifecycle, solving the problems of existing technologies where parameters cannot be transferred or adapted to thermal characteristic drift.

[0023] This invention significantly improves the signal-to-noise ratio of the identification input through active excitation. The measured RMSE of ARX dynamic identification is as low as 0.077℃, with R²=0.990, demonstrating extremely high identification accuracy. Simultaneously, a confidence assessment and gating mechanism based on prediction residuals is designed, allowing prediction and control output only when the confidence level is ≥0.80. This fundamentally avoids erroneous control during periods of identification instability, significantly improving the operational reliability of the thermal management system. It achieves high-precision, high-reliability online identification, solving the problems of low identification accuracy and erroneous control easily caused by the lack of confidence gating in existing technologies.

[0024] This invention automatically generates a structured evidence chain report containing seven core blocks, fully recording the entire process from stimulus input, data acquisition, parameter identification, risk prediction to control execution. This forms a complete causal chain document that is auditable and reproducible. Engineers can quickly distinguish whether performance anomalies are due to thermal issues or model / hardware problems, significantly shortening the troubleshooting cycle. It achieves fully traceable evidence chain recording, solving the problems of black-box operation, lack of auditability, and high difficulty in troubleshooting in existing technologies.

[0025] This invention employs an identification system combining an ARX dynamic model and a steady-state regression method. The ARX model verifies the dynamic accuracy of online identification, while the steady-state regression method yields inherent thermal resistance parameters of the equipment with clear physical meaning. These two methods complement each other, ensuring both real-time accuracy and reliable thermal risk prediction, thus providing a solid model foundation for active control. This achieves a balance between dynamic identification accuracy and steady-state prediction reliability, solving the problem of unclear physical meaning and unreliable predictive power in existing technologies.

[0026] This invention designs a hierarchical closed-loop control strategy based on the principle of minimum performance cost, prioritizing control actions that have the least impact on inference performance. The measured performance adjustment margin reaches up to 49.6%, avoiding thermal throttling while keeping throughput loss below 10%. It is expected to reduce thermal throttling events by more than 80% and improve inference latency jitter by more than 50%, perfectly balancing the requirements of thermal stability and inference performance. This achieves closed-loop control with minimum performance cost, maximizing the preservation of device inference performance while ensuring thermal stability. Attached Figure Description

[0027] Figure 1 This is a flowchart of the edge AI thermal parameter online identification and stability prediction control method of the present invention; Figure 2 This is a flowchart of the active excitation steps in an embodiment of the present invention; Figure 3 This is a flowchart of the standardized alignment steps in an embodiment of the present invention; Figure 4 This is a flowchart of the online identification and confidence assessment steps in an embodiment of the present invention; Figure 5 This is a flowchart of the drift detection and parameter update steps in an embodiment of the present invention; Figure 6 This is a flowchart of the thermal failure time prediction steps in an embodiment of the present invention; Figure 7 This is a flowchart of the closed-loop control steps in an embodiment of the present invention; Figure 8 This is a flowchart of the evidence chain generation steps in an embodiment of the present invention; Figure 9 This is a timing diagram of the active excitation input and system temperature response in an embodiment of the present invention; Figure 10 This is a graph showing the online convergence curve of the steady-state thermal parameter RLS in an embodiment of the present invention; Figure 11 This is a comparison chart of measured and online predicted temperature values ​​under different operating conditions in this embodiment of the invention. Figure 12 This is a monitoring diagram of the entire process of thermal parameter drift detection and local update in an embodiment of the present invention.

[0028] Figure 13 This is a comparison and verification diagram of the effectiveness of the closed-loop control strategy in the embodiments of the present invention. Detailed Implementation

[0029] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0030] This invention provides a method for online identification and stability prediction control of edge AI thermal parameters, such as... Figure 1 As shown, it includes the following steps: S1. Proactive motivation steps, such as Figure 2 As shown: Using the controllable inference load of the edge AI device as the active excitation source, a reproducible excitation sequence is generated according to the preset multi-condition excitation matrix, and multi-source telemetry data of the device are collected.

[0031] Furthermore, the controllable inference load is a YOLO series object detection inference task running on an edge AI device. By adjusting the three core parameters of the inference task—GPU frequency, inference accuracy, and batch size—a multi-dimensional and reproducible stimulus sequence is constructed. The preset multi-condition stimulus matrix consists of multiple reproducible combinations of conditions covering the three dimensions of GPU frequency, inference accuracy, and batch size. The stimulus duration is no less than twice the device's thermal time constant to ensure that the device temperature can fully respond to changes in the stimulus. The stimulus power consumption range is no less than 2W to provide sufficient signal-to-noise ratio for subsequent parameter identification.

[0032] Preferably, the multi-condition excitation matrix comprises 12 excitation sequences formed by combining 3 GPU frequencies, 2 inference accuracies, and 2 batch sizes. The GPU frequency levels include low frequency (305MHz), mid frequency (713MHz), and high frequency (914MHz). The inference accuracies include FP16 half-precision and FP32 single-precision. The batch sizes include single-frame bs1 and four-frame bs4. The duration of each excitation sequence is 180s, satisfying a requirement of not less than twice the thermal time constant (measured thermal time constant). =77.1s, The requirement of 154.2s was met, and the measured excitation power consumption span ΔP≈4W, which provided sufficient excitation amplitude for steady-state regression.

[0033] Furthermore, the multi-source telemetry data is collected through the telemetry interface built into the edge AI device system, without the need to connect any external hardware sensors. Specific collection interfaces include tegrastats and jtop. The collected telemetry data fields include timestamp, GPU temperature, GPU power consumption, GPU frequency, GPU utilization, and end-to-end inference latency.

[0034] S2. Standardize alignment steps, such as Figure 3 As shown: The collected telemetry data is time-series aligned and validity filtered to generate a standardized telemetry dataset.

[0035] Furthermore, the collected multi-source telemetry data is binned and time-aligned at a 1Hz frequency to eliminate time-series offsets and high-frequency noise from different data sources. Standardized telemetry fields are defined, including timestamp, gpu_temp_C (GPU temperature in °C), power_mW (GPU power consumption in mW), gpu_freq_MHz (GPU frequency in MHz), gpu_util_percent (GPU utilization in percentage), and end_to_end_ms (end-to-end inference latency in ms). A valid_flag is added to each data set to filter data validity. Only data with a temperature between 20°C and 90°C, power consumption greater than 0, and inference latency greater than 0 are set to valid (1) and deemed invalid (0). Only valid data is used for subsequent parameter identification, ensuring the quality of the input data.

[0036] S3. Online identification and confidence assessment steps, such as Figure 4 As shown: Online parameter identification is performed on a standardized telemetry dataset based on a heat dissipation dynamics model, and the real-time confidence level of the identification results is calculated and a confidence level gating threshold is set.

[0037] Furthermore, an autoregressive ARX thermal dynamics model with external input is used for online parameter identification. The expression of the ARX model is as follows: T_t = a·T_{t-1} + b·P_t + c; Where T_t is the GPU temperature at the current moment, T_{t-1} is the GPU temperature at the previous moment, P_t is the GPU power consumption at the current moment, and a, b, and c are the model parameters to be identified.

[0038] Furthermore, by using the identified ARX model parameters, thermodynamic parameters with clear physical meaning are calculated, including: (1) The steady-state thermal sensitivity coefficient k, in °C / W, reflects the inherent thermal resistance of the equipment and is the core parameter for predicting thermal failure time. It is obtained by steady-state linear regression across 12 operating conditions. The regression formula is: T_ss = k·P_ss + T_amb; Where T_ss is the steady-state temperature under operating conditions, P_ss is the steady-state power consumption under operating conditions, and T_amb is the equivalent ambient temperature. Linear regression was performed using 12 sets of steady-state observation points under operating conditions, and the measured values ​​were k=1.293℃ / W, R²=0.952, and the 95% confidence interval was ±0.180℃ / W.

[0039] (2) Thermal time constant The unit is seconds (s), which reflects the speed of the equipment's temperature response and determines the magnitude of thermal inertia. The calculation formula is: ; The mean value was obtained by fitting the first-order exponential curve of 5 independent working conditions, and the result was obtained by actual measurement. =77.1s, 95% confidence interval is ±14.3s.

[0040] Furthermore, the quality of ARX dynamic identification was verified. The average identification results of 12 sets of working conditions were: RMSE=0.077℃, R²=0.990, which proved the high accuracy of online identification.

[0041] It is important to note that the dynamic gain k_ARX=b / (1-a)≈6.469℃ / W obtained by ARX model conversion differs from k=1.293℃ / W obtained by steady-state regression by approximately 5 times. The core reason for this difference lies in the fact that the two methods measure different physical quantities: the steady-state regression method fits a large power consumption span (ΔP≈4W) across operating conditions, reflecting the inherent thermal resistance R_th of the device, which has a clear physical meaning and is suitable for long-term thermal failure time prediction; while the ARX dynamic identification uses small power consumption fluctuations (σ_P≈0.08W) within the same operating condition in a short window. The identification results contain transient effects, and small excitations lead to a higher condition number in the identification matrix. Therefore, k_ARX is only used to prove the dynamic accuracy of online identification and is not used for thermal failure time prediction. The two methods complement each other and together constitute the parameter identification system of this invention.

[0042] Furthermore, a real-time confidence calculation method for the identification results is designed, and the formula for calculating the confidence_t is: confidence_t=max(0, 1−RMSE_t / σ_T); Where RMSE_t is the root mean square error between the predicted and measured temperature values ​​within the current sliding window, and σ_T is the standard deviation of the temperature signal within the current sliding window. A confidence threshold c_gate = 0.80 is set. Subsequent thermal failure time prediction results and closed-loop control actions are only allowed to be output and executed when the real-time confidence_t ≥ 0.80, preventing erroneous control during periods of unstable identification and ensuring the reliability of equipment operation.

[0043] S4. Drift detection and parameter update steps, as follows Figure 5 As shown: Parameter drift detection is performed based on the identification prediction residual. When parameter drift is detected, local parameter update is triggered. During the update, risk prediction output is blocked by confidence gating.

[0044] Furthermore, exponentially weighted moving variance (EWMV) is used to statistically analyze the identified prediction residuals to achieve online detection of parameter drift. Specifically, the EWMV of the prediction residuals is calculated; when the EWMV exceeds three times the benchmark steady-state variance threshold, parameter drift is identified, triggering a local parameter update. Experimental results show that the drift detection delay of this invention is approximately 14 seconds, enabling rapid response to changes in equipment thermal characteristics.

[0045] Furthermore, the local parameter update strategy is as follows: Maintaining the overall structure of the thermal parameter map, only updating the parameters of nodes and edges affected by drift, without needing to re-identify all parameters. The measured parameter reconvergence time is approximately 38 seconds, significantly shortening the parameter update time. During parameter updates, the real-time confidence level of the identification results automatically drops below the 0.80 gating threshold, actively blocking the output of thermal failure time prediction and preventing erroneous control caused by unstable identification results during parameter updates.

[0046] Specifically, through perturbation experiments of nvpmodel power mode switching, it was verified that the convergence values ​​of the steady-state thermal sensitivity coefficient k of the device in the three power modes of 25W, 15W, and 25W are 1.293℃ / W, 1.296℃ / W, and 1.284℃ / W, respectively. The difference between the three is less than 1.5%, which proves that k is the inherent thermal resistance of the device and is independent of the operating conditions and configuration. This directly supports the core technical effect of the present invention that the parameters can be migrated across peripherals, operating conditions, and hardware versions.

[0047] S5. Steps for predicting thermal failure time, such as... Figure 6 As shown: Based on the identified thermal parameters, the remaining time for the device to reach the thermal failure threshold is predicted by a thermodynamic model, and the prediction results with confidence intervals are output.

[0048] Furthermore, a first-order thermodynamic model is used to predict temperature evolution. The expression for the first-order thermodynamic model is as follows: ; Where T(t) is the predicted temperature after time t, T_ss is the steady-state temperature under the current operating conditions, and T_now is the measured temperature at the current time. To identify the obtained thermal time constant, the formula for calculating the steady-state temperature T_ss is as follows: T_ss = T_amb + k·P_now; Where T_amb is the identified equivalent ambient temperature, k is the identified steady-state thermal sensitivity coefficient, and P_now is the measured GPU power consumption at the current moment.

[0049] Furthermore, based on the above temperature evolution model, the remaining time for the device to reach the thermal failure threshold T_thresh, i.e., the thermal failure time (TTF), is calculated using the following formula: ; Based on thermal time constant The 95% confidence interval is ±14.3s, and the confidence interval of the TTF prediction result is estimated. The mean absolute error (MAE) of the measured TTF prediction is <15s. When the current temperature of the equipment is at 75% of the steady-state temperature, the prediction lead is about 30s, which can reserve a buffer window of 15~30s for subsequent closed-loop control, and realize the early prediction of thermal throttling risk.

[0050] S6. Closed-loop control steps, such as Figure 7 As shown: when the predicted remaining time is less than the preset warning window and the confidence level meets the gating requirements, the hierarchical closed-loop control strategy is executed according to the principle of minimum performance cost.

[0051] Furthermore, the recommended value for the preset warning window is 30 seconds, which can be adjusted according to the real-time requirements of edge AI services. Closed-loop control is only triggered when TTF < 30 seconds and real-time confidence_t ≥ 0.80, thus avoiding unnecessary performance loss and miscontrol.

[0052] Furthermore, the hierarchical closed-loop control strategy sets execution priorities according to the principle of minimum performance cost, from low to high: GPU frequency soft limiting, reducing batch size, limiting inference FPS, and switching NVPModel power consumption mode. Control actions with the least impact on inference performance are executed first, and the next level of control action is only executed if the previous level of control action cannot effectively suppress thermal risks. This avoids thermal throttling while maximizing the performance of edge AI inference services.

[0053] Furthermore, each control action is recorded with complete information, including the action execution timestamp, triggering conditions, device status before execution, device status after execution, and action execution result, for use in subsequent evidence chain report generation and post-event auditing. Experimental results show that the control strategy of this invention has ample performance adjustment space. By increasing the GPU frequency from 305MHz to 914MHz, the inference latency in the FP32bs4 operating condition can be reduced by 49.6%, providing sufficient adjustment range for closed-loop control.

[0054] S7. Steps for generating the chain of evidence, such as Figure 8 As shown: Automatically generates a structured chain of evidence report with full traceability throughout the entire process.

[0055] Furthermore, the structured evidence chain report comprises seven core blocks: an incentive configuration block, recording the operating condition matrix and execution parameters of active incentives; a telemetry summary block, recording a statistical summary of the collected multi-source telemetry data; a data quality block, recording the validity screening results and quality assessment of the standardized telemetry dataset; an identification result block, recording the thermal parameter identification results, identification accuracy indicators, and confidence level changes; a thermal risk prediction block, recording the TTF prediction results, confidence intervals, and early warning trigger records; a control action log block, recording the complete execution information of all closed-loop control actions; and an effect verification block, recording the thermal management effect and performance impact assessment after the control strategy is implemented. The report is auditable and reproducible, providing a complete causal chain document for equipment health status assessment, performance anomaly investigation, and thermal management effect verification.

[0056] This application also provides an embodiment of an electronic device. The electronic device is manifested in the form of a general-purpose computing device. The components of the electronic device may include, but are not limited to: one or more processors or processing units, memory, and buses connecting different components (including memory and processing units).

[0057] Electronic devices typically include a variety of computer-readable media. These media can be any available media that can be accessed by the electronic device, including volatile and non-volatile media, and removable and non-removable media.

[0058] The memory may include computer-readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory. Electronic devices may further include other removable / non-removable, volatile / non-volatile computer device storage media. By way of example only, the storage system may be used to read and write non-removable, non-volatile magnetic media.

[0059] The electronic device can also communicate with one or more external devices (e.g., keyboard, pointing device, camera, etc.), may include a display, and may communicate with one or more devices that enable a user to interact with the electronic device, and / or with any device that enables the electronic device to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via an input / output (I / O) interface. Furthermore, the electronic device can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via a network adapter. The network adapter communicates with other modules of the electronic device via a bus. The processor executes various functional applications and data processing by running programs stored in memory, such as implementing the edge AI thermal parameter online identification and stability prediction control method provided in the above embodiments of the present invention.

[0060] The present invention will be further described in detail below with reference to specific embodiments: The edge AI thermal parameter online identification and stability prediction control method provided in this embodiment is deployed on the NVIDIA Jetson Orin NX edge AI platform, and the basic configuration of the platform is as follows: Hardware platform: NVIDIA Jetson Orin NX 16GB; Operating system: Linux 5.15.148-tegra-aarch64; Inference framework: TensorRT 10.3, CUDA 12.2, cuDNN 8.9; AI inference model: YOLOv8n object detection model, optimized with TensorRT engine using FP16 half-precision and FP32 single-precision respectively; GPU frequency levels: locked at 3 levels, namely low frequency 305MHz, mid frequency 713MHz, and high frequency 914MHz. CPU and EMC frequencies are locked throughout to eliminate interference from other units. DVFS mode: OFF locked throughout, eliminating DVFS automatic frequency reduction interference and ensuring the purity of the identified signal; Baseline power consumption mode: NVIDIA NVP Model 25W mode; Telemetry acquisition method: It is implemented by using the tegrastats tool and a self-developed Python acquisition script. The acquisition frequency is 10Hz, the frame-level timestamp is aligned, and subsequent 1Hz median binning preprocessing is performed. Experimental data scale: A total of 12 independent operating condition tests (R01~R12) were completed, each lasting 180 seconds, with a total of 115,246 frames of raw telemetry data collected, and a total test duration of approximately 3,130 seconds.

[0061] The specific implementation steps of the edge AI thermal parameter online identification and stability prediction control method in this embodiment are as follows: The active excitation and telemetry data acquisition step aims to generate a reproducible, high signal-to-noise ratio active excitation sequence through a controllable inference load in pure software, and to acquire the thermal response telemetry data of the device without any external hardware sensors.

[0062] In this embodiment, 12 reproducible stimulus conditions covering three dimensions—GPU frequency level, inference accuracy, and batch size—are designed to form a complete stimulus matrix. The specific combination of conditions is shown in Table 1.

[0063] Table 1: Parameters for 12 Groups of Active Excitation Conditions

[0064] The above 12 operating conditions cover the typical configuration of edge AI inference services. Among them, the R12 operating condition (FP32 bs4914MHz High) has the highest power consumption (10.24W) and the largest temperature rise, providing the best signal-to-noise ratio for model identification. The power consumption range of the entire operating condition is ΔP≈4W (from 6.29W in R01 to 10.24W in R12), which provides sufficient excitation amplitude for subsequent steady-state linear regression.

[0065] The execution flow for each set of operating conditions is as follows: First, the GPU frequency is locked to the target level, the YOLOv8n TensorRT inference engine with the corresponding precision and batch size is loaded, and a 30-second warm-up run is performed. After the device status stabilizes, the continuous inference task is started, and telemetry data acquisition is initiated simultaneously, running continuously for 180 seconds. The excitation duration of 180 seconds is ≥ twice the measured thermal time constant. This ensures that the device temperature can fully respond to changes in power consumption and reach a steady state.

[0066] Telemetry data is collected through the device’s built-in tegrastats interface and a self-developed Python script, without the need for any external sensors. The raw data fields collected include: timestamp, GPU temperature, GPU core voltage, GPU frequency, GPU utilization, GPU power consumption, CPU frequency, CPU utilization, system power consumption, end-to-end inference latency, and inference frame rate.

[0067] like Figure 9As shown in the figure, this graph illustrates the time-series relationship between the active excitation input and the system temperature response under typical R12 operating conditions. The horizontal axis represents time (in seconds), the left vertical axis represents the GPU core temperature (in °C), and the right vertical axis represents the GPU real-time power consumption (in W). The blue solid line represents the measured GPU temperature, the orange dashed line represents the temperature fitting curve based on the first-order thermodynamic model, and the gray solid line represents the GPU power consumption excitation input. This graph visually verifies the effectiveness of the active excitation sequence of this invention: the GPU temperature changes synchronously with the power consumption excitation, with a temperature lag of approximately 8 seconds; the first-order thermodynamic model fitting curve highly matches the measured temperature; and the measured thermal time constant is marked. The steady-state thermal sensitivity coefficient k = 1.293℃ / W verifies the applicability of the first-order thermodynamic model of this invention, and that active excitation provides an effective input with a high signal-to-noise ratio for parameter identification.

[0068] The standardization, alignment, and validity screening steps for telemetry data are primarily aimed at preprocessing the raw telemetry data to eliminate temporal offsets and high-frequency noise, filter valid data, and generate a standardized Golden CSV dataset to ensure the quality of input data for subsequent parameter identification.

[0069] Timing alignment and noise reduction processing: The raw telemetry data collected at 10Hz is binned at a frequency of 1Hz, that is, the median of 10 data points in each second is taken as the effective value of that second, so as to achieve timing alignment and high-frequency noise suppression and eliminate timing offset of different data sources.

[0070] Standardized field definitions: Define the standardized fields for the Golden CSV dataset, specifically including: timestamp: timestamp, in seconds, Unix time format; gpu_temp_C: GPU core temperature, in °C; power_mW: Real-time power consumption of the GPU, in mW; gpu_freq_MHz: GPU real-time frequency, in MHz; gpu_util_percent: Real-time GPU utilization, in percentage (%) end_to_end_ms: End-to-end latency for single-frame inference, in milliseconds; valid_flag: Data validity flag, 1 for valid, 0 for invalid.

[0071] Validity screening involves determining the validity of each binned data set. Data is considered valid only if all three of the following conditions are met simultaneously: `valid_flag` is set to 1. (1) GPU temperature gpu_temp_C∈[20, 90]℃, excluding data with abnormal temperature acquisition; (2) GPU power consumption power_mW>0, excluding data with abnormal power consumption acquisition; (3) End-to-end delay end_to_end_ms>0, exclude data that is abnormal in the inference task.

[0072] Only valid data with valid_flag=1 is used for subsequent parameter identification, and the proportion of valid data exceeds 99.5%, ensuring the high quality of the identification input data.

[0073] The online identification and confidence assessment of thermal parameters is a key step. The core objective of this step is to identify thermodynamic parameters with clear physical meaning online using standardized telemetry data, calculate the real-time confidence of the identification results, and set a gating threshold to ensure the reliability of the identification results.

[0074] This embodiment adopts an identification system that combines the ARX thermal dynamics model with steady-state linear regression, taking into account both dynamic identification accuracy and the physical reliability of steady-state prediction.

[0075] The online identification of the ARX dynamic model uses the following expression for the ARX thermal dynamics model: T_t = a·T_{t-1} + b·P_t + c; Where T_t is the GPU temperature at the current moment, T_{t-1} is the GPU temperature at the previous moment, P_t is the GPU power consumption at the current moment, and a, b, and c are the model parameters to be identified.

[0076] The recursive least squares (RLS) algorithm was used for online parameter identification to achieve real-time parameter updates, with a sliding window size of 30 seconds. The average values ​​were taken after identifying 12 sets of working conditions, and the identification results are shown in Table 2.

[0077] Table 2: Mean values ​​of ARX dynamic recognition results:

[0078] As can be seen, the RMSE of ARX dynamic identification is as low as 0.077℃, and R²=0.990, which proves that the online identification of the present invention has extremely high accuracy.

[0079] Steady-state thermal parameter regression calculation: In order to obtain thermal parameters with clear physical meaning and applicable to TTF prediction, this embodiment adopts the steady-state linear regression method across 12 operating conditions to calculate the inherent thermal resistance (steady-state thermal sensitivity coefficient k) of the equipment and the equivalent ambient temperature T_amb.

[0080] First, the median of the effective data in the last 100 seconds for each operating condition was taken as the steady-state observation point for that condition, resulting in 12 steady-state observation points (P_ss, T_ss). Then, the linear regression formula was used for fitting. T_ss = k·P_ss + T_amb; The regression results obtained are as follows: Steady-state thermal sensitivity coefficient k = 1.293℃ / W, 95% confidence interval ±0.180℃ / W, R² = 0.952; The equivalent ambient temperature T_amb = 41.96℃.

[0081] Meanwhile, the thermal time constant was obtained by first-order exponential fitting of five independent operating conditions. The mean value is 77.1s, the 95% confidence interval is ±14.3s, and the fitted RMSE is <0.2℃.

[0082] It is important to note that the k_ARX obtained from ARX dynamic identification (6.469℃ / W) differs from the k=1.293℃ / W obtained from steady-state regression by approximately 5 times. The core reason for this difference lies in the fact that the two methods measure different physical quantities. Steady-state regression fits a large power consumption span (ΔP≈4W) across operating conditions, reflecting the inherent thermal resistance R_th of the device. This is unaffected by changes in operating conditions and has a clear physical meaning, making it a core parameter for TTF prediction. In contrast, ARX dynamic identification uses small power consumption fluctuations (σ_P≈0.08W) within a short window under the same operating condition. The identification results include transient thermal effects, and the small excitation leads to a higher condition number in the identification matrix. Therefore, k_ARX is only used to demonstrate the dynamic accuracy of online identification and is not used for TTF prediction. The two methods complement each other and together constitute the complete identification system of this invention.

[0083] like Figure 10 As shown in the figure, this diagram illustrates the online convergence process of steady-state thermal parameters using the Recursive Least Squares (RLS) algorithm. It is divided into two subplots: the horizontal axis represents the cumulative observation condition number; the vertical axis of the upper subplot represents the steady-state thermal sensitivity coefficient k (unit: ℃ / W); and the vertical axis of the lower subplot represents the equivalent ambient temperature T_amb (unit: ℃). The solid black line represents the online estimated value of the parameter, the gray area represents the 95% confidence interval of the parameter, and the dashed black line represents the reference value for batch regression. This figure verifies the stable convergence of the online parameter identification method of this invention: the thermal sensitivity coefficient k and the equivalent ambient temperature T_amb monotonically converge from their initial values ​​with the accumulation of observation data, the 95% confidence interval continuously narrows, and the final converged value is completely consistent with the batch offline regression benchmark value. The measured steady-state thermal sensitivity coefficient k converges to 1.293 ℃ / W, with a deviation from the reference value of less than 1.5%, proving the accuracy and feasibility of the online identification method of this invention.

[0084] Confidence assessment and gating mechanism: This embodiment designs a real-time confidence calculation method based on prediction residuals. The formula for calculating confidence_t is as follows: confidence_t=max(0, 1−RMSE_t / σ_T); Where RMSE_t is the root mean square error between the predicted temperature and the measured temperature within the current 30-second sliding window; σ_T is the standard deviation of the measured temperature signal within the current 30-second sliding window.

[0085] The confidence threshold c_gate=0.80 is set so that subsequent TTF prediction results and closed-loop control actions are only allowed when the real-time confidence_t≥0.80; when confidence_t<0.80, the TTF prediction output is automatically blocked to prevent false control during the identification instability period, thus ensuring the reliability of the system operation from the mechanism.

[0086] like Figure 11 As shown in the figure, this graph compares the measured GPU temperature with the online predicted value of the sliding window OLS under typical operating conditions of low, medium, and high GPU frequencies. It includes three sub-graphs corresponding to the FP32bs1 operating conditions at 305MHz, 713MHz, and 914MHz frequencies, respectively. The horizontal axis represents the load running time (in seconds), and the vertical axis represents the GPU core temperature (in degrees Celsius). The solid blue line represents the measured GPU temperature, the dashed orange line represents the predicted value of the online identification model, and the shaded area represents the ±RMSE prediction error band. This graph verifies the high prediction accuracy of the online identification model and the effectiveness of the confidence gating mechanism of this invention: the predicted RMSE for all three typical operating conditions is below 0.25℃, far below the standard deviation of the temperature signal. The real-time confidence_t calculated based on the formula of this invention is greater than 0.80, fully meeting the confidence gating threshold requirements set by this invention, and can prevent miscontrol during periods of unstable identification.

[0087] The parameter drift detection and local update step aims to detect thermal characteristic drift throughout the entire life cycle of the equipment online, trigger local parameter updates, ensure the long-term validity of thermal parameters, and avoid miscontrol during the update period through confidence gating.

[0088] In this embodiment, the drift detection method employs exponentially weighted moving variance (EWMV) to statistically analyze the identification prediction residuals, thereby achieving online detection of parameter drift. The specific method is as follows: (1) Calculate the prediction residual of the identification model e_t = T_t_measured - T_t_predicted; (2) Calculate the exponentially weighted moving variance EWMV_t of the residuals, and set the smoothing coefficient to 0.95; (3) Set the drift detection threshold to 3 times the benchmark steady-state variance σ²_benchmark. When EWMV_t>3σ²_benchmark, it is determined that parameter drift has occurred and local parameter update is triggered.

[0089] A local parameter update strategy is employed when parameter drift is detected. This strategy maintains the overall structure of the heatmap, updating only the parameters of nodes and edges affected by the drift, eliminating the need for full re-identification and significantly reducing parameter reconvergence time. During parameter updates, the real-time confidence level of the identification results automatically drops below a gating threshold of 0.80, actively blocking TTF prediction output and preventing erroneous control caused by unstable results during the update period.

[0090] This embodiment verifies the effectiveness of drift detection and parameter updating through a perturbation experiment involving switching the NVPModel power consumption mode. The experiment process is as follows: the device's NVPModel power consumption mode is switched from 25W to 15W, run continuously for 300 seconds, and then switched back to 25W mode. Throughout the process, changes in thermal parameters, predicted residual EWMV, and confidence level are monitored. The changes in the core monitored indicators are as follows: (1) GPU temperature track: When the power consumption mode is switched from 25W to 15W, the steady-state temperature T_ss drops from 56.2℃ to 51.0℃, ΔT=5.2℃, which is completely consistent with the calculation results of the first-order thermal model, verifying the correctness of the model; (2) Real-time thermal sensitivity coefficient k_t orbit: The convergence values ​​of the three stages are 1.293℃ / W, 1.296℃ / W and 1.284℃ / W respectively. The difference between the three is less than 1.5%, which proves that k is the inherent thermal resistance of the device and is not related to the operating conditions and power consumption mode. This directly supports the core effect of the present invention that the parameters can be migrated across operating conditions, peripherals and hardware versions. (3) Predicting residual EWMV trajectory: 14s after the power consumption mode is switched, the EWMV exceeds the threshold of 3σ², triggering drift detection, which verifies the fast response capability of the drift detection of the present invention. (4) Real-time confidence level c_t orbit: During the drift and parameter update, the confidence level automatically drops below 0.80, actively blocking the TTF output, ensuring the reliability of the system. The parameter reconvergence time is about 38s. After convergence, the confidence level automatically recovers to above 0.80.

[0091] like Figure 12As shown in the figure, this figure is a full-process monitoring diagram of thermal parameter drift detection and local update in the nvpmodel power mode switching perturbation experiment. The horizontal axis represents time (unit: s) and includes 4 monitoring tracks: track 1 is the measured value of GPU temperature, track 2 is the estimated value of real-time thermal sensitivity coefficient k_t, track 3 is the exponentially weighted moving variance EWMV of the prediction residual, and track 4 is the real-time confidence level c_t of the identification result. The figure marks the power mode switching nodes from 25W to 15W to 25W, the drift detection trigger threshold (3 times the baseline steady-state variance), and the confidence level gating threshold (0.80). This figure fully verifies the effectiveness of the drift detection and local update strategy of this invention: parameter drift can be identified and local update can be triggered within 14s after power mode switching, and parameter reconvergence can be completed within 38s; the convergence value difference of steady-state thermal sensitivity coefficient k in the three power modes is less than 1.5%, proving that k is the inherent thermal resistance of the device and has cross-condition portability; at the same time, during drift and parameter update, the real-time confidence level automatically drops below the 0.80 gate threshold, actively blocking the thermal risk prediction output and avoiding miscontrol during the update period.

[0092] The Thermal Failure Time (TTF) prediction step aims to predict the remaining time for the equipment to reach the thermal failure threshold based on the identified thermal parameters using a first-order thermodynamic model, thereby enabling early prediction of thermal throttling risks.

[0093] The temperature evolution prediction model uses a first-order thermodynamic model for temperature evolution prediction. The model expression is as follows: ; Where T(t) is the predicted GPU temperature after time t; T_ss is the steady-state temperature under the current operating conditions, calculated by the formula T_ss=T_amb+k·P_now; T_now is the measured GPU temperature at the current time; and k is the identified steady-state thermal sensitivity coefficient, 1.293℃ / W. The thermal time constant obtained is 77.1s; P_now is the measured GPU power consumption at the current moment; T_amb is the equivalent ambient temperature obtained is 41.96℃.

[0094] The TTF calculation method, based on the temperature evolution model described above, calculates the remaining time for the device to reach the thermal failure threshold T_thresh, i.e., the TTF. The calculation formula is as follows: ; In this embodiment, the thermal failure threshold T_thresh is set to 85°C, the DVFS downclocking trigger threshold of the NVIDIA Jetson Orin NX platform, and can be adjusted according to actual needs. This is based on the thermal time constant. The 95% confidence interval is ±14.3s. The confidence interval of the TTF prediction result is estimated, and the mean absolute error (MAE) of the measured TTF prediction is <15s.

[0095] When the current temperature of the equipment is at 75% of the steady-state temperature, the prediction lead time is about 30 seconds, which can reserve a buffer window of 15 to 30 seconds for subsequent closed-loop control, realize the early prediction of thermal throttling risk, and fundamentally change the passive response mode of existing technology.

[0096] The hierarchical closed-loop control step aims to perform active closed-loop control according to the principle of minimum performance cost before thermal throttling occurs, thereby avoiding inference performance jitter caused by thermal throttling and maximizing the performance of inference services.

[0097] The control trigger condition is set so that the closed-loop control action is triggered only when both of the following conditions are met simultaneously: (1) The predicted TTF is less than the preset warning window. In this embodiment, the warning window is set to 30s, which can be adjusted according to the real-time requirements of the business. (2) The real-time confidence level confidence_t ≥ 0.80, which meets the confidence gating requirements.

[0098] The above triggering conditions avoid unnecessary performance loss and prevent accidental control from occurring.

[0099] The hierarchical control strategy, based on the principle of minimum performance cost, sets the execution priority of control actions from low to high (i.e., from smallest to largest performance impact): Level 1: Soft limit on GPU frequency. Without affecting inference performance, the upper limit of GPU frequency is soft-limited to the optimal frequency for the current operating conditions, avoiding the increase in power consumption caused by meaningless frequency increases, and has almost no impact on inference performance.

[0100] Level 2: Reduce batch size. While maintaining the same inference frame rate, appropriately reducing the batch size reduces the GPU load per inference cycle, lowers power consumption, and has minimal impact on inference latency.

[0101] Level 3: Limit inference FPS. While meeting the minimum frame rate requirements of the business, appropriately limit the maximum frame rate of inference to reduce the average load on the GPU and reduce heat generation.

[0102] Level 4: Switch NVPModel power consumption mode. When the current three-level control actions cannot effectively suppress thermal risks, switch to the lower power consumption NVPModel mode to fundamentally limit the maximum power consumption of the device and ensure that the temperature does not exceed the threshold.

[0103] Prioritize executing control actions that have the least impact on inference performance. Only execute the next level of control action if the previous level of control action cannot restore the TTF to more than 30 seconds. This avoids hot throttling while maximizing the performance of inference services.

[0104] Control action logs record the following information for each control action: action execution timestamp, TTF and confidence level at the time of triggering, device status before execution (temperature, power consumption, frequency, delay), type of control action executed, device status after execution, and action execution effect. All records are stored in the control action log for subsequent evidence chain report generation and post-event auditing.

[0105] This embodiment verifies through actual testing that the control strategy has sufficient performance adjustment range: when the GPU frequency is increased from 305MHz to 914MHz, the inference latency of FP16 bs1 is reduced by 36.8%, the inference latency of FP32 bs1 is reduced by 44.9%, the inference latency of FP16 bs4 is reduced by 38.5%, and the inference latency of FP32 bs4 is reduced by 49.6%, providing sufficient adjustment range for closed-loop control and achieving an optimal balance between thermal stability and inference performance.

[0106] like Figure 13 As shown, this figure is an A / B comparison verification of the effectiveness of the closed-loop control strategy. The experimental platform was an NVIDIA Jetson Orin NX, simultaneously running three concurrent inference tasks: YOLOv8 object detection, ACT robot arm control, and SuperPoint SLAM. The experiment set up two control groups: Group A (no EdgeTwin intervention) and Group B (EdgeTwin closed-loop control enabled), collecting 1197 frames of telemetry data. The figure contains three subplots: the upper subplot compares GPU temperatures, with the horizontal axis representing time and the vertical axis representing GPU core temperature (unit: °C). The red curve represents the temperature trajectory of Group A, and the green curve represents the temperature trajectory of Group B, with the EdgeTwin soft threshold T_soft=62℃ intervention point marked; the middle subplot compares YOLOv8 inference latency (unit: ms); and the lower subplot shows the ACT robot arm control latency (unit: ms), which is a key indicator of the task.

[0107] This figure fully verifies the actual effect of the closed-loop control strategy of this invention: At the GPU temperature level, Group A (no intervention) experienced an uncontrollable drift trend with its temperature continuously rising above approximately 64°C, while Group B (with EdgeTwin enabled) automatically triggered closed-loop intervention when its temperature approached the soft threshold of 62°C, stabilizing the temperature within the 60-63°C range, thus achieving proactive suppression of thermal risks. At the YOLOv8 inference latency level, Group A experienced severe inference latency fluctuations due to thermal throttling, while Group B maintained stable latency. At the ACT robot arm control latency level, enabling EdgeTwin stabilized the robot arm control latency at around 20ms with zero frame drops, requiring no manual intervention throughout the process, verifying the actual deployment effect of this invention in mission-critical, high-load concurrent scenarios. This experiment is based on the ARX thermal model of this invention (RMSE=0.077°C, R²=0.990), with a CUSUM drift detection trigger value of 1054. The measured results directly support the technical effects of this invention regarding reducing thermal throttling events and improving inference latency jitter.

[0108] The core purpose of this step is to automatically generate a structured evidence chain report that is traceable throughout the entire process, enabling the thermal management process to be auditable and reproducible, and providing a complete causal chain document for equipment health assessment and problem investigation.

[0109] The evidence chain report automatically generated in this embodiment is in HTML format, and also supports JSON machine-readable format. It contains 7 core blocks, the details of which are as follows: Incentive Configuration Block: Records complete configuration information for 12 sets of operating condition matrices for active incentives, execution parameters for each set of operating conditions, execution time, and incentive sequence, ensuring the reproducibility of the incentive process.

[0110] Telemetry Summary Block: Records statistical summaries of the collected multi-source telemetry data, including the maximum, minimum, mean, and standard deviation of temperature, power consumption, frequency, and delay, as well as information such as the total duration and total number of frames of data acquisition.

[0111] Data Quality Block: Records the generation process of the standardized Golden CSV dataset, the results of validity screening, the percentage of valid data, and data quality assessment indicators to ensure the traceability of the input data.

[0112] Identification Results Block: Records the complete results of thermal parameter identification, including steady-state thermal sensitivity coefficient k and thermal time constant. The estimated value and 95% confidence interval of the equivalent ambient temperature T_amb, the accuracy indicators such as RMSE and R² of ARX identification, the full range of confidence curves, and a complete record of the identification process.

[0113] Thermal Risk Prediction Block: Records the entire process of TTF prediction, confidence interval, early warning trigger time and trigger conditions, and the evolution of thermal risk, enabling traceability of risk prediction.

[0114] Control Action Log Block: Records complete execution information for all closed-loop control actions, including action execution time, triggering conditions, equipment status before and after execution, and execution effect, enabling full-process auditability of the control process.

[0115] Effect Verification Block: Records the thermal management effect after the control strategy is implemented, including the reduction rate of thermal throttling events, the improvement of inference latency jitter, the loss of throughput, and a quantitative evaluation of the thermal management effect.

[0116] The report fully records the entire process information from stimulus input to control execution, forming a complete causal chain document. Engineers can use the report to quickly distinguish whether the performance anomaly is a thermal issue or a problem with the model or hardware itself, reducing the troubleshooting time for thermal issues from hours to minutes.

[0117] Additionally, an edge AI thermal parameter online identification and stability prediction control system was disclosed, including: The active excitation module is used to generate a reproducible excitation sequence based on a preset multi-condition excitation matrix, using the controllable inference load of the edge AI device as the active excitation source, and to collect multi-source telemetry data of the device. The telemetry alignment module is used to perform time-series alignment and validity screening on the collected telemetry data to generate a standardized telemetry dataset. The online identification and confidence assessment module is used to identify parameters of a standardized dataset online based on a heat dissipation dynamics model, calculate the real-time confidence of the identification results, and set a gating threshold. The drift detection and parameter update module is used to detect parameter drift based on the identified prediction residual. When drift is detected, a local parameter update is triggered. During the update, the risk prediction output is blocked by confidence gating. The thermal failure time prediction module is used to predict the remaining time for the device to reach the thermal failure threshold based on the identified thermal parameters and a thermodynamic model, and output the prediction results with confidence intervals. The closed-loop control module is used to execute a hierarchical closed-loop control strategy according to the principle of minimum performance cost when the predicted remaining time is less than the preset warning window and the confidence level meets the gating requirements. The evidence chain generation module is used to automatically generate a structured evidence chain report that is traceable throughout the entire process.

[0118] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for online identification and stability prediction control of edge AI thermal parameters, characterized in that, Includes the following steps: S1. Using the controllable inference load of the edge AI device as the active excitation source, generate a reproducible excitation sequence according to the preset multi-condition excitation matrix, and collect multi-source telemetry data of the device operation. S2. Perform time-series alignment and validity screening on the collected telemetry data to generate a standardized telemetry dataset; S3. Based on the heat dissipation dynamics model, perform online parameter identification on the standardized dataset, calculate the real-time confidence of the identification results, and set the gating threshold; S4. Parameter drift detection is performed based on the identification prediction residual. When drift is detected, local parameter update is triggered. During the update, risk prediction output is blocked by confidence gating. S5. Based on the identified thermal parameters, predict the remaining time for the equipment to reach the thermal failure threshold through a thermodynamic model, and output the prediction results with confidence intervals; S6. When the predicted remaining time is less than the preset warning window and the confidence level meets the gating requirements, the hierarchical closed-loop control strategy is executed according to the principle of minimum performance cost. S7. Automatically generate a fully traceable, structured chain of evidence report.

2. The edge AI thermal parameter online identification and stability prediction control method according to claim 1, characterized in that, In step S1, the controllable inference load is a YOLO series target detection inference task, and the preset multi-condition excitation matrix is ​​a combination of multiple conditions covering three dimensions: GPU frequency level, inference accuracy, and batch size. The excitation duration is not less than twice the device thermal time constant, and the excitation power consumption span is not less than 2W.

3. The edge AI thermal parameter online identification and stability prediction control method according to claim 2, characterized in that, The multi-condition stimulus matrix contains 12 stimulus sequences formed by combining 3 GPU frequencies, 2 inference precisions, and 2 batch sizes. The GPU frequencies include 305MHz, 713MHz, and 914MHz. The inference precisions include FP16 and FP32. The batch sizes include bs1 and bs4. The duration of each stimulus sequence is 180s.

4. The edge AI thermal parameter online identification and stability prediction control method according to claim 1, characterized in that, In step S1, the multi-source telemetry data is collected through the telemetry interface built into the edge AI device system, without the need for external hardware sensors. The collection interface includes tegrastats and jtop, and the collection fields include timestamp, GPU temperature, GPU power consumption, GPU frequency, GPU utilization, and end-to-end inference latency.

5. The edge AI thermal parameter online identification and stability prediction control method according to claim 1, characterized in that, In step S2, the telemetry data is binned and time-aligned at a frequency of 1Hz. Standardized fields are defined, including timestamp, GPU temperature, GPU power consumption, GPU frequency, GPU utilization, and end-to-end inference latency. At the same time, a validity flag is set to retain only valid data with a temperature between 20°C and 90°C, power consumption greater than 0, and inference latency greater than 0.

6. The edge AI thermal parameter online identification and stability prediction control method according to claim 1, characterized in that, In step S3, an autoregressive ARX thermal dynamics model with external input is used for online parameter identification. The ARX model expression is: T_t = a·T_{t-1} + b·P_t + c, where T_t is the current GPU temperature, T_{t-1} is the previous GPU temperature, P_t is the current GPU power consumption, and a, b, and c are the parameters to be identified. Simultaneously, the inherent thermal sensitivity coefficient k and the equivalent ambient temperature T_amb are calculated through cross-condition steady-state linear regression, with the regression formula: T_ss = k·P_ss + T _amb, where T_ss is the steady-state operating temperature and P_ss is the steady-state operating power consumption; in step S3, the real-time confidence is calculated as: confidence_t=max(0,1−RMSE_t / σ_T), where RMSE_t is the root mean square error between the identified predicted value and the measured temperature within the current sliding window, and σ_T is the standard deviation of the temperature signal within the current sliding window; the confidence threshold is set to 0.80, and thermal failure time prediction output and closed-loop control execution are only allowed when confidence_t≥0.

80.

7. The edge AI thermal parameter online identification and stability prediction control method according to claim 1, characterized in that, In step S4, the exponentially weighted moving variance (EWMV) is used to detect drift in the identified prediction residuals. When the EWMV exceeds three times the threshold of the benchmark steady-state variance, it is determined to be a parameter drift, triggering a local parameter update. The local parameter update strategy is to keep the structure of the heat map unchanged and only update the parameters affected by the drift, with the parameter reconvergence time not exceeding 40 seconds.

8. The edge AI thermal parameter online identification and stability prediction control method according to claim 1, characterized in that, In step S5, a first-order thermodynamic model is used to predict temperature evolution. The model expression is as follows: Where T(t) is the predicted temperature after time t, T_ss is the steady-state temperature under current operating conditions, and T_now is the current measured temperature. The thermal time constant; The formula for calculating the thermal failure time (TTF) is: , where T_thresh is the thermal failure threshold.

9. The edge AI thermal parameter online identification and stability prediction control method according to claim 1, characterized in that, In step S6, the preset warning window is 30 seconds. The hierarchical closed-loop control strategy is set according to the principle of minimum performance cost, with the execution priority from low to high as follows: GPU frequency soft limit, reduce batch size, limit inference FPS, and switch NVPModel power consumption mode. Each control action records the execution timestamp, triggering condition, device status before and after execution, and execution result.

10. A control system using the edge AI thermal parameter online identification and stability prediction control method as described in any one of claims 1-9, characterized in that, include: The active excitation module is used to generate a reproducible excitation sequence based on a preset multi-condition excitation matrix, using the controllable inference load of the edge AI device as the active excitation source, and to collect multi-source telemetry data of the device. The telemetry alignment module is used to perform time-series alignment and validity screening on the collected telemetry data to generate a standardized telemetry dataset. The online identification and confidence assessment module is used to identify parameters of a standardized dataset online based on a heat dissipation dynamics model, calculate the real-time confidence of the identification results, and set a gating threshold. The drift detection and parameter update module is used to detect parameter drift based on the identified prediction residual. When drift is detected, a local parameter update is triggered. During the update, the risk prediction output is blocked by confidence gating. The thermal failure time prediction module is used to predict the remaining time for the device to reach the thermal failure threshold based on the identified thermal parameters and a thermodynamic model, and output the prediction results with confidence intervals. The closed-loop control module is used to execute a hierarchical closed-loop control strategy according to the principle of minimum performance cost when the predicted remaining time is less than the preset warning window and the confidence level meets the gating requirements. The evidence chain generation module is used to automatically generate a structured evidence chain report that is traceable throughout the entire process.