A motor temperature online monitoring real-time data analysis and prediction method and system

CN122824079APending Publication Date: 2026-09-25HARBOUR GREAT (WUXI) MOTOR CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610793303.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

该类方法无法区分由负载增加、环境温度升高等正常工况波动引起的温升,与散热风扇效率衰减、轴承摩擦加剧等早期故障造成的异常热积累,导致误报率高且预警滞后

Benefits of technology

本发明提出的电机温度在线监测实时数据分析和预测方法及系统,针对工业现场电机冲击性负载、频繁变载等非平稳工况下温度监测与故障诊断的场景痛点,对选用的物理信息神经网络与时序因果卷积网络进行了场景适配的改进,构建起物理机理与深度学习相配合的诊断架构;通过将电机等效热路方程约束结构性嵌入神经网络,让物理守恒规律的泛化能力与数据驱动的特征挖掘能力形成协同,对纯数据模型在未知工况下的泛化情况进行了优化,改善了传统物理模型对复杂工况适配性不足的情况,通过将诊断对象从传统温度数值偏差拓展至热力学过程物理一致性,以区分正常工况波动与早期故障引发的异常热积累,来降低误报风险;通过不确定性感知门控机制,提升了对不同工况的适配能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824079A_ABST
    Figure CN122824079A_ABST
Patent Text Reader

Abstract

The application discloses a motor temperature online monitoring real-time data analysis and prediction method and system, relates to the motor state monitoring technical field, and comprises the following steps: S1, collecting the multi-physical quantity historical data under the motor healthy operation state, constructing a physical information neural network model embedded with the motor equivalent thermal circuit equation constraint and training, and adding a thermal balance residual constraint term of the lumped heat capacity and thermal resistance parameter on the basis of the temperature reconstruction error term of the loss function. The motor temperature online monitoring real-time data analysis and prediction method and system are provided; the motor equivalent thermal circuit equation constraint structure is structurally embedded into the neural network, the generalization ability of the physical conservation law is cooperated with the feature mining ability of the data driving, the generalization of the pure data model under unknown working conditions is optimized, the situation that the traditional physical model is insufficient in the adaptability to complex working conditions is improved, the false alarm risk is reduced, and the adaptability to different working conditions is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of motor condition monitoring technology, specifically to a method and system for real-time data analysis and prediction of motor temperature online monitoring. Background Technology

[0002] During motor operation, winding temperature is a key thermal parameter reflecting the motor's health and insulation life. Traditional temperature monitoring generally employs a fixed threshold alarm strategy, triggering an alarm when the measured temperature exceeds a preset absolute threshold. This method cannot distinguish between temperature rises caused by normal operating condition fluctuations such as increased load and ambient temperature rises, and abnormal heat accumulation caused by early faults such as decreased cooling fan efficiency and increased bearing friction, resulting in a high false alarm rate and delayed warnings. Some improved solutions introduce dynamic thresholds or temperature rise rate criteria, but threshold correction still relies on manual experience, making it difficult to achieve adaptive discrimination in multivariate coupled industrial environments with millisecond-level impact loads, such as steel rolling and stamping.

[0003] Existing data-driven intelligent diagnostic methods employ a sequential architecture that combines a healthy temperature baseline with residual classification. The healthy temperature baseline is established using algorithms such as Gaussian process regression, and then the temperature residuals are classified for faults. In this architecture, the predictive uncertainty output by the baseline model is only used for residual normalization and is discarded in subsequent classification stages. This fails to achieve cross-module collaborative utilization, and the diagnostic object is limited to temperature numerical deviations rather than thermodynamic process anomalies, thus failing to fundamentally sever the coupling between operating condition fluctuations and fault signals. Furthermore, while directly modeling the original temperature and current time series using deep learning networks can uncover the statistical correlation between operating conditions and temperature, it lacks the structural constraints of the physical laws governing motor heat transfer. Such models are prone to generalization instability under new operating conditions not covered by the training data, misclassifying normal load changes as faults. Simultaneously, when the model issues an alarm, it cannot provide interpretable physical attributions related to root causes such as heat dissipation degradation, mechanical friction, or sensor drift, limiting its practical usability in preventative maintenance decisions.

[0004] To address this, we propose a method and system for real-time data analysis and prediction of motor temperature during online monitoring. Summary of the Invention

[0005] To address the problems raised in the background art, this invention provides a method for real-time data analysis and prediction of motor temperature online monitoring, comprising: Historical data of multiple physical quantities under the healthy operating state of the motor are collected, and a physical information neural network model with constraints of the motor's equivalent thermal circuit equation is constructed and trained. The loss function adds a thermal balance residual constraint term of lumped heat capacity and thermal resistance parameters to the temperature reconstruction error term. After training, the model outputs a set of diagnostic information online, including winding temperature reconstruction value, thermal balance residual, latent space state offset, and prediction uncertainty. The latent space state offset is the distance between the mapping point of the current input in the model's latent space and the center point of the latent space in the healthy state. The center point is determined by the mean of the latent space samples of the healthy training data. Real-time acquisition of multi-physical quantity data streams; time-series data segmentation and preprocessing based on sliding window statistics of current effective value; The preprocessed real-time data is input into the physical information neural network model to obtain temperature reconstruction error, thermal balance residual, latent space state shift and prediction uncertainty, and spliced ​​together to form a multi-source physical consistency tensor. Using the enhanced features of the multi-source physical consistency tensor as input, an uncertainty-aware gating temporal causal convolutional network is constructed; its channel attention module uses the predicted uncertainty as an external guiding signal to dynamically adjust the feature weights of each channel and outputs a fault occurrence probability vector. The comprehensive fault confidence level is calculated based on the fault occurrence probability vector, the fault warning level is classified, and the diagnostic results are output.

[0006] Preferably, the healthy operating state of the motor is defined, and historical data of multiple physical quantities under typical operating conditions of the motor's healthy state are collected. The multiple physical quantities include winding temperature, effective value of current, speed, and ambient temperature. Based on the rated parameters of the motor and the specifications of the operating environment, the effective value range of each physical quantity is preset, invalid data outside the range is eliminated, pulse interference noise in the winding temperature sequence is identified and eliminated, and missing data is filled by linear interpolation or discarded. A first-order exponential smoothing filter is applied to the winding temperature sequence, and a sliding median filter and a first-order hysteresis filter are applied to the current RMS value sequence in sequence. The time alignment of the time series of multiple physical quantities is completed with the winding temperature sampling time sequence as the reference. The input feature vector is constructed using the effective value of current, the rate of change of current, the rotational speed, the ambient temperature, and the temperature hysteresis term. The winding temperature at the corresponding time is used as the output target to construct the model training dataset.

[0007] Preferably, a physical information neural network model is constructed, which includes an encoder and a decoder. The encoder maps the input feature vector to a low-dimensional physical property latent space, and the decoder reconstructs the winding temperature from the latent space. The model's total loss function is constructed, which is composed of three weighted components: temperature reconstruction error term, thermal balance residual constraint term, and latent space regularization term. The thermal balance residual constraint term is established based on the first-order equivalent thermal path differential equation of the motor, and the lumped heat capacity and lumped thermal resistance parameters of the motor are used as trainable parameters of the network. Model training is completed by minimizing the total loss function through an optimization algorithm. During training, a Dropout layer is added after the fully connected layer. In the online inference stage, multiple random forward propagations are performed on the same input. The output mean is used as the winding temperature reconstruction value, and the output variance is used as the prediction uncertainty. The output thermal balance residual and latent space state offset are calculated simultaneously. The latent space state offset is calculated based on the distance between the mean of the latent variable in the current propagation and the center point of the latent space in the healthy state.

[0008] Preferably, a sliding analysis window and a sliding step size are set, and within each sliding window, the mean, variance, and first rate of change of the current effective value sequence are calculated; Define quantitative discrimination criteria for the start-up phase, steady-state phase, variable load phase, and shutdown phase, and complete the initial segmentation of the operating condition phase of the time-series data stream based on the discrimination criteria; Work condition segments with a duration shorter than the preset minimum length are merged, and transition intervals are set for the boundary points of adjacent work condition segments to correct the work condition segmentation results. Within each working condition segment, data cleaning, filtering, and time alignment are completed according to preprocessing rules.

[0009] Preferably, it is determined whether there is a frequent switching interval in the initial segmentation result of the working condition stage where the number of switching times of the working condition segment exceeds a preset threshold within a unit time. If the frequently switching interval exists, the Bayesian online change point detection algorithm is used to estimate the posterior probability of the working condition change online based on the Bayesian recursive framework, determine the working condition change point, and re-adaptively re-segment the interval. If, after processing by the Bayesian online change point detection algorithm, the interval still cannot form a stable segment with a continuous length exceeding the preset minimum length, then a fixed-length sliding window is used to perform temporal slicing on the interval. Adjacent windows maintain a preset overlap rate, and window segments with inconsistent lengths are completed or truncated to ensure the consistency of the input dimension.

[0010] Preferably, for the preprocessed real-time data, a physical information neural network model is used to synchronously output four types of diagnostic information at each time step, including: temperature reconstruction error, thermal balance residual, latent space state shift, and prediction uncertainty. The temperature reconstruction error is the difference between the measured winding temperature and the model reconstruction value; the thermal balance residual is the residual value after substituting real-time data into the thermal circuit equation; the latent space state offset is the Euclidean distance between the latent variable after the current input mapping and the center point of the latent space in the healthy state; and the prediction uncertainty is the variance of the output results of multiple random forward propagation. Using the winding temperature sampling time sequence as the time alignment master reference, the four types of diagnostic information are mapped to the corresponding timestamps of the master reference grid and spliced ​​to form a multi-source physical consistency tensor.

[0011] Preferably, for each channel of the multi-source physical consistency tensor, statistical enhancement features are extracted within the corresponding window. The statistical enhancement features include: local mean, variance, maximum value, minimum value, trend slope, kurtosis, and skewness. The extracted statistical enhancement features are concatenated with the original time series to obtain the enhancement feature matrix.

[0012] Preferably, a temporal causal convolutional network model is constructed, which is composed of an input layer, multiple sets of causal convolutional modules, an uncertainty-aware channel attention module, a multi-scale pyramid pooling layer, a global average pooling layer, a fully connected layer, and a Softmax output layer connected in series. The input layer receives the enhanced feature matrix and normalizes the sequence of each channel; the causal convolution module uses stacked one-dimensional causal convolution blocks, follows the temporal causality constraint, and completes feature extraction using the input of the current time and historical time. The uncertainty perception channel attention module takes the predicted value of the uncertainty channel as input, calculates the correction coefficient of each channel through learnable parameters, and multiplies the original channel attention weights by the corresponding correction coefficients to obtain the adjusted final attention weights. Multi-scale temporal features are extracted and fused through a multi-scale pyramid pooling layer, and then reduced in dimensionality through a global average pooling layer. Finally, a fully connected layer and a Softmax activation function are used to output the fault probability vector corresponding to various operating states of the motor.

[0013] Preferably, based on the failure occurrence probability vector, the maximum probability value of all non-healthy failure categories is taken, and combined with the probability of the healthy category, the comprehensive failure confidence is calculated according to the following formula. :

[0014] In the formula, For the probability of the health category, This represents the maximum probability value for all non-healthy fault categories. Based on the preset interval threshold of the comprehensive fault confidence, the motor operating status is divided into three levels: normal state, early warning state, and clear fault state. The system synchronously outputs the fault category corresponding to the highest probability, as well as the physical consistency quantification anomaly description associated with the fault, and continuously stores the full monitoring data and multi-source physical consistency tensor sequence before and after the fault occurs.

[0015] This invention also provides a real-time data analysis and prediction system for online monitoring of motor temperature, comprising: The physical information model training module is used to collect historical data of multiple physical quantities of the motor's healthy operating status, construct and train a physical information neural network model with constraints of the motor's equivalent thermal circuit equation, and output a set of diagnostic information including winding temperature reconstruction values, thermal balance residuals, hidden space state shifts and prediction uncertainties. The online operating condition segmentation module is used to collect multi-physical quantity data streams in real time and complete automatic segmentation and preprocessing of operating condition stages based on the sliding window statistics of the effective current value. The multi-source consistency generation module is connected to the physical information model training module and the working condition online segmentation module respectively. It is used to input the preprocessed data into the physical information neural network model, obtain the temperature reconstruction error, thermal balance residual, latent space state shift and prediction uncertainty, and splice them to form a multi-source physical consistency tensor. The uncertainty-gated classification module is connected to the multi-source consistency generation module. It is used to take the enhanced features of the multi-source physical consistency tensor as input, introduce the prediction uncertainty as a guiding signal through the channel attention module, dynamically adjust the feature weights of each channel, and output the failure probability vector through the temporal causal convolutional network. The diagnostic result output module is connected to the uncertainty gating classification module and is used to calculate the comprehensive fault confidence, classify the fault warning level and output the diagnostic result.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention proposes a method and system for real-time data analysis and prediction of motor temperature online monitoring. Addressing the pain points of temperature monitoring and fault diagnosis under non-stationary operating conditions such as impulsive loads and frequent load changes in industrial motors, the invention improves the selected physical information neural network and temporal causal convolutional network to adapt to different scenarios, constructing a diagnostic architecture that combines physical mechanisms with deep learning. By structurally embedding the motor's equivalent thermal circuit equation into the neural network, the generalization ability of physical conservation laws and data-driven feature mining capabilities are synergistically combined, optimizing the generalization of pure data models under unknown operating conditions and improving the adaptability of traditional physical models to complex operating conditions. By expanding the diagnostic object from traditional temperature numerical deviations to the physical consistency of thermodynamic processes, the invention distinguishes between normal operating condition fluctuations and abnormal heat accumulation caused by early faults, thereby reducing the risk of false alarms. Furthermore, the invention enhances the adaptability to different operating conditions through an uncertainty-aware gating mechanism. Attached Figure Description

[0017] Figure 1 This is a diagram illustrating the real-time data analysis and prediction method for online monitoring of motor temperature according to the present invention; Figure 2 This is a block diagram of the online monitoring, real-time data analysis, and prediction system for motor temperature according to the present invention. Detailed Implementation

[0018] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.

[0019] Reference Figure 1 As shown, a method for real-time data analysis and prediction of motor temperature online monitoring includes the following steps: Step 1: Collect historical data of multiple physical quantities under the healthy operating state of the motor, construct and train a physical information neural network model with constraints of the motor's equivalent thermal circuit equation. The loss function adds a thermal balance residual constraint term of lumped heat capacity and thermal resistance parameters to the temperature reconstruction error term. After training, the model outputs a set of diagnostic information online, including winding temperature reconstruction value, thermal balance residual, latent space state offset, and prediction uncertainty. The latent space state offset is the distance between the current input mapping point in the model's latent space and the center point of the healthy state latent space. The center point is determined by the mean of the latent space samples of the healthy training data. Step one includes the following: The so-called healthy operating status of a motor is defined as the operating status of the motor within 3 hours to 7 days after it has passed the factory test or been confirmed to be defect-free after maintenance or repair, and during this period, no traditional temperature, vibration or overcurrent threshold alarms occur.

[0020] Historical data on multiple physical quantities of the motor under healthy conditions and typical operating conditions are collected, including: winding temperature, RMS current, speed, and ambient temperature. Specifically, the RMS current is acquired through a current transformer or Hall sensor, representing the three-phase instantaneous current value. After anti-aliasing low-pass filtering, it undergoes analog-to-digital conversion, and the root mean square value over one power frequency cycle is calculated in the processing unit. For example, the cutoff frequency is 500Hz, and the sampling rate for analog-to-digital conversion is no less than 1kHz. The winding temperature is acquired through a Pt100 temperature sensor embedded in the winding end. The speed is acquired through an encoder or Hall sensor. The ambient temperature is acquired through a temperature sensor installed at the motor's ventilation inlet.

[0021] The collected data is preprocessed as follows: Based on the rated parameters and operating environment specifications of the monitored motor, the effective value ranges for each physical quantity are preset. For example, the ambient temperature is limited to between -20°C and 60°C; the current is limited to an effective current reading that must be ≥0 and ≤1.5 times the rated current value; data points exceeding the boundaries are directly marked as null values, i.e., invalid data points.

[0022] For the winding temperature sequence, based on its continuously changing, slowly varying characteristics due to thermal inertia, outlier identification is performed using three times the standard deviation of the rate of change. Specifically, the absolute value of the first-order difference of the sequence is calculated:

[0023] In the formula, in the formula, For the winding temperature sequence in The absolute value of the first-order forward difference at time t represents the absolute change in temperature within a single sampling period; For winding temperature at The sampled value at time; For the winding temperature at the previous adjacent sampling time The sampled value; if Then determine the current sample value. This is pulse interference noise. This represents the standard deviation of the entire historical temperature change series.

[0024] For isolated missing data points generated after the above processing, linear interpolation is used for estimation and filling. For segments with consecutive missing lengths exceeding a preset threshold (10 sampling points for example), the entire segment is discarded and not included in subsequent model training. Given that winding temperature changes are large-inertia, hysteresis thermodynamic processes, the winding temperature sequence... Using a first-order exponential smoothing filter, its recursive expression is:

[0025] In the formula, in the formula, for The winding temperature value output after passing through a first-order exponential smoothing filter at any time; This is a smoothing coefficient, exemplified by a value of 0.2; After data cleaning, The winding temperature value of the input filter is constantly being measured; For the previous moment The filtered output value.

[0026] Considering that the RMS current value sequence in a frequency converter drive may be superimposed with pulse interference from the secondary switching circuit, a sliding median filter with a window width of 5 sampling points and a first-order hysteresis filter are sequentially applied to the RMS current value sequence. The sliding median filter is used to eliminate isolated pulse glitches, and the first-order hysteresis filter is used to further smooth the waveform profile.

[0027] For time alignment, the winding temperature sampling time series is selected as the primary reference grid. For other physical quantities such as RMS current, rotational speed, and ambient temperature, a linear interpolation algorithm is used to map the observations of each series to the corresponding timestamps of the primary reference grid. After data cleaning, filtering, and time alignment, the input feature vector is constructed:

[0028] in, For feature vectors; This is the effective value of the current; Rotational speed; This is the rate of change of current, used to capture sudden load changes; The ambient temperature; This is a temperature lag term. The output target is the winding temperature at the current moment. .

[0029] The temperature hysteresis term is a pre-selected fixed positive integer, ranging from 3 to 5. Once selected, this value remains constant throughout the entire process of model training, validation, and online inference. Its construction method is as follows: The measured winding temperature values ​​of m consecutive historical moments before the current time t are used as independent scalar feature dimensions. These are then concatenated with four basic feature dimensions: effective current value, current change rate, rotational speed, and ambient temperature, to form an input feature vector with a fixed dimension of 4+m. This vector is adapted to the input requirements of the physical information neural network model of the subsequent fully connected structure.

[0030] For sampling points at sequence boundaries where m historical temperature values ​​cannot be fully obtained, standardization and padding are performed according to the following rules: Training dataset processing scenarios, for the beginning of time series sequences The sampling points, missing The temperature values ​​at each time point are uniformly filled using the ambient temperature at the beginning of the sequence; to reduce the impact of boundary filling bias on model training, the temperature values ​​at the beginning of the sequence are preferentially selected. The sampling points are used to construct training samples; In online inference scenarios, for sampling points at the beginning of the data stream or the beginning of a segment after the operation condition is divided, if the corresponding operation condition is the start-up end, the missing hysteresis term is filled with the steady-state ambient temperature of the motor before start-up; if the corresponding operation condition is the steady-state segment, variable load segment, or shutdown segment, the missing hysteresis term is filled with the measured winding temperature value at the adjacent time before the start of the segment.

[0031] The specific steps for calculating the rate of change of current are as follows: To avoid the amplification effect of measurement noise in differential operations, a window with a width of is applied to the original current RMS value sequence. A sliding median filter is used to eliminate isolated pulse glitches introduced by inverter switching harmonics or electromagnetic coupling. A first-order low-pass filter is applied to the filtered sequence to smooth the curve profile.

[0032] On the equally spaced time series formed by time alignment and resampling, the rate of change at the current moment is calculated using the five-point central difference formula with second-order accuracy. Its expression is:

[0033] Among them, the sampling time corresponds to the time interval of the time series. The sampling period is ; To calculate The derivative of the effective value of the current with respect to time at any given moment; for The sampled value of the effective current at time t; , These are the effective current values ​​at the first and second sampling points after the current time n; , These are the effective current values ​​at the first and second sampling points before the current time n; the central difference scheme can effectively suppress the contamination of the differential result by high-frequency noise; For the boundary sampling points at the beginning and end of the sequence, the above central difference scheme cannot be directly applied. Instead, forward difference or backward difference schemes should be used:

[0034]

[0035] In the formula, in the formula, This is an approximation of the first-order forward derivative of the effective current value with respect to time at the start of the sequence. This is the first sampling moment of the current RMS value resampling sequence; For a moment The effective value of the current at the location; for The next sampling time The effective value of the current at the location; This is an approximation of the first-order backward derivative of the effective current value with respect to time at the end of the sequence. This is the last sampling moment in the current RMS value resampling sequence; For a moment The effective value of the current at the location; for Previous sampling time The effective value of the current at the location; The calculated rate of change sequence can be further lightly filtered using first-order exponential smoothing to eliminate the small sawtooth oscillations introduced by the difference operation.

[0036] A physical information neural network model with constraints from the equivalent thermal circuit equation of the motor is constructed and trained. The specific construction and training methods of the model are as follows: The model employs an autoencoder structure, consisting of an encoder and a decoder. The encoder comprises three fully connected layers, with the hidden layer dimension progressively shrinking from the input feature dimension to a 2D physical property latent space. The activation function used is ReLU. The decoder is a symmetrical structure of the encoder, also consisting of three fully connected layers, progressively recovering the input feature dimension from the 2D latent space. The final output layer is 1D, corresponding to the reconstructed winding temperature value. For example, the input feature dimension is d, i.e., the effective current value, current change rate, rotational speed, and ambient temperature, plus m temperature hysteresis terms. The encoder's hidden layer dimensions are as follows: The hidden layer dimensions of the decoder are as follows: .

[0037] The first dimension of this 2D implicit space represents the inherent thermal properties of the motor, mapping the combined characteristics of the motor's lumped thermal capacity and lumped thermal resistance. It characterizes the motor's heat dissipation capacity and thermal inertia, and is an inherent property of the motor. Under healthy conditions, the value of this dimension is only related to the motor's structure and materials, unaffected by fluctuations in operating conditions, and remains stable within a fixed range. When the motor experiences faults such as heat dissipation degradation or winding insulation aging, which alter its inherent properties, the value of this dimension will exhibit a systematic shift. The second dimension of the implicit space represents the motor's operating condition characteristics, mapping the combined characteristics of external operating conditions such as the effective value of current, speed, and ambient temperature. It characterizes the motor's current load level and thermal boundary conditions, and is independent of the motor's inherent properties. The value of this dimension fluctuates synchronously with changes in operating conditions. Normal operating condition fluctuations only cause changes in the value of this dimension and do not trigger a shift in the first dimension.

[0038] During model training, the correlation between the two dimensions is constrained by the latent space regularization term, thereby achieving physical decoupling of the two dimensions and ensuring the interpretability of subsequent latent space state shift calculations.

[0039] The model's total loss function is constructed, which is composed of a weighted average of three parts: a temperature reconstruction error term, a thermal equilibrium residual constraint term, and a latent space regularization term. Its mathematical expression is as follows:

[0040] in, The temperature reconstruction error term is expressed as follows:

[0041] In the formula, This represents the total number of historical health status data samples obtained after preprocessing. This is the measured value of the winding temperature. The winding temperature values ​​reconstructed for the model; As the thermal balance residual constraint term, it is established based on the first-order equivalent thermal path differential equation of the motor, and the lumped heat capacity of the motor is considered. and lumped thermal resistance As trainable parameters of the network, their discretized expression is:

[0042] In the formula, Δt is the sampling interval. For the total heat capacity of the motor, The total thermal resistance of the motor is used as a parameter, and both are jointly optimized during the training process as trainable parameters of the network.

[0043] The equivalent lumped thermal resistance is the combined equivalent thermal resistance of the entire heat dissipation path from the motor windings to the cooling medium. It is the series equivalent total thermal resistance of the winding thermal conduction resistance, the stator core thermal conduction resistance, and the convective heat transfer resistance between the casing and the cooling medium, satisfying the following conditions: ,in For motors with natural cooling, the radiation thermal resistance of the casing surface must be included in the calculation. For motors with forced air cooling or water cooling, convective heat transfer is dominant. It can be ignored.

[0044] For forced air-cooled motors, the convective heat transfer thermal resistance is negatively correlated with the motor speed n(t), so a speed correction coefficient is introduced. , The rated speed of the motor, adjusted for real-time convective thermal resistance. , The reference convective thermal resistance at rated speed; For water-cooled motors, the convective heat transfer resistance is a fixed value, determined by the coolant flow rate, heat transfer area, and flow channel structure. This value remains constant during training; only the conductive heat resistance component is optimized. The initial values ​​of the lumped heat capacity and equivalent lumped thermal resistance are estimated operationally according to standard procedures. The initial value of the equivalent lumped thermal resistance is first obtained from the motor's nameplate rated parameters, including rated power. Rated voltage Rated current Rated speed Rated efficiency Maximum allowable winding temperature corresponding to insulation class Then calculate the total loss of the motor under rated operating conditions. The winding copper loss ratio is taken as the industry-standard experience value of 0.6~0.7, which is the rated copper loss. Based on the rated operating condition thermal balance steady-state condition, the steady-state value of winding temperature rise is... satisfy ,in The initial value of the reference thermal resistance under rated operating conditions is calculated based on the steady-state temperature difference between the winding and the ambient temperature, taken as 60% to 70% of the allowable temperature rise for the insulation class. The unit is K / W, and this value is used as a trainable parameter of the network. The initial values ​​are set, and parameter constraints are set during training. The range of values ​​is limited to This is to avoid parameter training divergence and loss of physical meaning.

[0045] The initial value of the lumped heat capacity is calculated based on the total mass of the motor core and windings and the specific heat capacity of the materials. ,in For the quality of the stator core, The specific heat capacity of silicon steel sheets is 480 J / (kg·K). This refers to the total mass of the copper material in the winding. The specific heat capacity of copper is 385 J / (kg·K). During the training process... The range of values ​​is limited to . Let t be the ambient temperature. Let be the effective value of the current at time t. To represent the resistance of the winding as a function of temperature, the commonly used linear resistance-temperature model for copper windings in the motor industry is adopted, and its mathematical expression is as follows:

[0046] In the formula, Reference temperature The DC resistance value of the winding below, The industry standard reference temperature is 20℃. The measured value of the stator winding cold-state DC resistance in the motor's factory test report should be used first. If no measured value is available, it can be calculated from the rated parameters. , The temperature coefficient of resistance for copper is taken from the International Electrotechnical Commission (IEC) standard value. , The reconstructed winding temperature value output by the neural network model decoder at time t is consistent with the temperature term in the thermal balance equation, ensuring the self-consistency of physical constraints. During model training and online inference, the parameters of this resistance model... , , All values ​​are fixed and do not participate in network training optimization; they are only used through... To achieve real-time updates of resistance values.

[0047] Here, is the latent space regularization term, used to constrain the clustered distribution of latent space variables under healthy conditions and maintain a smooth and continuous manifold structure. Its expression is:

[0048] In the formula, For the first Latent space variables of each sample after being mapped by the encoder; is the mean vector of the latent space variables in the current training batch; this regularization term ensures that the latent variables of the healthy state form a compact and clustered healthy manifold in the low-dimensional space, avoiding latent space divergence.

[0049] To address the issues of data source consistency and training stability in thermal equilibrium residual calculation, a closed-loop optimization scheme is adopted. The calculation of the thermal equilibrium residual constraint term uses the model output throughout the process. As the sole temperature reference, the winding resistance calculation of the heat source term, the temperature rise term of the heat balance equation, and the heat dissipation term are all based on... Calculation complete, only RMS current value. Ambient temperature Using measured values ​​as model input features, there is no prediction bias, eliminating the systematic residual bias caused by inconsistent temperature references. At the same time, during the training dataset construction stage, the input features and output targets are synchronously normalized, mapping the winding temperature, current RMS value, and ambient temperature to the [0,1] interval, eliminating the influence of dimensional differences on residual calculation and gradient backpropagation.

[0050] During model training, the weighting coefficients of the thermal equilibrium residual constraint term are... A dynamic adaptive adjustment strategy is adopted, and in the initial training phase, i.e., the first 30% of training rounds, the settings are... The initial value is 0.01, prioritizing the reduction of temperature reconstruction error to allow the model to quickly learn the temperature response pattern under healthy operating conditions. This is used in the middle stage of training, i.e., 30% to 70% of the training epochs. The values ​​are linearly increased to a maximum of 0.1, gradually strengthening the physical conservation constraints. This process continues until the later stages of training, after 70% of the training rounds. Keeping the value constant at 0.1, a temperature reconstruction error threshold is introduced. When the average relative error of temperature reconstruction for a single batch of samples exceeds 10%, The value is automatically reduced to 50% of the current value, and restored to the original value when the relative error drops to within 5%, ensuring the stability of the training process.

[0051] During gradient backpropagation, the gradient corresponding to the thermal equilibrium residual constraint term is clipped using the L2 norm, with a clipping threshold set to 1.0 to avoid gradient explosion. Simultaneously, the trainable parameters... , Setting hard boundary constraints ensures that parameters exceeding the boundary are automatically pulled back to their boundary values ​​after gradient updates, guaranteeing that the parameters always have reasonable physical meaning. (Latent space regularization term weight coefficients) For example, 0.01 is used.

[0052] The model training hyperparameters were set as follows: the optimizer was Adam, the initial learning rate was 1×10−3, the batch size was 64, the number of training epochs was 200, and an early stopping strategy on the validation set was adopted. During training, a Dropout layer was added after the fully connected layer with a Dropout rate of 0.1 to obtain prediction uncertainty during the online inference stage.

[0053] After training, all healthy training data is input into the encoder to obtain the latent space variables corresponding to all healthy samples. The latent space center points of the healthy states are calculated and stored. The main scheme for calculating the latent space center points of the healthy states uses the global arithmetic mean of the latent vectors of all healthy training samples. The mathematical expression is:

[0054] In the formula, The center point of the latent space in the healthy state is consistent with the dimension of the latent space output by the encoder. This represents the total number of all healthy training samples after preprocessing. For the first The latent space vector of a healthy training sample after being mapped by the encoder.

[0055] If batch training is used, a double-weighted average of the batch-level means can be used as an equivalent calculation method. The result is equivalent to the global arithmetic mean, and the expression is:

[0056] In the formula This represents the total number of batches in the training set. For the first Number of valid samples in each batch For the first The arithmetic mean of the latent vectors of samples within each batch.

[0057] When the sample size of health training data is too large, that is When the memory and computing power costs of loading and calculating the latent vectors of the full sample exceed the deployment hardware threshold, two unbiased equivalent alternative methods are allowed to replace the full global mean calculation. The first is the incremental recursive mean calculation method, which uses a streaming incremental recursive approach to calculate the latent space sample mean in batches. It does not require storing the full latent vectors, has a constant memory footprint, and is suitable for lightweight deployment requirements at the edge. The calculation formula is as follows:

[0058] In the formula Before completion The recursive mean vector calculated from each batch, with initial values... , For the first The arithmetic mean of the latent vectors of samples within each batch. The current batch number has been calculated. After traversing all training batches, the final recursive result is the center point of the hidden space of the healthy state.

[0059] The second method is the stratified random sampling unbiased estimation method. It adopts a stratified random sampling method to extract a representative sample set from the full amount of health training data and calculate the mean. The data is stratified according to the motor operating condition type, namely the starting section, steady-state section, and variable load section. The sampling ratio of each stratum is consistent with the proportion of the sample size of that stratum to the total sample size, ensuring that the sampled samples cover all health operating conditions and have no distribution bias. The arithmetic mean of the latent vectors of all the extracted samples is calculated, and the result is used as the center point of the latent space of the health state.

[0060] Once the center point is calculated, it is permanently stored in the model inference file and remains unchanged throughout the online inference phase, without being dynamically updated based on real-time monitoring data. Simultaneously, distribution representativeness verification is required by calculating the Euclidean distance between the latent vectors of all healthy training samples and the center point. This distance distribution should conform to a normal distribution, and the sample distances should fall within the range of... Within the interval, The mean distance The distance is the standard deviation. If the verification fails, the encoder training effect and the health status and validity of the training data need to be rechecked until the verification is successful.

[0061] In the online inference phase, multiple random forward propagations are performed on the same input using the Monte Carlo dropout method. The randomness originates from the Dropout layer added during training, and its physical meaning is the cognitive uncertainty at the model structure level. During online inference, the Dropout layer remains enabled, and the random dropout mechanism is not disabled. The dropout rate is the same as in training to ensure network distribution matching between training and inference, avoiding inaccurate uncertainty estimation caused by distribution shift. During each forward propagation, the Dropout layer independently generates a random mask conforming to a Bernoulli distribution, randomly zeroing the outputs of neurons in the fully connected layers. The random masks generated in each propagation are independent and uncorrelated, ensuring that the outputs of multiple propagations meet the requirement of independent and identically distributed distribution. For the input feature vector at the same time, a corresponding number of independent random forward propagations are performed to obtain multiple temperature reconstruction values. The arithmetic mean of all outputs is used as the winding temperature reconstruction value, and the unbiased sample variance of all outputs is used as the prediction uncertainty.

[0062] The number of random forward propagation iterations is determined based on the statistical convergence of Monte Carlo uncertainty estimation, while also meeting the mandatory real-time requirements of industrial online monitoring. The specific setting logic is as follows: During the offline model training and validation phase, the number of random forward propagations must meet the convergence requirements of Monte Carlo uncertainty estimation, and the value should not be lower than the industry-recognized critical number of convergences for uncertainty estimation. This ensures that the output prediction uncertainty results converge, providing unbiased benchmark data for the subsequent training of the uncertainty perception module.

[0063] In online conventional inference scenarios, for steady-state operation of motors, the number of random forward propagation iterations must simultaneously consider the accuracy and stability of uncertainty estimation and the real-time performance of online inference. While ensuring that the relative deviation of uncertainty estimation is within an acceptable range for engineering purposes, the time consumed in a single inference iteration should be compressed to the maximum extent possible to meet the millisecond-level inference time budget requirements under the standard sampling frequency in industrial settings.

[0064] In high-risk online scenarios, including the switching phases of motor start-up, load change, and shutdown, as well as critical states for fault warning, the number of random forward propagation attempts must prioritize ensuring the accuracy and conservatism of uncertainty estimation. The value must be increased to meet the high confidence convergence requirements of uncertainty estimation to avoid misjudgment and missed judgment of faults due to estimation bias.

[0065] For scenarios where edge embedded devices have limited computing power and a mandatory upper limit on inference budget, a matching lightweight optimization scheme is adopted to reduce inference time without sacrificing core diagnostic accuracy.

[0066] Before online inference, a sufficient number of fixed random mask groups matching the dimensions of the Dropout layer are pre-generated. During inference, the pre-generated masks are called cyclically to avoid the additional time overhead of generating random numbers in real time during each forward propagation. Based on the online segmentation results and fault warning status, the number of random forward propagations is adaptively and dynamically adjusted: for steady-state operating conditions, a baseline number that balances accuracy and real-time performance is adopted. For non-steady-state operating conditions such as startup, load change, and shutdown, the number of propagations is automatically increased to adapt to the increased uncertainty caused by operating condition fluctuations. For operational scenarios that enter the critical state of fault warning, the system automatically switches to the number of propagation attempts that meet the high confidence convergence requirements, thereby achieving a dynamic optimal balance between estimation accuracy and inference real-time performance.

[0067] At the same time, a maximum inference time threshold matching the industrial field control cycle is set. When the time for a single inference exceeds this threshold, the number of random forward propagations is automatically reduced within the range above the statistically valid critical value to ensure that the inference time always meets the mandatory real-time requirements of industrial online monitoring. The statistically valid critical value is the minimum sample size required for the Monte Carlo uncertainty estimate to be statistically significant. The number of propagations must not be less than this critical value. At the same time, a conservative correction term is added to the output prediction uncertainty to ensure that the uncertainty estimate is not underestimated and to avoid the failure risk being masked.

[0068] The latent space variables corresponding to the current input are obtained through forward propagation of the encoder, and the real-time thermal balance residuals are calculated by substituting them into the thermal circuit equation. The Euclidean distance between the current latent space variables and the center point of the pre-stored health state latent space is calculated through the post-processing step to obtain the latent space state offset, forming a complete diagnostic information group.

[0069] Step 2: Real-time acquisition of multi-physical quantity data streams, and segmentation and preprocessing of time-series data based on sliding window statistics of current effective values; Step two includes the following: Real-time acquisition of multi-physical data streams of motor operation from the same source as in step one. The acquired physical quantities include winding temperature, effective value of three-phase current, speed and ambient temperature. The hardware link, anti-aliasing filter parameters and analog-to-digital conversion sampling rate of the data acquisition are completely consistent with those in step one. The exemplary sampling frequency is not less than 1Hz to ensure the homogeneity and comparability of training data and online monitoring data.

[0070] Using the effective value of current as the core discriminative feature, combined with sliding window statistics, the real-time acquired time-series data stream is automatically segmented into operating condition stages. First, a sliding analysis window is set; an example window width is... With one sampling point and a sliding step size of one sampling point, the mean, variance, and first-order rate of change of the current effective value sequence are calculated simultaneously within each sliding window, serving as the core discrimination index for operating condition segmentation.

[0071] Based on the above statistics, quantitative discrimination criteria for each working condition stage are defined; The starting phase is defined as follows: the effective current value rises from a near-zero no-load value to more than 90% of the steady-state value, and the winding temperature change rate exceeds a preset threshold, with an example temperature change rate threshold of 0.5℃ / min; the steady-state phase is defined as: the fluctuation range of the effective current value within the window is less than ±5% of the motor's rated current, and the duration of this steady state exceeds 30 seconds; the variable load phase is defined as: the absolute value of the current change rate within the window exceeds a preset threshold, with an example threshold of ±20% of the rated current per second; and the shutdown phase is defined as: the effective current value continuously drops to below 5% of the rated current, and the duration exceeds 10 seconds.

[0072] Correction processing is performed on boundary points and short segments of the working condition segmentation. For working condition segments with a continuous length less than the preset minimum length (exemplary minimum length is 10 sampling points), they are merged into adjacent working condition segments with the same attribute. For boundary points of adjacent working condition segments, a transition interval is set for smoothing.

[0073] The number of sampling points in the transition interval follows three principles: smoothness effectiveness, integrity of operating condition characteristics, and adaptability of sampling frequency. There is no mandatory fixed value; the specific value is dynamically determined according to clear technical rules.

[0074] The minimum number of sampling points in the transition interval must meet the effectiveness requirements of the linear weighted smoothing algorithm and must not be less than the minimum sample size required for linear smoothing calculation. The maximum number of sampling points in the transition interval must not exceed 1 / 10 of the effective length of the shortest operating segment in the adjacent operating segments. The reason for choosing 1 / 10 as the upper limit is that in the segmented smoothing processing of industrial time-series signals, when the length of the transition interval does not exceed 1 / 10 of the effective length of the target segment, the smoothing operation will not change the essential operating characteristics of the entire operating segment and will not interfere with the subsequent calculation of the physical consistency tensor, statistical feature extraction, and fault classification.

[0075] The physical time length of the transition interval is matched to the dynamic response time of the motor operating condition switching. The number of sampling points is adjusted synchronously with the data sampling frequency. When the sampling frequency increases, the number of sampling points in the transition interval increases accordingly to ensure that the physical time length of the transition interval remains constant and to ensure consistent smoothing effects at different sampling frequencies. The transition interval is divided into two segments along the time axis. The first half belongs to the previous operating condition segment, and the second half belongs to the next operating condition segment, ensuring that a single sampling point belongs to only one operating condition segment and that no single sample spans two operating condition segments simultaneously.

[0076] The time-series characteristics of multiple physical quantities within the transition interval are processed using a linear weighted smoothing algorithm to eliminate feature jumps during operating condition switching. The specific calculation formula is as follows:

[0077] In the formula, These are the smoothed eigenvalues ​​at time t within the transition interval. These are the characteristic measured values ​​of the previous working condition at time t. This represents the characteristic measured value of the next working condition at time t. The weighting coefficient for the previous working condition section. The weighting coefficient for the next working condition section satisfies... The weighting coefficients change linearly with the time step of the transition interval, and the starting point of the transition interval... , End point of transition interval , The intermediate sampling points are weighted by linear interpolation.

[0078] All preprocessing and feature extraction operations within the transition interval use smoothed feature values ​​as input. Direct calculation using original measured values ​​across working conditions is prohibited. Preprocessing of segments after working condition segmentation is limited to non-transition intervals of the working condition segment to avoid data aliasing across working conditions.

[0079] The algorithm determines whether there exists an interval in the initial segmentation results of the operating condition stage where the number of operating condition segment switching times per unit time exceeds a preset threshold. If so, it uses a Bayesian online change point detection algorithm to adaptively re-segment this interval. This algorithm is based on a Bayesian recursive framework and estimates the posterior probability of changes in operating conditions in the time series online. Its core recursive formula is:

[0080] In the formula, for The running length from the previous point of change. As of The observation sequence at time, For the observation likelihood under a given runtime, a Gaussian distribution is used for modeling. For the hazard function, an example is a constant hazard rate. , for The posterior probability of the running length at each time point is used to determine the point of change in operating conditions when the peak of the posterior probability changes abruptly, thus completing the adaptive segmentation of the time series. If, after processing by the Bayesian online change point detection algorithm, a stable segment with a continuous length exceeding the preset minimum length still cannot be formed, then a fixed-length sliding window is directly used to perform temporal slicing on that interval. For example, the window length L = 128 sampling points, the sliding step size is 64 sampling points, and adjacent windows maintain a 50% overlap rate to ensure the continuity of temporal features. For segments with a window length less than L, linear interpolation is used to complete the segment to a fixed length L. For segments with a window length exceeding L, L sampling points in the middle of the window are truncated to ensure the consistency of the input dimension of the subsequent model.

[0081] Within each segment or sliding window of the completed work condition, data cleaning, filtering, and time alignment are performed according to the preprocessing rules in step one, providing standardized input for subsequent feature extraction and model inference.

[0082] Step 3: Input the preprocessed real-time data into the physical information neural network model to obtain temperature reconstruction error, thermal balance residual, latent space state shift and prediction uncertainty, and splice them to form a multi-source physical consistency tensor. Step three includes the following: For each segment of the operating condition or the time-series data within the sliding window that has been segmented, after data cleaning, filtering, and time alignment are completed according to the preprocessing rules in step one, the input feature vector that has been preprocessed in real time is then... Input the physical information neural network model trained in step one. The model synchronously outputs four types of diagnostic information at each time step, as follows: The first type is temperature reconstruction error, and its expression is: In the formula, in the formula, for The temperature reconstruction error at any given time reflects the deviation between the actual temperature and the predicted value of the health model; for The measured value of the winding temperature is collected in real time. Output of the physical information neural network model The reconstructed winding temperature value at any given time; The second category is thermal balance residuals. That is, the residual value after real-time data is substituted into the first-order equivalent thermal circuit equation of the motor embedded in the model, which reflects the degree of deviation of the energy conservation of the current system. The third type is latent space state shift, and its expression is:

[0083] In the formula, in the formula, for The latent space state shift at any given time reflects the overall degree of shift of the system in the physical property characteristic space; The latent variable is the current input mapped by the encoder; The center point of the hidden space for the healthy state is a fixed baseline value that is not updated during the online inference process; The fourth category is forecast uncertainty. The variance of the output results of T random forward propagations is taken to reflect the adequacy of the model's understanding of the current working conditions.

[0084] Using the winding temperature sampling time sequence as the time-aligned master reference grid, a linear interpolation algorithm is used to map the above four types of outputs to the corresponding timestamps of the master reference grid, and the resulting data is spliced ​​to form a multi-source physical consistency tensor. Where C is the number of channels, including at least temperature error channels, physical residual channels, state offset channels and uncertainty channels, and C=4 is taken as an example. L is the fixed sampling length of the operating condition segment or sliding window, and L=128 is taken as an example.

[0085] For each channel of the multi-source physical consistency tensor, statistical enhancement features are extracted within the corresponding window. These statistical enhancement features include the local mean, variance, maximum value, minimum value, first-order linear fitting trend slope, kurtosis, and skewness of the sequence within the window.

[0086] The above indicators respectively characterize the central tendency, dispersion, extreme value boundary, trend of change and distribution pattern of time series, and can fully cover the full-dimensional statistical characteristics of early motor faults and anomalies.

[0087] The statistical feature extraction window is divided into two categories: global statistical window and local sliding statistical window. The window division and applicable scenarios are bound to the working condition segmentation results in step two. The core constraint is to prohibit cross-working condition and cross-slice window calculations to avoid feature distortion caused by data aliasing in different operating states.

[0088] The global statistical window is applicable to stable operating condition segments whose duration meets the statistical feature convergence requirements after operating condition segmentation. The so-called statistical feature convergence requirement means that the length of the operating condition segment is sufficient to allow the time-series statistical features within the window to stably represent the motor operating characteristics under the corresponding operating condition, and will not cause irregular random fluctuations in the statistical results due to insufficient sample size. The range of the global statistical window is limited to the complete effective time range of the current operating condition segment and must not cross the transition interval and boundary points of adjacent operating conditions.

[0089] Local sliding statistical windows are suitable for short working condition segments whose length does not meet the convergence requirements of global statistical windows, transition intervals between adjacent working conditions, and fixed-length sliding slice windows for intervals with frequent working condition switching. The window length must meet the physical continuity of time series data within a single window. The sliding step size is consistent with the data sampling step size. The calculation range of each window is limited to the current working condition segment or slice window. At the working condition boundary points and the boundary of the transition interval, the window must be automatically truncated to the corresponding boundary and must not introduce the running data of adjacent working conditions or adjacent slices. For the frequently changing operating conditions, a fixed-length sliding slice window is used. The statistical feature extraction window is consistent with the range of the slice window. The statistical feature calculation is only completed within the time range of the current slice and must not cross the boundary of adjacent slices. The purpose of the fixed-length slice design is to provide a standardized input with uniform dimensions for scenarios where the operating conditions frequently change and it is impossible to form stable long segments, so as to ensure the consistency of the dimensions of subsequent feature splicing with the network input.

[0090] The number of sampling points in the statistical window does not change with the data sampling frequency, while the physical time length corresponding to the window changes synchronously and proportionally with the sampling frequency. This ensures that the time scale and physical meaning of statistical feature extraction are consistent under different sampling frequencies, guaranteeing feature comparability under different hardware acquisition configurations. Statistical features at operating condition boundary points and transition intervals are calculated using a local sliding statistical window.

[0091] The extracted statistical enhancement features are concatenated with the original time series along the channel dimension, i.e. the feature dimension. The concatenation process follows the core principle of not changing the length of the time dimension of the original time series.

[0092] The original multi-source physical consistency tensor has the dimension of the number of channels × the time series length, where the number of channels corresponds to the number of diagnostic information categories output by the physical information neural network, and the time series length is the time series length of the current working condition segment or slice window.

[0093] For each channel of the original tensor, a corresponding number of statistical enhancement features are extracted. For each extracted statistical feature, it is expanded into a time series of the same length as the original time series. The values ​​of all time steps in the sequence are the calculation results of the statistical feature, forming an independent statistical feature channel with the same dimension as the original channel. This processing method can ensure that the statistical features are aligned with the original time series in the time dimension, without destroying the causal structure of the time series, and adapting to the requirements of subsequent causal convolutional networks for the input time series structure.

[0094] Along the channel dimension, the original temporal channels are concatenated with all statistical feature channels. During the concatenation process, the temporal length of all channels remains consistent. Any form of truncation, zero-padding, or expansion of the temporal dimension is strictly prohibited to ensure that the causality, integrity, and temporal resolution of the temporal sequence remain unchanged. The enhanced feature matrix formed after concatenation is used as the standardized input for subsequent temporal causal convolutional networks.

[0095] The entire process of statistically enhanced feature extraction and dimension concatenation follows a fixed normalization process. The order of each step is based on clear technical principles and cannot be arbitrarily changed. The core objective is to eliminate the dimensional differences between different physical quantities and features, ensuring that the feature distribution is consistent between the training and inference phases, and avoiding a decrease in model generalization ability and diagnostic inaccuracies caused by distribution shifts.

[0096] The first step is to perform channel-level pre-normalization of the original time series. For each original channel of the multi-source physical consistency tensor, within the current statistical window, min-max normalization is performed using the global extrema of the corresponding channel pre-stored in the healthy training dataset. This maps the values ​​to a unified standardized interval. The reason for prioritizing this step is that different channels of the original time series data correspond to different physical quantities, with significant differences in dimensions and numerical ranges. Performing channel-level pre-normalization first can eliminate the interference of dimensional differences on the subsequent statistical feature calculation results. At the same time, the extreme value benchmark used for normalization is fixed after the model training is completed and should not be dynamically updated with real-time monitoring data during the online inference stage to avoid the feature distribution shift during the inference stage.

[0097] The second step is to extract statistical enhancement features. For each channel time series after pre-normalization, extract each statistical enhancement feature within the corresponding window. Extracting statistical features after pre-normalization ensures that the calculation of statistical features is performed on dimensionless standardized values. This prevents the statistical features of a certain channel from having an unreasonable dominant position in subsequent feature fusion due to an excessively large range of values, thus ensuring a balanced weight of features in each channel.

[0098] The third step is to perform independent normalization of the statistical features. For each extracted statistical feature, the global extremum of the corresponding statistical feature is pre-stored in the healthy training dataset, and independent min-max normalization is performed to map it to a standardized interval consistent with the original time series. The design of this step is based on the fact that the numerical distribution ranges of different statistical features are naturally different. Independent normalization can ensure that the numerical ranges of all statistical features are aligned with the original time series, providing a unified numerical basis for subsequent channel dimension splicing and avoiding the problem of unbalanced feature numerical distribution after splicing.

[0099] The fourth step completes the dimensional concatenation and final input normalization. The normalized statistical feature channels and the original temporal channels are concatenated along the channel dimension to form an enhanced feature matrix. When inputting into the input layer of the temporal causal convolutional network, a uniform second-order min-max normalization is performed on all channels of the enhanced feature matrix, which is finally mapped to the standardized interval, thus completing the input preprocessing.

[0100] By using the standardized feature extraction, splicing, and normalization process described above, the temporal characteristics and statistical anomalies of motor faults can be preserved, thereby improving the ability of subsequent classification models to capture early fault features and their diagnostic robustness under varying operating conditions.

[0101] Step 4: Using the enhanced features of the multi-source physical consistency tensor as input, construct an uncertainty-aware gating temporal causal convolutional network; its channel attention module uses the predicted uncertainty as an external guiding signal to dynamically adjust the feature weights of each channel and output a fault occurrence probability vector. Step four includes the following: Using the enhanced features of the multi-source physical consistency tensor generated in step three as input, an uncertainty-aware gating temporal causal convolutional network model is constructed to achieve end-to-end fault classification of motor operating status and output the probability vector of each type of fault.

[0102] The model adopts an end-to-end architecture that conforms to temporal causality, consisting of an input layer, multiple sets of causal convolutional modules, an uncertainty-aware channel attention module, a multi-scale pyramid pooling layer, a global average pooling layer, a fully connected layer, and a softmax output layer connected in series. The input layer receives enhanced features from a multi-source physically consistent tensor. For example, the input dimension is 4×128. The input layer first performs min-max normalization on the sequence of each channel, mapping the values ​​to the interval [−1,1], which improves the stability and convergence speed of model training.

[0103] The causal convolution module employs three stacked causal convolutional blocks to ensure that the model's temporal receptive field covers the entire length of the input sequence, while adhering to temporal causality constraints to prevent future information leakage and meet the real-time requirements of online monitoring. Each convolutional block contains two layers of one-dimensional causal convolutions, with the kernel size exemplarily set to... The inflation coefficients are set to 1, 2, and 4 respectively. The ReLU linear rectified activation function is used, and a Dropout layer is added after each convolutional block. The Dropout rate is set to 0.5 for example to effectively suppress overfitting during model training. The output feature map calculation formula for causal convolution is:

[0104] In the formula, The convolution output feature value at time t; The kernel size; The i-th weight coefficient of the convolution kernel; For the input feature map in The time values ​​are obtained by using only the current time and historical time inputs to ensure causality; This is the bias term for the convolutional layer.

[0105] After the causal convolution module, an uncertainty-aware channel attention module is connected. This module consists of two parts connected in series: a basic channel attention subunit and an uncertainty-aware gating subunit. The entire calculation is completed within the module without any external input. The prediction uncertainty output by the physical information neural network model is used as an external guiding signal to adjust the weights of the channel features.

[0106] The basic channel attention subunit uses a compression-activation structure to generate the original channel attention weights, and the input feature map is the feature map output by the causal convolution module. Where C=4, corresponding to the four channels of temperature reconstruction error, thermal equilibrium residual, latent space state shift, and prediction uncertainty, and L is the temporal length. Global average pooling is performed on the feature map along the temporal dimension to obtain the channel-level global feature vector. The expression is:

[0107] In the formula, For the first Global compression features of each channel, For the first Each channel is in The feature values ​​at each time step are mapped to the global feature vector through two fully connected layers to learn the dependencies between channels and output the original channel attention weight vector. The expression is:

[0108] In the formula , For the learnable weight matrix of the fully connected layer, , For learnable bias terms, a channel compression ratio of 4 is recommended. The Sigmoid activation function outputs... For the first The original attention weights for each channel.

[0109] The uncertainty-aware gating subunit uses the prediction uncertainty as the sole external guiding signal. It generates correction coefficients for each channel through a channel-specific constrained Sigmoid gating function. First, it obtains the mean value of the prediction uncertainty channel in the multi-source physical consistency tensor within the current window. and compared with the preset uncertainty benchmark value Compare and calculate the uncertainty ratio. Uncertainty benchmark The mean of the prediction uncertainty of the physical information neural network model is statistically output on the training data of health status, and is used to characterize the cognitive confidence benchmark of the model for known health conditions.

[0110] With uncertainty ratio As input to the learnable gating function, correction coefficients for the original attention weights of each channel are generated. For the four channels, two branches are defined, each with a learnable sigmoid gating function to generate the corresponding channel's correction coefficients. The expression for the correction coefficient of the temperature reconstruction error channel, i.e., c=1, is:

[0111] The correction coefficient expressions for the thermal equilibrium residual channel (c=2) and the latent space state offset channel (c=3) are as follows:

[0112] The correction factor for the uncertainty prediction channel, i.e., c=4, is fixed at 4. In the formula , , The learnable slope parameter of the gating function is initially set to 1.0. Together with other parameters of the temporal causal convolutional network, it is jointly trained end-to-end using health data and fault injection test data to automatically learn the optimal weighting strategy for each channel under different uncertainty levels.

[0113] During training, the learnable slope parameter , , Apply nonnegative boundary constraints to limit the range of values. Negative values ​​are prohibited. The monotonicity direction of the Sigmoid gating function is fixed to ensure that the monotonicity trend of the correction coefficient conforms to the physical logic of the design. When all channel correction coefficients approach 1, the original channel attention weights remain unchanged. At that time, the correction coefficient of the temperature reconstruction error channel varies with The correction coefficients for the thermal equilibrium residual channel and the hidden space state offset channel increase and decrease monotonically. The configuration increases monotonically and is based on the following physical principle: increased prediction uncertainty indicates that the current operating condition deviates from the known operating condition range covered by the training data, and the temperature reconstruction error based on data statistics is prone to generalization instability; while the thermal balance residual is constructed based on the law of conservation of energy in motor thermodynamics, and the latent space state offset is calculated based on the baseline distribution of the manifold of the physical properties of the healthy state. Both are far less affected by the fluctuation of unknown operating conditions than the temperature reconstruction error, and have higher diagnostic reliability in scenarios where the model's understanding is insufficient.

[0114] The original attention weights of each channel are multiplied element-wise with the corresponding correction coefficients to obtain the adjusted final attention weights. The final attention weights are then multiplied channel-wise with the original input feature map to complete the feature weighting. The weighted feature map is then output to the subsequent multi-scale pyramid pooling layer.

[0115] Multi-scale temporal features are extracted and fused through a multi-scale pyramid pooling layer, and then reduced in dimensionality through a global average pooling layer. Finally, a fully connected layer and a Softmax activation function are used to output the fault probability vector corresponding to various operating states of the motor.

[0116] The multi-scale pyramid pooling layer is set with pooling scales of 1, 2, 4 and 8. It performs average pooling on the input feature map at different scales to extract multi-scale temporal features. Then, through upsampling and feature concatenation operations, it restores the temporal dimension of the original feature map, realizes multi-scale feature fusion, and improves the robustness of the model to changing working conditions.

[0117] The output features of the pyramid pooling layer are fed into a global average pooling layer. Global average pooling is performed along the temporal dimension on the feature map, converting the two-dimensional feature map into a one-dimensional feature vector, significantly reducing the number of model parameters and further suppressing the risk of overfitting. The output of the global average pooling layer is fed into two fully connected layers. The first fully connected layer has 64 neurons (exemplarily) and uses the ReLU activation function. The second fully connected layer has the same number of neurons as the preset number of fault categories (exemplarily 5). Finally, the probability vectors for each fault category are output through the Softmax activation function. ,in Let be the probability of occurrence of the k-th type of fault, satisfying .

[0118] The preset fault categories and corresponding label definitions are as follows: Category 0 represents the healthy state, where all physical consistency indicators are within the random fluctuation range of the healthy state training data statistics, with no systematic bias or continuous unidirectional trend, corresponding to label 0; Category 1 is heat dissipation anomaly, the thermal balance residual is consistently positive and shows an accumulative growth trend, the hidden space state offset increases slowly in sync, and the temperature reconstruction error gradually increases with the heat accumulation process, corresponding to label 1; Category 2 is an overload fault, where the temperature reconstruction error and thermal balance residual jump synchronously with the sudden change in load, and the hidden space state shift shows a violent instantaneous shift, corresponding to label 2; Category 3 is a local overheating fault in the bearing. The hidden space state deviation exhibits periodic fluctuation characteristics synchronized with the rotational speed. The temperature reconstruction error shows a small but continuous abnormality, and the thermal balance residual does not show a significant systematic deviation, corresponding to label 3. Category 4 is sensor drift fault, where the temperature reconstruction error exhibits a systematic fixed bias unrelated to changes in operating conditions, the thermal balance residual shows no significant systematic deviation, and the latent space state shift remains within the healthy fluctuation range. The formation mechanism of this feature is as follows: the physical information neural network, through constraints on non-temperature input features such as current and rotational speed and the thermal circuit equation, has learned the temperature response law under healthy conditions. The fixed drift of the temperature sensor will only cause a fixed deviation between the measured temperature and the reconstructed temperature, without disrupting the consistency of the thermal balance equation or causing a shift in the latent space manifold, corresponding to label 4.

[0119] The model uses cross-entropy loss as the total loss function, and its mathematical expression is as follows:

[0120] In the formula, The total number of samples in the training batch; This is the one-hot encoded value of the true label of the i-th sample, which is 1 when the sample belongs to the k-th class and 0 otherwise. This represents the predicted probability that the i-th sample in the model output belongs to the k-th class. For imbalanced scenarios, a class weight coefficient is introduced to assign higher weights to fault classes with fewer samples, thereby improving the model's ability to identify faults with small sample sizes.

[0121] The model training dataset consists of two parts: One part consists of historical health status data collected, which is processed to generate multi-source physically consistent tensor enhancement features. The labels are uniformly set to 0, and the number of samples accounts for no less than 70% of the total training samples. The other part consists of full-cycle fault condition data collected from fault injection tests. By artificially blocking the cooling fan, applying excessive load, simulating bearing friction, and setting a fixed bias for the temperature sensor, multi-physical quantity data of various faults from their early budding to their clear occurrence are obtained. For the data in the fault injection process, the execution time of the fault injection action is used as the boundary, and the physical process of fault development is combined to label the data. After processing, multi-source physical consistency tensor enhancement features are generated, and corresponding labels 1 to 4 are labeled according to the fault type.

[0122] The training hyperparameters are set as follows: the optimizer is Adam optimizer, the initial learning rate is set to 1×10−3 for example, the weight decay coefficient is set to 1×10−4, the batch size is 32, the maximum number of training rounds is set to 100, and the validation set early stopping strategy is adopted. When the validation set loss does not decrease for 10 consecutive rounds, the training is terminated early, and the model weights with the lowest validation set loss are saved as the final inference model.

[0123] Step 5: Calculate the comprehensive fault confidence based on the fault occurrence probability vector, classify the fault warning level, and output the diagnostic results; Step five includes the following: Based on the various fault probability vectors output by the model, the comprehensive fault confidence is calculated, the fault warning level is classified, and interpretable fault location results are output, providing a quantitative and implementable basis for motor operation and maintenance decisions.

[0124] To suppress false alarms caused by random fluctuations in single-class probabilities, a comprehensive fault confidence calculation model that considers both fault class probabilities and health probabilities is constructed. Its mathematical expression is as follows:

[0125] In the formula, The overall fault confidence level is set to a value in the range of [0,1]. The maximum probability value among all non-healthy fault categories (categories 1-4); This represents the probability of the healthy category output by the model. This calculation ensures that the overall fault confidence is high only when the maximum probability of the unhealthy category is high and the probability of the healthy category is low, effectively reducing the false alarm rate and improving the engineering reliability of the diagnostic results.

[0126] Based on the overall fault confidence The value of is used to divide the motor's operating status into three levels, and corresponding differentiated operation and maintenance strategies are set. The specific division rules are as follows: when When the motor is in a healthy operating condition, it is considered to be in a normal state and requires no maintenance intervention. It is sufficient to continuously output real-time monitoring data and physical consistency index curves. when When the motor has an early warning status, it is determined that there is an early abnormality in the motor, triggering an early warning prompt. It is recommended to increase the monitoring frequency, for example, increase the data sampling frequency to twice the rated sampling rate, shorten the equipment inspection cycle, and focus on the changing trend of abnormal physical quantities. when When the fault is identified as a definite fault, indicating a definitive fault in the motor, a fault alarm is triggered. It is recommended to immediately develop a planned maintenance program and, if necessary, perform a shutdown inspection to prevent the fault from escalating and causing irreversible damage to the motor.

[0127] While outputting the fault confidence level and grade, the system also simultaneously outputs interpretable fault location results, including the most likely fault category and a quantitative description of the associated physical consistency anomaly. An example output format is as follows: "[Fault Category], Confidence Level XX.XX%, Anomaly Characteristics: [Quantitative Anomaly Description of Related Physical Consistency Indicators]", for example, "Heat dissipation anomaly, confidence level 82.5%, anomaly characteristics: thermal balance residual continuously +2.8σ and exponentially accumulating, latent space offset trend +0.15σ / min, temperature residual +1.2σ, ​​current and speed characteristics are not abnormal."

[0128] Meanwhile, it continuously stores and updates full monitoring data and multi-source physical consistency tensor sequences before and after the failure, providing complete data support for subsequent failure tracing, root cause analysis and operation and maintenance optimization.

[0129] Reference Figure 2 As shown, a real-time data analysis and prediction system for online monitoring of motor temperature includes: The physical information model training module is used to collect historical data of multiple physical quantities of the motor's healthy operating status, construct and train a physical information neural network model with constraints of the motor's equivalent thermal circuit equation, and output a set of diagnostic information including winding temperature reconstruction values, thermal balance residuals, hidden space state shifts and prediction uncertainties. The online operating condition segmentation module is used to collect multi-physical quantity data streams in real time and complete automatic segmentation and preprocessing of operating condition stages based on the sliding window statistics of the effective current value. The multi-source consistency generation module is connected to the physical information model training module and the working condition online segmentation module respectively. It is used to input the preprocessed data into the physical information neural network model, obtain the temperature reconstruction error, thermal balance residual, latent space state shift and prediction uncertainty, and splice them to form a multi-source physical consistency tensor. The uncertainty-gated classification module is connected to the multi-source consistency generation module. It is used to take the enhanced features of the multi-source physical consistency tensor as input, introduce the prediction uncertainty as a guiding signal through the channel attention module, dynamically adjust the feature weights of each channel, and output the failure probability vector through the temporal causal convolutional network. The diagnostic result output module is connected to the uncertainty gating classification module and is used to calculate the comprehensive fault confidence, classify the fault warning level and output the diagnostic result.

[0130] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for real-time data analysis and prediction of motor temperature through online monitoring, characterized in that, include: S1. Collect historical data of multiple physical quantities under the healthy operating state of the motor, construct and train a physical information neural network model with constraints of the motor's equivalent thermal circuit equation. The loss function adds a thermal balance residual constraint term of lumped heat capacity and thermal resistance parameters to the temperature reconstruction error term. After training, the model outputs a set of diagnostic information online, including winding temperature reconstruction value, thermal balance residual, latent space state offset, and prediction uncertainty. The latent space state offset is the distance between the current input mapping point in the model's latent space and the center point of the healthy state latent space. The center point is determined by the mean of the latent space samples of the healthy training data. S2. Real-time acquisition of multi-physical quantity data streams, and segmentation and preprocessing of time-series data based on sliding window statistics of current effective value; S3. Input the preprocessed real-time data into the physical information neural network model to obtain temperature reconstruction error, thermal balance residual, latent space state shift and prediction uncertainty, and splice them to form a multi-source physical consistency tensor. S4. Using the enhanced features of the multi-source physical consistency tensor as input, construct an uncertainty-aware gating temporal causal convolutional network; its channel attention module uses the predicted uncertainty as an external guiding signal to dynamically adjust the feature weights of each channel and output a fault occurrence probability vector. S5. Calculate the comprehensive fault confidence based on the fault occurrence probability vector, classify the fault warning level, and output the diagnostic results.

2. The method for real-time data analysis and prediction of motor temperature online monitoring according to claim 1, characterized in that, S1 includes: The healthy operating state of the motor is defined, and historical data of multiple physical quantities under typical operating conditions of the motor's healthy state are collected. The multiple physical quantities include winding temperature, effective value of current, speed, and ambient temperature. Based on the rated parameters of the motor and the specifications of the operating environment, the effective value range of each physical quantity is preset, invalid data outside the range is eliminated, pulse interference noise in the winding temperature sequence is identified and eliminated, and missing data is filled by linear interpolation or discarded. A first-order exponential smoothing filter is applied to the winding temperature sequence, and a sliding median filter and a first-order hysteresis filter are applied to the current RMS value sequence in sequence. The time alignment of the time series of multiple physical quantities is completed with the winding temperature sampling time sequence as the reference. The input feature vector is constructed using the effective value of current, the rate of change of current, the rotational speed, the ambient temperature, and the temperature hysteresis term. The winding temperature at the corresponding time is used as the output target to construct the model training dataset.

3. The method for real-time data analysis and prediction of motor temperature online monitoring according to claim 2, characterized in that, S1 further includes: A physical information neural network model containing an encoder and a decoder is constructed. The encoder maps the input feature vector to a low-dimensional physical property latent space, and the decoder reconstructs the winding temperature from the latent space. The model's total loss function is constructed, which is composed of three weighted components: temperature reconstruction error term, thermal balance residual constraint term, and latent space regularization term. The thermal balance residual constraint term is established based on the first-order equivalent thermal path differential equation of the motor, and the lumped heat capacity and lumped thermal resistance parameters of the motor are used as trainable parameters of the network. Model training is completed by minimizing the total loss function through an optimization algorithm. During training, a Dropout layer is added after the fully connected layer. In the online inference stage, multiple random forward propagations are performed on the same input. The output mean is used as the winding temperature reconstruction value, and the output variance is used as the prediction uncertainty. The output thermal balance residual and latent space state offset are calculated simultaneously. The latent space state offset is calculated based on the distance between the mean of the latent variable in the current propagation and the center point of the latent space in the healthy state.

4. The method for real-time data analysis and prediction of motor temperature online monitoring according to claim 1, characterized in that, S2 includes: Set the sliding analysis window and sliding step size, and calculate the mean, variance and first rate of change of the current effective value sequence within each sliding window; Define quantitative discrimination criteria for the start-up phase, steady-state phase, variable load phase, and shutdown phase, and complete the initial segmentation of the operating condition phase of the time-series data stream based on the discrimination criteria; Work condition segments with a duration shorter than the preset minimum length are merged, and transition intervals are set for the boundary points of adjacent work condition segments to correct the work condition segmentation results. Within each working condition segment, data cleaning, filtering, and time alignment are completed according to preprocessing rules.

5. The method for real-time data analysis and prediction of motor temperature online monitoring according to claim 4, characterized in that, Also includes: Determine whether there is a frequent switching interval in the initial segmentation result of the working condition stage where the number of switching times of the working condition segment exceeds a preset threshold within a unit time. If the frequently switching interval exists, the Bayesian online change point detection algorithm is used to estimate the posterior probability of the working condition change online based on the Bayesian recursive framework, determine the working condition change point, and re-adaptively re-segment the interval. If, after processing by the Bayesian online change point detection algorithm, the interval still cannot form a stable segment with a continuous length exceeding the preset minimum length, then a fixed-length sliding window is used to perform temporal slicing on the interval. Adjacent windows maintain a preset overlap rate, and window segments with inconsistent lengths are completed or truncated to ensure the consistency of the input dimension.

6. The method for real-time data analysis and prediction of motor temperature online monitoring according to claim 1, characterized in that, S3 includes: For the preprocessed real-time data, the physical information neural network model synchronously outputs four types of diagnostic information at each moment, including: temperature reconstruction error, thermal balance residual, latent space state shift, and prediction uncertainty. The temperature reconstruction error is the difference between the measured winding temperature and the model reconstruction value; the thermal balance residual is the residual value after substituting real-time data into the thermal circuit equation; the latent space state offset is the Euclidean distance between the latent variable after the current input mapping and the center point of the latent space in the healthy state; and the prediction uncertainty is the variance of the output results of multiple random forward propagation. Using the winding temperature sampling time sequence as the time alignment master reference, the four types of diagnostic information are mapped to the corresponding timestamps of the master reference grid and spliced ​​to form a multi-source physical consistency tensor.

7. The method for real-time data analysis and prediction of motor temperature online monitoring according to claim 6, characterized in that, S3 further includes: For each channel of the multi-source physical consistency tensor, statistical enhancement features are extracted within the corresponding window. These statistical enhancement features include: local mean, variance, maximum value, minimum value, trend slope, kurtosis, and skewness. The extracted statistical enhancement features are concatenated with the original time series to obtain the enhancement feature matrix.

8. The method for real-time data analysis and prediction of motor temperature online monitoring according to claim 7, characterized in that, S4 includes: A temporal causal convolutional network model is constructed, which consists of an input layer, multiple sets of causal convolutional modules, an uncertainty-aware channel attention module, a multi-scale pyramid pooling layer, a global average pooling layer, a fully connected layer, and a softmax output layer connected in series. The input layer receives the enhanced feature matrix and normalizes the sequence of each channel; the causal convolution module uses stacked one-dimensional causal convolution blocks, follows the temporal causality constraint, and completes feature extraction using the input of the current time and historical time. The uncertainty perception channel attention module takes the predicted value of the uncertainty channel as input, calculates the correction coefficient of each channel through learnable parameters, and multiplies the original channel attention weights by the corresponding correction coefficients to obtain the adjusted final attention weights. Multi-scale temporal features are extracted and fused through a multi-scale pyramid pooling layer, and then reduced in dimensionality through a global average pooling layer. Finally, a fully connected layer and a Softmax activation function are used to output the fault probability vector corresponding to various operating states of the motor.

9. The method for real-time data analysis and prediction of motor temperature online monitoring according to claim 1, characterized in that, S5 includes: Based on the failure occurrence probability vector, the maximum probability value of all non-healthy failure categories is taken, and combined with the probability of the healthy category, the comprehensive failure confidence is calculated using the following formula. : ; In the formula, For the probability of the health category, This represents the maximum probability value for all non-healthy fault categories. Based on the preset interval threshold of the comprehensive fault confidence, the motor operating status is divided into three levels: normal state, early warning state, and clear fault state. The system synchronously outputs the fault category corresponding to the highest probability, as well as the physical consistency quantification anomaly description associated with the fault, and continuously stores the full monitoring data and multi-source physical consistency tensor sequence before and after the fault occurs.

10. A real-time data analysis and prediction system for online monitoring of motor temperature, characterized in that, include: The physical information model training module is used to collect historical data of multiple physical quantities of the motor's healthy operating status, construct and train a physical information neural network model with constraints of the motor's equivalent thermal circuit equation, and output a set of diagnostic information including winding temperature reconstruction values, thermal balance residuals, hidden space state shifts and prediction uncertainties. The online operating condition segmentation module is used to collect multi-physical quantity data streams in real time and complete automatic segmentation and preprocessing of operating condition stages based on the sliding window statistics of the effective current value. The multi-source consistency generation module is connected to the physical information model training module and the working condition online segmentation module respectively. It is used to input the preprocessed data into the physical information neural network model, obtain the temperature reconstruction error, thermal balance residual, latent space state shift and prediction uncertainty, and splice them to form a multi-source physical consistency tensor. The uncertainty-gated classification module is connected to the multi-source consistency generation module. It is used to take the enhanced features of the multi-source physical consistency tensor as input, introduce the prediction uncertainty as a guiding signal through the channel attention module, dynamically adjust the feature weights of each channel, and output the failure probability vector through the temporal causal convolutional network. The diagnostic result output module is connected to the uncertainty gating classification module and is used to calculate the comprehensive fault confidence, classify the fault warning level and output the diagnostic result.