Small sample motor residual service life prediction method based on pre-training transfer learning and adaptive calibration fusion
By employing a method that integrates pre-trained transfer learning with adaptive calibration in predicting the remaining service life of motors, a dual-path network is constructed in the source and target domains. Combining multi-channel gated attention and covariance matrix adaptive calibration, the problems of insufficient model generalization ability and poor cross-domain adaptability under small sample conditions are solved, achieving high-precision and reliable prediction results and uncertainty calibration.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TIANJIN UNIV
- Filing Date
- 2026-06-02
- Publication Date
- 2026-07-31
AI Technical Summary
Existing methods for predicting the remaining service life of motors suffer from problems such as insufficient model generalization ability, poor cross-domain adaptability, lack of uncertainty calibration and catastrophic forgetting in prediction results under small sample conditions, making it difficult to output reliable prediction results when the target equipment has only a small amount of initial operating data.
A method based on the fusion of pre-trained transfer learning and adaptive calibration is adopted. By constructing a source domain degradation knowledge path and a target domain adaptive path, and combining a multi-channel gated attention network and a covariance matrix adaptive calibration mechanism, dynamic fusion and online calibration of motor degradation features are achieved.
It significantly improves prediction accuracy and cross-domain generalization ability under small sample conditions, provides reliable uncertainty quantification and interpretable prediction process, effectively alleviates catastrophic forgetting, and improves the credibility of prediction results and application value in engineering fields.
Smart Images

Figure CN122491056A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of predictive maintenance of motors, industrial artificial intelligence, and remaining service life prediction technology. Specifically, it relates to a method for predicting the remaining service life of a motor by using a fusion mechanism of source domain degradation knowledge, transfer learning, meta-learning, and adaptive calibration under the condition of scarce full-life samples of the target motor. Background Technology
[0002] Remaining Useful Life (RUL) prediction is a key technology for enabling predictive maintenance of equipment and ensuring the safe and reliable operation of industrial systems. Accurate RUL prediction can optimize maintenance resource allocation, reduce total lifecycle costs, and mitigate risks such as unplanned downtime through early warning.
[0003] Existing RUL prediction methods are mainly divided into two categories: one is physical model-based methods, such as the Wiener process and the Gamma process. These methods have clear physical interpretation, but the actual operating conditions of motors are complex, making it difficult to construct accurate mathematical models. The lack of parameter identification and adaptability seriously restricts their engineering applications. The second is data-driven methods, represented by deep learning, such as Convolutional Neural Networks (CNN) and Long Short-Term Memory (LSTM). These methods can automatically capture degradation features in the data, but their effectiveness is highly dependent on high-quality full-lifecycle data. In actual industrial scenarios, acquiring full-lifecycle data of equipment from health to failure is costly and time-consuming, resulting in insufficient model generalization ability under small sample conditions [1,2].
[0004] Data-driven methods have been widely used in recent years. The basic idea is to learn health indicators or degradation characteristics from sensor data and then establish a mapping relationship between degradation characteristics and RUL. Traditional machine learning methods include support vector regression, random forest, correlation vector machine, Gaussian process regression and hidden Markov model, etc.; deep learning methods include CNN, LSTM, gated recurrent unit (GRU), temporal convolutional network, autoencoder and Transformer (Transformer, attention mechanism model), etc. The above methods can reduce manual feature design and capture nonlinear degradation relationships. For example, LSTM can utilize long-term dependencies in multi-sensor sequences [3], and CNN can extract local degradation features from sliding window data [4]. However, these methods usually require a large amount of target equipment data with lifespan labels, and when there is a distribution difference between the training samples and the field equipment, the model prediction performance will decrease significantly [5].
[0005] To alleviate the problem of insufficient full-life data for target equipment, existing studies have introduced transfer learning, domain adversarial training, and meta-learning. For example, in transfer learning, the maximum mean difference constraint can bring the feature distributions of the source domain and the target domain closer together; domain adversarial training reduces inter-domain differences by learning representations that are difficult to distinguish between the source domain and the target domain [6]; meta-learning methods enable the model to quickly adapt to new tasks on a small number of samples [7]. However, the above-mentioned existing methods still have three shortcomings in motor RUL prediction: First, the pre-training targets mostly borrow from general reconstruction or classification tasks and are not fully designed to address degradation trends, temporal consistency, and changes in health status; Second, the source domain knowledge and new features of the target domain after transfer are usually fused using fixed weights or single-path fine-tuning, which can easily lead to negative transfer or catastrophic forgetting; Third, the prediction results mostly give a single point value and lack a prediction interval that is calibrated with changes in online residuals, making it difficult to support on-site maintenance decisions.
[0006] Therefore, a method for predicting the RUL of a small sample of target motors is needed, which can complete data preprocessing, pre-training of general degradation characterization, rapid adaptation of the target domain, adaptive calibration and fusion of source domain knowledge and target domain features, and online uncertainty calibration in an executable process, so as to output RUL prediction results that can be used for maintenance decisions even when the target device has only a small amount of initial operating data.
[0007] [References]
[0008] [1] Y. Ren, J. Lyu, M. Jebali and B. Zhang, "An AC ContactorRemaining Useful Life Prediction Method based on Degradation Event Analysis," 2023 6th International Symposium on Autonomous Systems (ISAS), Nanjing,China, 2023, pp. 1-6, doi: 10.1109 / ISAS59543. 2023.10164499.
[0009] [2] L. Xia, Y. Chen, H. Wang, L. Liang, Q. Huang and B. Zhang, "Remaining Useful Life Prediction of Lithium-Ion Batteries Based on FusionNeural Network CCLAtt," 2025 International Conference on Digital Analysis andProcessing, Intelligent Computation (DAPIC), Incheon, Korea, Republic of,2025, pp. 15-21, doi: 10.1109 / DAPIC66097.2025.00009.
[0010] [3] Yi, F., Shu, X., Zhou, J., Zhang, J., Feng, C., Gong, H., Zhang,C., & Yu, W. (2025). Remaining useful life prediction of PEMFC based onmatrix long short-term memory. INTERN- ATIONAL JOURNAL OF HYDROGEN ENERGY,111, 228–237.https: / / doi.org / 10.1016 / j.ijhydene.2025.02.302.
[0011] [4] Zhang, J., Tian, J., Li, M., Leon, J., Franquelo, L., Luo, H., &Yin, S. (2023). A Parallel Hybrid Neural Network With Integration of Spatialand Temporal Features for Remaining Useful Life Prediction in Prognostics.IEEE TRANSACTIONS ON INSTRUMENTATION AND MEASUREMENT, 72. https: / / doi.org / 10.1109 / TIM.2022.3227956.
[0012] [5] Li, H., Zhang, Z., Li, T., & Si, X. (2024). A review on physics-informed data-driven remaining useful life prediction: Challenges andopportunities. MECHANICAL SYSTEMS AND SIGNAL PROCESSING, 209. https: / / doi.org / 10.1016 / j.ymssp.2024.111120.
[0013] [6] Liu Jun Bridge damage identification method based on pre trainingand domain confrontation. South China University of technology, 2024.doi:10.27151 / d.cnki.ghnlu.2024.001421.
[0014] [7] Pang, H., Chen, K., Geng, Y., Wu, L., Wang, F., & Liu, J. (2024).Accurate capacity and remaining useful life prediction of lithium-ionbatteries based on improved particle swarm optimization and particle filter.ENERGY, 293. https: / / doi.org / 10.1016 / j.energy.2024.130555. Summary of the Invention
[0015] The purpose of this invention is to provide a small-sample method for predicting the remaining service life of motors based on the fusion of pre-trained transfer learning and adaptive calibration. This method does not simply train a regression model, but executes the steps sequentially: multi-source data construction, degradation characterization pre-training, target domain small-sample adaptation, dynamic gating fusion, and online calibration output. This allows subsequent steps to directly utilize the parameters, features, or calibration results generated in the previous step. Specifically, the pre-trained encoder parameters, standardized parameters, and source domain health state reference distribution obtained in the source domain pre-training stage serve as the foundation for the source domain knowledge path; the target domain adaptation features obtained in the target domain small-sample adaptation stage are input into the target domain adaptive path; the target motor degradation features output by dynamic gating fusion serve as input to the online prediction framework; and the covariance matrix and prediction interval coverage generated in the online calibration stage further serve as the state basis for prediction interval adjustment. Thus, this invention achieves closed-loop optimization from source domain degradation knowledge transfer, rapid target domain adaptation, dual-path feature fusion to dynamic calibration of the prediction interval.
[0016] To address the aforementioned technical problems, this invention proposes a small-sample method for predicting the remaining service life of motors based on the fusion of pre-trained transfer learning and adaptive calibration, comprising the following steps:
[0017] S1. Collect the full lifecycle operation data of the source domain device and the initial operation data of a small sample of the target motor, and perform signal layer, feature layer and sequence layer preprocessing on the data; the items of the full lifecycle operation data of the source domain device include: timestamp, vibration, temperature, current and remaining service life value; the items of the initial operation data of the small sample of the target motor are the same as the items of the full lifecycle operation data of the source domain device.
[0018] S2, based on the preprocessed source domain data, train a pre-trained model of general degradation representation to obtain a pre-trained encoder for representing the degradation state or health state of the source domain device, and the pre-trained encoder constitutes the source domain knowledge path.
[0019] S3, Based on the pre-processed small sample initial running data of the target motor, the pre-trained encoder is subjected to soft alignment and meta-learning for rapid adaptation to obtain the target domain adapted encoder; the target domain adapted encoder constitutes the target domain adaptive path.
[0020] S4, the source domain knowledge path and the target domain adaptive path constitute a dual-path network. The preprocessed small sample initial running data of the target motor is used as the input of the dual-path network. The output of the dual-path network is fused through dynamic gating vector to obtain the degradation features of the target motor.
[0021] S5, input the degradation features of the target motor into the online prediction framework. The online prediction framework includes a multi-channel gated attention network and a covariance matrix adaptive calibration mechanism. The fused degradation features are used to perform modal weighting calculation through the multi-channel gated attention network to obtain the remaining service life point estimate of the current target motor.
[0022] S6, the current remaining useful life prediction interval is calculated using the covariance matrix adaptive calibration mechanism, and the coverage rate of the current prediction interval is calculated. ; The current forecast interval coverage Coverage rate with preset target The comparison is made, wherein the current predicted interval coverage is... It refers to the proportion of times the actual remaining useful life value falls within the prediction range out of all predicted times from the initial time to the current time during the online prediction process.
[0023] S7, if According to the above and The difference is used to update the covariance matrix adaptive calibration mechanism, and the updated covariance matrix adaptive calibration mechanism is used to return to step S6; otherwise, the current remaining service life prediction interval is identified as the remaining service life confidence interval; the online prediction framework outputs the current target motor's remaining service life point estimate and its confidence interval.
[0024] Compared with the prior art, the beneficial effects of the present invention are:
[0025] (1) Significantly improves prediction accuracy under small sample conditions: By pre-training a general degradation representation model on large-scale heterogeneous source domain data, the model acquires prior knowledge of device degradation; combined with a meta-learning strategy, the model can quickly infer the appropriate degradation trajectory from a very small number of target domain samples (the first 10% of the lifetime data). Experiments show that, under the condition that the training data is only 10% of the entire lifetime, the mean absolute error (MAE) of the method of this invention is as low as 8.3%, which is about 9% lower than the existing state-of-the-art (SOTA) methods and more than 30% lower than traditional deep learning methods.
[0026] (2) Excellent cross-domain generalization ability: The pre-trained model learns a general degradation representation based on health status, rather than specific features of a particular device. t-SNE (t-distributed stochastic neighbor embedding) visualization demonstrates that data points from different devices and operating conditions are distributed along a continuous degradation trajectory according to health status, rather than clustered by device type. In the cross-domain task of migrating from publicly available bearing data to autonomously collected motor data, the method of this invention achieves a MAE of 8.3%, significantly outperforming Domain-Adversarial Neural Network (DANN) at 11.3% and Model-Agnostic Meta-Learning (MAML) at 10.7%.
[0027] (3) Reliable uncertainty quantification and online calibration: An innovative deep reinforcement learning covariance matrix adaptive calibration mechanism enables the model to dynamically adjust the uncertainty estimate based on real-time prediction errors, achieving online calibration of the prediction interval. Experiments show that the prediction interval coverage probability (ICP) of the method in this invention reaches 93.4%, and the negative log-likelihood (NLL) is as low as 1.91, which is significantly better than existing methods, providing a reliable confidence interval for engineering decision-making.
[0028] (4) Interpretable prediction process: The modal importance weights of the multi-channel gated attention network output are highly consistent with the physical process of motor aging—early stage temperature weights dominate (thermal aging), mid-stage current harmonic weights increase (intensified partial discharge), and late-stage vibration weights rise (mechanical imbalance). This weight allocation, which is synchronized with the physical failure mode, verifies the rationality of the model's reasoning logic and significantly improves the credibility of the prediction results in the engineering field.
[0029] (5) Effective mitigation of catastrophic forgetting: The innovative dual-path network and dynamic gating mechanism enable the model to retain the general degradation knowledge acquired during pre-training in the source domain while learning new knowledge in the target domain. Ablation experiments demonstrate that this mechanism improves prediction accuracy by approximately 12% and prediction interval coverage probability by approximately 5% under small sample conditions. Attached Figure Description
[0030] Figure 1 This is a flowchart of the main method of the present invention;
[0031] Figure 2 This is a flowchart of the multi-source data preprocessing and general degradation characterization pretraining of the present invention;
[0032] Figure 3 This is a flowchart of the small-sample migration fusion and online uncertainty calibration of the present invention;
[0033] Figure 4 Comparison of cross-domain RUL prediction results and error distribution for cross-domain tasks. Detailed Implementation
[0034] This invention proposes a small-sample method for predicting the remaining service life of motors based on the fusion of pre-trained transfer learning and adaptive calibration. The design concept is as follows: Figure 1 As shown, this method first collects the full lifecycle data of source domain devices and the operating data of the target motor, and completes preprocessing of the signal layer, feature layer, and sequence layer. Second, based on the full lifecycle data of multiple source domain devices and multiple operating conditions, a general degradation representation pre-trained model is trained to obtain a pre-trained encoder. This pre-trained encoder is then used to construct the source domain knowledge path. Third, using small sample data of the target motor, the pre-trained encoder undergoes target domain soft alignment, health indicator construction, and meta-learning for rapid adaptation, resulting in a target domain adapted encoder. This target domain adapted encoder is then used to construct the target domain adaptive path. Next, a dual-path network is constructed based on the source domain knowledge path and the target domain adaptive path. The outputs of the source domain knowledge path and the target domain adaptive path are fused using dynamic gating vectors to obtain the degradation features of the target motor. Finally, the degradation features of the target motor are input into an online prediction framework, where modal weighting is performed using a multi-channel gating attention network, and the prediction interval is calibrated online using a covariance matrix adaptive calibration mechanism. If the predicted interval coverage does not reach the preset target coverage, return to the covariance matrix calibration step to continue updating; if the predicted interval coverage reaches the preset target coverage, output the RUL point estimate of the target motor and its confidence interval.
[0035] like Figure 1 As shown, the present invention proposes a small-sample method for predicting the remaining service life of motors based on the fusion of pre-trained transfer learning and adaptive calibration, comprising the following steps:
[0036] Step 1: Collect the full life cycle operation data of the source domain equipment and the initial operation data of the target motor in a small sample, and preprocess the data in the signal layer, feature layer and sequence layer.
[0037] Specifically, the source domain equipment lifecycle operation data is selected from source domain equipment data related to the degradation mechanism of the target motor. The source domain equipment includes bearings, gearboxes, motors, or other rotating equipment. The source domain data includes vibration, current, and temperature data. Each source domain sample includes at least a timestamp, equipment number, sensor sequence, and remaining service life value. The target motor small sample initial operation data consists of continuous monitoring data from the initial operation phase of the target motor to be predicted, preferably early-life operation data. The items in the source domain equipment lifecycle operation data include: timestamp, vibration, temperature, current, and remaining service life value; the items in the target motor small sample initial operation data are the same as those in the source domain equipment lifecycle operation data.
[0038] The preprocessing of the data at the signal layer includes: segmenting the vibration, current, and temperature signals into sliding windows; obtaining the noise floor for each frequency band using the minimum statistics adaptive noise estimation method; determining the dynamic wavelet threshold based on the noise floor; and performing soft thresholding on the wavelet coefficients of each decomposition scale to obtain the denoised multimodal signal.
[0039] Specifically, let the noise standard deviation of the j-th wavelet decomposition scale at time t be . The scale window length is N, where N represents the number of sampling points within the current window. The dynamic threshold... (The wavelet denoising threshold for the j-th decomposition scale at time t) is determined according to the following formula:
[0040] (1)
[0041] in, This is the global threshold adjustment coefficient. This is the local noise adjustment coefficient. The local noise factor is determined by the ratio of the current window noise energy to the historical noise energy. Soft thresholding is applied to the wavelet coefficients w_{j,c}, where c represents the wavelet coefficient index at that scale, yielding the denoised coefficients:
[0042] (2)
[0043] Equations (1) and (2) automatically increase the threshold when the noise is enhanced and automatically decrease the threshold when the noise is reduced, thus taking into account both the noise reduction intensity and the preservation of degradation features.
[0044] The feature layer preprocessing of the data includes: extracting the mean, root mean square, peak value, peak-to-peak value, standard deviation, skewness, kurtosis, waveform factor, peak factor, impulse factor, margin factor, and variance from the vibration signal; extracting transient power fluctuation, negative sequence current ratio, total harmonic distortion rate, harmonic amplitude, and sideband energy ratio from the electrical signal; extracting the temperature rise rate, temperature fluctuation range, and steady-state temperature mean from the temperature signal; and standardizing the above features according to source domain statistics after concatenation.
[0045] The sequence-level preprocessing of the data includes: resampling and aligning each modal feature according to the timestamp; performing window-level alignment of sequences with different sampling rates using dynamic time warping; calculating the correlation index for cross-modal feature pairs and removing features that are irrelevant to degradation or are dominated by noise; and saving the standardized mean, variance, and feature order for reuse in the online prediction stage of the target domain.
[0046] The items in the source domain device's full lifespan operation data include: timestamp, vibration, temperature, current, and remaining service life value;
[0047] The items in the initial operating data of the target motor sample are the same as the items in the full life-cycle operating data of the source domain equipment.
[0048] Step 2: Train a pre-trained model of general degradation representation based on the pre-processed source domain data to obtain a pre-trained encoder for representing the degradation state or health state of the source domain device. The pre-trained encoder constitutes the source domain knowledge path.
[0049] The pre-trained model adopts a multi-task self-supervised training objective, which includes a multi-scale masking reconstruction task, a temporal consistency comparison task, and a degradation gradient prediction task. The loss of each task is jointly optimized by a weight adjustment method based on task uncertainty.
[0050] The pre-trained encoder includes a modality-specific encoder, a cross-modal interaction module, and a shared representation layer; wherein the modality-specific encoder processes vibration, electrical signal, and temperature features respectively, the cross-modal interaction module fuses the temporal semantics of different modalities, and the shared representation layer outputs a degradation representation vector of a unified dimension.
[0051] Specifically, the joint loss function during the pre-training phase (The total loss for the three types of self-supervised tasks) is defined as:
[0052] (3)
[0053] Where r represents the self-supervised task number. These represent the multi-scale masking reconstruction loss, the temporal consistency contrast loss, and the degradation gradient prediction loss, respectively. Let be the learnable uncertainty parameter for the r-th task. The weights of different self-supervised tasks can be automatically adjusted during training using Equation (3). After training until the validation loss converges, the encoder parameters, normalized parameters, and source domain health state reference distribution are saved as the initial model for target domain transfer.
[0054] Specifically, such as Figure 2 As shown, the pre-training sub-process includes the following steps: inputting source domain data from multiple devices, operating conditions, and failure modes; adaptive noise estimation and dynamic wavelet thresholding; extraction of vibration, electrical signal, temperature, and operating condition features; DTW (Dynamic Time Warping) alignment and correlation filtering for sequences at different sampling rates; inputting these data into modality-specific encoders; obtaining shared degradation representations through a cross-modal interaction module; performing masking reconstruction, temporal consistency comparison, and degradation gradient prediction tasks; checking for convergence of the self-supervised loss; solidifying the pre-trained encoder parameters and standardized parameters; and ending the process. It can be seen that the source domain data first undergoes noise suppression, feature extraction, sequence alignment, and feature filtering to form a unified input, which then enters the modality coding branches corresponding to vibration, electrical signal, temperature, and operating conditions, i.e., the modality-specific encoders. The cross-modal interaction module performs shared representation learning on the outputs of each branch, learning the coupling relationships between different modes. The pre-training process employs three self-supervised tasks: multi-scale masking reconstruction, temporal consistency comparison, and degradation gradient prediction, enabling the encoder to learn general degradation knowledge related to changes in device health status. After the task is executed, if the self-supervised loss does not converge, the training returns to the modality-specific encoder. If it converges, the pre-trained encoder parameters and standardized parameters are saved for reuse in the target domain transfer stage.
[0055] Step 3: Based on the pre-processed initial running data of the target motor (small sample), the pre-trained encoder is subjected to soft alignment and meta-learning for rapid adaptation, resulting in a target domain-adaptive encoder. This target domain-adaptive encoder constitutes the target domain adaptive path. The process involves: fixing the underlying parameters of the pre-trained encoder, adjusting the shared representation layer parameters, and constraining the distribution difference between source domain features and target domain features with the maximum mean difference; dividing the initial target domain samples into a support set and a query set; updating the trajectory mapper parameters through the support set; and evaluating the updated trajectory prediction error, physical monotonicity error, and temporal smoothness error through the query set. The target domain-adaptive encoder outputs target domain adaptation features, and the trajectory mapper generates the target motor degenerate trajectory based on these features.
[0056] Specifically, the small sample data in the target domain is processed using the feature order and standardized parameters saved in step 1, and then input into the pre-trained encoder saved in step 2. The low-level parameters of the pre-trained encoder are frozen, and only the parameters of the shared representation layer and a few adaptation layers are adjusted. The feature distributions of the source and target domains are constrained by the maximum mean difference, correlation alignment, or other distribution distances, so that the features of the target domain fall into the adjacent region of the healthy state reference distribution of the source domain. The feature distribution difference between the source and target domains is constrained by the maximum mean difference, and its alignment loss... (The loss representing the difference in distribution between the source and target domains) is:
[0057] (4)
[0058] in, Let be the degradation characterization of the i-th source domain sample. This represents the degradation characterization of the p-th target domain sample. (·) represents the kernel mapping function, and H represents the reproducing kernel Hilbert space. and These represent the number of samples in the source and target domains, respectively.
[0059] Furthermore, a hybrid autoencoder (CNN-KAN, or equivalent nonlinear autoencoder) is used to reconstruct early healthy samples in the target domain. The reconstruction error is then converted into a health index score via a monotonic mapping, which serves as the input to the degradation trajectory mapper. For the t-th window of the target domain, the reconstruction error of the hybrid autoencoder is... (representing the reconstruction error of the t-th window) and health indicators (Representing the health indicators for the t-th window) are as follows:
[0060] (5)
[0061] in, Let t be the input features of the t-th window. For the reconstruction features of the t-th window, This represents the scaling factor for health indicators. A larger reconstruction error indicates a more significant deviation between the window and the initial healthy sample. The smaller.
[0062] The target domain finite sequence is divided into a support set and a query set. The trajectory mapper parameters are updated on the support set, and the trajectory prediction error, physical monotonicity error, and temporal smoothness error are calculated on the query set. This process is repeated until the query set error stabilizes, ultimately yielding a degenerate trajectory mapper for the target domain. Specifically, the degenerate trajectory mapper models discrete health indicators as a continuous dynamic system.
[0063] (6)
[0064] in, For continuous time The following degenerate state, Input for operating conditions. Let be the state evolution function represented by a neural network, and θ be the trainable parameters of this state evolution function. The meta-learning optimization objective is:
[0065] (7)
[0066] in, To support the collection of losses, To query set loss, To support set update step size, For physical monotonicity constraints, For time series smoothness constraints, and For the corresponding weights.
[0067] Specifically, such as Figure 3 As shown, a small sample of initial operating data of the target motor is input into the pre-trained encoder. The bottom-level encoder is frozen, and soft alignment is performed on the shared representation layer and the target domain to ensure that the degradation features of the target domain are consistent with the health status reference distribution of the source domain. Subsequently, the target domain samples are divided into a support set and a query set. Meta-learning adaptation is completed through the support set and the query set. Then, health indicators are calculated through a hybrid autoencoder or an equivalent nonlinear autoencoder. The degradation trajectory mapping parameters are updated through a meta-learning method, i.e., neural differential equations, enabling the model to quickly form a degradation trajectory description of the target motor under a small amount of target domain data. Features are output by the source domain knowledge path and the target domain adaptive path, respectively. That is, the source domain knowledge path calls the frozen pre-trained encoder to output general degradation features; the target domain adaptive path calls the target domain adapted encoder to output target motor-specific degradation features. After dynamic gating feature fusion, i.e., the dynamic gating vector adaptively determines the fusion ratio of the two paths according to the current input data, thereby absorbing new degradation features of the target motor while retaining the degradation knowledge of the source domain. The fused degradation features are weighted and fused using a multi-channel attention mechanism to integrate vibration, electrical signals, temperature, and operating conditions. The covariance matrix is updated by the policy network based on the residual state. That is, the prediction uncertainty is updated based on the online prediction residual through the covariance matrix adaptive calibration mechanism. If the interval coverage or interval width does not meet the requirements, the calibration is returned to the covariance matrix update node until the prediction interval meets the maintenance decision requirements.
[0068] Step 4: The source domain knowledge path and the target domain adaptive path constitute a dual-path network. The preprocessed small sample initial running data of the target motor is used as the input of the dual-path network. The output of the dual-path network is fused through a dynamic gating vector to obtain the degradation features of the target motor. The dynamic gating vector is jointly determined by the preprocessed small sample initial running data of the target motor, the output of the source domain knowledge path, and the output of the target domain adaptive path.
[0069] Specifically, the degradation characteristics of the target motor can be expressed as:
[0070] (8)
[0071] in Output the knowledge path from the source domain. The target domain adaptive path output is given by z, which is the dynamic gating vector, and ⊙ represents element-wise multiplication.
[0072] The calculation process of the dynamic gating vector can be represented as follows:
[0073] (9)
[0074] Where x is the current input feature, Output the knowledge path from the source domain. The target domain adaptive path output is z, which is the dynamic gating vector. and Here, represents the weights and biases of the fully connected layers in the gated network, and sigmoid(·) is the sigmoid activation function. When the target domain has little data, high noise, or unstable distribution drift, the weights of the source domain knowledge path are increased; when new samples in the target domain gradually accumulate and the prediction residuals stabilize, the weights of the target domain adaptive path are increased. In the dual-path feature fusion stage, a source domain knowledge path and a target domain adaptive path are constructed. The source domain knowledge path calls the frozen pre-trained encoder to output general degradation features; the target domain adaptive path calls the target domain adapted encoder to output target motor-specific degradation features. The dynamic gating vector adaptively determines the fusion ratio of the two paths based on the current input data, thereby absorbing new degradation features of the target motor while retaining the source domain degradation knowledge.
[0075] Step 5: Input the degradation features of the target motor into the online prediction framework. The online prediction framework includes a multi-channel gated attention network and a covariance matrix adaptive calibration mechanism. The fused degradation features are used to perform modal weighting calculation through the multi-channel gated attention network to obtain the remaining service life point estimate of the current target motor.
[0076] The covariance matrix adaptive calibration mechanism uses the predicted residual characteristics, the previous time step covariance matrix, and the previous time step predicted interval coverage as states, and the covariance matrix scaling factor as an action; the update of the scaling factor follows the following reward criterion:
[0077] First, when the predicted interval coverage is less than the preset target coverage, a penalty value negatively correlated with the coverage difference is given, and the wider the interval, the greater the penalty.
[0078] Second, when the predicted interval coverage is greater than or equal to the preset target coverage, the coverage penalty is zero, and the reward value is only negatively correlated with the predicted interval width, so as to drive the covariance matrix adaptive calibration mechanism to output a scaling factor that minimizes the interval width under the premise of satisfying the coverage constraint.
[0079] Specifically, the scaling factor of the output is represented as the action vector. (This represents the scaling action of the covariance matrix at time t), used to update the predicted covariance matrix:
[0080] (10)
[0081] in, Let be the covariance matrix at time t. Let be the covariance matrix of the previous time step. The regularization coefficient is . It is the identity matrix. For the action vector The resulting diagonal matrix. Reward function. (The calibration reward at time t) is defined as:
[0082] (11)
[0083] in, This represents the current forecast interval coverage. To achieve the preset target coverage rate, the current prediction interval coverage rate refers to the proportion of times the actual remaining useful life value falls within the prediction interval out of all predicted times from the initial time to the current time during the online prediction process. The preferred value is 0.9.
[0084] (Prediction Interval Normalized Average Width) is the width of the normalized average interval. and is the penalty coefficient. Equations (10) and (11) are used to strike a balance between achieving coverage targets and having a smaller interval width.
[0085] In the online prediction phase, the fused degradation features are input into the remaining service life prediction framework. This online prediction framework recursively updates the health status and degradation trajectory based on real-time collected target motor data, and performs weighted fusion of vibration, electrical signals, temperature, and operating mode through a multi-channel gated attention network. Simultaneously, an adaptive calibration mechanism of the covariance matrix updates the prediction uncertainty based on the online prediction residuals, ultimately outputting a point estimate and confidence interval for the remaining service life of the target motor.
[0086] Step 6: Calculate the current remaining useful life prediction interval using the covariance matrix adaptive calibration mechanism, and calculate the coverage of the current prediction interval. ; The current forecast interval coverage Coverage rate with preset target The comparison is made, wherein the current predicted interval coverage is... In online forecasting, the percentage of times the actual remaining useful life falls within the forecast range out of all forecasted moments from the initial moment to the current moment; preset target coverage. That is, the preset threshold, preferably 0.9.
[0087] Step 7: If According to the above and The difference is used to update the covariance matrix adaptive calibration mechanism, and the updated covariance matrix adaptive calibration mechanism is used to return to step S6; otherwise, the current remaining service life prediction interval is identified as the remaining service life confidence interval; the online prediction framework outputs the current target motor's remaining service life point estimate and its confidence interval.
[0088] Finally, to verify the prediction effect of the method of the present invention in cross-domain small sample scenarios, this embodiment migrates the source domain device data to the target motor data for cross-domain RUL prediction, and compares it with SVR (Support Vector Regression), XGBoost (eXtreme Gradient Boosting), DANN and SOTA. Figure 4 This shows the RUL prediction results and error distribution of different methods in cross-domain tasks. Figure 4 The scatter plot in (a) compares the predicted and true labels, with the horizontal axis representing the true labels, the vertical axis representing the predicted labels, the dashed line representing the ideal case line, and circles, squares, triangles, diamonds, and stars representing the prediction results of SVR, XGBoost, DANN, SOTA, and the present method, respectively. The predicted points of the present invention are closer to the ideal case line, indicating a higher consistency between the predicted values and the true labels. Figure 4In the middle (b), the prediction error box plot is shown. The horizontal axis represents SVR, XGBoost, DANN, SOTA and the present method, and the vertical axis represents the prediction error. The error box of the present invention is lower and less discrete, indicating that the present invention has higher prediction accuracy and more stable error distribution under cross-domain small sample conditions.
[0089] Although the present invention has been described above in conjunction with the accompanying drawings, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are preferred application examples that demonstrate the core technical ideas of the present invention, and are merely illustrative and not restrictive. Those skilled in the art can make many improvements and changes under the guidance of the present invention without departing from the spirit of the present invention, and these all fall within the protection scope of the present invention.
Claims
1. A method for predicting the remaining service life of a motor based on a small sample size using a fusion of pre-trained transfer learning and adaptive calibration, characterized in that, Includes the following steps: S1, collect the full life cycle operation data of the source domain device and the initial operation data of the target motor in a small sample, and preprocess the data in the signal layer, feature layer and sequence layer. The items in the source domain device's full lifespan operation data include: timestamp, vibration, temperature, current, and remaining service life value; the items in the target motor's small sample initial operation data are the same as those in the source domain device's full lifespan operation data. S2, based on the preprocessed source domain data, train a pre-trained model of general degradation representation to obtain a pre-trained encoder for representing the degradation state or health state of the source domain device, and the pre-trained encoder constitutes the source domain knowledge path. S3, Based on the pre-processed small sample initial running data of the target motor, the pre-trained encoder is subjected to soft alignment and meta-learning for rapid adaptation to obtain the target domain adapted encoder; the target domain adapted encoder constitutes the target domain adaptive path. S4, the source domain knowledge path and the target domain adaptive path constitute a dual-path network. The preprocessed small sample initial running data of the target motor is used as the input of the dual-path network. The output of the dual-path network is fused through dynamic gating vector to obtain the degradation features of the target motor. S5, input the degradation features of the target motor into the online prediction framework. The online prediction framework includes a multi-channel gated attention network and a covariance matrix adaptive calibration mechanism. The fused degradation features are used to perform modal weighting calculation through the multi-channel gated attention network to obtain the remaining service life point estimate of the current target motor. S6, the current remaining useful life prediction interval is calculated using the covariance matrix adaptive calibration mechanism, and the coverage rate of the current prediction interval is calculated. ; The current forecast interval coverage Coverage rate with preset target The comparison is made, wherein the current predicted interval coverage is... It refers to the proportion of times the actual remaining useful life value falls within the prediction range out of all predicted times from the initial time to the current time during the online prediction process. S7, if According to the above and The difference is used to update the covariance matrix adaptive calibration mechanism, and the updated covariance matrix adaptive calibration mechanism is used to return to step S6; otherwise, the current remaining service life prediction interval is identified as the remaining service life confidence interval; the online prediction framework outputs the current target motor's remaining service life point estimate and its confidence interval.
2. The method for predicting the remaining service life of a small-sample motor according to claim 1, characterized in that, In step S1, the preprocessing of the data at the signal layer includes: performing sliding window segmentation on the vibration, current, and temperature signals respectively; obtaining the noise floor of each frequency band using the minimum statistics adaptive noise estimation method; determining the dynamic wavelet threshold based on the noise floor, and performing soft thresholding on the wavelet coefficients of each decomposition scale to obtain the denoised multimodal signal.
3. The method for predicting the remaining service life of a small-sample motor according to claim 1, characterized in that, In step S1, the feature layer preprocessing of the data includes: extracting the mean, root mean square, peak value, peak-to-peak value, standard deviation, skewness, kurtosis, waveform factor, peak factor, impulse factor, margin factor, and variance from the vibration signal; extracting transient power fluctuation, negative sequence current ratio, total harmonic distortion rate, harmonic amplitude, and sideband energy ratio from the electrical signal; extracting the temperature rise rate, temperature fluctuation range, and steady-state temperature mean from the temperature signal; and standardizing the above features according to source domain statistics after splicing them together.
4. The method for predicting the remaining service life of a small-sample motor according to claim 1, characterized in that, In step S1, the sequence layer preprocessing of the data includes: resampling and aligning each modal feature according to the timestamp; using dynamic time warping to align sequences with different sampling rates at the window level; calculating the correlation index for cross-modal feature pairs, and removing features that are irrelevant to degradation or are dominated by noise; and saving the standardized mean, variance, and feature order for reuse in the online prediction stage of the target domain.
5. The method for predicting the remaining service life of a small-sample motor according to claim 1, characterized in that, In step S2, the pre-trained model adopts a multi-task self-supervised training objective, which includes a multi-scale masking reconstruction task, a temporal consistency comparison task, and a degradation gradient prediction task. The loss of each task is jointly optimized by a weight adjustment method based on task uncertainty.
6. The method for predicting the remaining service life of a small-sample motor according to claim 1, characterized in that, In step S2, the pre-trained encoder includes a modality-specific encoder, a cross-modal interaction module, and a shared representation layer; wherein the modality-specific encoder processes vibration, electrical signal, and temperature features respectively, the cross-modal interaction module fuses the temporal semantics of different modalities, and the shared representation layer outputs a degradation representation vector of a unified dimension.
7. The method for predicting the remaining service life of a small-sample motor according to claim 1, characterized in that, Step S3 involves fixing the underlying parameters of the pre-trained encoder, adjusting the parameters of the shared representation layer, constraining the distribution difference between source domain features and target domain features with the maximum mean difference, dividing the initial samples of the target domain into a support set and a query set, updating the trajectory mapper parameters through the support set, and evaluating the updated trajectory prediction error, physical monotonicity error, and temporal smoothness error through the query set.
8. The method for predicting the remaining service life of a small-sample motor according to claim 1, characterized in that, In step S4, the dynamic gating vector is determined by the preprocessed small sample initial running data of the target motor, the source domain knowledge path output, and the target domain adaptive path output.
9. The method for predicting the remaining service life of a small-sample motor according to claim 1, characterized in that, In step S5, the covariance matrix adaptive calibration mechanism uses the predicted residual characteristics, the covariance matrix of the previous time step, and the predicted interval coverage of the previous time step as states, and the covariance matrix scaling factor as an action; the update of the scaling factor follows the following reward criterion: First, when the predicted interval coverage is less than the preset target coverage, a penalty value negatively correlated with the coverage difference is given, and the wider the interval, the greater the penalty. Second, when the predicted interval coverage is greater than or equal to the preset target coverage, the coverage penalty is zero, and the reward value is only negatively correlated with the predicted interval width, so as to drive the covariance matrix adaptive calibration mechanism to output a scaling factor that minimizes the interval width under the premise of satisfying the coverage constraint.