A Multi-Objective Domain Bearing Remaining Useful Life Prediction Method and System

Through the ModernTCN network and dynamic time regular DTW combined with the maximum mean difference MMD method, the problem of data distribution differences and high computing resource consumption in the residual service life prediction of bearings is solved, and feature extraction and transfer learning across target domains is realized, and prediction accuracy and efficiency are improved.

CN119939259BActive Publication Date: 2025-07-11WUHAN TEXTILE UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510433006.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-11
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

The prior art faces the problems of large data distribution differences, in the prediction of residual service life of bearings, the inability to adjust the model online, and the consumption of computing resources, and it is difficult to adapt to industrial application scenarios in multi-target domains.

Method used

The multi-objective domain bearing residual service life prediction method is used, offline training is carried out through the ModernTCN network, combining dynamic time regularization and maximum mean difference measurement, feature extraction and transfer learning across the target domain are achieved, and the dynamic time regularization DTW is used to allocate weights, combined with maximum mean difference MMD is used to measure distribution differences, and the model is fine-tuned in the online stage to improve prediction accuracy and efficiency.

Benefits of technology

It realizes more accurate residual service life prediction in diversified industrial application scenarios, reduces computing resource occupation, and improves the generalization ability and prediction efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939259B_ABST
    Figure CN119939259B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for predicting the remaining useful life of multi-objective domain bearings, including: randomly selecting data under different working conditions from relevant datasets as the source domain, target domain, and test data respectively, and then pre-training the processed source domain data using ModernTCN; next, calculating distances and assigning weights using DTW, and calculating the weighted maximum mean difference value with the help of MMD to achieve feature-level transfer learning between the test data and the multi-objective domain; further, inputting the weighted mixed target domain into the RUL predictor to obtain a prediction error, which is used as a loss function together with WMMD to fine-tune some parameters of the ModernTCN part to obtain a prediction model; finally, inputting the test data into the prediction model to obtain the prediction result of the remaining useful life of the online data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of predicting the remaining useful life (RUL) of mechanical equipment, and particularly to a method for predicting the RUL of bearings in multiple target domains. Background Art

[0002] Under the background of the rapid development of industrial big data and industrial Internet of Things, the technology of prognostics and health management (PHM) in mechanical reliability engineering is becoming increasingly important, and the prediction of the remaining useful life (RUL) of equipment has become a key task. As a crucial rotating component in mechanical equipment, the performance degradation of bearings directly affects the overall performance of the mechanical system. Therefore, the prediction of bearing RUL is of great significance for improving the level of PHM, optimizing maintenance strategies, and enhancing the reliability of rotating machinery.

[0003] Currently, many studies on bearing RUL prediction focus on data-driven methods based on deep learning technology, but traditional deep learning models face many challenges in practical applications. On the one hand, when predicting RUL, it is required that the training and test dataset distributions be consistent. However, in the actual industrial environment, the operating conditions of equipment are complex and variable, and the data distributions are quite different, making it difficult to meet this requirement, resulting in limited model performance. On the other hand, most traditional methods use offline modeling, and the model parameters cannot be adjusted online, making it difficult to adapt to new working conditions. Although online modeling methods can handle new working conditions, they consume a large amount of computing resources and have low model application efficiency.

[0004] In addition, most existing studies assume that the test data only comes from a single target domain. However, in the actual industrial process, it is impossible to ensure that the domain where the test data is located has complete life cycle data, and it is also difficult to determine its source domain or target domain. It is difficult to obtain good prediction results by directly using existing methods. Summary of the Invention

[0005] To solve the problems existing in the prior art, the present invention provides a method for predicting the remaining useful life of bearings in multiple target domains, which includes the following steps:

[0006] S1, randomly select data under different working conditions from the industrial process bearing operation-related dataset as source domain data 、 、……、 , n >= 1, is the number of source domains, target domain data 、 、……、 , m >= 1, is the number of target domains, and test data Test, ensuring that the test data Test is in a different domain from the source domain data and in the same domain as some target domain training data;

[0007] S2, in the offline stage, for the source domain data The patches obtained by padding and splitting are input into a 1D convolutional layer to obtain input embeddings ;

[0008] S3. Use ModernTCN to train to obtain model model_1; the ModernTCN includes stacked ModernTCN blocks, i.e., the backbone network, and a linear head with a flatten layer;

[0009] S4. In the online phase, for the test data Test and the target domain data , , ……, Use dynamic time warping (DTW) to align the sequences respectively and calculate the distances, and assign the initial weights of the target domain according to the distances;

[0010] S5. Use the maximum mean discrepancy (MMD) to measure the distribution difference between the target domain processed by DTW and the test data, and calculate the value of the initial weighted maximum mean discrepancy (WMMD);

[0011] S6. Take the last two layers of model model_1 as the remaining useful life predictor model_2, input the weighted mixed target domain, and obtain the predicted remaining useful life;

[0012] S7. Take the sum of the value of WMMD and the mean square error between the predicted remaining useful life obtained by model_2 and the true remaining useful life as the loss function, fine-tune some model parameters of model_1, and then repeat S4, S5, and S6 in sequence to obtain the final remaining useful life prediction model model_1_final. Finally, pass the test data Test through model_1_final to obtain the corresponding remaining useful life prediction target value.

[0013] Furthermore, the ModernTCN block includes a depthwise convolution (DWConv) and two convolutional feed-forward networks (ConvFFN), namely ConvFFN1 and ConvFFN2. DWConv is responsible for independently learning the temporal information between tokens on each feature channel. The role of ConvFFN1 is to generate new feature representations for each variable, while the task of ConvFFN2 is to capture the interdependencies between variables of each feature.

[0014] Furthermore, the specific implementation method of step S3 is as follows:

[0015] (3.1) Input into the backbone network to capture the cross-time and cross-variable interdependencies, so as to learn rich information representations , and the formula is as follows:

[0016]

[0017] Among them, M represents the number of variables included in the source domain data, N represents the number of patches, and D represents the number of channels. is the stacked ModernTCN blocks;

[0018]

[0019] Among them, , {1, ..., K} is the input of the i-th ModernTCN block, represents the ModernTCN block. After passing through the backbone network, the final ;

[0020] (3.2) Use a linear head with a flattening layer to obtain the final prediction result:

[0021]

[0022] Among them, is the prediction result with a freely set length of T, which contains M variables; The operation refers to the layer that reshapes the final representation into a tensor Z, and the dimension of the tensor Z is M×(D×N); represents the linear projection layer, which converts the final representation into the final prediction result;

[0023] (3.3) In the offline stage, the calculation formula of the loss function of the model model_1 is as follows:

[0024]

[0025] Among them, n represents the number of source domains, represents the true label.

[0026] Furthermore, the process of using dynamic time warping DTW to assign weights in step S4 is as follows:

[0027] (4.1) Dynamic time warping DTW is used to measure the similarity between two sequences. Given two time series, which are ) and ), DTW(P, Q) is defined as the optimal path among all possible paths T; a path T is a sequence ), where , , , a and b represent sequence numbers, and for each step k∈ belongs to the set , the DTW formula is as follows:

[0028]

[0029] Among them, \(t\) represents the number of nodes on the optimal path, represents the value of the \(k\)-th dimension of the \(i\)-th data point in the time series \(P\), represents the value of the \(k\)-th dimension of the \(j\)-th data point in the time series \(Q\), is the Euclidean distance, which is a distance metric and is calculated as follows:

[0030]

[0031] (4.2) The weighted coefficient is obtained by normalizing the DTW distance, and the calculation formula is as follows:

[0032]

[0033] Among them, represents the number of labeled target domains, \(r\) represents the \(r\)-th labeled target domain, represents the labeled target domain data embedded through the backbone network, represents the test data after being embedded through the backbone network.

[0034] Furthermore, in step S5, WMMD is used to measure the dissimilarity of the domain adaptation distribution between the target domain and the test set using DTW, so as to achieve feature-level transfer learning between the test data and multiple target domains. The calculation formula is as follows:

[0035]

[0036] Among them, represents the value of WMMD, () is a mapping function that maps data to a high-dimensional feature space, represents the norm of the Hilbert space.

[0037] Furthermore, in step S6, the mixed target domain is defined as:

[0038]

[0039] Among them, represents the mixed target domain.

[0040] Furthermore, the loss function formula between the true label and the target domain data label in the remaining useful life predictor model_2 is as follows:

[0041]

[0042] Among them, \(m\) represents the number of target domain data, Represents the true label, indicating the predicted value obtained by the predictor model_2.

[0043] Furthermore, in step S7, based on randomly freezing some model parameters in step S3, other parameters adopt pre-trained values, and then the model model_1 is fine-tuned to complete the model-based transfer learning between the second-level hierarchical source domain and multiple target domains. The formula of the overall optimization function for multi-level domain transfer learning is as follows:

[0044]

[0045] where λ and γ are regularization parameters.

[0046] Furthermore, it also includes using the root mean square error RMSE and the mean absolute error MAE as evaluation indicators to evaluate the final prediction results.

[0047] The present invention also provides a multi-target domain bearing remaining useful life prediction system, including the following modules:

[0048] A data domain acquisition module, configured to randomly select data under different working conditions from the industrial process bearing operation-related data set as source domain data , , ……, , n >= 1, is the number of source domains, and the target domain data , , ……, , m >= 1, is the number of target domains, and the test data Test, ensuring that the test data Test is in a different domain from the source domain data and in the same domain as some of the target domain training data;

[0049] A source domain data preprocessing module, configured to, in the offline stage, fill and segment the source domain data to obtain the input embedding by feeding the obtained patches into a 1D convolutional layer ;

[0050] An offline feature extraction module, using ModernTCN to train and obtain the model model_1; the ModernTCN includes stacked ModernTCN blocks, i.e., the backbone network and a linear head with a flattening layer;

[0051] A data relationship measurement module, configured to, in the online stage, use dynamic time warping DTW to , , ……, align the sequences of the test data Test and the target domain data

[0052] A distribution difference calculation module, which is used to measure the distribution difference between the target domain after DTW processing and the test data by using the Maximum Mean Discrepancy (MMD), and calculate the value of the initial Weighted Maximum Mean Discrepancy (WMMD).

[0053] A remaining useful life prediction module, which is used to take the last two layers of the model model_1 as the remaining useful life predictor model_2, and input the weighted hybrid target domain to obtain the predicted remaining useful life.

[0054] A final prediction module, which is used to take the sum of the value of WMMD and the mean square error between the predicted remaining useful life obtained by model_2 and the true remaining useful life as the loss function, fine-tune some model parameters of model_1, and then repeat the data relationship measurement module, the distribution difference calculation module and the remaining useful life prediction module in sequence to obtain the final remaining useful life prediction model model_1_final. Finally, the test data Test is passed through model_1_final to obtain the corresponding predicted target value of the remaining useful life.

[0055] Through the above technical solutions conceived by the present invention, compared with the prior art, the following beneficial effects can be achieved:

[0056] The present invention proposes a novel RUL transfer learning method across multiple target domains. In the current research field, the literature on related topics seems to be scarce.

[0057] A multi-level domain adaptation strategy is designed, in which the improved weighted online MMD completes the transfer of feature extraction and degradation knowledge from the source domain to the target domain and the test set, improves the generalization ability of the RUL prediction model, and enables it to more accurately adapt to diverse industrial application scenarios.

[0058] The strategy of placing the main training task of the model in the offline stage and then performing online fine-tuning according to the test set not only ensures the prediction accuracy, avoids over-consuming online computing resources, but also greatly improves the operation efficiency of the entire prediction system. Description of the Drawings

[0059] Figure 1 is the flowchart provided by the embodiment of the present invention;

[0060] Figure 2 is a schematic diagram of the RUL prediction fitting effect of four tasks provided by the embodiment of the present invention;

[0061] Figure 3 is an ablation experiment on Task1 provided by the embodiment of the present invention, where (a) the method proposed by the present invention; (b) without fine-tuning; (c) without WMMD; (d) only one target domain. Detailed Embodiments

[0062] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0063] To solve the problem of bearing multi-target domain RUL prediction and realize the creation of a transferable model that can dynamically adapt to online test data from different target domains in real time, the present invention provides a method for predicting the remaining useful life of bearings based on multi-level domain transfer lightweight. First, in the offline stage, different working conditions are randomly selected from the industrial process bearing operation-related data set as the source domain, target domain, and test data. The source domain data is filled and segmented and input into a one-dimensional convolutional layer, and then this embedding vector is used to preliminarily model the ModernTCN network. In the subsequent online stage, DTW is used to calculate the distance between the test data and the target domain data and assign weights, and then MMD is used to calculate the distribution difference between the target domain and the test set after DTW processing to obtain the WMMD value. Then, the last two layers of ModernTCN are taken to construct an RUL predictor, and the weighted mixed target domain data is input to obtain the predicted remaining useful life. The preliminary model ModernTCN is fine-tuned with the sum of WMMD and the mean square deviation between the weighted true RUL and the predicted RUL as the loss function. After iterative optimization, finally, Test is input into the model to obtain the prediction of the remaining useful life. Figure 1 The flowchart showing the present invention is as follows, and the following are specific implementation examples.

[0064] (1) Data collection and processing:

[0065] Randomly select different working conditions from the industrial process bearing operation-related data set as the source domain , , ……, , n>=1, is the number of source domains, the target domain , , ……, , m>=1, is the number of target domains, and the test data Test, ensuring that the test data Test is from a different domain from the source domain data and from the same domain as part of the target domain training data. Specifically:

[0066] Obtain the industrial process bearing operation-related data , where M represents the number of variable features and L represents the sequence length;

[0067] For the source domain data The patches obtained by filling and segmenting are input into a one-dimensional convolutional layer to obtain the input embedding , specifically:

[0068] (1.2.1) Uncompress the source domain input data sequence of M variables with length L into ;

[0069] (1.2.2) The step size during the chunking process is S, repeat the last value (P - S) times, and add it to the end of to keep N = L / / S.

[0070] (1.2.3) Divide the padded into N patches of size P. Then input these patches into a 1D convolutional layer, which maps 1 input channel to D output channels, to obtain the embedding vector , the formula is as follows:

[0071]

[0072] (2) Model pre-training and online domain adaptation

[0073] (2.1) Use ModernTCN to train to obtain the model model_1. The ModernTCN includes stacked ModernTCN blocks, i.e., the backbone network and a linear head with a flatten layer; specifically:

[0074] (2.1.1) Input into the backbone network to capture the interdependencies across time and variables, so as to learn an information-rich representation , the formula is as follows:

[0075]

[0076] is the stacked ModernTCN blocks.

[0077]

[0078] Among them, , {1,...,K} is the input of the i-th ModernTCN block, represents the ModernTCN block. After passing through the backbone network, we get the final .

[0079] (2.1.2) The ModernTCN block includes a depthwise convolution DWConv and a convolutional feed-forward network ConvFFN. The following are their respective operations and functions:

[0080] (2.1.2.1)Depthwise Convolution DWConv: Responsible for independently learning the temporal information between tokens on each feature channel, mixing information in the temporal dimension, and mapping input channels to output channels to achieve the independence of variables and feature dimensions.

[0081] (2.1.2.2)Convolutional Feed-Forward Network ConvFFN: After decoupling ConvFFN based on the group convolution idea, ConvFFN1 and ConvFFN2 are formed. By setting the number of groups to M, ConvFFN1 can independently learn new feature representations for each variable, while ConvFFN2 captures cross-variable dependencies in each feature dimension by setting the number of groups to D.

[0082] (2.1.3)Using a linear head with a flattening layer to obtain the final prediction result:

[0083]

[0084] where, is the prediction result of length T freely set by us, containing M variables. The operation refers to the layer that reshapes the final representation into tensor Z, and the dimension of tensor Z is M×(D×N). represents the linear projection layer, which converts the final representation into the final prediction result.

[0085] (2.1.3)In the offline phase, the calculation formula of the loss function is as follows:

[0086]

[0087] where, n represents the number of source domains, represents the true label. When the loss function converges, the pre-training ends, and when the loss is minimized, the best model is saved. In this study, the best pre-trained model parameters of the hidden layer in the source domain are frozen and copied, and used as the fixed parameters of the model in the subsequent target domain.

[0088] (2.2)In the online phase, for the test data Test and the target domain data 、 、……、 Dynamic Time Warping DTW is used to align the sequences and calculate the distances respectively, and the initial weights of the target domain are assigned according to the distances. Specifically:

[0089] (2.2.1)Dynamic Time Warping DTW is used to measure the similarity between two sequences. Given two time series, which are ) and ), DTW(P, Q) is defined as the optimal path among all possible paths T. A path T is a sequence ), where , , , a and b represent sequence numbers, and for each step k ∈ belongs to the set , the DTW formula is as follows:

[0090]

[0091] where t represents the number of nodes on the optimal path, represents the value of the k-th dimension of the i-th data point in the time series P, represents the value of the k-th dimension of the j-th data point in the time series Q, is the Euclidean distance, which is a distance metric and is calculated as follows:

[0092]

[0093] (2.2.2) The weighted coefficient is obtained by normalizing the DTW distance, and the calculation formula is as follows:

[0094]

[0095] where, represents the number of labeled target domains, r represents the r-th labeled target domain, represents the data of the labeled target domain embedded through the backbone network, represents the test data after being embedded through the backbone network.

[0096] (2.3) Use the maximum mean discrepancy MMD to measure the distribution difference between the target domain processed by DTW and the test set, and calculate the value of the initial weighted maximum mean discrepancy WMMD. Specifically:

[0097] WMMD is used to measure the dissimilarity of the domain adaptation distribution between the target domain and the test set using DTW, so as to achieve feature-level transfer learning between the test data and multiple target domains. The calculation formula is as follows:

[0098]

[0099] where, represents the value of WMMD, () is the mapping function that maps the data to the high-dimensional feature space, represents the norm of the Hilbert space. By minimizing , it enables the model to learn how to adjust parameters to reduce the difference in data distributions between the test domain and multiple target domains, and achieve feature-level transfer learning between the test data and multiple target domains.

[0100] (3) RUL Prediction and Model Iterative Optimization

[0101] (3.1) Use the last two flattening layers and linear layer of ModernTCN as the RUL predictor model_2, retain the model parameters trained in the source domain, and input the weighted mixed target domain to obtain the predicted remaining useful life. Specifically:

[0102] (3.1.1) The mixed target domain is defined as follows:

[0103]

[0104] (3.1.2) To obtain the degradation information contained in the target data labels, use the RUL prediction results in model_2 to predict the mixed target domain. The loss function formula between the true label and the target data label is as follows:

[0105]

[0106] where m represents the number of target domains, represents the true label.

[0107] (3.2) Use the sum of the value of WMMD and the mean square error between the predicted remaining useful life obtained by model_2 and the true remaining useful life as the loss function to fine-tune some model parameters of model_1. After the fine-tuning is completed, the updated model_1 is obtained, and then repeat (2.2), (2.3) and (3.1) in sequence to obtain the final RUL prediction model model_1_final. Finally, pass test through model_1_final to obtain the corresponding predicted target value of the remaining useful life. Specifically:

[0108] (3.2.1) Based on randomly freezing some model parameters in step S3, other parameters use the pre-trained values, and then fine-tune the model to complete the model-based transfer learning between the second-level hierarchical source domain and multiple target domains. Integrate the above loss function, and the formula of the overall optimization function for multi-level domain transfer learning is as follows:

[0109]

[0110] where λ and γ are regularization parameters. During this process, the transfer learning between the target domain and the test data, and between the source domain and the target domain continues.

[0111] (3.2.2)When the loss function converges, input Test into the trained model to calculate the final predicted value of the remaining useful life, and the formula is as follows:

[0112]

[0113] Where represents the input embedding after padding and segmentation, represents the trained model, represents the input embedding After The predicted value of the remaining useful life.

[0114] The second aspect of the present invention provides a multi-objective domain bearing remaining useful life prediction system, including the following modules:

[0115] The data domain acquisition module is used to randomly select data under different working conditions from the industrial process bearing operation-related data set as the source domain data , ,……, , n>=1, is the number of source domains, and the target domain data , ,……, , m>=1, is the number of target domains, and the test data Test, ensuring that the test data Test is in a different domain from the source domain data and in the same domain as part of the target domain training data;

[0116] The source domain data preprocessing module is used to, in the offline stage, perform padding and segmentation on the source domain data to obtain the input embedding through the patch input 1D convolutional layer ;

[0117] The offline feature extraction module uses ModernTCN to train to obtain the model model_1; the ModernTCN includes a stacked ModernTCN block, i.e., the backbone network and a linear head with a flattening layer;

[0118] The data relationship measurement module is used to, in the online stage, perform dynamic time warping DTW on the test data Test and the target domain data , ,……, to align the sequences respectively and calculate the distance, and allocate the initial weights of the target domain according to the distance;

[0119] The distribution difference calculation module is used to measure the distribution difference between the target domain and the test data after DTW processing by using the maximum mean difference MMD, and calculate the value of the initial weighted maximum mean difference WMMD;

[0120] The remaining useful life prediction module is used to take the last two layers of the model model_1 as the remaining useful life predictor model_2, input the weighted mixed target domain, and obtain the predicted remaining useful life;

[0121] The final prediction module is used to take the sum of the value of WMMD and the mean square error between the predicted remaining useful life obtained by model_2 and the true remaining useful life as the loss function, fine-tune some model parameters of model_1, and then repeat the data relationship measurement module, the distribution difference calculation module, and the remaining useful life prediction module in sequence to obtain the final remaining useful life prediction model model_1_final. Finally, the test data Test is passed through model_1_final to obtain the corresponding remaining useful life prediction target value.

[0122] The specific implementation methods of each module correspond to the respective steps, and are not described in this invention.

[0123] The following further explains the beneficial effects that can be achieved by this invention in combination with specific application scenarios.

[0124] To verify the effectiveness of the proposed method in predicting the RUL under multiple working conditions using online insufficient unseen data, we conducted a series of experiments on the IEEE PHM Challenge 2012 bearing dataset and the XJTU-SY bearing dataset. The IEEE PHM Challenge 2012 bearing dataset (hereinafter referred to as PHM) was collected from the PRONOSTIA platform. The vibration signals of the bearings were measured by two accelerometers in the horizontal and vertical directions, and the sampling frequency was 25.6 kHz. The dataset contains 17 running failed bearings under three different working conditions, with different rotational speeds and radial force magnitudes. The XJTU-SY bearing dataset (hereinafter referred to as XJTU) was jointly collected by Xi'an Jiaotong University and Changxing Sumitomo Technology. The sampling content and frequency of this dataset are the same as those of PHM. This dataset contains 15 misaligned bearings under three different working conditions. Table 1 lists the detailed test conditions of these data.

[0125] Table 1 Experimental conditions of the IEEE PHM Challenge 2012 and XJTU - SY bearing datasets

[0126]

[0127] To illustrate the generality of the proposed method, we randomly select a working condition from the PHM dataset as source domain 1 (S1) and a working condition from the XJTU dataset as source domain 2 (S2) to perform the transfer prediction task. The target domain is another working condition randomly selected from the PHM and XJTU datasets. The test data does not belong to the same domain as the training data of the source domain and belongs to the same domain as part of the training data of the target domain. For the same task, data from two datasets is cross-selected for training, and these two datasets come from different devices. The differences between devices are generally larger than those within a device. Therefore, it is more challenging to verify the effectiveness of the model in this case. Table 2 lists the detailed experimental settings. In each task, we use two target domains and consider the multi-target domain problem to extract features. The data of the test set comes from a device with the same target domain but different bearings. Experiments have shown that effective information can be extracted from data of different devices and different bearings.

[0128] Table 2 Experimental Settings

[0129]

[0130] For the feature extractor, i.e., the backbone network, it adapts all the same parameter settings. The patch size is 8, the patch stride is 4, the dropout rate is 0.3, and the learning rate is the dynamic learning rate of OneCycleLR with an initial value of 0.0001. To verify the prediction accuracy of the proposed method, the root mean square error (RMSE) and the mean absolute error (MAE) are selected as evaluation metrics.

[0131]

[0132]

[0133] where n represents the number of samples, represents the true value of the i-th sample, is the predicted value of the i-th sample.

[0134] Figure 2 shows the RUL predicted values and their 95% confidence intervals for the four tasks in Table 2. The RUL prediction results obtained by the proposed method are very good and are very close to the actual RUL decline trajectory. From the perspective of the 95% confidence interval, the prediction uncertainty at each moment is very small. These results verify the effectiveness of the method in extracting domain-invariant features and promoting the transfer of fault degradation diversity, which is crucial for industrial maintenance.

[0135] This method is compared with three representative RUL prediction methods to analyze the tasks listed in Table 2. These methods include LSTM, Transferable CNN+MMDs, and TPDAN.

[0136] LSTM: A method that uses deep feature representation and adopts a Long Short-Term Memory (LSTM) neural network to predict the remaining useful life.

[0137] CNN+MMD: Transferable Convolutional Neural Network with Multicore Maximum Mean Discrepancy

[0138] TPDAN: A multi-source ensemble adversarial adaptation strategy with stage-weighted maximum mean discrepancy and ranking-based feature regularization.

[0139] All comparison methods used the best parameters provided in the original papers. The performance comparison is shown in Table 3. It should be noted that the RUL transfer prediction in this paper is a supervised prediction task, and the full-life cycle labeled data in the target domain is used for training. As shown in Table 3, the method proposed in this paper obtained the minimum RMSE and MAE in four tasks, indicating that this method significantly improved the prediction performance.

[0140] Table 3 Error result data

[0141]

[0142] As can be seen from the table, the proposed method has the smallest prediction error in all tasks. The error of our method is less than 50% of that of LSTM, which may be because this method cannot obtain satisfactory prediction results when there are differences in the data distributions of different machines. Although CNN+MMD adopts domain adaptation to extract domain-invariant features, different working conditions have different degradation trends, and using only a single prediction model is not sufficient to mine the specific degradation information of each source domain. Therefore, the performance of CNN+MMD is inferior to the proposed method, and the error is more than four times that of ours. TPDAN performs the worst in most tasks. The reason is that it trains multiple discriminators without considering multi-source and multi-target conditions.

[0143] Furthermore, this paper discusses the influence of some important hyperparameters in the model on the prediction performance through experiments. There are three main hyperparameters in the model we proposed: , and the weights in

[0144] and the learning rate. They will all have a certain impact on the results and are adjusted according to each task in this experiment. Hyperparameters with smaller errors are selected for testing. Here, we take Task 1 as an example to illustrate how to obtain the parameters. and The hyperparameter tuning results show that under combination 1, the RMSE and MAE of the model reach the optimal values, and our model finally selects this combination.

[0145] Table 4 Hyperparameter Tuning

[0146]

[0147] In addition, to verify the importance of certain modules in this method, ablation experiments were designed in the instance, including the entire model, the model without fine-tuning, the model without WMMD, and the model using only one target domain. To further verify the functions and importance of multi-level transfer learning and multiple objectives in the proposed method, we conducted three ablation experiments. The detailed description of the ablation experiments is shown in Table 5.

[0148] Table 5 Detailed Description of Ablation Experiments

[0149]

[0150] Still taking Task1 as an example, the remaining useful life prediction results are shown in Table 6, and each component in the proposed method is beneficial to RUL prediction.

[0151] Table 6 Performance Metrics of Task1 in Different Ablation Experiments

[0152]

[0153] Multi-level domain adaptation plays an important role in the proposed method, improving the ability of ordinary feature extraction. To verify this contribution, we calculated and visualized the experimental results of multi-source and multi-target domains. Due to space limitations, only the experimental results of Task 1 are plotted, as shown in Figure 3.

[0154] As shown in (a) of Figure 3, they are the experimental results of the features of two source domains and one mixed target domain after passing through the feature extractor and before fine-tuning, and the differences between these experimental results are very obvious. (b) in Figure 3 shows the situation after network fine-tuning. Now, the experimental result distributions of the source domain and mixed target domain features are more consistent.

[0155] As shown in (c) of Figure 3, they are the experimental results of the features from two target domains and the WMMD-free test set, and the differences between these experimental results are quite significant. (d) in Figure 3 is the feature distribution diagram of the mixed target domain and the test set, showing the consistency of its distribution.

[0156] Multi-level domain adaptation reduces the differences in the data distributions of different domains, and the extracted features are more universal, which is also the reason why the model achieves excellent results.

[0157] In this patent document, the provided embodiments and detailed descriptions are intended to illustrate the concepts, principles, and applications of the present invention, rather than to limit the scope of protection of the present invention. The scope of protection of the present invention should be determined according to the claims, and it should be understood that any technical improvements, substitutions, or changes that conform to the innovative spirit and essence of the present invention, as long as they do not exceed the scope of protection defined by the claims, should be regarded as within the scope of protection of the present invention. Those skilled in the art can make various forms of substitutions or adjustments under the inspiration of the present invention to adapt to different application scenarios and requirements without departing from the core idea and scope of protection of the present invention.

Claims

1. A method for predicting the remaining service life of a multi-objective domain bearing, characterized in that, It includes the following steps: S1. Randomly select data under different working conditions from the industrial process bearing operation-related dataset as the source domain data S Total , the target domain data T m , and the test data Test, ensuring that the test data is from a different domain from the source domain data and from the same domain as some of the target domain data; S2. In the offline stage, the source domain data S Total is filled and segmented to obtain patches, which are input into a 1D convolutional layer to obtain the input embedding X emb ; S3, Use ModernTCN to train X emb to obtain the model model_1; the ModernTCN includes stacked ModernTCN blocks, i.e., the backbone network, and a linear head with a flatten layer; S4. In the online stage, for the test data Test and the target domain data T m Use dynamic time warping (DTW) to align the sequences respectively and calculate the distances, and assign the initial weights of the target domain according to the distances; S5. Use the Maximum Mean Discrepancy (MMD) to measure the distribution difference between the target domain processed by DTW and the test data, and calculate the value of the initial Weighted Maximum Mean Discrepancy (WMMD). S6. Take the last two layers of the model model_1 as the remaining useful life predictor model_2, input the weighted mixed target domain, and obtain the predicted remaining useful life. S7. Use the sum of the value of WMMD and the mean square error between the predicted remaining useful life obtained by model_2 and the true remaining useful life as the loss function to fine-tune some model parameters of model_1, then repeat S4, S5, and S6 in sequence to obtain the final remaining useful life prediction model model_1_final. Finally, pass the test data Test through model_1_final to obtain the corresponding predicted target value of the remaining useful life.

2. The multi-objective domain bearing remaining service life prediction method according to claim 1, wherein: The Modern Temporal Convolutional Network (ModernTCN) block includes a Depthwise Convolution (DWConv) and two Convolutional Feed-Forward Networks (ConvFFN), namely ConvFFN1 and ConvFFN2. DWConv is responsible for independently learning the temporal information between tokens on each feature channel. The role of ConvFFN1 is to generate new feature representations for each variable, while the task of ConvFFN2 is to capture the inter-variable dependencies of each feature.

3. A method for predicting the remaining service life of a multi-target domain bearing according to claim 1, characterized in that: The specific implementation method of step S3 is as follows: (3.1) Input X emb into the backbone network to capture the interdependencies across time and variables, so as to learn informative representations The formula is as follows: Among them, M represents the number of variables included in the source domain data, N represents the number of patches, D represents the number of channels, and Backbone(·) is a stacked ModernTCN block. Among them, is the input of the i-th ModernTCN block. Block(·) represents the ModernTCN block. After passing through the backbone network, the final (3.2) Use a linear head with a flattening layer to obtain the final prediction result: where, is a predicted result of length T freely set, including M variables; the Flatten(·) operation refers to the layer that reshapes the final representation into a tensor Z with dimensions M×(D×N); Head(·) represents a linear projection layer that converts the final representation into the final predicted result; (3.3) In the offline stage, the calculation formula of the loss function of the model model_1 is as follows: where n represents the number of source domains, represents the true label.

4. A method for predicting the remaining service life of a multi-target domain bearing according to claim 1, characterized in that: The process of allocating weights using Dynamic Time Warping (DTW) in step S4 is as follows: (4.1) Dynamic Time Warping (DTW) is used to measure the similarity between two sequences. Given two time series, namely P = (p1, p2,..., p a ), and Q = (q1, q2,..., q b ), DTW(P, Q) is defined as the optimal path among all possible paths T; a path T is a sequence ((i1, j1), (i2, j2),..., (i t , j t )) where i1 = j1 = 1, i t = a, j t = b, a and b represent sequence numbers, and for each step k' ∈ {1,..., t - 1}, (i k'+1 , j k′+1 ) belongs to the set {(i k′ + 1, j k′ ), (i k′ , j k′ + 1), (i k′ + 1, j k′ + 1)}. The DTW formula is as follows: Among them, t represents the number of nodes on the optimal path, and d(·,·) is the Euclidean distance, which is a distance metric, and its calculation method is as follows: (4.2) Obtain the weighting coefficient by normalizing the DTW distance, and the calculation formula is as follows: Among them, TNs represents the number of labeled target domains, and r represents the r-th labeled target domain. represents the data of the labeled target domain embedded through the backbone network. represents the test data after being embedded through the backbone network.

5. The multi-objective domain bearing remaining service life prediction method according to claim 4, wherein: In step S5, WMMD is used to measure the dissimilarity of the domain adaptation distribution between the target domain and the test set using DTW to achieve feature-level transfer learning between the test data and multiple target domains. The calculation formula is as follows: Among them, L WMMD represents the value of WMMD, φ() is a mapping function that maps data to a high-dimensional feature space, and ||·|| H represents the norm of the Hilbert space.

6. The multi-objective domain bearing remaining service life prediction method according to claim 4, characterized in that: The mixed target domain in step S6 is defined as: Among them, represents the mixed target domain.

7. A multi-objective domain bearing remaining service life prediction method according to claim 5, characterized in that: The formula of the loss function between the true label and the target domain data label in the remaining useful life predictor model_2 is as follows: Among them, m represents the number of target domain data, represents the true label, represents the predicted value obtained by the predictor model_2.

8. A method for predicting the remaining service life of a multi-objective domain bearing according to claim 7, characterized in that: In step S7, based on randomly freezing some model parameters in step S3, other parameters use the pre-trained values, and then fine-tune the model model_1 to complete the model-based transfer learning between the second-level hierarchical source domain and multiple target domains. The formula of the overall optimization function for multi-level domain transfer learning is as follows: L Total = λL RUL + γL WMMD Among them, λ and γ are regularization parameters.

9. The method for predicting the remaining service life of a multi-target domain bearing according to claim 1, characterized in that: It also includes using the Root Mean Square Error (RMSE) and Mean Absolute Error (MAE) as evaluation indicators to evaluate the final prediction result.

10. A multi-objective domain bearing remaining useful life prediction system, characterized in that It includes the following modules: A data domain acquisition module, which is used to randomly select data under different working conditions from an industrial process bearing operation-related dataset as source domain data S Total Target domain data T m , and test data Test, ensuring that the test data is from a different domain from the source domain data and from the same domain as some of the target domain training data; Source domain data preprocessing module, which is used to preprocess the source domain data S in the offline stage Total The patches obtained by filling and segmenting are input into the 1D convolutional layer to obtain the input embedding X emb ; Offline feature extraction module, using ModernTCN to train X emb to obtain the model model_1; the ModernTCN includes stacked ModernTCN blocks, namely the backbone network, and a linear head with a flatten layer; A data relationship measurement module, which is used in the online stage to process the test data Test and the target domain data T m Using Dynamic Time Warping (DTW) to align sequences respectively and calculate distances, and assigning initial weights to the target domain according to the distances; The distribution difference calculation module is used to measure the distribution difference between the target domain after DTW processing and the test data by using the Maximum Mean Discrepancy (MMD), and calculate the value of the initial Weighted Maximum Mean Discrepancy (WMMD). The remaining useful life prediction module is used to use the last two layers of the model model_1 as the remaining useful life predictor model_2, and input the weighted mixed target domain to obtain the predicted remaining useful life. The final prediction module is used to take the sum of the value of WMMD and the mean square error between the predicted remaining useful life obtained by model_2 and the true remaining useful life as the loss function, fine-tune some model parameters of model_1, and then repeat the data relationship measurement module, the distribution difference calculation module and the remaining useful life prediction module in sequence to obtain the final remaining useful life prediction model model_1_final. Finally, the test data Test is passed through model_1_final to obtain the corresponding remaining useful life prediction target value.

Citation Information

Patent Citations

  • Gas consumption behavior anomaly detection method and system based on transfer learning

    CN115017970A

  • Method for predicting residual service life of rotary machine

    CN119129364A