Remaining life prediction method based on multi-source feature dynamic weighting sub-domain self-adaption

Through the multi-source feature dynamic weighted subdomain adaptive method, the problems of low prediction accuracy and insufficient generalization ability of the single-source domain adaptive method in variable operating condition systems are solved. Through cross-domain difference perception and joint distribution adaptive modules, a system remaining life prediction network is constructed, achieving higher prediction accuracy and generalization ability.

CN120632487APending Publication Date: 2025-09-12TAIYUAN UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510789343.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing single-source domain adaptive methods have problems with low prediction accuracy and insufficient model generalization ability in the remaining life prediction of variable operating condition systems. This is mainly because the source domain data cannot cover the diversity of the target domain and the target domain data is insufficient, resulting in the model being unable to make accurate predictions in the target domain.

Method used

A method based on dynamic weighted subdomain adaptation of multi-source features is adopted. Multi-source domain and target domain features are extracted through a deep feature fusion extraction network. The similarity in degradation trends is calculated and feature weights are assigned using the cross-domain difference perception layer. Feature alignment is performed in combination with the joint distribution adaptation module to construct a system remaining life prediction network.

Benefits of technology

The prediction accuracy of the target domain and the generalization ability of the model are improved. When facing multi-source domain data, it can balance the characteristics between different domains, retain diversity and rich information, adapt to the characteristics of the target domain, and achieve more accurate remaining life prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632487A_ABST
    Figure CN120632487A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of residual life prediction, and discloses a residual life prediction method based on multi-source feature dynamic weighting subdomain self-adaption, and the specific technical scheme is as follows: constructing a system residual life prediction network, extracting key information from input data of a plurality of source domains and target domains by a multi-domain feature extractor, the cross-domain difference sensing layer measures the similarity of the multi-source domain features and the target domain features in the degradation trend, dynamic weighting is carried out on the multi-source features to form weighted features containing the multi-source domain features, and the weighted features and the target domain features are sent to the joint distribution self-adaption module and used for distribution difference processing in a cross-domain feature alignment task. A joint loss function of a system residual life prediction network is constructed, a residual life prediction module is formed by using a bidirectional gating cycle unit network, features output by a test set feature extractor are processed, a final prediction result is output, and the prediction accuracy of a target domain is effectively improved by using multi-source domain data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remaining life prediction, and in particular relates to a remaining life prediction method based on multi-source feature dynamic weighted subdomain self-adaptation. Background Art

[0002] In recent years, domain adaptation methods, a core area of ​​transfer learning, have been widely used to predict the remaining life of systems with variable operating conditions. However, existing methods mostly rely on single-source domain analysis, which is inconsistent with actual operating conditions. In real-world conditions, remaining life prediction often involves data from multiple devices or operating environments. These devices may have different operating conditions, failure modes, and service lives, resulting in significant differences in their degradation characteristics and data distribution.

[0003] Single-source domain system remaining life prediction models face problems of low prediction accuracy and insufficient model generalization in actual industrial conditions. These problems mainly stem from two aspects:

[0004] First, single-source domain adaptation models usually rely on data from the source domain to learn features and knowledge. If the source domain data cannot cover the diversity of the target domain, the model may not be able to fully learn the features of the target domain, resulting in the model being unable to make accurate predictions in the target domain.

[0005] Secondly, in many practical applications, labeled data in the target domain is often scarce, especially when data from the target device or environment is difficult to obtain or label. The effectiveness of single-source domain adaptation usually relies on the existence of target domain data, but when target domain data is insufficient, model training and adaptation will be limited. Similarly, although single-source domain adaptation methods can make certain adjustments to the target domain using data from the source domain, their ability to adapt to different target domain characteristics is still limited. In particular, when the difference between the target and source domains is large, the model's generalization ability is significantly restricted. Summary of the Invention

[0006] In order to solve the technical problems existing in the prior art, the present invention provides a remaining life prediction method based on multi-source feature dynamic weighted sub-domain adaptation, which effectively utilizes multi-source domain data to improve the prediction accuracy of the target domain.

[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows: a method for predicting remaining life based on dynamic weighted subdomain adaptation of multi-source features is used to construct a system remaining life prediction network: a deep feature fusion extraction network is used to extract multi-source domain and target domain features, and the extracted features are sent to the cross-domain difference perception layer. Multiple source domain features are sequentially compared with target domain features to calculate the similarity in degradation trends. Feature weights of different source domains are assigned based on the similarity measurement results calculated by the cross-domain difference perception layer to form a new weighted feature that integrates multi-source domain features. A joint distribution adaptation module is used as a constraint to achieve alignment of conditional distribution and marginal distribution.

[0008] The specific prediction steps are as follows:

[0009] Step 1: A multi-domain feature extractor extracts key information from input data of multiple source domains and target domains;

[0010] Step 2: The cross-domain difference perception layer measures the similarity of the degradation trends between the multi-source domain features and the target domain features, and dynamically weights the multi-source features to form a weighted feature that includes the multi-source domain features;

[0011] Step 3: After fusing the multi-source domain features into dynamic weighted features, the weighted features and target domain features are sent to the joint distribution adaptation module to handle distribution differences in the cross-domain feature alignment task.

[0012] Step 4: Construct a joint loss function for the system's remaining life prediction network, use a bidirectional gated recurrent unit network to form a remaining life prediction module, process the features output from the test set feature extractor, and output the final prediction results.

[0013] In step 1, the multi-domain feature extractor consists of two parts: a multi-scale convolutional neural network and a multi-head self-attention mechanism. The multi-scale convolutional network mainly extracts features in local areas, while the multi-head self-attention mechanism models the whole world. It enhances feature expression by considering the global context, and complements each other by their respective areas of expertise to better complete the feature extraction task. The preprocessed frequency domain data is input into three convolutional networks arranged at different scales to extract local features in different dimensions. The efficiency of feature extraction is improved by parallel convolution operations. The feature map after convolution operation can be expressed as:

[0014]

[0015] Where, Represents the weight in the current convolutional layer operation, Represents the bias in the current convolutional layer operation, represents the feature mapping function, Represents the feature map output by the current convolutional layer.

[0016] pass The operation performs additive fusion on features extracted at different scales, comprehensively retaining the unique features extracted by each convolutional network at its own scale, providing richer prediction features. The fused features contain more comprehensive degradation information, which helps improve the accuracy of prediction. The fusion calculation is shown below:

[0017]

[0018] Where, represents the output of the fusion layer, Indicates the number of channels, 、 、 Represent the outputs of the three convolutional networks respectively.

[0019] The features that will be obtained Input to the multi-head attention encoder layer to capture global features. The multi-head attention is composed of multiple single-head attentions, each of which focuses on different relevance, including the query vector , key vector Sum value vector Three inputs. and Perform scaled dot product to get the attention matrix, and then Multiply to get the weighted data sample. The calculation formula of scaled dot product is as follows:

[0020]

[0021] Where, Finally, the multi-head attention mechanism is calculated as follows:

[0022]

[0023]

[0024] Where, represents the number of heads in the multi-head self-attention mechanism, 、 、 、 They are all learnable parameters, and because they focus on different correlations, the weights obtained during training are also different.

[0025] In step 2, the cross-domain difference perception layer includes two parts: local similarity measurement and global similarity measurement. The local similarity measurement measures the similarity between source domain and target domain samples in the feature space. Cosine similarity shows good performance in processing features with large differences in data feature distribution. Cosine similarity is selected to measure the local similarity of source domain and target domain samples. Its formula is expressed as follows:

[0026]

[0027] Where, represents the dot product operation, Indicates the The modulus length of the source domain, The modulus of the target domain feature;

[0028] Local similarity metrics usually focus on the short-distance similarity between data points, focusing on the similarity between adjacent data points or the feature matching of local areas. However, local metrics cannot capture the differences in global distribution. If the data distribution of the source domain and the target domain is significantly different overall, the local metrics may not accurately reflect this difference. While considering local similarity, KL divergence is used to measure the source domain. The characteristic distribution of and target domain feature distribution The differences between:

[0029]

[0030] Each source domain feature and target domain features The KL divergence of is expressed as:

[0031] (8)

[0032] Different source domains may contribute unevenly to the target domain task. Some source domains may be more similar to the target domain in some features, while other source domains may contribute less. The adaptive weight mechanism can automatically learn the contribution of each source domain and dynamically adjust the weight of the source domain features according to the characteristics of the target domain. During the training process, the model adjusts the weight of each feature according to its contribution. It will be automatically adjusted by the gradient descent method, and the weight update formula is:

[0033] (9)

[0034] in, represents the learning rate, is the loss function About weight The gradient, Indicates the The source domain is The weight at the iteration;

[0035] Through the adaptive weight mechanism, the model can automatically adjust the importance of features based on feedback during training, thereby allocating weights more accurately.

[0036] After weight optimization, each source domain feature The calculated weight The weighted source domain features are expressed as follows:

[0037] (10)

[0038] Where, represents the weighted source domain features, Source domain The weight coefficient of

[0039] Finally, all the weighted features of the source domain are fused into a new feature :

[0040] (11)

[0041] Through the cross-domain difference perception layer, the multi-source domain features are weighted and fused, which can combine multiple source domains The dynamic weighting mechanism adjusts weights in real time based on the similarity between source and target domain features. This allows the model to maintain a balance between domains when dealing with multi-source data, preventing performance degradation caused by overly prominent features in a particular source domain or a significant divergence from the target domain. Therefore, the final fused features not only preserve the diversity and richness of the source domain information, but also prioritize those source domain features that contribute most to the target task.

[0042] In step three, the joint distribution adaptation module effectively extracts the feature relationship between the source domain and the target domain by fusing the convolutional feature enhancement strategy, and introduces label discretization in the feature alignment process. Specifically, the true RUL labels of the source domain and the pseudo labels of the target domain are first discretized to generate a discrete label space as a constraint for feature alignment. In this way, the marginal distribution and the conditional distribution are effectively aligned at the same time, and a more robust joint distribution alignment mechanism is established between the source domain and the target domain. First, the features extracted by the deep feature fusion network are weighted using the CBAM attention mechanism. CBAM is a hybrid attention mechanism module that combines space and channel attention mechanisms.

[0043] In the channel attention part, firstly perform global average pooling and global maximum pooling operations on the input data, and output two sizes of The vectors are input into two layers of MLP, the first layer has neurons, of which is the reduction rate, the second layer has The parameters of these two MLP layers are shared. Next, the outputs of the MLPs are element-wise summed. To ensure that the attention weights are between 0 and 1, a sigmoid activation function is applied to generate channel attention weights. The activation function is expressed as follows:

[0044] (12)

[0045] The spatial attention part then concatenates the features after maximum pooling and average pooling along the channel dimension to obtain a feature with contextual information of different scales. Its operation part includes a convolution layer and a Softmax activation function, which helps to optimize the distribution of features along the one-dimensional space and more effectively capture the structural features in the data.

[0046] After the features extracted by CBAM attention are compressed into 3C through the maximum pooling layer, the extracted features are mapped to the corresponding remaining lifespan through the regressor. The mean square error between the predicted RUL and the true RUL is calculated as the loss function, and the prediction loss is described as:

[0047] (13)

[0048] Where, represents the number of samples in the source domain, represents the predicted lifespan, Indicates the real lifespan.

[0049] RUL prediction is a typical regression problem, and its label ranges from This is a continuous variable, which poses a great challenge to characterizing the differences between the conditional distributions of remaining life. Therefore, before performing conditional distribution adaptation, the present invention discretizes the regression labels using formula (13) to simplify the modeling complexity of the conditional distribution. Through this formula, the discretized source domain RUL labels and the RUL pseudo-labels generated in the target domain together constitute the conditional distribution basis of the features, thus laying a solid foundation for conditional distribution adaptation.

[0050] (14)

[0051] Where, Indicates the number of degradation stages into which the remaining lifetime label is divided, .

[0052] Multilinear multiplication mapping is used to simulate the interaction between different features and labels to capture the interaction between features and life prediction results and the multimodal structure of data distribution. The joint distribution adaptive training process can be expressed as:

[0053] (15)

[0054] Where, Represents the data in the source domain, represents the data of the target domain, represents the marginal distribution of the source domain, represents the marginal distribution of the target domain, Representation domain discriminator, used to distinguish data, through resistance training, effectively realizes adversarial learning of variable working condition tasks.

[0055] In step 4, the joint loss function is composed of the weighted fusion loss of multi-source domain features and the regression prediction error loss. , joint distribution adaptive training loss, where the multi-source feature weighted fusion loss is defined as:

[0056] (16)

[0057] Where, represents the cosine similarity measure between source domain features and target domain features, Represents the KL divergence measure between source domain features and target domain features;

[0058] The regression loss is used to minimize the error between the predicted life expectancy and the actual life expectancy, assuming To predict life expectancy, Represents the true RUL value, and the regression loss is expressed as:

[0059] (17)

[0060] Joint distribution adaptation training is expressed as:

[0061] (18)

[0062] The total loss of the remaining life prediction network model of the subdomain adaptive system based on multi-source characteristics dynamic weighting is:

[0063] (19)

[0064] in, , , Represents the weight coefficient of each part.

[0065] In step 4, the bidirectional gated recurrent unit network is composed of two independent gated recurrent units stacked together. By extracting information from both the forward and reverse time axes and integrating them through mutual interaction, a more complex and powerful network structure is constructed. This bidirectional structure enables the bidirectional gated recurrent unit network to simultaneously consider both forward and backward information in the sequence, thereby more comprehensively capturing the contextual dependencies in the sequence. Compared with traditional unidirectional gated recurrent units, the bidirectional gated recurrent unit network can better handle long-term dependencies and is particularly suitable for variable and complex time series tasks.

[0066] The core of the gated recurrent unit is to control the update gate through two gating mechanisms and reset gate The internal operation is as follows:

[0067] (20)

[0068] (twenty one)

[0069] (twenty two)

[0070] (twenty three)

[0071] Where, represents the output of the update gate, represents the output of the reset gate, represents the weight matrix of the update gate, represents the weight matrix of the reset gate, The weight matrix representing the candidate hidden state, is the hidden state at the previous moment, represents the activation function, tanh is the hyperbolic tangent activation function, Indicates the output at the current moment;

[0072] The output of the remaining life prediction module is expressed as:

[0073] (twenty four)

[0074] in, represents the weight of the output layer, represents the bias of the output layer.

[0075] Negative transfer typically occurs when the data distributions of the source and target domains differ significantly. Simply using transfer learning methods cannot effectively improve the prediction accuracy of target domain tasks. In this case, the source domain knowledge borrowed by transfer learning does not have a positive impact on the learning tasks in the target domain and may even cause interference that degrades the model's performance. This paper proposes a system remaining useful life prediction based on dynamic weighted subdomain adaptation of multi-source features. This method optimizes and adjusts features from multiple source domains through a dynamic weighting strategy to adapt to the characteristics of the target domain, achieving more accurate RUL prediction.

[0076] Based on the system remaining life prediction method based on a multi-constraint deep feature fusion network under variable working conditions, in order to effectively improve the utilization rate of common features in multiple source domains, the present invention proposes a system remaining life prediction network based on dynamic weighted multi-source sub-domain adaptation. The network first uses a deep feature fusion extraction network to extract multi-source domain and target domain features, and then sends the extracted features to the cross-domain difference perception layer. In this layer, multiple source domain features and target domain features are compared in sequence to calculate the similarity in degradation trends. Then, based on the similarity measurement results calculated by the cross-domain difference perception layer, feature weights are assigned to the features of different source domains to form new weighted features that fused multiple source domain features. Finally, the joint distribution adaptation module is used as a constraint to achieve alignment of the conditional distribution and the marginal distribution, reduce distribution differences, and construct a multi-source feature dynamic weighted sub-domain adaptive remaining life prediction model to improve the generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1 Schematic diagram of system remaining life prediction based on dynamic weighted sub-domain adaptation of multi-source characteristics.

[0078] Figure 2 Schematic diagram of the feature extraction module based on deep feature fusion.

[0079] Figure 3 Schematic diagram of the cross-domain difference perception layer.

[0080] Figure 4 Schematic diagram of the joint distribution adaptation module based on convolution enhancement.

[0081] Figure 5 Schematic diagram of the Bi-GRU network structure.

[0082] Figure 6 Schematic diagram of the GRU network structure.

[0083] Figure 7 This is the structural diagram of the experimental platform.

[0084] Figure 8 This is the original vibration signal data diagram of bearings 1-5. Figure 8 (a) is the horizontal vibration signal data diagram, Figure 8 (b) is the vertical vibration signal data diagram.

[0085] Figure 9 This is the original vibration signal data diagram of bearing 2-2. Figure 9 (a) is the horizontal vibration signal data diagram, Figure 9 (b) is the vertical vibration signal data diagram.

[0086] Figure 10 This is a schematic diagram of the feature distribution visualization results. Figure 10 (a) is the probability density distribution diagram before multi-domain feature migration, Figure 10 (b) is the probability density distribution diagram after weighted multi-source domain feature migration.

[0087] Figure 11 This is the prediction result curve of bearing 1-1.

[0088] Figure 12 This is the predicted result curve of bearing 2-1.

[0089] Figure 13 This is the predicted result curve of bearing 3-2. DETAILED DESCRIPTION

[0090] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0091] Based on a multi-source feature dynamic weighted subdomain adaptive RLS prediction method, a system RLS prediction network is constructed. This network first extracts features from multiple source and target domains using a deep feature fusion extraction network. The extracted features are then fed into a cross-domain difference perception layer, where multiple source and target domain features are sequentially compared to calculate the similarity in their degradation trends. Next, weights are assigned to features from different source domains based on the similarity metric calculated by the cross-domain difference perception layer, forming a new weighted feature that fuses features from multiple source domains. Finally, a joint distribution adaptation module is used as a constraint to align the conditional distribution and marginal distribution, reducing distribution differences. This results in a multi-source feature dynamic weighted subdomain adaptive RLS prediction model, improving its generalization capability.

[0092] like Figure 1As shown, this method extracts features from multiple source and target domains and performs feature similarity learning in the cross-domain difference perception layer, effectively fusing them into weighted source domain features. Furthermore, a convolution-enhanced joint distribution adaptation module simultaneously reduces the marginal and conditional distributions between the source and target domains, enhancing the model's predictive ability for target domain data and its overall prediction accuracy and generalization performance. This method consists of three core components: data acquisition and processing, offline training and online testing, and evaluation of remaining life prediction results.

[0093] Among them, the adversarial training network contains four core parts, namely, source domain feature extractor based on deep feature fusion , target domain feature extractor , domain discriminator and the remaining life prediction module R, which work together to achieve feature extraction of variable operating condition data, minimization of domain differences, and accurate estimation of remaining life.

[0094] The multi-domain feature extractor is responsible for extracting key information from input data in multiple source and target domains. The feature extractor consists of two parts: a multi-scale convolutional neural network and a multi-head self-attention mechanism. The multi-scale convolutional network mainly extracts features from local areas, while the multi-head self-attention mechanism performs modeling on a global scale and enhances feature expression by considering the global context. They complement each other in their respective areas of expertise to better complete the feature extraction task.

[0095] like Figure 2 As shown in the figure, by designing a feature extraction module based on deep feature fusion, we not only avoid the slow convergence caused by excessive computation in deep networks, but also improve the integrity of input information and the ability to communicate information across channels. At the same time, we successfully achieve the fusion of local and global features of the data, fully tapping into the interactive capabilities of features while retaining the original information, thus improving model performance. The preprocessed frequency domain data is input into three convolutional networks arranged at different scales to extract local features in different dimensions. The efficiency of feature extraction is improved by parallel convolution operations. The feature map after convolution operation can be expressed as:

[0096]

[0097] Where, Represents the weight in the current convolutional layer operation, Represents the bias in the current convolutional layer operation, represents the feature mapping function, Represents the feature map output by the current convolutional layer.

[0098] pass The operation performs additive fusion on features extracted at different scales, comprehensively retaining the unique features extracted by each convolutional network at its own scale, providing richer prediction features. The fused features contain more comprehensive degradation information, which helps improve the accuracy of prediction. The fusion calculation is shown below:

[0099]

[0100] Where, represents the output of the fusion layer, Indicates the number of channels, 、 、 Represent the outputs of the three convolutional networks respectively.

[0101] The features that will be obtained Input to the multi-head attention encoder layer to capture global features. The multi-head attention is composed of multiple single-head attentions, each of which focuses on different relevance, including the query vector , key vector Sum value vector Three inputs. and Perform scaled dot product to get the attention matrix, and then Multiply to get the weighted data samples, and the calculation formula of the scaled dot product is as follows:

[0102]

[0103] Where, Finally, the multi-head attention mechanism is calculated as follows:

[0104]

[0105]

[0106] Where, represents the number of heads in the multi-head self-attention mechanism, 、 、 、 They are all learnable parameters, and because they focus on different correlations, the weights obtained during training are also different.

[0107] like Figure 3 It is shown that the purpose of the cross-domain difference perception layer is to measure the similarity between the degradation trends of multi-source domain features and target domain features, and to dynamically weight the multi-source features to form a new weighted feature that includes multi-source domain features.

[0108] The cross-domain difference perception layer includes two parts: local similarity measurement and global similarity measurement. The local similarity measurement measures the similarity between source domain and target domain samples in the feature space. Cosine similarity shows good performance in processing features with large differences in data feature distribution. Therefore, the present invention selects cosine similarity to measure the local similarity of source domain and target domain samples. Its formula is expressed as follows:

[0109]

[0110] Where, represents the dot product operation, Indicates the The modulus length of the source domain, Indicates the modulus of the target domain feature.

[0111] Local similarity metrics usually focus on the short-distance similarity between data points. They focus on the similarity between adjacent data points or the feature matching of local areas. However, local metrics cannot capture the differences in global distribution. If the data distribution of the source domain and the target domain is significantly different overall (i.e., the global distribution is different), local metrics may not accurately reflect this difference. Therefore, while considering local similarity, KL divergence is used to measure the source domain. The characteristic distribution of and target domain feature distribution The differences between:

[0112]

[0113] Then for each source domain feature , which is consistent with the target domain characteristics The KL divergence of can be expressed as:

[0114]

[0115] Different source domains may contribute unevenly to the target domain task. Some source domains may be more similar to the target domain in some features, while other source domains may contribute less. The adaptive weight mechanism can automatically learn the contribution of each source domain and dynamically adjust the weight of the source domain features according to the characteristics of the target domain. During the training process, the model adjusts the weight of each feature according to its contribution. It will be automatically adjusted by the gradient descent method, and the weight update formula is:

[0116]

[0117] in, represents the learning rate, is the loss function About weight The gradient, Indicates the The source domain is The weight at the iteration.

[0118] Through the adaptive weight mechanism, the model can automatically adjust the importance of features based on feedback during training, thereby allocating weights more accurately.

[0119] After weight optimization, each source domain feature The calculated weight The weighted source domain features can be expressed by the following formula:

[0120]

[0121] In the formula, represents the weighted source domain features, Source domain The weight coefficient of .

[0122] Finally, all the weighted features of the source domain are fused into a new feature :

[0123]

[0124] Through the cross-domain difference perception layer, the multi-source domain features are weighted and fused, which can combine multiple source domains The dynamic weighting mechanism fully combines the useful information of the source domain and the target domain, adjusting the weights in real time based on the similarity between the source domain features and the target domain. This allows the model to balance the different domains when faced with multi-source domain data, preventing performance degradation caused by overly prominent features in one source domain or excessive differences from the target domain. Therefore, the final fused features not only retain the diversity and rich information of the source domains, but also enable the model to place greater emphasis on those source domain features that contribute most to the target task.

[0125] After fusing the multi-source domain features into dynamic weighted features, the weighted features and the target domain features are sent to the joint distribution adaptation module for distribution difference processing in the cross-domain feature alignment task. This module effectively extracts the feature relationship between the source and target domains by fusing the convolutional feature enhancement strategy, and introduces label discretization in the feature alignment process. Specifically, the true RUL labels of the source domain and the pseudo labels of the target domain are first discretized to generate a discrete label space as a constraint for feature alignment. In this way, the marginal distribution and the conditional distribution are effectively aligned at the same time, and a more robust joint distribution alignment mechanism is established between the source and target domains.

[0126] The joint loss function of the remaining life prediction network model of the adaptive system based on the dynamic weighted sub-domain of multi-source features consists of three parts: multi-source domain feature weighted fusion loss, regression prediction error loss , joint distribution adaptive training loss. Among them, the multi-source feature weighted fusion loss is defined as:

[0127]

[0128] Where, and They represent the cosine similarity metric and KL divergence metric between source domain features and target domain features, respectively.

[0129] The regression loss is used to minimize the error between the predicted life expectancy and the actual life expectancy, assuming To predict life expectancy, Represents the true RUL value, then the regression loss is expressed as:

[0130]

[0131] The training process of the joint distribution adaptation training loss is expressed as:

[0132]

[0133] In summary, the total loss of the remaining life prediction network model of the subdomain adaptive system based on multi-source features dynamic weighting is:

[0134] (15)

[0135] in, , , Represents the weight coefficient of each part, which is used to control the contribution of each loss function to the final loss. Figure 3As shown in the figure, in the online remaining life prediction stage, a bidirectional gated recurrent unit network is used to form a remaining life prediction module to process the features output from the test set feature extractor and output the final prediction results.

[0136] The Bi-GRU (Bidirectional Gated Recurrent Unit) network model is not simply composed of two independent GRU networks (Gated Recurrent Units). Instead, it extracts information from both the forward and reverse time axes and integrates them through mutual interaction, thereby constructing a more complex and powerful network structure. This bidirectional structure enables the Bi-GRU to simultaneously consider both forward and backward information in the sequence, thereby more comprehensively capturing the contextual dependencies in the sequence. Compared with the traditional unidirectional GRU, the Bi-GRU can better handle long-term dependencies and is particularly suitable for variable and complex time series tasks.

[0137] like Figure 4 As shown in the figure, the present invention performs joint distribution alignment on marginal distribution and conditional distribution based on the convolution-enhanced joint distribution adaptation module. The joint distribution adaptation CBAM module, regressor, discretization module, dropout layer and fully connected layer are composed. First, the features extracted by the deep feature fusion network are weighted using the CBAM attention mechanism. CBAM is a hybrid attention mechanism module that combines spatial and channel attention mechanisms. In the channel attention part, the input data is first subjected to global average pooling and global maximum pooling operations, and two outputs of size The vectors are input into two layers of MLP, the first layer has neurons, of which is the reduction rate, the second layer has The parameters of these two layers of MLP are shared. Next, the outputs of the MLP are summed element-wise. To ensure that the attention weights are between 0 and 1, a Sigmoid activation function is applied to generate channel attention weights. The activation function is expressed as follows:

[0138] (16)

[0139] The spatial attention part then concatenates the features after maximum pooling and average pooling along the channel dimension to obtain a feature with contextual information of different scales. Its operation part includes a convolution layer and a Softmax activation function, which helps to optimize the distribution of features along the one-dimensional space and more effectively capture the structural features in the data.

[0140] After the features extracted by CBAM attention are compressed into 3C through the maximum pooling layer, the extracted features are mapped to the corresponding remaining life through the regressor, and the mean square error between the predicted RUL and the true RUL is calculated as the loss function. The prediction loss is described as:

[0141] (17)

[0142] Where, represents the number of samples in the source domain, represents the predicted lifespan, Indicates the real lifespan.

[0143] RUL prediction is a typical regression problem. Its label ranges from This is a continuous variable, which poses a great challenge to characterizing the differences between the conditional distributions of remaining life. Therefore, before performing conditional distribution adaptation, the present invention discretizes the regression labels using formula (13) to simplify the modeling complexity of the conditional distribution. Through this formula, the discretized source domain RUL labels and the RUL pseudo-labels generated in the target domain together constitute the conditional distribution basis of the features, thus laying a solid foundation for conditional distribution adaptation.

[0144] (18)

[0145] Where, Indicates the number of degradation stages into which the remaining lifetime label is divided, .

[0146] Multilinear multiplication mapping is used to simulate the interaction between different features and labels to capture the interaction between features and life prediction results and the multimodal structure of data distribution. The joint distribution adaptive training process can be expressed as:

[0147] (19)

[0148] Where, and Represent the data of the source domain and the target domain respectively, represents the features extracted from the source domain, represents the features extracted by the target domain feature extractor, represents the marginal distribution of the source domain, represents the marginal distribution of the target domain, Representation domain discriminator, used to distinguish data, through resistance training, effectively realizes adversarial learning of variable working condition tasks.

[0149] like Figure 5 The core of GRU is to control the flow of information through two gating mechanisms: update gate and reset gate , GRU decides how much past information to retain and when to forget information through the control of these two gates.

[0150] Its internal operation is as follows:

[0151] (20)

[0152] (twenty one)

[0153] (twenty two)

[0154] (twenty three)

[0155] Where, represents the output of the update gate, represents the output of the reset gate, represents the weight matrix of the update gate, represents the weight matrix of the reset gate, The weight matrix representing the candidate hidden state, is the hidden state at the previous moment, represents the activation function, tanh is the hyperbolic tangent activation function, Indicates the output at the current moment.

[0156] Finally, the output of the remaining life prediction module can be expressed as:

[0157] (twenty four)

[0158] in, represents the weight of the output layer, represents the bias of the output layer.

[0159] Negative transfer typically occurs when the data distributions of the source and target domains differ significantly. Simply using transfer learning methods cannot effectively improve the prediction accuracy of target domain tasks. In this case, the source domain knowledge borrowed by transfer learning fails to have a positive impact on the learning task in the target domain and may even cause interference, degrading model performance. This paper proposes a system remaining useful life prediction method based on dynamic weighted subdomain adaptation of multi-source features. This method optimizes and adjusts features from multiple source domains through a dynamic weighting strategy to adapt to the characteristics of the target domain, achieving more accurate RUL prediction. The training process of the proposed method is shown in Table 1.

[0160] Table 1 Multi-source feature dynamic weighted subdomain adaptive network model algorithm

[0161] In order to evaluate the effectiveness of the system remaining life prediction method based on multi-source feature dynamic weighted subdomain adaptation, the XJTU-SY bearing dataset was selected for prediction analysis. Figure 7 As shown in the figure, the experimental platform consists of several key components including an AC motor, a hydraulic loading system, a speed controller and test bearings. The platform conducts accelerated degradation experiments on bearings under various operating conditions and fully obtains data on bearings from operation to failure.

[0162] The platform conducted 15 degradation experiments under three operating conditions, obtaining full-lifecycle vibration data for 15 LDK UER204 rolling bearings, as shown in Table 2. The experiments were conducted under three different operating conditions: 35 Hz, 12 kN; 37.5 Hz, 11 kN; and 40 Hz, 10 kN. The experiments involved four types of faults under three operating conditions: inner ring wear, cage fracture, and outer ring fracture and wear. This diversity of fault types simulates a wide range of potential failure scenarios in real industrial environments, thereby comprehensively verifying the effectiveness and robustness of the proposed method.

[0163] Table 2 XJTU-SY bearing full life cycle information

[0164] To verify the domain generalization ability and transfer performance of the model under different working conditions, three transfer tasks were designed, covering multi-working condition transfer between the source domain and the target domain: Task 1 is to migrate data from working conditions 2 and 3 to working condition 1;

[0165] Task 2 involves migrating data from working conditions 1 and 3 to working condition 2; task 3 involves migrating data from working conditions 1 and 2 to working condition 3. These tasks aim to evaluate the model's generalization capabilities across different working conditions, as well as its ability to improve target domain prediction performance when there are differences in the distribution of the source and target domains. In each task, bearing data from multiple source domains is selected for training, and some bearing data from the target domain is used for auxiliary training. Finally, the model performance is verified on test bearings in the target domain. This design simulates complex real-world scenarios under multiple working conditions, providing a comprehensive experimental basis for evaluating the model's generalization capabilities and robustness. The specific details are shown in Table 3.

[0166] Table 3 Three domain generalization tasks

[0167] at last, Figure 8 and 9 The following are the original vibration signal data of bearings under some working conditions. Bearing 1-5 under 35Hz12kN working condition and bearing 2-2 under 37.5Hz11kN working condition are taken as examples. These data will provide important basis for subsequent analysis.

[0168] This paper uses the root mean square error (RMSE) and mean absolute error (MAE) as the main evaluation indicators to comprehensively measure the accuracy and stability of the prediction method, thereby ensuring the reliability of the prediction results. These two indicators are evaluated by calculating the error between the predicted remaining life point estimate and the true value. The specific calculation formula is as follows:

[0169] (25)

[0170] (26)

[0171] Where, represents the point estimate of the forecast, represents the true value, is the total number of test set samples; when the values ​​of these two evaluation indicators decrease, the error between the model's predicted point estimate and the true value also decreases, further reflecting the model's higher prediction accuracy.

[0172] When constructing a network model for predicting the remaining life of a system under variable operating conditions based on adversarial training, the model parameters were set to ensure that the model can demonstrate excellent performance in the remaining life prediction task, and the optimizer was selected based on evaluation indicators. The specific model parameters are shown in Table 4.

[0173] Table 4 Model parameter information

[0174] Model parameter selection:

[0175] A. Feature extraction module parameter settings:

[0176] The multi-scale convolution kernels are 1×1, 3×3, and 5×5 with a stride of 1 and 32, 64, and 128 channels, respectively. The activation function is ReLU, and max pooling is used in the pooling layer. The multi-head self-attention mechanism uses six attention heads, each with a dimension of 64 and a dropout rate of 0.3. By deeply fusion-enhancing feature extraction modules, the model can extract information from multiple scales and simultaneously learn the internal relationships of the data from multiple perspectives. This helps enhance the comprehensiveness and diversity of feature representation while avoiding overfitting, thereby improving overall predictive and generalization capabilities.

[0177] B. Cross-domain difference perception layer:

[0178] The cross-domain difference perception layer calculates the similarity between multiple source domain features and target domain features. The local similarity is calculated by cosine similarity, and the global similarity is calculated by KL divergence. Based on the calculated similarity, the gradient descent method is used to adaptively assign weights to the source domain features, with a learning rate of 0.001. Finally, the weighted features are fused, and the weight range is .

[0179] C. Joint distribution adaptation module:

[0180] The joint distribution adaptation module discretizes the remaining lifetime into 12 degradation stages and multiplies them with the weighted features and target domain features into the domain discriminator. The domain discriminator consists of two fully connected layers, with the first layer having 128 nodes and the second layer having 64 nodes. It uses the ReLU activation function, and the output layer is a binary sigmoid output with a dropout rate of 0.5. The Adam optimizer is used, with a learning rate of 0.0002 and a batch size of 32.

[0181] (2) Online life prediction:

[0182] In the online remaining life prediction, two Bi-GRUs are used to construct the remaining life prediction module. The hidden layer of each Bi-GRU is set as follows:

[0183] The first hidden layer consists of 64 BiGRU units that return the entire sequence and use a dropout rate of 0.3 to introduce randomness and prevent overfitting.

[0184] The second hidden layer consists of 64 BiGRU units, which return the entire sequence and continue to use a dropout rate of 0.3 to further enhance the robustness of the model.

[0185] The third hidden layer: the flattening layer flattens the output of the BiGRU layer in order to integrate the temporal features into a form suitable for subsequent processing.

[0186] The fourth hidden layer has 128 nodes and a dropout rate of 0.5, which helps to further reduce overfitting and enhance the generalization ability of the model.

[0187] Finally, a one-dimensional output layer is used to predict the remaining life and output a continuous value representing the remaining life of the system.

[0188] (3) Feature distribution visualization:

[0189] Taking Task 1 as an example, we also use t-SNE technology to show the changes in the feature probability density distribution before and after migration. Figure 10As shown in the figure, before feature migration, there are significant differences in the feature distributions between the multi-source domain and the target domain. The peak distribution is dispersed and the overlap area is small, indicating that there is a significant distribution shift between the domains. After weighted multi-source domain feature migration, the feature distributions of each domain tend to be consistent, the overlap of density curves is significantly enhanced, and the feature alignment effect is significant, effectively alleviating the problem of inconsistent distribution between domains. This shows that the adopted multi-domain feature dynamic weighted migration strategy can improve the feature consistency between the multi-source domain and the target domain, providing a good foundation for improving the performance of the model in the target domain.

[0190] After dynamic weighted fusion of multi-source features, the differences between feature distributions are significantly reduced, and the fused features learned in the source domain tend to be similar to the features of the target domain. The method proposed in this invention effectively eliminates the differences in cross-domain feature distributions.

[0191] In order to verify the effectiveness of the system remaining life prediction network based on multi-source feature dynamic weighted sub-domain adaptation, the results of the three migration tasks proposed in this invention and the results of the single-source domain migration model were visualized and analyzed. The visualization results are shown in Figure 2. Figure 11 、 Figure 12 and Figure 13 shown.

[0192] from Figure 11 、 Figure 12 and Figure 13 It can be seen that the model proposed in the present invention is overall superior to the single-source domain migration model in the remaining life prediction task, and is closer to the trend of the real life. Specifically, the model of the present invention can more accurately fit the real life curve in the early and middle stages, with a smaller prediction error; in long-term prediction tasks, it shows higher stability and robustness, and can better cope with changes under complex working conditions. In contrast, the single-source domain migration model has a larger prediction error in the middle and late stages, and the fluctuation is significant, indicating that the model of the present invention effectively improves the accuracy and generalization ability of cross-domain remaining life prediction through dynamic weighting of multi-source domain features and joint distribution adaptation.

[0193] In order to demonstrate the superiority of the present invention, the present invention was compared with other classic network models, including the DCAE+CNN network proposed by Wang et al., the CAE+RNN model proposed by Xu et al., and representative single-source domain adaptation methods, such as DANN and the BiLSTM and attention mechanism-based method proposed by Dong et al. In addition, it was also compared with multi-source domain adaptation methods, including the multi-source domain adaptation method proposed by Liu et al. and the MDAN method proposed by Ding et al. Table 5 shows the average error comparison results of the present invention and other network models under the two indicators of RMSE and SCORE on the XJTU-SY bearing dataset. It can be seen from the data in the table that the method proposed by the present invention shows better prediction effect under the same test bearing conditions, which further verifies the effectiveness and reliability of the method.

[0194] Table 5 Comparison of RMSE values ​​and SCORE functions with other models

[0195] As can be seen from Table 5, the single-source domain to single-target domain method has certain limitations when dealing with equipment remaining life prediction under complex working conditions. It is relatively inadequate when dealing with differences in cross-domain feature distributions, and both prediction accuracy and robustness need to be improved. The method proposed in this paper uses a dynamic weighted multi-source feature fusion method to fully utilize multi-source domain features, adaptively adjust feature weights, and effectively align cross-domain feature distributions. Experimental results show that this method exhibits significant advantages in both RMSE and SCORE indicators, and demonstrates excellent adaptability and stability under complex working conditions.

[0196] In response to the problems of low model training accuracy and poor generalization ability of the single-source domain cross-domain remaining life prediction method when facing the diversity of target domain features and insufficient target domain data, the present invention proposes a dynamic weighted sub-domain adaptive remaining life prediction method based on multi-source features. This method dynamically adjusts the weights of multi-source domain features by designing a strategy for perceiving the difference between multi-source domain and target domain features, thereby improving the generalization ability of the model. In order to improve the prediction accuracy, a bidirectional gated recurrent unit is introduced into the remaining life prediction module. Comparative experiments were conducted on the XJTU-SY bearing dataset, and the results showed that the method proposed in the present invention showed significant advantages in both indicators. Compared with traditional single-source domain and multi-source domain methods, the model proposed in the present invention not only improves the prediction accuracy, but also has a significant improvement in robustness, and can more effectively deal with cross-domain feature distribution differences under complex working conditions.

[0197] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of the present invention.

Claims

1. A method for predicting remaining useful life based on dynamic weighted subdomain adaptation of multi-source features, characterized by: Constructing a system remaining life prediction network: A deep feature fusion extraction network is used to extract features from multiple source and target domains. These features are then sent to a cross-domain difference perception layer. Multiple source and target domain features are sequentially compared to calculate similarities in degradation trends. Weights are assigned to features from different source domains based on the similarity metrics calculated by the cross-domain difference perception layer to form new weighted features that fuse features from multiple source domains. A joint distribution adaptation module is used as a constraint to align the conditional distribution with the marginal distribution. The specific prediction steps are as follows: Step 1: A multi-domain feature extractor extracts key information from input data of multiple source domains and target domains; Step 2: The cross-domain difference perception layer measures the similarity of the degradation trends between the multi-source domain features and the target domain features, and dynamically weights the multi-source features to form a weighted feature that includes the multi-source domain features; Step 3: After fusing the multi-source domain features into dynamic weighted features, the weighted features and target domain features are sent to the joint distribution adaptation module to handle distribution differences in the cross-domain feature alignment task. Step 4: Construct a joint loss function for the system's remaining life prediction network, use a bidirectional gated recurrent unit network to form a remaining life prediction module, process the features output from the test set feature extractor, and output the final prediction results.

2. The method for predicting remaining useful life based on dynamic weighted subdomain adaptation of multi-source features according to claim 1 is characterized in that: In step 1, the multi-domain feature extractor includes two parts: a multi-scale convolutional neural network and a multi-head self-attention mechanism. The multi-scale convolutional network extracts local area features, and the multi-head self-attention mechanism performs modeling on a global scale. The pre-processed frequency domain data is input into three convolutional networks arranged at different scales. The feature map after convolution operation is represented as: ; Where, Represents the weight in the current convolutional layer operation, Represents the bias in the current convolutional layer operation, represents the feature mapping function, Represents the feature map output by the current convolutional layer; pass The operation performs additive fusion on the features extracted at different scales. The fusion calculation is shown as follows: ; Where, represents the output of the fusion layer, Indicates the number of channels, 、 、 Represent the outputs of three convolutional networks respectively; The features that will be obtained Input into the multi-head attention encoder layer to capture global features. The multi-head attention is composed of multiple single-head attentions, including the query vector , key vector Sum value vector Three inputs, through and Perform scaled dot product to get the attention matrix, and then Multiply to get the weighted data samples, and the calculation formula of the scaled dot product is as follows: ; Where, Representing the dimension, the multi-head attention mechanism is calculated as follows: ; ; Where, represents the number of heads in the multi-head self-attention mechanism, 、 、 、 are all learnable parameters.

3. The method for predicting remaining useful life based on dynamic weighted subdomain adaptation of multi-source features according to claim 2, characterized in that: In step 2, the cross-domain difference perception layer includes two parts: local similarity measurement and global similarity measurement. The local similarity measurement measures the similarity between source domain and target domain samples in the feature space. The cosine similarity is selected to measure the local similarity of source domain and target domain samples. Its formula is expressed as follows: ; Where, represents the dot product operation, and Indicates the The modulus of the source and target domain features; Use KL divergence to measure the source domain The characteristic distribution of and target domain feature distribution The differences between: ; Each source domain feature and target domain features The KL divergence of is expressed as: (8); Adjust the weight according to the contribution of each feature, weight Automatically adjusted by gradient descent method, the weight update formula is: (9); in, represents the learning rate, is the loss function About weight The gradient, Indicates the The source domain is The weight at the iteration; Through the adaptive weight mechanism, the model automatically assigns weights based on the importance of features adjusted by feedback during training; After weight optimization, each source domain feature The calculated weight The weighted source domain features are expressed as follows: (10); Where, represents the weighted source domain features, Source domain The weight coefficient of Fusion of all source domain weighted features into a new feature : (11); Multi-source domain features are weighted and fused through the cross-domain difference perception layer, and the dynamic weighting mechanism adjusts the weights in real time according to the similarity between the source domain features and the target domain.

4. The method for predicting remaining useful life based on dynamic weighted subdomain adaptation of multi-source features according to claim 3, characterized in that: In step 3, the joint distribution adaptation module extracts the feature relationship between the source and target domains by fusing convolutional feature enhancement strategies and introduces label discretization in the feature alignment process; First, the features extracted by the deep feature fusion network are weighted using the CBAM attention mechanism. In the channel attention part, global average pooling and global maximum pooling operations are performed on the input data, and two sizes of The vectors are input into two layers of MLP, the first layer has neurons, of which is the reduction rate, the second layer has neurons; Add the output of MLP element by element and use Sigmoid activation function to generate channel attention weights. The activation function is expressed as follows: (12); The spatial attention part performs a Concat operation along the channel dimension to concatenate the features after maximum pooling and average pooling; After the features extracted by CBAM attention are compressed into 3C through the maximum pooling layer, the extracted features are mapped to the corresponding remaining lifespan through the regressor. The mean square error between the predicted RUL and the true RUL is calculated as the loss function, and the prediction loss is described as: (13); Where, represents the number of samples in the source domain, represents the predicted lifespan, Indicates real lifespan; The label predicted by RUL is in the range of Before the conditional distribution adaptation is performed on the continuous variable, the regression label is discretized using formula (13). The discretized source domain RUL label and the RUL pseudo label generated in the target domain together constitute the conditional distribution basis of the feature; (14); Where, Indicates the number of degradation stages into which the remaining lifetime label is divided, ; Multilinear multiplication mapping is used to simulate the interaction between different features and labels to capture the interaction between features and life prediction results and the multimodal structure of data distribution. The joint distribution adaptive training process is expressed as: (15); Where, and Represent the data of the source domain and the target domain respectively, and represents the marginal distribution of the source domain and the target domain, represents the domain discriminator.

5. The method for predicting remaining useful life based on dynamic weighted subdomain adaptation of multi-source features according to claim 4, characterized in that: The specific steps of label discretization are: discretize the real RUL labels of the source domain and the pseudo labels of the target domain, generate a discrete label space as a constraint for feature alignment, align the marginal distribution and conditional distribution, and establish a joint distribution alignment mechanism between the source domain and the target domain.

6. The method for predicting remaining useful life based on dynamic weighted subdomain adaptation of multi-source features according to claim 5, characterized in that: In step 4, the joint loss function is composed of the weighted fusion loss of multi-source domain features and the regression prediction error loss. , joint distribution adaptive training loss, where the multi-source feature weighted fusion loss is defined as: (16); Where, represents the cosine similarity measure between source domain features and target domain features, Represents the KL divergence measure between source domain features and target domain features; The regression loss is used to minimize the error between the predicted life expectancy and the actual life expectancy, assuming To predict life expectancy, Represents the true RUL value, and the regression loss is expressed as: (17); Joint distribution adaptation training is expressed as: (18); The total loss of the remaining life prediction network model of the subdomain adaptive system based on multi-source characteristics dynamic weighting is: (19); in, , , Represents the weight coefficient of each part.

7. The method for predicting remaining useful life based on dynamic weighted subdomain adaptation of multi-source features according to claim 6, characterized in that: In step 4, a bidirectional gated recurrent unit network is constructed by superimposing two independent gated recurrent units, extracting information in the forward and reverse time axes respectively; The core of the gated recurrent unit is to control the update gate through two gating mechanisms and reset gate The internal operation is as follows: (20); (21); (22); (23); Where, and Represent the outputs of the update gate and reset gate respectively, 、 and represents the weight matrix of the update gate, reset gate, and candidate hidden state, is the hidden state at the previous moment, represents the activation function, tanh is the hyperbolic tangent activation function, Indicates the output at the current moment; The output of the remaining life prediction module is expressed as: (24); Where, and Represents the weights and biases of the output layer.

Citation Information

Cited By

  • System multi-degradation-stage residual life prediction method, system and equipment based on dynamic domain adaptive network, and storage medium

    CN121580146A

  • System, device and storage medium for predicting residual life of multiple degradation stages of system based on dynamic domain adaptive network

    CN121580146B