Condition-aware meta-learning domain generalization method for bearing residual life prediction

CN122675409APending Publication Date: 2026-09-01乔先鹏
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610876185.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-17
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

[0008]有鉴于此,本发明提供一种面向轴承剩余寿命预测的条件感知元学习域泛化方法,以解决或缓解现有技术中存在的技术问题之一,至少提供一种有益的选择

Benefits of technology

1.首次提出函数偏移主导的域泛化框架,摒弃协变量偏移假设,精准匹配轴承运行条件变化导致的信号-RUL映射函数差异,从根本上解决跨域泛化瓶颈;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122675409A_ABST
    Figure CN122675409A_ABST
Patent Text Reader

Abstract

The application discloses a condition-aware meta-learning domain generalization method for bearing residual life prediction, discloses a function offset dominant domain generalization framework by abandoning the traditional covariate shift assumption, preserves the cross-domain life scale correlation through a global normalization strategy, and realizes absolute residual life prediction. The method builds a CTT-DAE condition-aware deep learning architecture, differentiates modeling of bearing health and degradation stage features, and realizes working condition adaptive feature alignment and life scaling by relying on a nonlinear projection meta-learner. Meanwhile, a CMLDG double-layer optimization mechanism is designed, data padding masking, position coding and early stopping strategies are combined, and model training is optimized. The application can accurately capture the degradation law of bearings depending on the working condition without the participation of target domain data, effectively improve the prediction accuracy and robustness of unseen working condition domains, and adapt to various industrial equipment bearing predictive maintenance scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bearing remaining life prediction technology, and in particular to a condition-aware meta-learning domain generalization method for bearing remaining life prediction. Background Technology

[0002] Predictive maintenance (PdM) plays a crucial role in achieving cost-effective and reliable maintenance strategies by preventing unexpected mechanical failures caused by wear, fatigue, and material degradation. The effectiveness of PdM largely depends on the accurate prediction of the remaining useful life (RUL) of mechanical components. Among these components, rolling element bearings are particularly important because they significantly reduce friction and ensure smooth, stable operation in a wide range of industrial applications.

[0003] However, in real-world industrial environments, rolling element bearings typically operate under a variety of conditions, such as variations in radial force and rotational speed. Compared to single-condition operation, the performance of traditional RUL prediction models degrades significantly when applied to multi-condition scenarios, with the root mean square error (RMSE) increasing by approximately 60% to 98%.

[0004] Variations in operating conditions introduce significant domain shifts between the source and target domains, severely hindering the ability of traditional RUL models to achieve high prediction accuracy. In this context, each source and target domain corresponds to a specific operating condition. Depending on whether target domain data is accessible during training, domain shift methods can be broadly categorized into Domain Adaptation (DA) and Domain Generalization (DG). DA aims to reduce the distributional differences between the source and target domains caused by changes in external operating conditions by explicitly aligning feature distributions. In contrast, DG focuses on learning domain-invariant representations from multiple source domains with heterogeneous distributions. This allows trained models to maintain robust prediction performance for previously unseen target domains without requiring access to any target domain data during training.

[0005] In contrast, representation learning-based DG methods focus on extracting domain-invariant features from multiple source domains, aiming to align the input feature distributions so that the learned representations remain consistent across unseen target domains. For example, Juan Xu et al. proposed an inter-domain-intra-domain normalized generalization (IDIDNG) network, applying adaptive normalization before feature extraction to reduce inter-domain differences. Xiaoqi Xiao et al. introduced a contrastive domain-invariant generalization (CDIG) method, which combines multi-head conditional awareness attention to extract invariant features across multiple running conditions, achieving a 32.19% improvement on the C-MAPSS dataset. Despite these advances, most existing DG methods still assume that covariate shift caused by multiple domains is the primary cause of domain shift. However, evidence suggests that the main factor caused by running conditions is function shift rather than covariate shift. Function shift refers to the change in the signal-to-RUL mapping function across different domains. Specifically, under running conditions 1 and 2, the input signals are highly similar, but the corresponding UL values ​​are significantly different. This gap in assumption limits the effectiveness of data manipulation-based and representation learning-based DG methods.

[0006] In summary, the existing technology has the following core defects: 1) The assumption that domain shift is mainly based on covariate shift ignores the function shift dominated by actual operating conditions, resulting in insufficient cross-domain generalization ability of the model. 2) Using local RUL normalization can only predict relative lifetimes and cannot output absolute RUL, making it difficult to support industrial decision-making; 3) The feature alignment method is singular, which cannot simultaneously achieve input-level feature alignment and output-level RUL scaling, making it difficult to capture the degradation patterns of conditional dependencies; 4) The meta-learning strategy is not deeply integrated with conditional awareness modeling, the two-layer optimization mechanism is missing, and the robustness in unseen domains is poor. 5) Lack of differentiated modeling between the health stage (HS) and the decline stage (DS), and insufficient extraction of temporal features; 6) Key components such as position encoding, masking attention, and multi-head attention are not specifically adapted to the bearing degradation timing characteristics, resulting in low accuracy of time-dependent modeling.

[0007] Therefore, there is an urgent need for a novel bearing remaining life prediction method that does not rely on existing technical frameworks, can directly solve function offset, supports absolute RUL prediction, integrates conditional awareness and meta-learning, and achieves strong cross-domain generalization. Summary of the Invention

[0008] In view of this, the present invention provides a condition-aware meta-learning domain generalization method for bearing remaining life prediction, in order to solve or alleviate one of the technical problems existing in the prior art, and at least provide a beneficial option.

[0009] The technical solution of this invention is implemented as follows: a condition-aware meta-learning domain generalization method for bearing remaining life prediction, comprising the following steps: 1) Data Acquisition and Domain Definition: Collect bearing vibration data throughout its entire lifespan, operating condition data, and remaining service life (RUL) labels to construct a source domain set. With the target domain set Each source domain is defined as ,in It is a multi-source vibration signal sequence. Number of signal channels For sequence length, For source domain Number of lower bearings For running condition variables, The number of channels required for operation. The target domain data is not accessible during the model training phase and is labeled as RUL. 2) Function offset modeling: Abandoning the covariate offset assumption, we define the function offset of the signal-to-RU mapping between different domains, as shown in the formula: ,in , Domains ,domain The parameterized prediction function, The domain generalization objective is to minimize the expected prediction error in the unseen target domain, as expressed in the formula:

[0010] Model parameters Based solely on source domain data, the optimization formula is:

[0011] 3) Global RUL Normalization: A global minimum-maximum scaling strategy is adopted to uniformly normalize the RUL labels, vibration signals, and operating condition variables of all source domains. The RUL normalization formula is:

[0012] in:

[0013] To find the global maximum RUL value in the source domain, the vibration signal normalization formula is:

[0014] The normalization formula for operating conditions is:

[0015] The target domain data is normalized using the maximum and minimum values ​​of the source domain statistics, while preserving the cross-domain absolute lifetime scale correlation. 4) Data segmentation and padding: The bearing life data is segmented into a healthy phase (HS) and a degradation phase (DS) according to the first prediction time (FPT). FPT is defined as the moment when the peak acceleration change rate exceeds 0.3. Zero padding is used for the HS segment and life end padding is used for the DS segment to unify the sequence length. During training, the padding values ​​are masked to avoid the padding values ​​interfering with the temporal feature learning. 5) Construct a condition-aware deep learning architecture: Build a condition-based time transformer and denoising autoencoder (CTT-DAE), which includes a denoising autoencoder (DAE) module, a condition-based time transformer (CTT) module, a positional encoding module, and a predictor; among them, the condition-based time transformer (CTT) module integrates three major modules: a nonlinear projection meta-learner (NPM), a health module, and a degradation module. 6) Conditional Meta-Learning Domain Generalization (CMLDG) Two-Layer Optimization: The source domain data is divided into non-overlapping meta-training datasets, meta-test datasets, and test datasets. The CTT-DAE model parameters are updated through inner loop optimization and the meta-learner parameters are updated through outer loop optimization, thereby achieving cross-domain transferable meta-knowledge extraction. 7) Model evaluation: The generalization performance of the model in the unseen target domain is quantitatively evaluated using the average prediction accuracy, worst case error (WCE), and percentage improvement index.

[0016] Preferably, the workflow of the DAE module in step 5) is as follows: white noise is added to the input vibration signal, the noisy signal is compressed into latent spatial features by the encoder module, and the original input signal is reconstructed by the decoder module. The reconstruction loss function formula is:

[0017] in For vibration signals reconstructed by DAE, robust spatial features are extracted by minimizing the reconstruction loss.

[0018] Preferably, the workflow of the nonlinear projection meta-learner (NPM) in step 5) is as follows: [The following text appears to be incomplete and requires further context: "with running condition variables..."] Spatial features extracted by DAE As input, input-level alignment coefficients are generated through a nonlinear projection function. With output level scaling factor The formula is Input-level feature alignment is achieved through element-wise multiplication, as shown in the formula: Output-level RUL scaling is achieved through element-wise multiplication, as shown in the formula: ,in The initial RUL prediction value output by the CTT module.

[0019] Preferably, the encoding function of the position encoding module in step 5) is:

[0020] in For time step Position encoding vector, For encoding dimensions, As a location index, temporal sequence information is injected into the model through location encoding, which makes up for the lack of explicit temporal awareness in the time transformer.

[0021] Preferably, the health module described in step 5) includes There are several encoder layers, each consisting of a multi-head attention layer and a feedforward network layer. The scaling dot product attention calculation process for the multi-head attention layer is as follows: mapping the input features to the query... ,key ,value The formula is:

[0022] in , , Given a learnable weight matrix, the scaling dot product attention formula is: For the key vector dimension, multi-head attention extracts multi-dimensional temporal features of the health phase and outputs them as memory features by concatenating the outputs of multiple parallel scaled dot product attention.

[0023] Preferably, the degradation module in step 5) includes a masked multi-head attention layer, an encoder-decoder attention layer, and a feedforward network layer. The masked multi-head attention layer prevents the model from focusing on future temporal information when predicting features of subsequent time steps through a masking mechanism. The encoder-decoder attention layer uses the memory features output by the health module as keys and values ​​and the degradation stage features as queries to achieve the fusion of health stage and degradation stage features. The feedforward network layer converts linear temporal features into nonlinear temporal features to accurately model the temporal evolution law of bearing degradation stage.

[0024] Preferably, the specific process of the CMLDG two-layer optimization in step 6) is as follows: 1) Inner loop optimization: Using the meta-training dataset as input, update the CTT-DAE model parameters by minimizing the source domain prediction error. Initial parameters of the meta-learner The optimized formula is:

[0025] 2) Outer loop optimization: Fix the parameters of the CTT-DAE model after inner loop optimization. Using the meta-test dataset as input, the meta-learner parameters are updated by minimizing the prediction error in the meta-test domain. The optimized formula is:

[0026] 3) Early stopping mechanism: The meta-training dataset and meta-test dataset are divided into support set and query set respectively. The support set is used for parameter updates and the query set is used for validation. When the validation loss of the query set does not decrease for 5 consecutive training cycles, the model training is terminated to avoid overfitting.

[0027] Preferably, the average prediction accuracy index in step 7) includes root mean square error (RMSE) and mean absolute error (MAE). The formula for calculating RMSE is:

[0028] The formula for calculating MAE is:

[0029] in For the target domain Each bearing at time step The predicted RUL value, This is the true RUL value. The total length of the target domain sequence. The number of bearings in the target domain.

[0030] Preferably, the worst-case error (WCE) mentioned in step 7) is the maximum prediction error of all tested bearings in the target domain. Before calculation, the globally normalized RUL value needs to be converted to a locally normalized scale. The conversion formula is as follows:

[0031] in For globally normalized RUL values, For locally normalized RUL values, For the target domain The total lifetime of each bearing, WCE, is calculated based on the locally normalized RUL value, reflecting the robustness of the model in the most challenging unseen domain scenarios.

[0032] Preferably, the percentage improvement index mentioned in step 7) is used to quantitatively evaluate the performance improvement of the method of the present invention compared to the baseline model, and the calculation formula is as follows:

[0033] in This represents the MAE value of the baseline deep learning model in the unseen target domain. This is the MAE value of the method of the present invention on an unseen target domain, and this index is used to achieve a fair performance comparison with existing domain generalization methods.

[0034] The embodiments of the present invention have the following advantages due to the adoption of the above technical solutions: 1. This paper proposes a domain generalization framework dominated by function offset for the first time, abandons the covariate offset assumption, accurately matches the difference in signal-RUL mapping function caused by changes in bearing operating conditions, and fundamentally solves the bottleneck of cross-domain generalization. 2. A unique global RUL normalization strategy breaks through the limitation of local normalization, which can only predict relative lifespan, and achieves absolute RUL prediction, directly supporting quantitative decision-making for industrial predictive maintenance, significantly improving practicality; 3. Construct a CTT-DAE condition-aware architecture that integrates DAE spatial feature extraction, meta-learner conditional adaptation, and time transformer temporal modeling, while simultaneously achieving input-level feature alignment and output-level RUL scaling to fully capture the degradation patterns of conditional dependencies. 4. Design a CMLDG two-layer meta-learning optimization mechanism, with the inner loop optimizing the basic model and the outer loop optimizing meta-knowledge extraction, which significantly improves the prediction robustness of unseen target domains under limited source domain diversity. 5. Differentiated modeling of the health and degradation stages, combined with zero-fill / end-of-life filling, masked attention, and position encoding, accurately captures the temporal degradation characteristics of the bearing throughout its entire life cycle, significantly improving the accuracy of time-dependent modeling; 6. The meta-learner uses a nonlinear projection design that directly links the operating conditions with the feature transformation coefficients, adaptively adapting to the degradation rate of different domains and avoiding the limitations of existing methods that fix transformation parameters; 7. It has strong industrial adaptability, requires no target domain data during the training phase, and is suitable for bearing RUL prediction needs in actual industrial scenarios such as turbines, coal mills, air compressors, and generators. Its generalization performance far exceeds that of existing DG methods.

[0035] The above overview is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features of the invention will become readily apparent from the accompanying drawings and the following detailed description. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0037] Figure 1 This is a schematic diagram of the domain adaptation (DA) and domain generalization (DG) of the present invention; Figure 2 The performance of the DA method in this invention is evaluated using the average RMSE score on the C-MAPSS dataset; Figure 3 This is a schematic diagram illustrating the function offset concept of the present invention; Figure 4 This is the architecture of the DAE of the present invention; Figure 5 This is the overall architecture of the condition-based time transformer (CTT) for RUL prediction according to the present invention; Figure 6 The overall architecture of the meta-learner of the present invention is used for: (a) input-level feature alignment, and (b) output-level RUL scaling; Figure 7 The (left) scaled dot product attention and (right) multi-head attention of this invention consist of multiple attention layers running in parallel; Figure 8 This is a schematic diagram of the parameter update process of the CMLDG method of the present invention. Detailed Implementation

[0038] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of the invention. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.

[0039] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0040] Example 1: Multi-operational-condition domain generalized RUL prediction based on the XJTU-SY bearing dataset

[0041] 1. Data Preparation

[0042] The XJTU-SY bearing life dataset was used, which includes vibration data of 15 bearings under 3 operating conditions (speed / radial force). The dataset was divided into a source domain (2 operating conditions, 10 bearings) and a target domain (1 unseen operating condition, 5 bearings). The source domain was further divided into a meta-training set (6 bearings), a meta-test set (2 bearings), and a test set (2 bearings). The data sampling frequency was 25.6kHz, and 1024 sampling points were taken for each sample.

[0043] 2. Global RUL Normalization

[0044] The maximum RUL value of all bearings in the statistical source domain h, maximum vibration signal g, minimum value g. Maximum operating conditions r / min, minimum value r / min; The source and target domain data are scaled uniformly according to the global normalization formula, and the target domain adopts the statistical values ​​of the source domain, while retaining the absolute lifetime scale correlation.

[0045] 3. Data splitting and filling

[0046] The HS and DS segments are split according to an FPT threshold of 0.3: HS represents the initial stage of operation with no obvious degradation, while DS represents the accelerated degradation stage after FPT; the HS segment is zero-padded to a length of 2048, and the DS segment is padded to a length of 2048 at the end of its lifespan; the padded values ​​are masked during training to avoid interference from the padded values ​​in the self-attention calculation.

[0047] 4. CTT-DAE Model Construction and Training

[0048] DAE module: The encoder contains 3 fully connected layers (dimension 1024→512→256), and the decoder contains 3 fully connected layers (256→512→1024); 20dB white noise is added, MAE is used for reconstruction loss, 50 training epochs are used, and the batch size is 32. Meta-learner: Non-linear projection layer with dimension 256, ReLU activation outputs projected features, and alignment coefficients are generated by taking the diagonal elements after matrix multiplication. With scaling factor The range of Tanh activation constraint coefficients; Location encoding: The encoding dimension is 256, matching the spatial feature dimension. A temporal location vector is generated according to the location encoding formula and concatenated with the spatial features. Health module: 2 encoder layers, 4 multi-head attention heads, 512 dimensions of feedforward layer, extracting HS temporal features; Degeneracy module: 2 decoder layers, 4 masked multi-head attention heads, encoder-decoder attention fusion healthy memory, feedforward layer dimension 512; Predictor: Fully connected layer (256→1), outputs RUL prediction value; CMLDG two-layer optimization: inner loop optimizes CTT-DAE parameters and learning rate. Rounds 30; outer loop optimizes meta-learner parameters, learning rate. Rounds 20; early stop threshold 5 rounds; MAE is used to verify loss.

[0049] 5. Result Evaluation

[0050] In the unseen target domain, the method of this invention achieves RMSE=8.2h and MAE=5.6h, representing a percentage improvement of 42.3% compared to the baseline model (CNN-LSTM); the worst-case error WCE (RMSE) is 12.5h, significantly outperforming existing DG methods.

[0051] Example 2: Cross-condition RUL prediction based on IEEE PHM 2012 bearing dataset

[0052] 1. Data Preparation

[0053] The IEEE PHM 2012 bearing dataset was used, which includes 4 operating conditions (speed / load) and a total of 12 bearing life data. The dataset was divided into a source domain (3 operating conditions, 9 bearings) and a target domain (1 unseen operating condition, 3 bearings). The source domain was further divided into a meta-training set (5 bearings), a meta-test set (2 bearings), and a test set (2 bearings). The data sampling frequency was 25kHz, and each sample had 2048 sampling points.

[0054] 2. Global RUL Normalization

[0055] Source domain statistics: h、 g、 g、 r / min r / min; All data are processed according to the global normalization formula, and the target domain reuses the source domain statistics.

[0056] 3. Data splitting and filling

[0057] The FPT threshold is 0.3, separating HS and DS; HS is zero-padded to 4096, and DS is padded to 4096 when the lifetime ends; masking padding values ​​are used to ensure the effectiveness of temporal feature learning.

[0058] 4. CTT-DAE Model Construction and Training

[0059] DAE module: encoder (2048→1024→512), decoder (512→1024→2048), signal-to-noise ratio 15dB white noise, reconstruction loss MAE, rounds 60, batch size 16; Meta-learner: Projection dimension 512, ReLU activation, matrix multiplication generation , Tanh constraint; Location encoding: 512 dimensions, time-series vector generated according to the formula; Health module: 3 encoder layers, 8 multi-head attention heads, 1024 feedforward layer dimensions; Degeneracy module: 3 decoder layers, 8 masking multi-head attention heads, encoder-decoder attention fusion memory; Predictor: 512→1 fully connected layer; CMLDG Optimization: Inner Loop Learning Rate Rounds 40; outer loop learning rate Round 25; 5 rounds stopped early.

[0060] 5. Result Evaluation

[0061] In the unseen target domain, RMSE=7.5h and MAE=4.9h, representing a 45.7% improvement over the baseline model; WCE(RMSE)=11.8h, demonstrating excellent generalization performance across operating conditions.

[0062] Example 3: Multi-condition RUL prediction based on industrial field bearing data

[0063] 1. Data Preparation

[0064] Vibration data of the bearings throughout the entire lifespan of a coal mill in a power plant were collected, including 5 operating conditions (speed 1500-3000 r / min, radial force 5-20 kN), totaling 20 bearing data points; source domain (4 conditions, 16 bearings), target domain (1 unseen condition, 4 bearings); source domain meta-training set (10 bearings), meta-test set (3 bearings), test set (3 bearings); sampling frequency 32 kHz, sample length 1024.

[0065] 2. Global RUL Normalization

[0066] Source domain statistics: h、 g、 g、 r / min r / min; global normalization processing, target domain reuses statistical values.

[0067] 3. Data splitting and filling

[0068] FPT threshold of 0.3, separating HS and DS; HS is zero-filled up to 2048, DS is filled up to 2048 at the end of its lifetime; masking fill value to ensure the accuracy of time series modeling.

[0069] 4. CTT-DAE Model Construction and Training

[0070] DAE module: encoder (1024→512→256), decoder (256→512→1024), signal-to-noise ratio 25dB white noise, reconstruction loss MAE, rounds 45, batch size 24; Meta-learner: Projection dimension 256, ReLU activation, generation , ; Location encoding: 256 dimensions, time-series vector generation; Health module: 2 encoder layers, 4 multi-head attention heads, 512 dimensions of feedforward layer; Degeneracy module: 2 decoder layers, masking multi-head attention with 4 heads, and incorporating healthy memories; Predictor: 256→1 fully connected layer; CMLDG Optimization: Inner Loop Learning Rate Rounds 35; outer loop learning rate Round 22; 5 rounds stopped early.

[0071] 5. Result Evaluation

[0072] In the industrial field without target domain, RMSE=9.1h and MAE=6.3h, representing a 40.5% improvement over the baseline model; WCE(RMSE)=13.2h, fully adapting to complex operating conditions in industrial settings and meeting actual predictive maintenance needs.

[0073] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope disclosed in the present invention, and these should all be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A condition-aware meta-learning domain generalization method for bearing remaining life prediction, characterized in that, Includes the following steps: 1) Data Acquisition and Domain Definition: Collect bearing vibration data throughout its entire lifespan, operating condition data, and remaining service life (RUL) labels to construct a source domain set. With the target domain set Each source domain is defined as ,in It is a multi-source vibration signal sequence. Number of signal channels For sequence length, For source domain Number of lower bearings For running condition variables, The number of channels required for operation. The target domain data is not accessible during the model training phase and is labeled as RUL. 2) Function offset modeling: Abandoning the covariate offset assumption, we define the function offset of the signal-to-RU mapping between different domains, as shown in the formula: ,in , Domains ,domain The parameterized prediction function, The domain generalization objective is to minimize the expected prediction error in the unseen target domain, as expressed in the formula: Model parameters Based solely on source domain data, the optimization formula is: 3) Global RUL Normalization: A global minimum-maximum scaling strategy is adopted to uniformly normalize the RUL labels, vibration signals, and operating condition variables of all source domains. The RUL normalization formula is: in: To find the global maximum RUL value in the source domain, the vibration signal normalization formula is: The normalization formula for operating conditions is: The target domain data is normalized using the maximum and minimum values ​​of the source domain statistics, while preserving the cross-domain absolute lifetime scale correlation. 4) Data segmentation and padding: The bearing life data is segmented into a healthy phase (HS) and a degradation phase (DS) according to the first prediction time (FPT). FPT is defined as the moment when the peak acceleration change rate exceeds 0.

3. Zero padding is used for the HS segment and life end padding is used for the DS segment to unify the sequence length. During training, the padding values ​​are masked to avoid the padding values ​​interfering with the temporal feature learning. 5) Construct a condition-aware deep learning architecture: Build a condition-based time transformer and denoising autoencoder (CTT-DAE), which includes a denoising autoencoder (DAE) module, a condition-based time transformer (CTT) module, a position encoding module, and a predictor; among them, the condition-based time transformer (CTT) module integrates three major modules: a nonlinear projection meta-learner (NPM), a health module, and a degradation module. 6) Conditional Meta-Learning Domain Generalization (CMLDG) Two-Layer Optimization: The source domain data is divided into non-overlapping meta-training datasets, meta-test datasets, and test datasets. The CTT-DAE model parameters are updated through inner loop optimization and the meta-learner parameters are updated through outer loop optimization, thereby achieving cross-domain transferable meta-knowledge extraction. 7) Model evaluation: The generalization performance of the model in the unseen target domain is quantitatively evaluated using the average prediction accuracy, worst case error (WCE), and percentage improvement index.

2. The conditional-aware meta-learning domain generalization method for bearing remaining life prediction according to claim 1, characterized in that: The workflow of the DAE module described in step 5) is as follows: white noise is added to the input vibration signal; the noisy signal is compressed into latent spatial features by the encoder module; and the original input signal is reconstructed by the decoder module. The reconstruction loss function formula is: in For vibration signals reconstructed by DAE, robust spatial features are extracted by minimizing the reconstruction loss.

3. The conditional-aware meta-learning domain generalization method for bearing remaining life prediction according to claim 1, characterized in that: The workflow of the nonlinear projection meta-learner (NPM) described in step 5) is as follows: It operates with condition variables... Spatial features extracted by DAE As input, input-level alignment coefficients are generated through a nonlinear projection function. With output level scaling factor The formula is Input-level feature alignment is achieved through element-wise multiplication, as shown in the formula: Output-level RUL scaling is achieved through element-wise multiplication, as shown in the formula: ,in The initial RUL prediction value output by the CTT module.

4. The conditional-aware meta-learning domain generalization method for bearing remaining life prediction according to claim 1, characterized in that: The encoding function of the position encoding module mentioned in step 5) is: in For time steps Position encoding vector, For encoding dimensions, As a location index, it injects temporal sequence information into the model through location encoding, making up for the lack of explicit temporal awareness in the time transformer.

5. The conditional-aware meta-learning domain generalization method for bearing remaining life prediction according to claim 1, characterized in that: The health module mentioned in step 5) includes There are several encoder layers, each consisting of a multi-head attention layer and a feedforward network layer. The scaling dot product attention calculation process for the multi-head attention layer is as follows: mapping the input features to the query... ,key ,value The formula is: in , , Given a learnable weight matrix, the scaling dot product attention formula is: For the key vector dimension, multi-head attention extracts multi-dimensional temporal features of the health phase and outputs them as memory features by concatenating the outputs of multiple parallel scaled dot product attention.

6. The conditional-aware meta-learning domain generalization method for bearing remaining life prediction according to claim 1, characterized in that: The degradation module described in step 5) includes a masked multi-head attention layer, an encoder-decoder attention layer, and a feedforward network layer. The masked multi-head attention layer prevents the model from focusing on future temporal information when predicting features of subsequent time steps through a masking mechanism. The encoder-decoder attention layer uses the memory features output by the health module as keys and values ​​and the degradation stage features as queries to achieve the fusion of health stage and degradation stage features. The feedforward network layer converts linear temporal features into nonlinear temporal features to accurately model the temporal evolution law of bearing degradation stage.

7. The conditional-aware meta-learning domain generalization method for bearing remaining life prediction according to claim 1, characterized in that: The specific process of CMLDG two-layer optimization described in step 6) is as follows: 1) Inner loop optimization: Using the meta-training dataset as input, update the CTT-DAE model parameters by minimizing the source domain prediction error. Initial parameters of the meta-learner The optimized formula is: 2) Outer loop optimization: Fix the parameters of the CTT-DAE model after inner loop optimization. Using the meta-test dataset as input, the meta-learner parameters are updated by minimizing the prediction error in the meta-test domain. The optimized formula is: 3) Early stopping mechanism: The meta-training dataset and meta-test dataset are divided into support set and query set respectively. The support set is used for parameter updates and the query set is used for validation. When the validation loss of the query set does not decrease for 5 consecutive training cycles, the model training is terminated to avoid overfitting.

8. The conditional-aware meta-learning domain generalization method for bearing remaining life prediction according to claim 1, characterized in that: The average prediction accuracy index mentioned in step 7) includes the root mean square error (RMSE) and the mean absolute error (MAE). The formula for calculating RMSE is: The formula for calculating MAE is: in For the target domain Each bearing at time step The predicted RUL value, This is the true RUL value. The total length of the target domain sequence. The number of bearings in the target domain.

9. A conditional-aware meta-learning domain generalization method for bearing remaining life prediction according to claim 1, characterized in that: The worst-case error (WCE) mentioned in step 7) is the maximum prediction error of all tested bearings in the target domain. Before calculation, the globally normalized RUL value needs to be converted to a locally normalized scale. The conversion formula is as follows: in For globally normalized RUL values, For locally normalized RUL values, For the target domain The total lifetime of each bearing, WCE, is calculated based on the locally normalized RUL value, reflecting the robustness of the model in the most challenging unseen domain scenarios.

10. A condition-aware meta-learning domain generalization method for bearing remaining life prediction according to claim 1, characterized in that: The percentage improvement metric mentioned in step 7) is used to quantitatively evaluate the performance improvement of the method of the present invention compared to the baseline model. The calculation formula is as follows: in This represents the MAE value of the baseline deep learning model in the unseen target domain. This is the MAE value of the method of the present invention on an unseen target domain, and this index is used to achieve a fair performance comparison with existing domain generalization methods.