Knowledge-guided dual-path multi-feature fusion machine remaining service life prediction method and system
By adopting a knowledge-guided dual-path multi-feature fusion method in aero engine life prediction, combining convolutional neural networks and GRU networks for feature extraction, and fusing them with manual features, the problems of insufficient feature extraction and time series information loss in the prior art are solved, significantly improving prediction performance and robustness.
Patent Information
- Application Number
- CN202510045542.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art has limitations in the life prediction of aero engines, loss of time series information and single path prediction, resulting in insufficient prediction performance and robustness.
A knowledge-guided dual-path multi-feature fusion method is adopted to combine domain knowledge with convolutional neural networks and self-focusing mechanism-driven GRU networks for spatial and time series feature extraction, and fuse them with manual features to form a dual-path multi-feature fusion framework.
It significantly improves the performance of life prediction and the robustness of the model, enhances the diversity of features and the interpretability of the model, and can more accurately capture complex engine degradation patterns.
Smart Images

Figure CN119962369A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of industrial production life (RUL) prediction, and specifically relates to a method and system for predicting the remaining useful life of a machine based on knowledge-guided dual-path multi-feature fusion. Background Art
[0002] Since Hinton et al. introduced the concept of deep learning in 2006, deep learning has become a key technology in many fields such as image processing, speech recognition, fault diagnosis, etc. Deep learning uses the construction of multi-layer neural networks to deeply mine hidden features in data through its powerful modeling and representation capabilities, thereby achieving high-precision prediction results.
[0003] In recent years, data-driven methods have attracted more and more attention in life prediction. These methods are divided into three categories: CNN-based methods, recurrent neural network (RNN)-based methods, and hybrid methods. The CNN-based remaining useful life prediction method automatically extracts local features of time series data through the convolution layer, integrates these features through the fully connected layer, and finally performs regression prediction of the RUL value. RNN is a neural network model specifically designed for processing sequence data. Unlike CNN-based methods, RNN captures time dependencies by passing previous information to current inputs through its cyclic structure. RNN has many variants, such as long short-term memory (LSTM) and gated recurrent unit (GRU), which further enhances its performance in processing long-term dependencies and sequence data. Although these methods can effectively complete the task of life prediction, there are still some challenges in practical applications.
[0004] First, due to the complexity and variability of the actual working conditions of aircraft engines, the existing CNN-based methods mainly extract features through local convolution operations, which makes them insufficient in modeling the global spatial relationship between sensors. This limitation may lead to an incomplete understanding of the complex internal structure of the engine and the interaction between components, thereby affecting the accuracy and comprehensiveness of the features and reducing the model's ability to predict potential failures and performance degradation. Secondly, the existing RNN-based methods usually rely on the output of the last time step for prediction, which may lead to the loss of key prediction information. In fact, the contributions of features at different time steps to life prediction are unequal, and the information of some time steps is more important than that of other time steps. Thirdly, current research usually uses a single path for RUL prediction, which makes it difficult to fully extract features when processing complex data tasks. If these shallow features are ignored, the accuracy of model prediction may be affected when performing RUL prediction tasks. Therefore, it is necessary to design a new life prediction method to solve the above technical problems. Summary of the invention
[0005] In view of this, the purpose of the present invention is to provide a machine remaining service life prediction method based on knowledge-guided dual-path multi-feature fusion, so as to reduce the dependence on a single feature, enhance the diversity of features and the interpretability of the model, significantly improve the prediction performance, and enhance the robustness of the model under different working conditions;
[0006] The technical solution of the present invention is: first, the present invention provides a machine remaining service life prediction method based on knowledge-guided dual-path multi-feature fusion, including:
[0007] S1: Data preprocessing: Use the min-max normalization method to process the data of the aircraft engine dataset in the RUL model;
[0008] S2: Feature extraction: In a self-supervised manner, a convolutional neural network (CNN) enhanced with domain knowledge is used to extract spatial features, a GRU network driven by a self-attention mechanism is used to extract time series features, and manual features are combined for feature fusion.
[0009] Specifically, S1 uses the min-max normalization method to map the data dimension of the RUL model to [-1,1], and uses the following specific equation to process the data:
[0010]
[0011] in, and denote the jth sensor in the ith period before and after normalization, respectively; and are the minimum and maximum values of the jth sensor of the same engine in the data;
[0012] Set the RUL threshold to 125;
[0013] A time window sequence of size D*W is used as the training sample, where D represents the data dimension and W represents the window width: the first sliding window sample is represented as T1=[X1,X2,...,X W ], after sliding with a step length of L, the next sequence sample data is T2 = [X1+L,X2+L,...,X W +L, the corresponding data label is the RU prediction label of the last time step.
[0014] Specifically, S2 includes:
[0015] S21: Modeling the physical relationship of engine components, the correlation ρ between sensor i and sensor j is defined as follows:
[0016]
[0017] in, represents the variance of the knowledge vector represented by sensor i, represents the covariance between sensor i and sensor j;
[0018] S22: clustering the sensors into 9 groups according to the physical relationship of the engine components;
[0019] S23: Use the CNN network to capture the spatial features in each group of sensors and obtain a concatenated feature map;
[0020] S24: The GRU network driven by the self-attention mechanism is used to extract time series features from the feature map. The formula is defined as follows:
[0021] r t =σ(U r x t +W r h t-1 +b r )
[0022] z t =σ(U z x t +W z h t-1 +b z )
[0023]
[0024] Among them, r t and z t Represent the outputs of the update gate and reset gate, respectively, X t is the input at time step t, h t-1 and h t are the hidden states at t-1 and t respectively, U r , W r , U Z , W Z , U h and W h is the weight, b r 、b z and b h is the bias, σ and tanh are sigmoid and hyperbolic tangent functions respectively;
[0025] S25: Extract the feature X from the deep space t Put it into the GRU network, X t As the spatial features extracted by the convolutional neural network combined with domain knowledge, after convolution and feature concatenation, the last dimension is the number of sensor groups 9, which is used as the input dimension of the GRU network, and the number of hidden layers is set to 50;
[0026] Assume that the feature representation learned by the GRU network for a sample is h = [h1,h2,...,h n ] T , T is the transposition operation, n is the number of neurons in the hidden layer of the GRU network, and the information obtained by each neuron in the GRU network is extracted;
[0027] Based on the self-attention mechanism, the input h i The importance of different sequential steps is expressed as:
[0028]
[0029] Where W and b are the weight matrix and bias vector respectively, and Φ(·) is the activation function sigmoid;
[0030] After obtaining the score of the i-th eigenvector, the weight is calculated using the softmax function as follows:
[0031]
[0032] Each obtained score α i and the output of the initial neuron h i Multiply them together to get the final feature output vector. The final output feature O of the self-attention mechanism is expressed as:
[0033]
[0034] Among them, α i represents the attention weight of the i-th time step, h i Represents the corresponding output of GRU at that time step.
[0035] Specifically, it also includes: evaluating the method on the publicly available aircraft engine dataset C-MAPSS in the RUL field.
[0036] The present invention also provides a machine remaining service life prediction system based on knowledge-guided dual-path multi-feature fusion, including:
[0037] Data preprocessing module: used to normalize the raw sensor data;
[0038] Feature extraction module: It is used to extract spatial features using a convolutional neural network combined with domain knowledge in a self-supervised manner, extract time series features using a GRU network driven by a self-attention mechanism, and realize feature fusion by combining traditional manual features;
[0039] Evaluation module: The effectiveness of the prediction method is evaluated on the publicly available aircraft engine dataset C-MAPSS in the RUL field.
[0040] Preferably, the feature extraction module includes a deep feature extraction module and a shallow feature extraction module;
[0041] The deep feature extraction module adopts a spatiotemporal progressive approach, using a CNN enhanced by domain knowledge to capture spatial features and a GRU driven by self-attention to capture temporal features;
[0042] The shallow feature extraction module adopts a feature fusion framework, which integrates the spatiotemporal features extracted by deep learning and the features extracted manually.
[0043] The present invention provides a machine remaining useful life prediction method based on knowledge-guided dual-path multi-feature fusion. The method is applied to the RUL prediction field. In the data preprocessing stage, the original sensor data is normalized, and the dimension effect is eliminated by label addition and sliding window interception, thereby improving the model convergence speed and reducing the influence of outliers. In the feature extraction stage, a convolutional neural network combined with domain knowledge is used in a self-supervised manner to extract spatial features and a GRU network driven by a self-attention mechanism is used to extract time series features, and feature fusion is achieved by combining traditional manual features. At the same time, the effectiveness of this method is comprehensively evaluated on the publicly available C-MAPSS dataset in the RUL field. Experimental results show that this method enhances the diversity of features and the interpretability of the model in dealing with RUL prediction tasks, and significantly improves the prediction performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0046] Figure 1 An overview diagram of the framework of the knowledge-guided dual-path multi-feature fusion method provided by the present invention;
[0047] Figure 2 A physical structure diagram of the aero-engine provided by the present invention;
[0048] Figure 3 A diagram showing the physical connection relationship of the engine components and the gas flow relationship provided by the present invention;
[0049] Figure 4A GRU-SAM network framework diagram driven by a self-attention mechanism provided by the present invention;
[0050] Figure 5 A manual characteristic diagram of the sensor mean and trend coefficient provided by the present invention;
[0051] Figure 6 The training and testing flow chart of FDMFusion provided by the present invention;
[0052] Figure 7 Visual analysis of the prediction results of FD001-FD004 and the test result diagram of a single engine provided by the present invention;
[0053] Figure 8 This is a diagram of the selection results of the sliding window size parameters provided by the present invention. DETAILED DESCRIPTION
[0054] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present invention. Instead, they are merely examples of systems consistent with some aspects of the present invention as detailed in the appended claims.
[0055] The present invention provides an application method of a knowledge-guided dual-path multi-feature fusion method in the field of life prediction, which reduces the dependence on a single feature, enhances the diversity of features and the interpretability of the model, significantly improves the prediction performance, and enhances the robustness of the model under different working conditions.
[0056] The technical solution of KDMFusion of the present invention includes the following steps:
[0057] S1: In the data preprocessing stage, the raw sensor data is normalized, labels are added, and sliding windows are intercepted to eliminate the dimension effect, improve the model convergence speed, and reduce the impact of outliers;
[0058] In the data preprocessing stage, since the aircraft engine dataset contains sensor degradation data of various units and magnitude ranges, different features may have different dimensions and value ranges, and such data will have an impact on some machine learning algorithms. At the same time, in real data, there may be outliers or abnormal values, which may have an adverse effect on the performance of the model. Therefore, the data normalization method is adopted to map the data dimension of the model to [-1,1] to eliminate the dimension effect, improve the model convergence speed, and reduce the impact of abnormal values. The following specific equation is used to process the data:
[0059]
[0060] in, and denote the jth sensor in the ith period before and after normalization, respectively. and are the minimum and maximum values of the j-th sensor of the same engine in the data.
[0061] In the context of aircraft engine RUL prediction, turbofan engines typically show minimal signs of performance degradation in the early stages of operation due to extensive industrial testing prior to deployment. Therefore, RUL prediction becomes more important in the middle and late stages of engine life. To this end, the RUL prediction process is divided into two stages: the unchanged state during early operation and the linear degradation stage in the middle and late stages. When the RUL exceeds a predetermined threshold, the engine is considered to be in a normal state and the RUL label remains unchanged. Once the RUL enters the linear degradation stage, the engine is classified as gradually degraded. Based on the analysis of turbofan engine data, the RUL threshold is set to 125.
[0062] Given that aircraft engine degradation information has a strong time dependency, selecting an appropriate time window is crucial to effectively capture relevant features. This application uses a time window sequence of size D*W as a training sample, where D represents the data dimension and W represents the window width. According to this rule, the first sliding window sample can be expressed as T1=[X1,X2,...,X W On this basis, after sliding with a step length of L, the next sequence sample data is T2 = [X1+L,X2+L,...,X W +L, the corresponding data label is the predicted label of the last time step. Selecting the optimal window length can improve the feature extraction from time series data, thereby improving the accuracy of life prediction;
[0063] S2: In the feature extraction stage, a convolutional neural network combined with domain knowledge is used in a self-supervised manner to extract spatial features and a GRU network driven by a self-attention mechanism is used to extract time series features, and feature fusion is achieved by combining traditional manual features;
[0064] This step first requires modeling the physical relationship of engine components and realizing sensor clustering, and then predicting through multi-feature fusion based on spatial-temporal features and manually designed features.
[0065] In the deep feature extraction stage: To improve accuracy, this paper proposes a spatiotemporal progressive approach, in which spatial features are captured using a CNN enhanced with domain knowledge, and temporal features are captured by a self-attention driven GRU. The knowledge-based CNN uses aeroengine expertise to better model spatial relationships, thereby improving interpretability. To address the limitations of CNN in terms of temporal data, the GRU adaptively weights the time steps, thereby effectively capturing long-term dependencies and reducing noise, improving prediction accuracy and model robustness.
[0066] The performance of an aircraft engine depends heavily on the state of its individual components. Understanding the interactions between these components allows for more efficient use of sensor data, leading to more accurate extraction of degradation features such as Figure 2 The engine in the C-MAPSS dataset consists of several key components, including the fan, low-pressure compressor (LPC), high-pressure compressor (HPC), high-pressure turbine (HPT), low-pressure turbine (LPT), and combustion chamber. Figure 2 In the combustion chamber, air enters the fan and is compressed by the LPC and HPC into high-temperature, high-pressure gas, which is then mixed with fuel in the combustion chamber. The resulting gas expands through the HPT and LPT, converting thermal energy into mechanical energy to drive the engine. The fan, LPC, and LPT are connected by the fan rotor, while the HPC and HPT are connected by the core rotor. To capture sensor dependencies, the components are modeled as follows: C1 (fan), C2 (LPC), C3 (bypass duct), C4 (HPC), C5 (combustion chamber), C6 (HPT), C7 (LPT), C8 (fan rotor), and C9 (core rotor), as shown in Figure 2. Figure 3 As shown. Figure 3 In the figure, the dashed lines indicate the gas flow direction, while the solid lines indicate the physical connections between components. Two key relationships are modeled: (1) physical connections, represented by (Ci, connect, Cj), where i and j represent components, for example, (C1, connect, C8) and (C7, connect, C8). (2) airflow relationships, represented by (Ci, airflow, Cj), for example, (C1, airflow, C3) and (C5, airflow, C6).
[0067] By analyzing the physical connections and airflow relationships of the engine, a total of 11 triplets are formed. These triplets help model the dependencies between sensors. On this basis, the expertise in the field of aerospace engines is used to establish sensor correlations based on the mechanical relationships of sensors (such as efficiency and mass flow). First, low-dimensional vectors are generated for each entity and relationship, and the engine component relationships and sensor data are converted into knowledge vectors. The correlation ρ between sensors i and j is calculated as follows:
[0068]
[0069] in, represents the variance of the knowledge vector represented by sensor i. It represents the covariance between sensor i and sensor j. Finally, the sensors are divided into 9 groups through clustering to guide the construction of deep learning model architecture.
[0070] The final sensor clustering results are as follows: Group 1 [S8, S13, S18, S19], Group 2 [S4, S10, S21], Group 3 [S1, S5], Group 4 [S2], Group 5 [S17], Group 6 [S9, S14], Group 7 [S6, S15], Group 8 [S2, S16, S20] and Group 9 [S3, S7, S11]. For example, in Group 1, S8 represents the physical fan speed, S13 represents the corrected fan speed, S18 represents the demand fan speed, and S19 corresponds to the demand-corrected fan speed. These sensors are clustered together and are all related to different aspects of the fan speed, so the clustering results are meaningful. The sensor groups effectively reflect the efficiency-related components of the physical engine entity, thereby affecting its operating performance. In order to capture the spatial features in these groups, a CNN network is used. After sliding window processing, the 3D data (batchsize, window length, sensor dimension) is first expanded to grayscale format (batchsize, 1, window length, sensor dimension). Then, 2D convolution is applied with the kernel size set to (4, the number of sensors in each cluster) and a sliding stride of 2. The output width of each convolution head is set to 1. This produces 9 feature maps, which are concatenated to ensure a uniform metric scale across features. These concatenated feature maps are then passed to the temporal feature extraction network, which further captures the temporal features associated with degradation.
[0071] Although existing RNN methods are effective for life prediction regression prediction, they usually rely on the last time step for prediction, which may lead to the loss of key information and thus reduce accuracy. In fact, not all time steps contribute equally to RUL prediction, and some steps carry more valuable information than others. To address this issue and ensure rich and comprehensive feature extraction, the present invention designs a GRU network with an enhanced self-attention mechanism, built on the basis of spatial feature extraction that contains domain knowledge. The method of the present invention adaptively assigns higher weights to more critical features and time steps, thereby enhancing the model's ability to capture long-term dependencies while reducing noise. GRU-SAM driven by the self-attention mechanism is shown in Figure 2. Figure 4 As shown, the gated recurrent unit (GRU) includes a reset gate and an update gate, which can be described as follows:
[0072] rt =σ(U r x t +W r h t-1 +b r )
[0073] z t =σ(U z x t +W z h t-1 +b z )
[0074]
[0075] Among them, r t and z t Represent the outputs of the update gate and reset gate respectively. t is the input at time step t, h t-1 and h t are the hidden states at time t-1 and time t respectively. r , W r , U Z , W Z , U h and W h is the weight, b r , b z and b h is the bias, σ and tanh are the sigmoid and hyperbolic tangent functions respectively.
[0076] First, the feature X extracted from the deep space t Put it into the GRU network, X t As the spatial features extracted by the convolutional neural network combined with domain knowledge, after convolution and feature concatenation, the last dimension is the number of sensor groups 9, which is used as the input dimension of the GRU network, and the number of hidden layers is set to 50. Assume that the features learned by the GRU network for a sample can be expressed as h = [h1,h2,...,h n ] T , T is the transposition operation, n is the number of neurons in the hidden layer of the GRU network, the information obtained by each neuron in the GRU network is extracted, and the extracted features are put into the self-attention mechanism. First, based on the self-attention mechanism, the importance of different sequential steps of the input hi is expressed as:
[0077]
[0078] Among them, W and b are the weight matrix and bias vector respectively, and Φ(·) is the activation function sigmoid. In this way, the contribution of each neuron can be obtained. After obtaining the score of the i-th eigenvector, the softmax function can be used to calculate the weight as follows:
[0079]
[0080] Finally, each obtained score α i and the output of the initial neuron h i Multiply them together to get the final feature output vector. The features obtained in this way realize time step self-attention and avoid the problem of early important time step information loss. The final output feature O of the attention mechanism can be expressed as:
[0081]
[0082] Among them, α i represents the attention weight of the i-th time step, h i Represents the corresponding output of GRU at that time step.
[0083] At the shallow feature extraction stage: Since current RUL prediction methods usually emphasize deep features extracted by neural networks, such as CNN networks for spatial feature extraction and GRU networks for time series feature extraction. However, these methods often ignore features manually derived by domain experts, which can provide key insights into physical degradation processes that deep learning methods may not be able to fully capture. For example, simple indicators such as the mean and trend coefficient of linear regression provide intuitive insights into sensor behavior: the mean represents the size of the sensor data, while the trend coefficient reveals the speed and direction of degradation, such as Figure 5 To address this gap, a feature fusion framework is proposed, which integrates the spatiotemporal features extracted by deep learning and the manually extracted features.
[0084] S3: The effectiveness of the proposed method is fully evaluated on the publicly available aero-engine dataset C-MAPSS in the RUL domain;
[0085] The C-MAPSS dataset has 4 sub-datasets FD001-FD004, each with a different number of operating conditions and fault conditions, and each sub-dataset is divided into a training subset and a test subset. The data contains a total of 26 columns: the first column represents the engine ID, the second column represents the current number of operating cycles, columns 3-5 represent three operation settings related to the operating status, and columns 6-26 represent 21 sensor values. The engine runs normally at the beginning of each time series, and the fault occurs at an unspecified time point. The training set records the fault progression until the system failure, while the test set contains sensor data before the fault point, aiming to predict the remaining operating cycles. All sub-datasets are used to evaluate the effectiveness of the method of the present invention.
[0086] First, the dataset is divided into a training set and a validation set in a ratio of 95:5. The validation set is used to fine-tune the model parameters by repeatedly testing and adjusting the network to achieve the best performance. Data preprocessing for both the training set and the test set involves normalization, label addition, and sliding window interception processing to ensure the consistency and relevance of the input data. This experiment uses Intel(R) Core(TM) i7-10510U CPU@1.80GHz, 64-bit Microsoft Windows operating system, and is implemented using the Python 3.7.0 framework. During the training phase, the Adam optimizer was used in the present invention, with a learning rate of 0.001 for the first 15 epochs and 0.0001 thereafter, and the batch size was set to 256. At the same time, an early stopping strategy was adopted, and training would stop if the validation loss did not improve for 10 consecutive epochs. The model integrates features extracted from sensors grouped according to domain knowledge. These spatial features are obtained using group convolutions, which capture the interdependencies between sensors within each group. The extracted spatial features are combined with the temporal features using a self-attention guided GRU network, which assigns different weights to time steps to effectively capture key temporal information related to engine degradation. After feature extraction, the model is trained on the preprocessed data, and a validation set is used to fine-tune the model parameters and improve its prediction accuracy. During the testing phase, the processed test data is fed into the trained neural network to generate real-time RUL predictions. The method of the present invention combines neural network features with manually extracted features to improve the accuracy of RUL predictions by capturing complex degradation patterns. The training process and testing process are as follows: Figure 6 For performance evaluation, the present invention uses the root mean square error RMSE and the scoring function Score, which are defined as follows:
[0087]
[0088] in, is the RUL prediction value of sample i (i = 1, 2, ..., N), yi is the true value of RUL of sample i. N is the total number of test samples.
[0089] To evaluate the effectiveness of the proposed KDMFusion, it is compared with several state-of-the-art RUL prediction methods using the same NASA dataset from the past eight years. For methods whose results are not available, they are marked with a backslash ('\'). The comparison results in Tables 1 and 2 highlight the performance of each method in terms of RMSE and Score metrics.
[0090] KDMFusion shows superior performance compared to traditional CNN-based models (CNN, DCNN, SSE-CNN, and MS-DCNN). While these models excel in spatial feature extraction, our approach enhances this capability by integrating domain-specific knowledge through grouped convolutions, which can capture degradation patterns in time series data in a more detailed manner. Compared to attention-based models that focus on improving temporal feature extraction through attention mechanisms (such as MCLSTM, Attention-LSTM, and ABGRU), KDMFusion stands out by incorporating a self-attention mechanism into the GRU network. The dynamic adjustment of attention weights at each time step enhances the model's ability to manage long-term dependencies and reduces noise interference. Among fusion-based methods such as EAGDE-SVM, VAE+RNN, and CATA-TCN, KDMFusion stands out by integrating spatiotemporal features with manually designed features through a dual-path multi-level feature extraction framework. This combination improves feature diversity, model interpretability, and reduces reliance on a single feature type, resulting in significant improvements in predictive performance.
[0091] Table 1 Comparison results of the KDMFusion method of the present invention with other methods in terms of RMSE, with the best result marked in bold.
[0092]
[0093]
[0094] Table 2: Comparison of the Score results of the KDMFusion method in this application and other methods. The best results are marked in bold.
[0095]
[0096]
[0097] In order to evaluate the contribution of different components of the proposed KDMFusion, ablation experiments were performed using the original GRU network as a baseline, and four configurations were evaluated: Model 1 combines traditional convolution with the original GRU network for spatial feature extraction (CNN+GRU); Model 2 uses grouped convolution with domain knowledge (knowledge-guided CNN+GRU); Model 3 adds a time-step self-attention mechanism to GRU (knowledge-guided CNN+GRU+Attention); Model 4 combines all previous components with a shallow hand-crafted feature extraction network (KDMFusion).
[0098] Table 3 shows the experimental results of various ablation methods applied to four sub-datasets. As observed, the baseline model provides basic temporal feature extraction but lacks advanced spatial feature extraction and domain-specific insights. Model 1 is introduced, which combines traditional convolutional layers for spatial feature extraction with GRU networks for temporal processing, thereby improving spatial feature capture. However, it does not integrate domain knowledge for sensor grouping, resulting in inconsistent performance. Specifically, Model 1 shows degraded performance on FD001 and FD004, and only slightly improves on FD002 and FD003. The lack of domain-specific knowledge leads to suboptimal sensor alignment, which affects the overall effectiveness of the model. Model 2 builds on this by integrating domain knowledge into the grouped convolution process, enhancing sensor grouping based on interdependencies. This adjustment leads to significant performance improvements on all datasets compared to Model 1. The significant enhancements on more complex datasets such as FD002 and FD004 show that domain knowledge greatly optimizes sensor alignment and improves spatial feature extraction. Model 3 further enhances the GRU network through the time-step self-attention mechanism, improving its ability to manage long-term dependencies and reduce noise, thereby improving the performance of all datasets. Finally, Model 4 of the proposed KDMFusion method integrates all previous components and adds a shallow hand-crafted feature extraction network. By integrating spatial and temporal feature extraction, domain knowledge, and shallow feature extraction, KDMFusion excels in capturing complex degradation patterns, especially on FD002 and FD004, resulting in excellent RUL prediction accuracy and enhanced interpretability.
[0099] In summary, the performance differences between the models emphasize the effectiveness of each component in our invention. Integrating domain knowledge significantly enhances sensor alignment and spatial feature extraction, while the time-step self-attention mechanism in the GRU network improves the handling of long-term dependencies and reduces noise. In addition, integrating shallow feature extraction provides complementary insights. Overall, these modules enable the KDMFusion method to greatly improve its ability to capture complex degradation patterns and achieve excellent RUL prediction accuracy.
[0100] Table 3: Impact results of ablation experiments (RMSE / SCORE)
[0101]
[0102] Since the sensor grouping is guided by integrating the knowledge of the engine component structure, choosing the right time window is crucial to accurately capture the sensor dependencies. The optimal window size ensures that the model training contains enough data to obtain the best prediction results. The present invention varies the size of the sliding window in the range of [30, 40, 50, 60, 70] of the four C-MAPSS sub-datasets. The effects of these window sizes on the RMSE and Score indicators are shown in Figure 2. Figure 7 As shown. Based on the experimental results, the window sizes selected for datasets FD001 to FD004 are 40, 60, 60, and 60, respectively. These selections are consistent with the increasing complexity and variability of operating and fault conditions from FD001 to FD004. Larger time windows integrated with the knowledge-guided dual-path multi-feature fusion framework of the present invention provide richer data for capturing extended feature series in complex data. They also help smooth time series data, reduce noise, and improve RUL prediction performance.
[0103] This implementation scheme also provides a knowledge-guided dual-path multi-feature fusion system KDMFusion for accurately predicting the remaining useful life RUL of aircraft engines, which consists of three main sub-modules: data preprocessing module, feature extraction module and prediction regression module. In the data preprocessing process, the original sensor data is normalized, label added and sliding window intercepted to ensure that the processed data is not limited by dimension and is representative enough. Secondly, in the feature extraction stage, a dual-branch structure is used to effectively capture multi-level features related to engine degradation. One branch combines the convolutional neural network CNN based on domain knowledge and the gated recurrent unit GRU network driven by the self-attention mechanism to extract the deep features of the data. The other branch uses intuitive manual features such as the mean and trend coefficient of linear regression to extract the shallow features of the data. Finally, in the prediction regression module, after the deep features are fused with the shallow features, a fully connected layer with an activation function is used to integrate and regress all the feature representations obtained after the fusion.
[0104] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These changes and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A machine remaining useful life prediction method based on knowledge-guided dual-path multi-feature fusion, characterized in that: include: S1: Data preprocessing: Use the min-max normalization method to process the data of the aircraft engine dataset in the RUL model; S2: Feature extraction: In a self-supervised manner, a convolutional neural network (CNN) enhanced with domain knowledge is used to extract spatial features, a GRU network driven by a self-attention mechanism is used to extract time series features, and manual features are combined for feature fusion.
2. The machine remaining useful life prediction method based on knowledge-guided dual-path multi-feature fusion according to claim 1 is characterized in that: In S1, the min-max normalization method is used to map the data dimension of the RUL model to [-1,1], and the following specific equation is used to process the data: in, and denote the jth sensor in the ith period before and after normalization, respectively; and are the minimum and maximum values of the jth sensor of the same engine in the data; Set the RUL threshold to 125; A time window sequence of size D*W is used as the training sample, where D represents the data dimension and W represents the window width: the first sliding window sample is represented as T1=[X1,X2,...,X W ], after sliding with a step length of L, the next sequence sample data is T2 = [X1+L,X2+L,...,X W +L, the corresponding data label is the RU prediction label of the last time step.
3. The machine remaining useful life prediction method based on knowledge-guided dual-path multi-feature fusion according to claim 1 is characterized in that S2 include: S21: Modeling the physical relationship of engine components, the correlation ρ between sensor i and sensor j is defined as follows: in, represents the variance of the knowledge vector represented by sensor i, represents the covariance between sensor i and sensor j; S22: clustering the sensors into 9 groups according to the physical relationship of the engine components; S23: Use the CNN network to capture the spatial features in each group of sensors and obtain a concatenated feature map; S24: The GRU network driven by the self-attention mechanism is used to extract time series features from the feature map. The formula is defined as follows: r t =σ(U r x t +W r h t-1 +b r ) z t =σ(U z x t +W z h t-1 +b z ) Among them, r t and z t Represent the outputs of the update gate and reset gate, respectively, X t is the input at time step t, h t-1 and h t are the hidden states at t-1 and t respectively, U r , W r , U Z , W Z , U h and W h is the weight, b r , b z and b h is the bias, σ and tanh are sigmoid and hyperbolic tangent functions respectively; S25: Extract the feature X from the deep space t Put it into the GRU network, X t As the spatial features extracted by the convolutional neural network combined with domain knowledge, after convolution and feature concatenation, the last dimension is the number of sensor groups 9, which is used as the input dimension of the GRU network, and the number of hidden layers is set to 50; Assume that the feature representation learned by the GRU network for a sample is h = [h1,h2,...,h n ] T , T is the transposition operation, n is the number of neurons in the hidden layer of the GRU network, and the information obtained by each neuron in the GRU network is extracted; Based on the self-attention mechanism, the input h i The importance of different sequential steps is expressed as: s i =Φ(W T h i +b) Where W and b are the weight matrix and bias vector respectively, and Φ(·) is the activation function sigmoid; After obtaining the score of the i-th eigenvector, the weight is calculated using the softmax function as follows: Each obtained score α i and the output of the initial neuron h i Multiply them together to get the final feature output vector. The final output feature O of the self-attention mechanism is expressed as: O=[h1α1h2α2...h n a n ] Among them, α i represents the attention weight of the i-th time step, h i Represents the corresponding output of GRU at this time step.
4. The method for predicting the remaining useful life of a machine based on knowledge-guided dual-path multi-feature fusion according to claim 1 is characterized in that: Also includes: The proposed method is evaluated on the publicly available aero-engine dataset C-MAPSS in the RUL domain.
5. A machine remaining useful life prediction system based on knowledge-guided dual-path multi-feature fusion, characterized in that: include: Data preprocessing module: used to normalize the raw sensor data; Feature extraction module: It is used to extract spatial features using a convolutional neural network combined with domain knowledge in a self-supervised manner, extract time series features using a GRU network driven by a self-attention mechanism, and realize feature fusion by combining traditional manual features; Evaluation module: The effectiveness of the prediction method is evaluated on the publicly available aircraft engine dataset C-MAPSS in the RUL field.
6. The machine remaining useful life prediction system based on knowledge-guided dual-path multi-feature fusion according to claim 5 is characterized in that: The feature extraction module includes a deep feature extraction module and a shallow feature extraction module; The deep feature extraction module adopts a spatiotemporal progressive approach, using a CNN enhanced by domain knowledge to capture spatial features and a GRU driven by self-attention to capture temporal features; The shallow feature extraction module adopts a feature fusion framework, which integrates the spatiotemporal features extracted by deep learning and the features extracted manually.
Citation Information
Cited By
Method for predicting residual service life of industrial equipment based on DMM-JA model
CN121167627A