Equipment life prediction method based on multi-source fusion model and application
Through the multi-source fusion model, the CNN-Informer model is used to extract device features and combine Informer self-attention mechanism to predict the device life, solving the problem of low prediction accuracy caused by the complexity of device degradation characteristics, and achieving higher prediction accuracy and stability.
Patent Information
- Application Number
- CN202510216682.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-20
AI Technical Summary
When dealing with the complexity of equipment degradation characteristics, existing equipment life prediction methods have low prediction accuracy and insufficient stability, especially long-term dependency problems and insufficient serial information capture capabilities.
The device life prediction method based on multi-source fusion model is adopted, and the device features are extracted through the CNN-Informer model, combined with the Informer's self-attention mechanism for global modeling, and the DCNN-Informer model is constructed to perform feature learning and prediction using the secondary training and interval probability prediction methods.
It improves the accuracy and stability of equipment life prediction, enhances the generalization ability of the model, and optimizes the prediction performance.
Smart Images

Figure CN120180869A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of fault prediction and health management, and particularly to a method, device, and electronic device for predicting the life of a device based on a multi-source fusion model, as well as a computer-readable storage medium. Background Technique
[0002] The working reliability of an aero-engine is an important factor in measuring flight safety. The working conditions of the engine are complex, and any fault may lead to a serious accident. Therefore, predicting the remaining useful life (RUL) of an aero-engine is crucial for evaluating the health status of the engine and ensuring the stable operation of the aircraft. Prognostics and Health Management (PHM) originated in developed countries in Europe and America in the 1980s and is a method for detecting, diagnosing, and predicting the state of a device or system. Predicting the remaining life is an important part of PHM. By predicting the RUL in the early stage of the degradation of a turbofan engine, corresponding maintenance or repair measures can be taken to prevent faults from occurring and save economic costs. There are many mature methods for RUL prediction, mainly divided into physical model-based methods and data-driven methods.
[0003] The method based on physical model prediction uses the physical characteristics and operating parameters of a device or system to estimate its remaining useful life. By understanding the working principle of the device and the mechanism of fault occurrence, a model is established to simulate the operating state and life consumption process of the device, thereby predicting the life of the device. Common physical model-based methods include state space algorithms such as Kalman filtering and particle filtering, as well as Boolean distributions. Cai et al. constructed a double non-linear hidden degradation model using a non-linear Wiener process, gradually updating the random coefficients and historical degradation conditions, thereby realizing the prediction of the remaining useful life. Wang et al. proposed a mechanical degradation state prediction method based on particle filtering, integrating physical knowledge and process measurements into the state space framework, improving the robustness and accuracy of the prediction of the fan bearing. Li Wei et al. introduced the Kalman filtering algorithm on the basis of support vector machines to correct the time series results, taking turning processing as the research object, improving the accuracy of tool wear state recognition, and at the same time solving the problems of slow convergence speed and easy to fall into local minimum of the conventional model.
[0004] Data-driven prediction methods directly use the historical operating parameters and real-time monitoring data of devices. Generally, this method first analyzes the data using signal processing techniques to extract information that can reflect system degradation and fault characteristics, then feeds this information into a model for training, and finally establishes a prediction model. Compared with physics-based model methods, data-driven methods do not require knowledge of the physical characteristics and working principles of the device. By simply analyzing the patterns and trends in the data, a prediction model can be quickly established to achieve the purpose of predicting the device's lifespan. For example, Liu et al. proposed an encoder-decoder model for RUL prognosis. Bi-LSTM and CNN were combined in the encoder to obtain temporal relationships and basic features from the time series, while a fully connected layer was used in the decoder to predict RUL by decoding the feature information. Hu et al. proposed an automatically expandable long short-term memory network model, which is based on a multi-level prediction method. Through a sub-module structure connected step by step, the output error of the previous level is used as the training value of the next level to form a multi-level error correction mechanism, thereby improving the prediction accuracy. In addition to LSTM, convolutional neural network (CNN) is another popular deep learning algorithm in the field of RUL. CNN has strong representation learning ability and can extract useful local features from data. Babu et al. used CNN to predict the RUL of aeroengines. In their study, the data was segmented using time windows, and convolutional operations were performed separately on the time dimension of the data. Qin et al. proposed a method for predicting the remaining useful life of aeroengines based on multi-scale feature fusion. The data noise interference was reduced through statistical denoising techniques, and a weighted spatio-temporal feature extraction module was designed by combining convolutional bidirectional long short-term memory network and multi-head attention mechanism to extract data features from multiple time scales. After fusing the manually extracted degradation information, it was input into a fully connected network to achieve high-precision prediction.
[0005] Although the above methods have achieved good performance, there are still some limitations. The RUL prediction method based on LSTM usually leads to long-term dependence problems, especially when historical information is needed to perform the current task. Since the gradient usually disappears after propagating through several stages, for long-term dependence, information accumulation requires several time steps to establish, and the farther the distance, the less likely it is to capture effective information. The Transformer model has achieved excellent results in object detection, traffic flow prediction, image segmentation, etc. However, the Transformer model ignores the importance of different characteristics in a single time step of the sequence, so it is also difficult to achieve high-accuracy lifespan prediction. Summary of the Invention
[0006] To overcome the defects of the above-mentioned existing technologies, embodiments of the present invention provide a device life prediction method and application based on a multi-source fusion model, which can solve the problem of low prediction accuracy of the remaining life of a device caused by the increasing complexity of the device degradation characteristics, and can improve the stability and accuracy of device life prediction compared with existing models.
[0007] On the one hand, embodiments of the present invention propose a device life prediction method based on a multi-source fusion model, including: obtaining device life data and performing feature normalization, setting RUL labels for the obtained normalized data, and extracting samples through a sliding window to obtain a preprocessed training set; establishing a CNN-Informer model, and inputting the data of the preprocessed training set into the CNN-Informer model for pre-training to obtain a pre-trained model; grouping devices according to device characteristics and respectively inputting them into the pre-trained model based on the groups to obtain secondary training models characterizing the device characteristics; testing the trained secondary training models on a test set to obtain an initial RUL prediction result and performing filtering, thereby constructing a DCNN-Informer model; processing the filtered initial RUL prediction result using an interval probability prediction method, thereby constructing a multi-source fusion model based on the DCNN-Informer model, and obtaining a final target RUL prediction result using the multi-source fusion model.
[0008] In an embodiment of the present invention, the feature normalization adopts min-max normalization, and the formula is: In the formula, is the j-th output of the i-th sensor of the engine, is the minimum value of all output values of the i-th sensor, is the maximum value of all output values of the i-th sensor.
[0009] In an embodiment of the present invention, the CNN-Informer model includes: a CNN feature extraction layer and an Informer Encoder layer, and the data processing flow includes: performing feature extraction operations through CNN; sending the features extracted by CNN to the Informer module, flattening, linearly mapping, adding classification markers to the extracted feature vectors, and then performing position encoding to complete serialization; inputting the serialized data into the Informer encoder, and completing feature learning through normalization, multi-head attention mechanism, and multi-layer perceptron; inputting the learned prediction features into a fully connected layer to predict the remaining service life of the device.
[0010] In one embodiment of the present invention, the CNN consists of an input layer, a convolutional layer, a pooling layer, a fully connected layer, and an output layer. Among them, the feature vector input by the input layer extracts features through the convolution operation of the convolutional layer, and then simplifies the complexity of network calculation through the pooling layer. All features are linked by the fully connected layer, and the obtained result is output by the output layer; the convolutional layer and the result of its output vector are expressed as: where σ is the sigmoid activation function, b j is the bias of the feature map, w is the weight of the kernel, f v is the filter index, and x is the input vector; the pooling layer represents dimensionality reduction, and its operation formula is: where R is the pooling size, T is the step size determining the moving distance of the input data region, y is the input size, and T < y. The fully connected layer connects each neuron in the pooling layer to each neuron in the output layer.
[0011] In one embodiment of the present invention, the Informer consists of an input layer, positional encoding, self-attention distillation, and a fully connected layer; among them, the input layer maps the input data to a d-dimensional vector Z ∈ R n×c ; the positional encoding accepts the input sequence and maps it to a high-dimensional vector, and then feeds it to the decoder to generate the output sequence; the self-attention distillation calculates the sparsity metric of any Q vector according to the query sparsity metric formula, and selects the Q vector that plays a dominant role in the attention calculation, and calculates the attention with this vector. The formula is: where Q is the query vector, K is the key vector, V is the value vector, represents calculating the similarity between the query and the key, d is the dimension of the key, and Softmax is the activation function; the fully connected layer adds the activation function, and the output passes through the residual connection and layer normalization again.
[0012] In one embodiment of the present invention, grouping the devices according to the device characteristics and respectively sending them into the pre-trained model based on the group to obtain a secondary training model representing the characteristics of each device includes: grouping the data in the training set according to each device, where the first engine is represented as y1, and the nth engine is represented as y n ; sending y1…,y n into the pre-trained model to obtain a model for each device; dividing the data X of each device into multiple sliding windows TW1…TWn to obtain subsequences X i1 …X im , where m represents the number of subsequences after division; for each subsequence X ij , use the pre-trained model for training to obtain sub-models AM1…AMn.
[0013] In one embodiment of the present invention, the prediction result adopts second-order exponential smoothing filtering, and the formula is: L t =αy t +(1-α)(L t-1 +T t-1 )t; T t =β(L t -L t-1 )+(1-β)T t-1 ; F t+m =L t +mT t ; where L t is the horizontal estimated value at time point t; T t is the trend estimated value at time point t; α is the smoothing coefficient of the level, which controls the response degree of the horizontal estimated value to the current observed value and the long-term trend law of the time series, and α = 0.1 - 0.3 is taken; β is the smoothing coefficient of the trend, which controls the response degree of the trend estimated value to the change of the horizontal estimated value; F t+m is the predicted value for the next m periods.
[0014] In one embodiment of the present invention, the interval probability prediction method includes: generating M independent sub-models {M1, M2,..., M M} based on the DCNN-Informer model with different hyperparameter combinations; performing K rounds of training using the M sub-models, where the predicted value of the k-th round of the i-th sub-model is: All predicted values form a set Introduce kernel density to estimate the probability distribution of the predicted values, and the formula is: where h is the bandwidth parameter, which controls the smoothing degree; K(·) is the kernel function; calculate the probability density of each prediction point according to the probability distribution, and perform weighted average on the predicted life values based on the probability density, and the formula is: where each predicted value is multiplied by its corresponding probability density and then added together to obtain the final predicted result of the target RUL
[0015] On the other hand, an embodiment of the present invention further provides a device life prediction device based on a multi-source fusion model, including: a data preprocessing module, configured to obtain device life data and perform feature normalization, set RUL labels for the obtained normalized data, and extract samples through a sliding window to obtain a preprocessed training set; a model pre-training module, configured to establish a CNN-Informer model, and input the preprocessed training set data into the CNN-Informer model for pre-training to obtain a pre-trained model; a model secondary training module, configured to group devices according to device characteristics, and separately input them into the pre-trained model based on the groups to obtain a secondary training model characterizing each device characteristic; an initial RUL prediction module, configured to test the trained secondary training model on a test set, obtain an initial RUL prediction result and perform filtering, so as to construct a DCNN-Informer model; a target RUL prediction module, configured to process the filtered initial RUL prediction result using an interval probability prediction method, thereby constructing a multi-source fusion model based on the DCNN-Informer model, and obtaining a final target RUL prediction result using the multi-source fusion model.
[0016] On yet another aspect, an embodiment of the present invention further provides an electronic device, including: a memory and one or more processors connected to the memory, the memory stores a computer program, and the processor is configured to execute the computer program to implement the device life prediction method based on a multi-source fusion model as described in any one of the above embodiments.
[0017] On still another aspect, an embodiment of the present invention further provides a computer-readable storage medium, the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are configured to execute the device life prediction method based on a multi-source fusion model as described in any one of the above embodiments.
[0018] As can be seen from the above, compared with the prior art, the above embodiments of the present invention can at least have one or more of the following beneficial effects:
[0019] The high-dimensional spatial features of time series data are extracted using a CNN convolutional neural network, and the self-attention mechanism of Informer is combined to globally model these features, so as to fully extract the information in the time dimension. In addition, an engine secondary training framework is designed. The dataset is grouped by engine, and the data of each group is sent into the CNN-Informer model for secondary training to obtain a personalized model for each engine, further improving the accuracy and generalization ability of the model. The second-order exponential smoothing filtering algorithm and the interval probability prediction method are used to process the prediction results of the model, so as to construct a multi-source fusion model to improve the stability and accuracy of the prediction. This prediction model has obvious advantages in RUL prediction and is superior in prediction performance compared with existing models. Description of the Drawings
[0020] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0021] Figure 1 is a flowchart of a device life prediction method based on a multi-source fusion model provided by an embodiment of the present invention;
[0022] Figure 2 is a schematic diagram of the specific execution logic of a device life prediction method based on a multi-source fusion model provided by an embodiment of the present invention;
[0023] Figure 3 is a schematic diagram of the CNN structure provided by an embodiment of the present invention;
[0024] Figure 4 is a schematic diagram of the Informer structure provided by an embodiment of the present invention;
[0025] Figure 5 is a schematic diagram of the sensor dataset provided by an embodiment of the present invention;
[0026] Figure 6 is a schematic diagram of the data trend after feature normalization provided by an embodiment of the present invention;
[0027] Figure 7 is a schematic diagram of the effect of pre-training of the CNN-Informer model provided by an embodiment of the present invention;
[0028] Figure 8 is a schematic diagram of the secondary training effect of the pre-trained model provided by an embodiment of the present invention;
[0029] Figure 9 is a schematic diagram of the prediction result of the secondary training model provided by an embodiment of the present invention;
[0030] Figure 10 Schematic diagram of the effect after the prediction result filtering operation provided by the embodiment of the present invention;
[0031] Figure 11 Schematic diagram for comparing the predicted values and true values of the pre-trained model, secondary training model, and smoothing filtering model provided by the embodiment of the present invention;
[0032] Figure 12 Error distribution diagram of the prediction results and actual values of several models provided by the embodiment of the present invention on the test set;
[0033] Figure 13 Box plot of the prediction metrics RMSE and Score of several models provided by the embodiment of the present invention;
[0034] Figure 14 Schematic diagram of the structure of a device life prediction device based on a multi-source fusion model provided by the embodiment of the present invention;
[0035] Figure 15 Schematic diagram of the structure of an electronic device provided by the embodiment of the present invention;
[0036] Figure 16 Schematic diagram of the structure of a computer-readable storage medium provided by the embodiment of the present invention. Detailed implementation manners
[0037] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other. The present invention will be described below with reference to the accompanying drawings and in conjunction with the embodiments.
[0038] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments, and all should fall within the protection scope of the present invention.
[0039] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0040] It should also be noted that the division of multiple embodiments in the present invention is only for the convenience of description and should not constitute a special limitation. The features in various embodiments can be combined and referenced to each other without contradiction.
[0041] As Figure 1 shown, the first embodiment of the present invention proposes a device life prediction method based on a multi-source fusion model, for example, including: Step S1, obtaining device life data and performing feature normalization, setting RUL labels for the obtained normalized data, and extracting samples through a sliding window to obtain a preprocessed training set; Step S2, establishing a CNN-Informer model, and inputting the data of the preprocessed training set into the CNN-Informer model for pre-training to obtain a pre-trained model; Step S3, grouping devices according to device characteristics and respectively inputting them into the pre-trained model based on the groups to obtain a secondary training model representing the characteristics of each device; Step S4, testing the trained secondary training model on a test set to obtain an initial RUL prediction result and performing filtering, thereby constructing a DCNN-Informer model; Step S5, using an interval probability prediction method to process the filtered initial RUL prediction result, thereby constructing a multi-source fusion model based on the DCNN-Informer model, and using the multi-source fusion model to obtain a final target RUL prediction result.
[0042] Specifically, as Figure 2 shown, in Step S1, since the measurement ranges of different sensors are different and the data ranges vary greatly, it will have a great impact on the final engine life prediction. Therefore, it is necessary to perform feature normalization on the original data. Commonly used normalization methods include standardization (z-score normalization) and min-max normalization. The formula for min-max normalization is as follows:
[0043]
[0044] In the formula, is the j-th output of the i-th sensor of the engine, is the minimum value of all output values of the i-th sensor, is the maximum value of all output values of the i-th sensor.
[0045] Since there are no labels in the original dataset, RUL labels are added for model training. The sliding window is a commonly used data processing method. After data normalization, sliding time windows can be used to extract samples to increase the number of samples during training. Its basic principle is to divide the dataset into windows of equal size and then process the data step by step by sliding the window. Sliding windows are often used for training models and predicting future values. By training and predicting within the sliding window, the time information and trends in the sequence data can be effectively utilized to improve the accuracy of prediction. Set the sliding window size to 32. Using this method can help the model better capture the time correlations and features in the data, thereby improving the model's ability to predict future values.
[0046] In step S2, the preprocessed data is passed into the CNN-Informer model for training. The purpose of pre-training is to enable the model to learn the general characteristics and patterns of the data. After pre-training is completed, a CNN-Informer pre-trained model is obtained, which serves as the basis for subsequent secondary training.
[0047] CNN is the most widely used neural network model currently. As Figure 3 shown, the overall structure of CNN is relatively simple, mainly consisting of an input layer, a convolutional layer, a pooling layer, a fully connected layer, and an output layer, etc. The input feature vector is used to extract features through convolutional operations, and then the complexity of network calculations is simplified by the pooling layer. Secondly, all features are linked by the fully connected layer to obtain the final output result.
[0048] The convolutional layer is the main block of CNN, which obtains local attributes from the higher-level input and passes all the information to the lower level to obtain more complex features. The result of the first convolutional layer and its output vector can be expressed by the following formula:
[0049]
[0050] where σ is the sigmoid activation function, b j is the bias of the feature map, w is the weight of the kernel, f v is the filter index, and x is the input vector.
[0051] The max pooling layer is used to reduce the dimensionality of the representation, thereby further reducing the computational burden of the model.
[0052] The operation formula of the max pooling layer is:
[0053]
[0054] Among them, R is the pooling size, T is the step size that determines the moving distance of the input data region, y is the input size, with T < y. The fully connected layer connects each neuron in the pooling layer to each neuron in the output layer.
[0055] The Informer model consists of an input layer, positional encoding, self-attention distillation, and a fully connected layer. The input layer maps the input data into a d-dimensional vector Z ∈ R n×c , preparing for the subsequent feature extraction process. In the Encoder part, it receives extremely long input data. The Encoder module improves the robustness of the algorithm by stacking the above two operations. The encoder receives the input sequence and maps it to a high-dimensional vector, and then feeds it to the decoder to generate the output sequence. In this paper, the encoder is used to learn the long-term correlations of degradation from the engine's operation data records. The self-attention layer is used to capture the dependencies between features. Different from the self-attention layer in Transformer, it does not calculate the dot product of each vector in the Q matrix and each vector in the K matrix. Instead, it calculates the sparsity metric of any Q vector according to the query sparsity metric formula, and selects the Q vector that plays a dominant role in the attention calculation, and calculates the attention with this vector. The specific formula is as follows:
[0056]
[0057] where Q is the query vector, K is the key vector, V is the value vector, represents calculating the similarity between the query and the key, d is the dimension of the key, and Softmax is the activation function.
[0058] The fully connected layer adds an activation function and passes the output through residual connection and layer normalization again.
[0059] The Informer model can extract long-term correlations and local features from the data, and can accurately capture the long-term dependencies between the output and the input. CNN performs well in processing local features, but is weak in processing global information. While Informer can handle the modeling and generation of sequence data, it has advantages in processing global information, but its ability to process local information is relatively weak. By combining CNN and Informer, the local and global information in the features can be effectively captured and processed, thereby improving the performance and effect of the model.
[0060] Through the above analysis, the CNN-Informer model proposed in this embodiment is as Figure 4 shown. This model consists of two parts: the CNN feature extraction layer and the Informer Encoder layer. The working process based on the CNN-Informer model is as follows:
[0061] 1) Feature extraction is performed through CNN.
[0062] 2) The features extracted by CNN are sent to the Informer module. After flattening, linearly mapping, adding classification tokens, and performing positional encoding on the extracted feature vectors, serialization is completed.
[0063] 3) After serialization, it is input into the Informer encoder, and feature learning is completed through normalization, multi-head attention mechanism, and multi-layer perceptron.
[0064] 4) The learned prediction features are input into the fully connected layer to predict the remaining useful life of the engine.
[0065] In step S3, since the operating characteristics of each engine are different, after pre-training through the CNN-Informer model, a pre-trained CNN-Informer model is obtained. Then, the data in the training set is grouped by each engine. The first engine is denoted as y1, and the nth engine is denoted as y n , and then y1…,y n are fed into the pre-trained model to obtain models for each engine. The data X of each engine is divided into multiple sliding windows TW1…TWn to obtain subsequences X i1 …X im , where m represents the number of subsequences after division. For each subsequence X ij , the pre-trained model is used for training to obtain sub-models AM1…AMn.
[0066] This targeted model training method can better consider the uniqueness of each engine, thereby improving the accuracy of the model. By training a separate model for each engine, we can better capture the differences between engines, and thus more accurately predict the remaining useful life of each engine.
[0067] In step S4, when there is a certain trend in the predicted data, in order to assign higher weights to longer observations, for example, second-order exponential smoothing filtering (DES) is adopted to construct the DCNN-Informer model. This method can make the difference between the prediction result and the actual data become smooth. Second-order exponential smoothing filtering has the advantages of simple calculation, less sample requirements, strong adaptability, and stable results. The formula is:
[0068] L t =αy t +(1-α)(L t-1 +T t-1 )t;
[0069] T t =β(L t-L t-1 )+(1 - β)T t-1 ;
[0070] F t+m = L t + mT t ;
[0071] where L t is the horizontal estimate at time point t, T t is the trend estimate at time point t, α is the smoothing coefficient of the horizontal, controlling the response degree of the horizontal estimate to the current observed value and the long-term trend law of the time series, taking α = 0.1 - 0.3, β is the smoothing coefficient of the trend, controlling the response degree of the trend estimate to the change of the horizontal estimate, and F t+m is the predicted value for the next m periods.
[0072] During the flight of an aircraft, the state of the engine is often affected by various factors, showing a high degree of uncertainty. Existing research usually takes point prediction as the task goal, estimating a definite RUL value for each device. However, deterministic estimation is difficult to meet the requirements of maintenance decision-making in engineering practice. Even though some scholars have explored introducing interval probability prediction methods, they often assume that the predicted values follow a fixed probability distribution. This premise assumption ignores the true distribution characteristics of the data itself, making the prediction interval likely to deviate from the actual situation and difficult to accurately reflect the uncertainty.
[0073] To overcome the above problems, in this embodiment, based on the prediction framework of frequency domain enhancement and multi-source fusion, an interval probability prediction method based on Ensemble Learning and Kernel Density Estimation (KDE) is further proposed. This method generates diverse prediction results through ensemble learning and uses KDE to perform non-parametric modeling on the distribution of the predicted values, generating the corresponding probability density function, thereby constructing an interval prediction range to improve the uncertainty estimation of the confidence level.
[0074] In step S5, for example, based on the DCNN-Informer model, different combinations of hyperparameters are used to generate M independent sub-models {M1, M2,..., M M}; the M sub-models are used for K rounds of training, where the predicted value of the k-th round of the i-th sub-model is: All the predicted values form a set The final predicted result set is defined as:
[0075]
[0076] After obtaining the predicted result set After that, the probability distribution of the predicted values is modeled by introducing kernel density estimation (KDE), and its formula is:
[0077]
[0078] where h is the bandwidth parameter that controls the smoothness; K(·) is the kernel function, and the definition of the commonly used Gaussian kernel function is:
[0079]
[0080] The bandwidth h directly affects the density estimation effect and needs to be optimized through cross-validation or empirical methods. The probability density function generated by KDE describes the probability distribution of the predicted values and provides a basis for uncertainty modeling and decision-making. For example, the prediction interval [y lower , y upper with a confidence level of α needs to satisfy:
[0081]
[0082] The prediction interval quantifies the uncertainty of the prediction result. KDE has the advantage of adapting to multimodal or asymmetric distributions and provides reliable decision support for complex equipment degradation prediction. To obtain the final average probability predicted lifetime value, calculate the probability density of each prediction point and perform a weighted sum of the predicted lifetime values based on these probability values, that is:
[0083]
[0084] where each predicted value is multiplied by its corresponding probability density and then added together. In this way, a DCIDMS (double training CNN-Informer with Double Exponential Smoothing Multi-Source Fusion) model is constructed to obtain a more robust final prediction result.
[0085] This method can effectively integrate the prediction information of different sub-models, reduce the influence of individual abnormal predicted values on the final result, improve the prediction accuracy, and provide a more reliable reference basis for RUL evaluation in engineering applications.
[0086] The following combines specific examples to elaborate on the solution and effect of this application in detail:
[0087] The following dataset is sourced from the Commercial Modular Aero-Propulsion System Simulation (C-MAPSS) dataset of the National Aeronautics and Space Administration (NASA). The table lists the characteristic descriptions of some sensors.
[0088]
[0089] In the C-MAPSS dataset, not all sensors reflect the degradation trend of the engine. Under the same operating mode, the data of some sensors remain almost unchanged throughout the entire lifespan. Therefore, it is necessary to discard the data of some sensor signals, and their trends can be roughly divided into four types: monotonically increasing, monotonically decreasing, unchanged, and irregular. As Figure 5 shows the characteristic trends of some sensors. Through analysis, it can be seen that the sensors of X 3 、X 4 、X 8 、X 13 、X 19 、X 21 、X 22 do not provide obvious change information. Therefore, the experiment selected the data of the remaining 17 sensors to train the model.
[0090] For example, using Min-Max normalization, the original data features are mapped to the interval [0, 1] to eliminate the dimensional differences between features and ensure that the model fairly considers each feature on the same scale. The data trends after feature normalization are as Figure 6 shown.
[0091] Figure 7 shows the prediction effects of Engine No. 63 at different learning rates. The test results show that on the CNN-Informer pre-trained model, the prediction effect is optimal when the learning rate is 0.001. As the learning rate increases, the prediction effect gradually deteriorates.
[0092] As Figure 8 shown, after sending the processed data into the pre-trained model and undergoing multiple rounds of training, a PCNN-Informer secondary training model is obtained. The model iteration process is shown in the figure. The abscissa is the number of training times, and the ordinate is the value of RMSE. The value of RMSE tends to level off at approximately 260 times.
[0093] After grouping each engine and sending them into the pre-trained model batch by batch according to the engine number, a model based on each engine is obtained to predict the test set data, and the obtained effect is as Figure 9 shown, Figure 9Subgraphs a and b respectively show the iteration processes of Engine No. 27 and Engine No. 63 on the CNN-Informer pre-trained model. The abscissa is the number of training times, and the ordinate is the RMSE value. From Figure 9 the trend, it can be seen that for Engine No. 63 during the first 20 rounds of training, the RMSE has a slightly decreasing trend and then remains flat. Engine No. 27 generally remains flat, with a slightly decreasing trend in the early stage.
[0094] Perform a filtering operation on the obtained prediction results. After using second-order exponential smoothing filtering, the predicted trend becomes smoother. Figure 10 Shows the overall life prediction comparison chart of 3 different engines. In subgraphs a and c, the overall predicted curves are closer to the true values after filtering. In subgraph b, the filtered curve is smoother, and the points predicting the last time series are closer to the true values, and the final prediction effect is better.
[0095] To verify the prediction effect of the DCNN-Informer model proposed in this paper, the basic model and the advanced model in the table are used for experimental comparison in this embodiment. The RMSE and Score values of the models in the table are the averages of the results of 20 experiments. It can be seen from the table that on the FD001 and FD003 datasets, DCNN-Informer is superior to the other comparison models listed. Taking the FD001 dataset as an example, compared with the XGBoost-DCNN model, the RMSE is reduced by 3.27% and the Score is reduced by 5.43%. This shows that the DCNN-Informer model has more excellent prediction performance.
[0096]
[0097] Figure 11 Is a comparison chart of the predicted values and the true values of the CNN-Informer pre-training, PCNN-Informer secondary training model, and DCNN-Informer smoothing filtering model. The Figure 11 Is a comparison of the predicted values and the true values of the last time series of 100 engines of the three models on the test set. Experiments show that CNN can better extract information related to engine features, and the prediction effect of the DCNN-Informer model after adding CNN convolution, secondary training, and filtering is better.
[0098] Figure 12 Gives the error distribution chart of the prediction results and the actual values of different models on the test set. The abscissa represents the range of errors, and the ordinate represents the frequency of errors in each range. From Figure 12It can be seen that the error distribution of the pre-trained model is relatively wide, indicating that the pre-trained model may have a large prediction bias in some cases. The error distribution of the secondarily trained model is relatively narrow, and the prediction error can be reduced through secondary training. The model with smoothed filtering shows a more concentrated error distribution, indicating that smoothed filtering can further improve the prediction accuracy.
[0099] Based on the above models, ten groups of experiments were conducted on the test set data, and the error metrics of the experimental results were recorded. Figure 13 It is a box plot of the prediction metrics RMSE and Score for several models proposed in this paper. Figure 13 It can be seen that after adding secondary training, the stability of the model is significantly improved. After adding smoothed filtering, the prediction performance of the model is slightly improved. Compared with the M6-M8 models, the DCNN-Informer proposed in this paper has the best effect.
[0100] In summary, the first embodiment of the present invention proposes a device life prediction method based on a multi-source fusion model, which uses a CNN convolutional neural network to extract high-dimensional spatial features of time series data and combines the self-attention mechanism of Informer to globally model these features, so as to fully extract the information in the time dimension; in addition, an engine secondary training framework is designed, the dataset is grouped by engine, and the data of each group is respectively sent into the CNN-Informer model for secondary training to obtain a personalized model for each engine, further improving the accuracy and generalization ability of the model; the second-order exponential smoothing filtering algorithm and the interval probability prediction method are used to process the prediction results of the model, thereby constructing the DCIDMS model to improve the stability and accuracy of the prediction. This prediction model has obvious advantages in RUL prediction and is more superior in prediction performance compared with existing models.
[0101] In addition, as Figure 14 shown, the second embodiment of the present invention also proposes a device life prediction device based on a multi-source fusion model, including: a data preprocessing module 201, a model pre-training module 202, a model secondary training module 203, an initial RUL prediction module 204, and a target RUL prediction module 205.
[0102] Among them, the data preprocessing module 201 is used to obtain the equipment life data and perform feature normalization, set the RUL label for the obtained normalized data, and extract samples through a sliding window to obtain a preprocessed training set; the model pre-training module 202 is used to establish a CNN-Informer model and input the data of the preprocessed training set into the CNN-Informer model for pre-training to obtain a pre-trained model; the model secondary training module 203 is used to group the equipment according to the equipment characteristics and send them into the pre-trained model based on the groups respectively to obtain a secondary training model representing the characteristics of each equipment; the initial RUL prediction module 204 is used to test the trained secondary training model on the test set, obtain an initial RUL prediction result and perform filtering, so as to construct a DCNN-Informer model; the target RUL prediction module 205 is used to process the filtered initial RUL prediction result using the interval probability prediction method, so as to construct a multi-source fusion model based on the DCNN-Informer model and obtain a final target RUL prediction result using the multi-source fusion model.
[0103] The device life prediction method based on the multi-source fusion model implemented by the device life prediction device disclosed in the second embodiment of the present invention is as described in the foregoing first embodiment, so it will not be elaborated in detail here. Optionally, each module and the above other operations or functions in the second embodiment are respectively for implementing the method described in the first embodiment, and the beneficial effects of the device life prediction device based on the multi-source fusion model provided in this embodiment are the same as those of the device life prediction method based on the multi-source fusion model provided in the foregoing first embodiment. For the sake of brevity, they will not be repeated here.
[0104] As Figure 15 shown, the third embodiment of the present invention also proposes an electronic device, for example, including: at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program, and when the computer program is executed by the processing unit, the processing unit is enabled to execute the method described in the first embodiment, and the beneficial effects of the electronic device provided in this embodiment are the same as those of the device life prediction method based on the multi-source fusion model provided in the first embodiment.
[0105] As Figure 16 shown, the fourth embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above method are implemented, and the beneficial effects of the computer-readable storage medium provided in this embodiment are the same as those of the device life prediction method based on the multi-source fusion model provided in the first embodiment.
[0106] Among them, the computer-readable storage medium may include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs, CD-ROMs, microdrives, and magneto-optical disks, ROMs, RAMs, EPROMs, EEPROMs, DRAMs, VRAMs, flash memory devices, magnetic or optical cards, nanosystems (including molecular memory ICs), or any type of medium or device suitable for storing instructions and / or data.
[0107] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0108] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0109] In the several embodiments provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some service interfaces. The indirect couplings or communication connections of the devices or units can be in electrical or other forms.
[0110] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0111] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0112] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned memory includes various media that can store program codes, such as USB flash drives, read-only memories (ROM), random access memories (RAM), external hard drives, magnetic disks, or optical discs.
[0113] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable memory, and the memory can include: flash drives, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs, etc.
[0114] The above are only exemplary embodiments of the present disclosure, and the scope of the present disclosure cannot be limited thereby. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. After considering the specification and practicing the present disclosure herein, those skilled in the art will readily think of other embodiments of the present disclosure. This application aims to cover any variations, uses, or adaptive changes of the present disclosure. These variations, uses, or adaptive changes follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not described in the present disclosure. The specification and embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.
[0115] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0116] Those skilled in the art can easily understand that the above are only preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for predicting equipment life based on a multi-source fusion model, characterized in that: include: Obtain equipment life data and perform feature normalization, set RUL labels for the normalized data, and extract samples through a sliding window to obtain a preprocessed training set; Establish a CNN-Informer model, and pass the preprocessed training set data into the CNN-Informer model for pre-training to obtain a pre-trained model; Grouping devices according to device characteristics, and sending the grouped groups to the pre-trained model to obtain a secondary training model that characterizes the characteristics of each device; The trained secondary training model is tested on the test set to obtain the initial RUL prediction result and filter it, thereby constructing a DCNN-Informer model; The initial RUL prediction result after filtering is processed using an interval probability prediction method, so as to construct a multi-source fusion model based on the DCNN-Informer model, and the final target RUL prediction result is obtained using the multi-source fusion model.
2. The equipment life prediction method based on the multi-source fusion model according to claim 1 is characterized in that: The feature normalization adopts minimum-maximum normalization, and the formula is: In the formula, is the jth output of the i-th sensor of the engine, is the minimum value of all output values of the i-th sensor, is the maximum value of all output values of the i-th sensor.
3. The equipment life prediction method based on the multi-source fusion model according to claim 1 is characterized in that: The CNN-Informer model includes: a CNN feature extraction layer and an Informer Encoder layer, and the data processing flow includes: Perform feature extraction operations through CNN; The features extracted by CNN are sent to the Informer module, and the extracted feature vectors are flattened, linearly mapped, and classified tags are added before position encoding to complete serialization; After serialization, it is input into the Informer encoder, and feature learning is completed through normalization, multi-head attention mechanism, and multi-layer perceptron; The learned prediction features are input into the fully connected layer to predict the remaining service life of the equipment.
4. The equipment life prediction method based on the multi-source fusion model according to claim 3 is characterized in that: The CNN is composed of an input layer, a convolution layer, a pooling layer, a fully connected layer and an output layer, wherein the feature vector input by the input layer is subjected to the convolution operation of the convolution layer to extract features, and then the complexity of network calculation is simplified by the pooling layer, and all features are linked by the fully connected layer, and the obtained results are output by the output layer; The result of the convolutional layer and its output vector is represented as: Among them, σ is the sigmoid activation function, b j is the bias of the feature map, w is the weight of the kernel, and f v is the filter index, x is the input vector; The pooling layer represents dimensionality reduction, and its calculation formula is: Among them, R is the pooling size, T is the step size that determines the moving distance of the input data area, y is the input size, T<y, and the fully connected layer connects each neuron in the pooling layer to each neuron in the output layer.
5. The equipment life prediction method based on multi-source fusion model according to claim 3 is characterized in that: The Informer consists of an input layer, a position encoding, a self-attention distillation, and a fully connected layer; in The input layer maps the input data into a d-dimensional vector Z∈R n×c ; The positional encoding takes an input sequence and maps it to a high-dimensional vector, which is then fed to the decoder to produce an output sequence; Self-attention distillation calculates the sparsity measure of any Q vector according to the query sparsity measure formula, and selects the Q vector that plays a dominant role in the attention calculation, and uses this vector to calculate the attention. The formula is: Among them, Q is the query vector, K is the key vector, and V is the value vector. It means calculating the similarity between the query and the key, d is the dimension of the key, and Softmax is the activation function; The fully connected layer adds an activation function and the output is normalized again through the residual connection and layer.
6. The equipment life prediction method based on multi-source fusion model according to claim 1 is characterized in that: The device is grouped according to the device characteristics, and the groups are respectively sent to the pre-trained model based on the grouping to obtain a secondary training model that characterizes the characteristics of each device, including: The data of the training set is grouped by each device, where the first engine is denoted as y1 and the nth engine is denoted as y n ; Set y1…,y n Input the pre-trained model to obtain a model for each device; Divide the data X of each device into multiple sliding windows TW1…TWn to obtain subsequence X i1 …X im , where m represents the number of subsequences after division; For each subsequence X ij , use the pre-trained model to train and obtain sub-models AM1…AMn.
7. The equipment life prediction method based on multi-source fusion model according to claim 1 is characterized in that: The initial RUL prediction result adopts second-order exponential smoothing filtering, and the formula is: L t ay t +(1-α)(L t-1 +T t-1 )t; T t =β(L t -L t-1 )+(1-β)T t-1 ; F t+m =L t +mT t ; Among them, L t is the estimated value of the level at time point t; T t is the trend estimate at time point t; α is the level smoothing coefficient, which controls the response of the level estimate to the current observation and the long-term trend of the time series, and takes α = 0.1 to 0.3; β is the trend smoothing coefficient, which controls the response of the trend estimate to the change of the level estimate; F t+m is the predicted value for the next m periods.
8. The equipment life prediction method based on multi-source fusion model according to claim 1 is characterized in that: The interval probability prediction method comprises: Based on the DCNN-Informer model, different hyperparameter combinations are used to generate M independent sub-models {M1, M2, ..., M M }; The M sub-models are used to perform K rounds of training, where the prediction value of the k-th round of the i-th sub-model is: All predicted values form a set The probability distribution of the predicted value is estimated by introducing kernel density, and the formula is: Where h is the bandwidth parameter, which is used to control the degree of smoothing; K(·) is the kernel function; The probability density of each prediction point is calculated according to the probability distribution, and the predicted life value is weighted averaged based on the probability density. The formula is: Among them, each predicted value Multiply it by its corresponding probability density Then add them together to get the final target RUL prediction result 9. A device for predicting equipment life based on a multi-source fusion model, characterized in that: include: The data preprocessing module is used to obtain equipment life data and perform feature normalization, set RUL labels on the normalized data, and extract samples through a sliding window to obtain a preprocessed training set; The model pre-training module is used to establish a CNN-Informer model and pass the pre-processed training set data into the CNN-Informer model for pre-training to obtain a pre-trained model; A model secondary training module, used for grouping devices according to device characteristics, and sending them to the pre-trained model based on the groups to obtain a secondary training model that characterizes the characteristics of each device; An initial RUL prediction module is used to test the trained secondary training model on a test set to obtain an initial RUL prediction result and filter it, thereby constructing a DCNN-Informer model; The target RUL prediction module is used to process the initial RUL prediction result after filtering using an interval probability prediction method, thereby building a multi-source fusion model based on the DCNN-Informer model, and using the multi-source fusion model to obtain the final target RUL prediction result.
10. An electronic device, characterized in that: include: A memory and one or more processors connected to the memory, the memory storing a computer program, the processor being used to execute the computer program to implement the equipment life prediction method based on a multi-source fusion model as described in any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable commands, and the computer-executable commands are used to execute the equipment life prediction method based on the multi-source fusion model as described in any one of claims 1-8.
Citation Information
Cited By
New energy equipment part aging prediction method and system based on digital twin and multi-source data fusion
CN121835240A
Photovoltaic power adaptive confidence interval prediction method and device, medium and equipment
CN122000865A