Equipment residual life prediction method based on improved convolutional neural network

By introducing technical means such as deep separation convolution and jump connection in the equipment residual life prediction, the convolution neural network structure is optimized, and the problems of high computational complexity and insufficient feature extraction capabilities in the existing technology are solved, and efficient and accurate equipment residual life prediction is achieved.

CN120234907APending Publication Date: 2025-07-01CHINA RONGTONG SCI RES INST GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510297564.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The prior art has problems such as high computational complexity, insufficient feature extraction capability and poor training stability in the prediction of equipment residual life, especially when processing high-dimensional time-series data, it is difficult to effectively capture time dependencies.

Method used

A device residual life prediction method based on an improved convolutional neural network is proposed. By introducing deep separation convolution, jump connection, batch normalization layer and LeakyReLU activation function, the network structure is optimized to improve feature extraction ability and training stability.

Benefits of technology

It significantly improves the accuracy and robustness of equipment residual life prediction, reduces calculation complexity and training costs, and can adapt to changes in various operating conditions, and has strong generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234907A_ABST
    Figure CN120234907A_ABST
Patent Text Reader

Abstract

The invention discloses a method for predicting the residual life of equipment based on an improved convolutional neural network, and aims to solve the problems of high calculation complexity, insufficient feature extraction capability and limited prediction precision in the prior art. The method comprises the following steps: based on a CMPASS data set, rejecting invalid features and performing normalization processing on residual features, and constructing multi-dimensional time sequence data with a fixed length by using a sliding window method; the method comprises the following steps: constructing an improved CNN network, extracting local and global features by adopting a depth separable convolution module, realizing feature transfer and relieving a gradient disappearance problem through jump connection, combining stable training of a batch normalization layer, enhancing a nonlinear modeling capability by using a Leaky ReLU activation function, and preventing overfitting through a Dropout layer; and finally, generating a predicted value of the residual life of the equipment through the full connection layer. Experimental results show that indexes such as MSE, MAE and R2 on a CMPASS data set are remarkably superior to those of traditional CNN, GRU and DCNN models, and higher prediction precision and generalization ability are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of equipment health management, and particularly to a method for predicting the remaining useful life of equipment based on an improved convolutional neural network. Background Art

[0002] With the development of industrial equipment towards large-scale and complexity, the monitoring and prediction of its operating status have become crucial. The prediction of the remaining useful life (RUL) of equipment, as an important part of predictive maintenance, can help predict equipment failures, reduce unplanned downtime, extend the equipment life, and optimize maintenance costs. Traditional maintenance methods often rely on periodic maintenance or corrective maintenance. This method is difficult to fully utilize the remaining value of the equipment and may lead to resource waste or significant losses caused by equipment failures. Therefore, data-driven prediction of the remaining useful life of equipment has gradually become a research hotspot.

[0003] In the aviation field, the turbofan engine is an important core component of an aircraft. Its complex operating environment and multi-condition characteristics make it difficult to predict faults. Once a fault occurs, it may lead to serious economic losses and even casualties. Therefore, the research on the Prognostics and Health Management (PHM) technology of turbofan engines is of great significance, and the prediction of the remaining useful life is an important part of PHM. The operating data of turbofan engines has the following characteristics:

[0004] High-dimensional multimodality: The data comes from different sensors and contains multiple physical variables (such as temperature, pressure, vibration, etc.), and the data dimension is relatively high;

[0005] Temporal correlation: The engine performance gradually degrades over time, and its characteristics show significant time dependence;

[0006] Multi-condition adaptability: The engine operates under various conditions such as different flight altitudes, speeds, and thrust lever angles, and the data distribution has heterogeneity.

[0007] To solve the above problems, traditional RUL prediction methods mainly include:

[0008] Physics-based methods: Establish a mathematical model using the operating mechanism of the equipment, such as a fluid dynamics model or a thermodynamics model. This method requires a large amount of prior knowledge and accurate physical parameters, and the modeling cost is high and it is not easy to be extended to complex equipment.

[0009] Statistics and traditional machine learning-based methods: Such as multiple regression analysis, support vector machine (SVM), random forest, etc. These methods have limited ability to model non-linear features and are difficult to process high-dimensional time series data.

[0010] Deep learning-based methods: In recent years, convolutional neural networks (CNNs), recurrent neural networks (RNNs), and their variants (such as GRUs and LSTMs) have shown good performance in RUL prediction. Deep learning methods can automatically extract features from high-dimensional data, but the optimization of their network structures and the capture of time-dependent relationships still pose challenges.

[0011] To this end, this application proposes a method and system for predicting the remaining useful life of equipment based on an improved convolutional neural network to solve the above problems. Summary of the Invention

[0012] The purpose of the present invention is to solve the above technical problems and provide a method and system for predicting the remaining useful life of equipment based on an improved convolutional neural network.

[0013] To achieve the above object, the present invention is implemented through the following technical solutions:

[0014] One aspect of the solution of the present invention provides a method for predicting the remaining useful life of equipment based on an improved convolutional neural network, including the following steps:

[0015] Input multi-dimensional time series data of equipment operation and preprocess the data;

[0016] Construct an improved convolutional neural network, which includes a convolutional block, a batch normalization layer, an activation function, a skip connection, a pooling layer, and a fully connected layer;

[0017] Use depthwise separable convolutions to extract local and global features from the time series data;

[0018] Introduce skip connections in the convolutional block to bypass some network layers through the residual path;

[0019] Add a batch normalization layer in the convolutional block to stabilize the training process;

[0020] Adopt the LeakyReLU activation function to prevent the problem of dead neurons in the network;

[0021] Reduce the dimension of the convolutional features through max pooling operation and input them into the fully connected layer to generate the remaining useful life prediction value.

[0022] Further, the data preprocessing includes the following steps:

[0023] Eliminate irrelevant features with small variation amplitudes in the time series data;

[0024] Normalize the data after eliminating irrelevant features and adjust all feature values to the range of 0 to 1;

[0025] Construct time series subsequences of a fixed length through a sliding window method, where the length of the time series subsequence is 31.

[0026] Furthermore, the depthwise separable convolution in the improved convolutional neural network includes the following steps:

[0027] Independently apply depthwise convolution to each input channel of the time series data to extract in-channel local features;

[0028] Use pointwise convolution to fuse the inter-channel feature maps output by the depthwise convolution through a 1×1 convolutional kernel to extract global features;

[0029] Add a batch normalization layer and a LeakyReLU activation function after the depthwise convolution and the pointwise convolution respectively to improve training stability and feature expression ability.

[0030] Furthermore, the skip connection includes the following steps:

[0031] Adjust the number of channels of the feature tensor of the residual path through a 1×1 convolution to make it consistent with the number of channels of the output of the main branch;

[0032] Use the bilinear interpolation method to adjust the size of the feature tensor of the residual path to match the size of the feature tensor output by the main branch;

[0033] Add the adjusted feature tensor of the residual path and the output of the main branch element-wise to form the final feature representation.

[0034] Furthermore, the optimization measures of the improved convolutional neural network include the following:

[0035] Use the LeakyReLU activation function, and its mathematical expression is:

[0036]

[0037] where α is a small positive number, which is used to prevent the weight update of negative input neurons from being interrupted;

[0038] Add a batch normalization layer after the convolutional layer, which is used to normalize the mean and standard deviation of each batch of input data, so as to stabilize the training process and accelerate the convergence of the model;

[0039] Add a Dropout layer, and the dropout probability of the Dropout layer is set to 0.2, which is used to randomly mask the outputs of some neurons to reduce overfitting.

[0040] Furthermore, the training process of the improved convolutional neural network includes the following steps:

[0041] The parameters of the network are optimized using the Adam optimizer, and the initial learning rate is set to 0.01;

[0042] The mean squared error loss function is used as the objective function to measure the error between the prediction result of the network and the true value;

[0043] The batch size is set to 64, and 64 samples are selected from the time series dataset as a batch for each iteration;

[0044] The early stopping mechanism is used to dynamically control the number of iterations, and the training stops when the performance on the validation set has not improved in 10 consecutive iterations.

[0045] Furthermore, the time series data is from the CMPASS dataset, and the characteristics of the dataset include:

[0046] It contains 3 operating condition parameters and 21 performance monitoring parameters;

[0047] After removing the ineffective features with relatively small changes, the remaining number of features is 17;

[0048] The tensor shape of the input data is 64, 1, 31, 17, where 64 is the batch size, 1 is the number of channels, 31 is the time series length, and 17 is the number of features.

[0049] Furthermore, the main hierarchical structure of the improved convolutional neural network includes:

[0050] Depthwise separable convolution blocks, which are used to efficiently extract local and global features in the time series data;

[0051] Skip connection modules, which are used to transfer information between the main branch and the residual path of the convolution block to alleviate the problem of gradient disappearance;

[0052] Max pooling layers, which are used to reduce the size of the output feature map of the convolution block and capture key features;

[0053] Fully connected layers, which are used to map the output features of the convolution block to the predicted value of the remaining useful life of the device.

[0054] On the other hand, the solution of the present invention also provides a device remaining useful life prediction system based on an improved convolutional neural network, and the system includes:

[0055] Data input module: used to receive and preprocess multi-dimensional time series data of device operation;

[0056] Feature extraction module: uses depthwise separable convolution to extract local and global features of the input data;

[0057] Skip connection module: transfers information through the residual path to alleviate the problem of gradient disappearance;

[0058] Output module: Generates the predicted value of the remaining useful life of the device through a fully connected layer.

[0059] Furthermore, the improved convolutional neural network includes depthwise separable convolution, LeakyReLU activation function, batch normalization layer, Dropout layer, skip connection module and max pooling layer, which are combined with the time series characteristics in the CMPASS dataset to efficiently predict the remaining useful life of the device.

[0060] Compared with the prior art, the present invention has the following beneficial effects:

[0061] 1. By improving the structure of the traditional convolutional neural network (CNN), the present invention proposes an efficient method for predicting the remaining useful life of a device (DP_CNN), significantly overcoming problems such as high computational complexity, insufficient feature extraction ability and poor training stability in the prior art. By introducing depthwise separable convolution, the number of network parameters and computational complexity are greatly reduced, and the training cost is reduced; the skip connection is used to alleviate the problem of gradient disappearance and improve the deep feature learning ability of the network; combined with batch normalization (BN) to stabilize the training process, the LeakyReLU activation function enhances the network's ability to model non-linear degradation features, and finally significantly improves the accuracy and robustness of the prediction. In the experiment, the MSE and MAE of the model of the present invention are respectively reduced by 3228.91 and 32.51 compared with the traditional CNN method, and the R 2 value is increased from 0.41 to 0.73, comprehensively superior to existing methods such as GRU and DCNN.

[0062] 2. At the application level, by optimizing the preprocessing process of the CMPASS dataset, including feature screening, normalization and time series division, the present invention successfully adapts to the complex multi-dimensional and multi-condition monitoring data of turbofan engines and realizes high-precision prediction of the remaining useful life. The present invention can not only capture the time-dependent pattern of the performance degradation of turbofan engines, but also has strong generalization ability and can adapt to various condition changes. In addition, while maintaining high prediction performance, the improved CNN network reduces the computational resource requirements of the model and can meet the requirements of real-time prediction. Therefore, the present invention is not only applicable to the health management of aeroengines, but also can be extended to fields such as wind power, rail transit and industrial manufacturing, providing technical support for intelligent operation and predictive maintenance. Description of the Drawings

[0063] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0064] Figure 1 It is a schematic diagram of a turbofan engine for the C-MPASS dataset;

[0065] Figure 2 It is a comparison chart of the loss curves of the original CNN network (left) and the improved network (right);

[0066] Figure 3 It is an architecture diagram of the improved convolutional neural network;

[0067] Figure 4 It is a flowchart of the method of the present invention. Detailed implementation manners

[0068] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0069] The following in conjunction with Figures 1-4 , will describe in detail the specific implementation manners of the present invention.

[0070] The implementation solution of the present invention provides a method for predicting the remaining useful life of a device based on an improved convolutional neural network. The method includes the following steps:

[0071] Input the multi-dimensional time series data of the device operation and preprocess the data. The data preprocessing includes the following steps:

[0072] Eliminate the irrelevant features with small variation amplitudes in the time series data;

[0073] Normalize the data after eliminating the irrelevant features and adjust all feature values to the interval of 0 and 1;

[0074] Construct time series subsequences with a fixed length through the sliding window method. The length of the time series subsequence is 31.

[0075] Construct an improved convolutional neural network, which includes a convolutional block, a batch normalization layer, an activation function, a skip connection, a pooling layer, and a fully connected layer;

[0076] Extract local and global features from the time series data using depthwise separable convolutions; the depthwise separable convolutions include the following steps:

[0077] Independently apply depthwise convolutions to each input channel of the time series data to extract in-channel local features;

[0078] Use pointwise convolutions to fuse the inter-channel feature maps output by the depthwise convolutions through 1×1 convolutional kernels to extract global features;

[0079] Add a batch normalization layer and a LeakyReLU activation function after the depthwise convolution and the pointwise convolution respectively to improve training stability and feature expression ability.

[0080] Introduce skip connections in the convolutional block to bypass some network layers through the residual path; the skip connections include the following steps:

[0081] Adjust the number of channels of the feature tensor of the residual path through 1×1 convolutions to make it consistent with the number of channels of the output of the main branch;

[0082] Use the bilinear interpolation method to adjust the size of the feature tensor of the residual path to match the size of the feature tensor output by the main branch;

[0083] Element-wise add the adjusted feature tensor of the residual path to the output of the main branch to form the final feature representation.

[0084] Add a batch normalization layer in the convolutional block to stabilize the training process;

[0085] Adopt the LeakyReLU activation function to prevent the problem of dead neurons in the network; use the LeakyReLU activation function, and its mathematical expression is:

[0086]

[0087] where α is a small positive number used to prevent the weight update of negative input neurons from being interrupted;

[0088] Add a batch normalization layer after the convolutional layer to normalize the mean and standard deviation of each batch of input data, thereby stabilizing the training process and accelerating model convergence;

[0089] Add a Dropout layer, and set the dropout probability of the Dropout layer to 0.2, which is used to randomly mask the output of some neurons to reduce overfitting.

[0090] Reduce the dimension of the convolutional features through max pooling operations and input them into a fully connected layer to generate the remaining useful life prediction value.

[0091] The training process of the improved convolutional neural network includes the following steps:

[0092] Use the Adam optimizer to optimize the parameters of the network, and set the initial learning rate to 0.01;

[0093] Adopt the mean squared error loss function (MSELoss) as the objective function to measure the error between the prediction result of the network and the true value;

[0094] Set the batch size to 64, and select 64 samples from the time series dataset as a batch for each iteration;

[0095] Use the early stopping mechanism to dynamically control the number of iterations, and stop training when the performance of the validation set has not improved in 10 consecutive iterations; The time series data is from the CMPASS dataset, and the characteristics of the dataset include:

[0096] It contains 3 operating condition parameters and 21 performance monitoring parameters;

[0097] After removing the ineffective features with small changes, the remaining number of features is 17;

[0098] The tensor shape of the input data is 64, 1, 31, 17, where 64 is the batch size, 1 is the number of channels, 31 is the time series length, and 17 is the number of features.

[0099] The main hierarchical structure of the improved convolutional neural network includes:

[0100] Depthwise separable convolution blocks, which are used to efficiently extract local and global features in the time series data;

[0101] Skip connection modules, which are used to transfer information between the main branch and the residual path of the convolution block to alleviate the problem of gradient disappearance;

[0102] Max pooling layers, which are used to reduce the size of the output feature map of the convolution block and capture key features;

[0103] Fully connected layers, which are used to map the output features of the convolution block to the predicted value of the remaining useful life of the device.

[0104] A system for predicting the remaining useful life of a device based on an improved convolutional neural network, the system includes:

[0105] Data input module: used to receive and preprocess multi-dimensional time series data of device operation;

[0106] Feature extraction module: uses depthwise separable convolution to extract local and global features of the input data;

[0107] Skip connection module: Transmits information through the residual path to alleviate the problem of vanishing gradients;

[0108] Output module: Generates the predicted remaining useful life of the device through a fully connected layer.

[0109] The improved convolutional neural network includes depthwise separable convolution, LeakyReLU activation function, batch normalization layer, Dropout layer, skip connection module and max pooling layer, and combines the time series characteristics in the CMPASS dataset for efficient prediction of the remaining useful life of the device.

[0110] The following makes specific descriptions of the above implementation plans in combination with specific embodiments.

[0111] Embodiment 1: A method for predicting the remaining useful life of a turbofan engine based on an improved convolutional neural network;

[0112] Dataset introduction and preprocessing:

[0113] Introduction: The CMPASS dataset is selected to train and compare the original CNN network and the improved network. This dataset is a dataset simulating a large commercial turbofan engine generated by NASA using the Commercial Modular Aero-Propulsion System Simulation software. The engine schematic diagram is as Figure 1 shown.

[0114] The simulation modeling mainly focuses on the gas path part of the engine, which mainly includes the fan blade (Fan), low-pressure compressor (LPC), high-pressure compressor (HPC), combustor (Combustor), low-pressure rotor (N1), high-pressure rotor (N2), high-pressure turbine (HPT), low-pressure turbine (LPT), and nozzle (Nozzle). The dataset is the monitoring data of the turbofan engine performance components at different times, so the data has temporality and reflects the degradation of the turbofan engine performance.

[0115] The dataset files are divided into three categories: training data train_FD00x.txt, test data test_FD00x.txt, and the remaining useful life of the turbofan engine at the last moment of each sample in the test data, corresponding to the file RUL_FD00x.txt. (The value of x is different, that is, the fault modes and conditions of the turbofan engine are different), specifically refer to Table 1.

[0116] Table 1 Data distribution of the C-MPASS dataset

[0117] data set FD001 FD002 FD003 FD004 training set 100 260 100 249 test set 1 00 259 100 248 operating conditions 1 6 1 6 fault status 1 1 2 2

[0118] The data format is an N×26-dimensional matrix. Here, N in the matrix represents different cycle periods of different samples, and the 26 dimensions respectively correspond to the sample number, time cycle, operation 1, operation 2, operation 3, sensor 1, sensor 2, and sensor 21. That is, each engine contains three working condition monitoring parameters (flight altitude, Mach number, and thrust lever angle) and 21 performance monitoring parameters, as shown in Table 2.

[0119] Table 2 Schematic diagram of the training set data of the C-MPASS dataset;

[0120]

[0121] Dataset preprocessing: The present invention selects the FD001 dataset, which includes input files train_FD001.txt, test_FD001.txt, and RUL_FD001.txt, representing the training file, test file, and RUL file (remaining useful life of the engine, used to verify the prediction results), respectively.

[0122] Among them, both the training and test files contain the unit number (unit), time cycle (time), three operation settings (os_1, os_2, os_3), and 21 sensor measurements (sm_1 to sm_21), as shown in Table 2.

[0123] First, the present invention conducts a preliminary check on the data to ensure there are no missing values. Further analysis shows that some features (such as os_3, sm_1, sm_5, sm_10, sm_16, sm_18, sm_19) do not change much throughout the life cycle and contribute less to the change from the healthy state to the faulty state. Therefore, it is decided to eliminate these features to reduce the unnecessary computational burden and improve the model efficiency. After processing, the number of features is 17.

[0124] To ensure that different features are compared on the same scale and to accelerate model convergence, the selected features are normalized. The normalization uses the min-max scaling method, which limits all values within the range of [0, 1]. This helps to stabilize the training process and improve the model performance.

[0125] By analyzing the maximum time cycle number of each unit in the test dataset, the minimum length of the sequence is determined to be 31. This means that each training sample is composed of data at 31 consecutive time points, thus constructing a time series input format suitable for the deep learning model.

[0126] For each training unit, starting from its earliest time point, consecutive subsequences of length 31 are sequentially selected as input samples. All these subsequences are combined into a large tensor x_train for training the model.

[0127] The corresponding label y_train represents the number of time steps until each input sequence fails at this unit, that is, the remaining life. The label data is also converted into a tensor form and paired with the input data.

[0128] After being processed through the above steps, the shape of the input data x_train is [64 (batch size), 1, 31, 17], while the shape of the label data y_train is [64, 1].

[0129] Improve the model deployment and implementation method;

[0130] Deployment settings: The algorithm programming environment of the present invention is Python 3.8.19 and runs on a computer equipped with a GTX1650 graphics card. The Adam optimizer is adopted, the mean squared error loss function is selected, the initial value of Epoch is set to 200, and an early stopping mechanism (patience value is 10 iterations) is adopted during the training process to dynamically control the number of iterations (stop training when the fitting degree reaches the best). The Batchsize is set to 64, and the learning rate is set to 0.01.

[0131] Implementation method:

[0132] Step 1: The shape of the input data is [64, 31, 17]. To meet the requirements of the CNN, the channel dimension is increased through the operation of torch.unsqueeze(x, dim=1) to ensure that the input is four-dimensional, that is, the shape becomes [64, 1, 31, 17].

[0133] Step 2: The input data is fed into the first convolutional layer self.conv1, where a 3x3 depthwise convolution is first applied, followed by batch normalization and the LeakyReLU activation function. Then a 1x1 pointwise convolution is performed, and after batch normalization and the LeakyReLU activation again, 2x2 max pooling is adopted.

[0134] At this time, the size of the output feature map changes from [64, 1, 31, 17] to [64, 20, 16, 9].

[0135] Step 3: To add skip connections, an additional 1x1 convolutional layer self.residual_conv is defined, which provides a shortcut path directly from the input to the last layer of the main branch. The residual branch transforms the input into a tensor with the same number of channels as the main branch through residual_conv, with a shape of [64, 20, 31, 17]. Since the output size of the main branch has changed, bilinear interpolation (F.interpolate) is needed to adjust the size of the residual branch to match the main branch output. Finally, the two are added together to form the final feature representation, with the shape remaining [64, 20, 16, 9].

[0136] Step 4: After the feature map with skip connections is flattened, the shape changes from [64, 20, 16, 9] to [64, 20*16*9]. Then, these flattened features are mapped to the desired output dimension, which is a single value representing the predicted RUL estimate, through a fully connected layer (self.fc).

[0137] Step 5: To prevent overfitting, a Dropout layer is added. Ensure that the Dropout mechanism is correctly included each time the model is instantiated, so that during the training process, a part of the neurons are randomly discarded, reducing the model complexity and improving the generalization ability.

[0138] Test results: Loss curve: Early stopping is used during training to dynamically control the number of iterations to achieve the best training effect. The loss curves plotted during the training process are compared as Figure 2 shown:

[0139] It can be seen from the comparison that the loss of the improved network drops more significantly and the curve is smoother during the training process.

[0140] Prediction results: The best models obtained after training using the original CNN network and the improved network respectively are used to perform prediction verification on the test set. The comparison of the scores of each index obtained by the verification is shown in Table 3:

[0141] Table 3 Comparison of prediction index scores of the original CNN network (left) and the improved network (right)

[0142]

[0143] It can be seen from the comparison that the prediction indexes MSE, MAE, RMSE and R2 scores of the improved network model are significantly better than those of the original CNN network. The improved network model has a better fitting effect on the data, higher accuracy, and smaller overall error. In addition, the improved network model better captures the change patterns in the data and has stronger interpretability.

[0144] Performance comparison test with common networks: The GRU (Gated Recurrent Unit) network is a simplified version of the long short-term memory network (LSTM), aiming to solve the vanishing gradient problem encountered by traditional RNNs when processing long time series data. By introducing an update gate and a reset gate, GRU can effectively control the information flow, thereby retaining long-term dependencies without adding excessive computational complexity. The update gate determines the proportion of past information passed to the future, while the reset gate allows the network to selectively ignore previous state information, enabling the model to better capture the dynamic changes in the time series. For the RUL prediction task, the advantage of GRU lies in its powerful time-dependent modeling ability, which is particularly suitable for continuous sensor readings generated during device operation and can accurately predict the remaining useful life and identify potential failure modes.

[0145] DCNN extracts complex features in images or other types of data by stacking multiple convolutional layers, pooling layers, and fully connected layers. Each convolutional layer operates on the input using a set of filters to generate a series of feature maps, and then the pooling layer reduces the spatial dimensions of these feature maps to reduce the computational amount and prevent overfitting. As the number of layers increases, DCNN can learn more abstract and higher-level feature representations. In the RUL prediction task, DCNN can automatically extract meaningful spatial features from multi-dimensional time series data, such as measurements from sensors at different locations. This method not only improves the feature expression ability but also enhances the generalization performance of the model, especially suitable for processing data with obvious spatial structures, such as simultaneous readings from multiple sensors, providing more accurate support for remaining life prediction. In addition, combining the time-dependent modeling ability of GRU, the joint application of the two can further improve the accuracy of RUL prediction.

[0146] Performance comparison test results:

[0147] The test metrics are as follows: MSE (Mean Squared Error): It is a commonly used regression evaluation metric. It measures the accuracy of the model by calculating the average of the squared differences between the predicted values and the true values. Generally, the lower the MSE value, the smaller the gap between the predicted value and the actual value, and the higher the prediction accuracy. MSE is very sensitive to large errors because the errors are squared before calculation, so a large error will significantly affect the value of MSE.

[0148] MAE (Mean Absolute Error): It is the average of the absolute values of the differences between the predicted values and the true values. Compared with MSE, MAE is less affected by outliers because it directly measures the average distance between the predicted value and the true value without considering the direction or magnitude of the error. Generally, the lower the MAE value, the more generally the prediction results are close to the actual value, reflecting that the model has good overall accuracy.

[0149] RMSE (Root Mean Square Error): The square root of MSE. RMSE inherits the characteristic of MSE being sensitive to large errors. At the same time, since its unit is the same as the input data, it provides a more intuitive measure of error. When the RMSE value is small, it means that most predictions are closely around the true value, reflecting the good fitting degree and reliability of the model.

[0150] R 2 (Coefficient of determination): Represents the proportion of variance explained by the model, that is, the proportion of the variability predicted by the model relative to the total variability. The value range of R2 is from negative infinity to 1, where 1 represents a perfect fit, 0 means the model is no better than a constant model, and a negative value means the model performs even worse than a simple average prediction. Generally speaking, the closer R 2 is to 1, the better the prediction performance of the model. It not only reflects the accuracy of the model but also demonstrates the model's ability to understand the trend of data changes.

[0151] The test results are shown in the following table (the improved CNN network designed in the present invention is temporarily called the DP_CNN network):

[0152] model category MAE MSE RMSE <![CDATA[R 2 > CNN 49.03 3688.43 60.73 0.41 DP_CNN 16.52 459.52 16.52 0.73 DCNN 28.76 1285.28 35.85 0.26 GRU 31.47 1482.11 38.50 0.14

[0153] From the comparison results, it can be seen that the DCNN network and GRU network models are superior to the CNN network model in prediction accuracy, but the R 2 values decrease by 0.15 and 0.27 respectively compared to the CNN network model, and their abilities to understand the trend of data changes are relatively poor compared to the CNN network.

[0154] The DP_CNN network designed in the present invention has a significantly improved prediction accuracy compared to the other three networks. The MSE decreases by 825.76 compared to the commonly used DCNN network today, and the performance is excellent. The R 2 score is also higher than the other three networks, increasing by 0.32 compared to the original CNN network model. The improved network has good generalization ability.

[0155] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. The method for predicting the remaining life of equipment based on an improved convolutional neural network is characterized in that: The following steps are involved: Input multi-dimensional time series data of equipment operation and pre-process the data; Construct an improved convolutional neural network, which includes convolution blocks, batch normalization layers, activation functions, skip connections, pooling layers, and fully connected layers; extracting local and global features from the time series data using depthwise separable convolutions; Introducing skip connections in convolutional blocks to bypass some network layers through residual paths; Add batch normalization layers in the convolutional blocks to stabilize the training process; Use LeakyReLU activation function to prevent the problem of dead neurons in the network; The convolutional features are reduced in dimension through the maximum pooling operation and input into the fully connected layer to generate the remaining life prediction value.

2. The method for predicting the remaining life of equipment based on an improved convolutional neural network according to claim 1, characterized in that: The data preprocessing comprises the following steps: Eliminate irrelevant features with small changes in the time series data; After removing irrelevant features, the data is normalized and all feature values ​​are adjusted to the range of 0 and 1; A time series subsequence with a fixed length is constructed by a sliding window method, and the length of the time series subsequence is 31.

3. The method for predicting the remaining life of equipment based on an improved convolutional neural network according to claim 1, characterized in that: The depth-wise separable convolution in the improved convolutional neural network comprises the following steps: Applying a depthwise convolution independently to each input channel of the time series data to extract local features within the channel; Using point-by-point convolution to fuse the inter-channel feature maps output by the depth convolution through a 1×1 convolution kernel to extract global features; Batch normalization layer and LeakyReLU activation function are added after depthwise convolution and pointwise convolution respectively to improve training stability and feature expression ability.

4. The method for predicting the remaining life of equipment based on an improved convolutional neural network according to claim 1, characterized in that: The jump connection comprises the following steps: The number of channels of the feature tensor of the residual path is adjusted through 1×1 convolution to make it consistent with the number of channels output by the main branch; Using a bilinear interpolation method, the feature tensor size of the residual path is adjusted to match the feature tensor size output by the main branch; The adjusted residual path feature tensor is added element-by-element to the main branch output to form the final feature representation.

5. The method for predicting the remaining life of equipment based on an improved convolutional neural network according to claim 1, characterized in that: The optimization measures of the improved convolutional neural network include the following: Using the LeakyReLU activation function, its mathematical expression is: Where α is a small positive number used to prevent the weight update of negative input neurons from being interrupted; Add a batch normalization layer after the convolutional layer to normalize the mean and standard deviation of each batch of input data, thereby stabilizing the training process and accelerating model convergence; A Dropout layer is added, and the dropout probability of the Dropout layer is set to 0.2 to randomly mask the outputs of some neurons to reduce overfitting.

6. The method for predicting the remaining life of equipment based on an improved convolutional neural network according to claim 1, characterized in that: The training process of the improved convolutional neural network includes the following steps: The parameters of the network were optimized using the Adam optimizer, and the initial learning rate was set to 0.01; A mean square error loss function is used as an objective function to measure the error between the prediction result of the network and the true value; Set the batch size to 64, and select 64 samples from the time series dataset as a batch in each iteration; The early stopping mechanism is used to dynamically control the number of iterations, and training is stopped when the performance of the validation set does not improve in 10 consecutive iterations.

7. The method for predicting the remaining life of equipment based on an improved convolutional neural network according to claim 6, characterized in that: The time series data is derived from the CMPASS dataset, and the characteristics of the dataset include: Contains 3 operating condition parameters and 21 performance monitoring parameters; After removing invalid features with small changes, the number of remaining features is 17; The tensor shape of the input data is 64, 1, 31, 17, where 64 is the batch size, 1 is the number of channels, 31 is the time series length, and 17 is the number of features.

8. The method for predicting the remaining life of equipment based on an improved convolutional neural network according to claim 1, characterized in that: The main hierarchical structure of the improved convolutional neural network includes: A depthwise separable convolutional block for efficiently extracting local and global features from the time series data; A skip connection module is used to transfer information between the main branch and the residual path of the convolution block to alleviate the gradient vanishing problem; A maximum pooling layer, used to reduce the size of the feature map output by the convolution block and capture key features; The fully connected layer is used to map the output features of the convolution block into a predicted value of the remaining life of the equipment.

9. A system for predicting the remaining life of an equipment based on an improved convolutional neural network, which is based on the method for predicting the remaining life of an equipment based on an improved convolutional neural network according to any one of claims 1 to 8, characterized in that: The system comprises: Data input module: used to receive and pre-process multi-dimensional time series data of equipment operation; Feature extraction module: uses depth-wise separable convolution to extract local and global features of input data; Skip connection module: transfers information through residual paths to alleviate the gradient vanishing problem; Output module: Generates the predicted value of the remaining life of the equipment through the fully connected layer.

10. The equipment remaining life prediction system based on improved convolutional neural network according to claim 9 is characterized in that: The improved convolutional neural network includes a depthwise separable convolution, a LeakyReLU activation function, a batch normalization layer, a Dropout layer, a skip connection module and a maximum pooling layer, and is combined with the time series characteristics in the CMPASS dataset to efficiently predict the remaining life of the equipment.