A fault prediction method and system

By adopting the MKCNN-LSTM data prediction model and MFSCNN fault diagnosis model based on attention mechanism in the fault prediction system, the existing fault prediction methods have solved the problem of low diagnostic accuracy and weak generalization ability, and achieved high accuracy and high generalization ability fault prediction effects.

CN117195953BActive Publication Date: 2025-05-30XIDIAN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202210593198.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-27
Publication Date
2025-05-30
Estimated Expiration
2042-05-27

AI Technical Summary

Technical Problem

The existing fault prediction methods have problems with low diagnostic accuracy and weak generalization ability.

Method used

The MKCNN-LSTM data prediction model and MFSCNN fault diagnosis model based on attention mechanism are adopted, and the fault prediction system of offline training and online inference is used to perform data preprocessing and various types of data prediction to improve the accuracy and generalization ability of fault diagnosis.

Benefits of technology

Improves the accuracy and generalization of fault diagnosis, ensuring efficient fault prediction with lower latency and fewer computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117195953B_ABST
    Figure CN117195953B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of communication technologies, and particularly relates to a fault prediction method and system. It overcomes the problems of low diagnostic accuracy and weak generalization ability existing in the existing methods. Before fault prediction, historical data in the production process is used to perform offline training on the fault prediction model, and the trained model parameters are saved. When performing fault prediction, first, the device status data in the production process is preprocessed according to a predetermined rule, and the preprocessed data is used to predict the device status data at a future moment through the trained data prediction model. Then, the trained fault diagnosis model is used to perform fault diagnosis on the predicted device status data at the future moment to obtain the result of fault prediction. The present invention has high fault diagnosis accuracy and strong generalization ability at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of communication technologies, and particularly relates to a fault prediction method and system. Background Art

[0002] Fault prediction refers to analyzing the operating state of a device through historical state data to predict the operating state of the device within a certain period in the future, including whether a fault will occur, the time and manner of the fault occurrence, etc. In the case of predicting a fault, maintenance personnel can take effective and reasonable maintenance guarantee measures to minimize the possibility of fault shutdown and reduce losses. Currently, there are mainly two fault prediction methods, namely the model-based fault prediction method and the data-driven fault prediction method. The former refers to using expert experience to construct a physical model with a higher degree of compliance with the actual usage scenario. The model constructed by this method has weak generalization ability and requires a large amount of artificial prior knowledge. The data-based fault prediction method, on the other hand, can construct a fault prediction model through real scenario data. The model constructed by this method has high accuracy, strong generalization ability, a simple structure, and does not require a large amount of artificial prior knowledge.

[0003] Fault diagnosis can be regarded as a part of fault prediction. It can master the types and causes of faults occurring in a device through the data during the operation of the device, and can help technicians quickly locate problems when a device fault occurs, reducing the maintenance time.

[0004] The main data-driven methods for fault prediction currently include: time series analysis, support vector machines, artificial neural networks, deep learning, etc. The two most commonly used time series analysis methods are autoregressive moving average models and grey models. In addition to time series analysis methods, machine learning is also a commonly used fault prediction method. For example, in the literature Hu Z, Xiao M. Research of Fault Prediction Based on Time Sequence Model[J]. Computer Measurement & Control, 2013, 21(6): 1421-1427., a new method for feature extraction was proposed. First, the Hilbert Huang transform was performed on the signal, and then SVM and SVR were used to detect degradation of the transformed signal and predict the degradation trend. The prediction error of this method was 0.6%. Artificial intelligence technologies used for fault prediction include artificial neural networks, deep learning, decision tree algorithms, random forests, etc. In the literature Soualhi A, Medjaher K, Zerhouni N. Bearing Health Monitoring Based on Hilbert-Huang Transform, Support Vector Machine and Regression.[J]. IEEE Transactions on Instrumentation & Measurement, 2014, 64(1): 52-62., it was proven that the BP neural network achieved good results in the fault prediction of unmanned aerial vehicle systems, with a prediction accuracy of 87%. In the literature: Su Xujun, Lü Xuezhi. Application Analysis of BP Neural Network Model in Fault Prediction of Unmanned Aerial Vehicle Systems[J]. Computer Applications and Software, 2019, 36(9): 6., it was proposed to use LSTM for bearing life prediction. This method can mine the deep information of bearing fault data based on time domain features, can better predict the degradation trend, and can be used for the real-time estimation of RUL. In the literature Chen Y, Han B. Prediction of Bearing Degradation Trend based on LSTM[C] / / Symposium Series on Computational Intelligence. Xiamen, China: IEEE, 2019: 1035-1040., the CNN-LSTM model was proposed, where CNN is used to automatically extract data features, and LSTM analyzes the feature sequence extracted by CNN. The prediction accuracy of this method can reach 94.65%.

[0005] In Chinese invention patent CN112633317A, a CNN-LSTM fan fault prediction method and system based on an attention mechanism is disclosed, and a fan fault prediction model is proposed. However, in this model, due to the simple structure of CNN-LSTM, when the amount of data is the same, the model cannot extract more effective information from the samples, which may lead to low accuracy of fault prediction. Moreover, since this invention is based on the equipment data of the fan, the generalization ability of this model for different equipment faults may not be strong. In Chinese invention patent CN114118586A, a motor fault prediction method and system based on CNN-BiLSTM is disclosed. The structure of this model is relatively simple, and the input data of this model are vibration data and current data. These two types of data are fused before being input into the model for analysis, which may result in insufficient accuracy of fault prediction. In Chinese invention patent CN113821875A, an intelligent vehicle fault real-time prediction method and system based on cloud-edge collaboration is disclosed. Through the CNN-LSTM prediction model deployed in the cloud for real-time prediction training, a single-system remaining service life prediction model corresponding to each system is obtained and sent to the vehicle end for information integration to obtain the prediction result. This invention has the advantage of cloud-edge collaboration in the fault prediction process. However, for the independent prediction methods used for various types of data, that is, similar to the fault prediction mode of a single sensor, the diagnostic accuracy of this method is relatively low, and the generalization ability is also relatively weak. Summary of the Invention

[0006] The purpose of the present invention is to provide a fault prediction method and system to overcome the problems of low diagnostic accuracy and weak generalization ability existing in the existing methods.

[0007] The technical solution of the present invention is to provide a fault prediction method, which is characterized in that it includes an offline training process of a fault prediction system and an online inference process of the fault prediction system;

[0008] Among them, the offline training process of the fault prediction system is as follows:

[0009] Training a data prediction model and a fault diagnosis model according to the historical data in the production process; the data prediction model is an MKCNN-LSTM model based on an attention mechanism, and the fault diagnosis model is an MFSCNN model;

[0010] Among them, the online inference process of the fault prediction system is as follows:

[0011] First, preprocess the equipment status data in the production process according to predetermined rules; then, use the trained data prediction model to predict the preprocessed equipment status data in the production process to obtain the equipment status data at future times; finally, use the trained fault diagnosis model to perform fault diagnosis on the equipment status data at future times to obtain the results of fault prediction.

[0012] Further, the historical data in the production process and the equipment status data in the production process include n types of data, namely vibration data, current data, voltage data, and / or temperature data, where n is an integer greater than or equal to 2;

[0013] The number of data prediction models corresponds to the number of data types;

[0014] In the offline training process of the fault prediction system, train the corresponding data prediction models according to different types of historical data in the production process; in the online inference process of the fault prediction system, use the trained data prediction models to predict the preprocessed equipment status data of the corresponding type in the production process.

[0015] Further, the MKCNN-LSTM model based on the attention mechanism includes two convolutional modules, an LSTM (Long Short-Term Memory Network) module based on the attention mechanism, and a fully connected layer. Each convolutional module convolves the input signal through convolutional kernels of different sizes, enabling the model to learn different information in the same sample with different receptive fields, rather than extracting a single piece of information from the sample. The LSTM module based on the attention mechanism is used to predict the features learned by the convolutional module; the fully connected layer is used to integrate the information output by the LSTM module based on the attention mechanism to obtain the final prediction result, that is, the equipment status data at future times.

[0016] Further, the MFSCNN model includes three types of modules, namely a hybrid convolutional module, a standard convolutional module, and a classification module. The number of hybrid convolutional modules corresponds to the number of data types; the different types of data output by each MKCNN-LSTM model based on the attention mechanism are respectively extracted with features by the corresponding hybrid convolutional modules and then the features are concatenated, and then the concatenated data features are used as the input of the standard convolutional module to extract the joint features of multiple data sources, and finally the joint features are classified by the softmax classification module.

[0017] Furthermore, each hybrid convolution module includes a softmax classifier and two types of convolution modules, namely a standard convolution module and a dilated convolution module; in the standard convolution module, a first standard convolution layer, a second standard convolution layer, and a max pooling layer are sequentially arranged; in the dilated convolution module, a first dilated convolution layer, a second dilated convolution layer, and a max pooling layer are sequentially arranged; the operation process is as follows:

[0018] (1) Respectively use the first standard convolution layer and the first dilated convolution layer to perform standard convolution and dilated convolution on the input to obtain outputs Z r and Z q ;

[0019] (2) The output feature map Z r of the first standard convolution layer is concatenated with the output feature map Z q of the first dilated convolution layer to obtain Z;

[0020] (3) Use the second dilated convolution layer to perform dilated convolution operation on the feature Z again, and at the same time use the second standard convolution layer to perform standard convolution on the output feature map Z r of the first standard convolution layer again;

[0021] (4) Pool and expand the two convolution output feature maps obtained in step (3) using the corresponding max pooling layer, and then concatenate and output to obtain the effective features learned from the original data;

[0022] (5) Input the learned effective features into the softmax classifier for classification to obtain the final classification result.

[0023] Furthermore, in the preprocessing of the device status data in the production process according to a predetermined rule, the wavelet packet transform method is used to achieve the preprocessing.

[0024] Furthermore, the data prediction model and the fault diagnosis model are trained according to the historical data in the production process, specifically:

[0025] At the central cloud node, the data prediction model and the fault diagnosis model are trained according to the historical data in the production process; and the trained data prediction model is sent to the edge computing node.

[0026] Furthermore, the preprocessing of the production data according to a predetermined rule is specifically:

[0027] The edge computing node receives the production data sent by the terminal device and preprocesses the production data according to a predetermined rule.

[0028] Furthermore, the data prediction of the device status data in the preprocessed production process is performed through the trained data prediction model to obtain the device status data at a future moment, specifically:

[0029] The edge computing node uses the trained data prediction model to perform data prediction on the preprocessed data, and obtains the device status data at future moments; then it sends the predicted results to the central cloud node.

[0030] Furthermore, the trained fault diagnosis model is used to perform fault diagnosis on the device status data at future moments, and the specific results of the fault prediction are as follows:

[0031] After receiving the device status data at future moments sent by the edge computing node, the central cloud node uses the trained fault diagnosis model to perform fault diagnosis and obtains the results of the fault prediction.

[0032] The present invention also provides a fault prediction system for implementing the above-mentioned fault prediction method, which is characterized in that it includes a terminal device, a data preprocessing and prediction module, and a fault diagnosis module;

[0033] The terminal device is used to collect the device status data during the production process and send it to the data preprocessing and prediction module;

[0034] The data preprocessing and prediction module is used to collect the device status data during the production process sent by the terminal device, preprocess the data according to a predetermined rule, and use the trained data prediction model to perform data prediction on the preprocessed data to obtain the device status data at future moments, and send the device status data at future moments to the fault diagnosis module;

[0035] The fault diagnosis module is used for training the data prediction model and the fault diagnosis model, and is used to receive the device status data at future moments, and perform fault diagnosis on the device status data at future moments through the trained fault diagnosis model to obtain the results of the fault prediction.

[0036] Furthermore, the terminal device includes n sensors, which are respectively used to collect the device status data during the production process, including vibration data, current data, voltage data, or temperature data, etc.; where n is an integer greater than or equal to 2.

[0037] Furthermore, the data preprocessing and prediction module is located at the edge computing node, and the fault diagnosis module is located at the central cloud node.

[0038] The beneficial effects of the present invention are:

[0039] 1. High diagnostic accuracy and strong generalization ability;

[0040] The present invention uses an MKCNN-LSTM data prediction model based on the attention mechanism to predict production data, which can make the predicted data closer to the real data. At the same time, the MFSCNN model is used to diagnose faults in the predicted data, which can improve the accuracy and generalization ability of fault diagnosis;

[0041] 2. High data prediction accuracy;

[0042] In the MKCNN-LSTM data prediction model based on the attention mechanism of the present invention, the LSTM module based on the attention mechanism is composed of an LSTM layer and an attention layer. The LSTM layer extracts features related to future data in the data. The attention layer can assign different weights to its inputs according to the degree of influence on the correct output value to learn features that have a greater impact on the output, making the predicted data closer to the real data. At the same time, the input data of the present invention can be vibration data, current data, voltage data, and temperature data. Different data prediction models are used to predict the equipment status data in the corresponding type of production process after preprocessing, further improving the accuracy of data prediction. Then, different types of predicted data are input into the fault diagnosis model for fault diagnosis, further improving the accuracy and generalization ability of fault prediction.

[0043] 3. High fault diagnosis accuracy;

[0044] There are two dilated convolution operations in the hybrid convolution model of the MFSCNN fault diagnosis model of the present invention. One is to perform dilated convolution on the original data, and the other is to perform dilated convolution on the features after fusing the standard convolution results. The second dilated convolution can not only learn the features obtained from the previous dilated convolution but also learn the features extracted by the standard convolution. The combination of dilated convolution and standard convolution enables the standard convolution to make up for the possible blind spot problems brought by dilated convolution and avoid losing continuous information. It can expand the receptive field of model learning without generating blind spots or increasing the number of parameters, and deeply learn the potential features in the data. The MFSCNN model considers the relationship between the features of various types of data and the fault state respectively, and a standard convolution module is added at the end. After learning the features of multi-source data respectively, the standard convolution module further explores the relationship between the features of various types of data and the implicit connection between the fused features and the fault, improving the accuracy of fault diagnosis.

[0045] 4. Can occupy less computing resources at a lower latency;

[0046] Due to the large amount of data generated by the industrial Internet of Things and the large computational workload of the fault prediction algorithm based on artificial intelligence, by deploying part of the model to the edge computing node and jointly performing model inference by the central cloud and the edge computing node, the computational pressure on the central cloud and the transmission pressure on the communication link can be effectively reduced, the security of the data can be guaranteed, and a new idea is provided for the practical application of the fault prediction model. For fault prediction in the industrial field, designing a high-precision fault prediction model for the industrial Internet of Things and implementing a fault prediction system based on the cloud-edge collaborative architecture to complete the prediction under the conditions of low latency and less resource occupancy can improve production efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a flowchart of the fault prediction method of the present invention;

[0048] Figure 2 It is a signal decomposition process based on wavelet packet transform;

[0049] Figure 3 It is a schematic diagram of the attention mechanism;

[0050] Figure 4 It is a schematic diagram of the MKCNN-LSTM network structure based on the attention mechanism of the present invention;

[0051] Figure 5 It is a schematic diagram of the hybrid convolutional network structure of the present invention;

[0052] Figure 6a It is a partial schematic diagram of the MFSCNN network structure of the present invention;

[0053] Figure 6b It is another partial schematic diagram of the MFSCNN network structure of the present invention;

[0054] Figure 7 It is a schematic diagram of the fault prediction model structure of the present invention;

[0055] Figure 8 It is a schematic diagram of the fault prediction system structure of the present invention;

[0056] Figure 9 It is a fault prediction flowchart implemented based on the fault prediction system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention is provided in conjunction with the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments.

[0058] The fault prediction model of the present invention is divided into two parts: a data prediction model and a fault diagnosis model. As Figure 1As shown, before fault prediction, historical data in the production process (i.e., the source domain data in Figure 1 , and the data can be vibration data, current data, voltage data, and / or temperature data of equipment in the production process, etc.) are used to perform offline training on the fault prediction model, and the trained model parameters are saved. The number of data prediction models corresponds to the number of data types, and the corresponding data prediction models are trained according to different types of historical data in the production process. During the training process, the training can be implemented using the process shown in Figure 1 : model pre-training - model training - error calculation and backpropagation optimization - determine whether to converge. If it converges, the training is completed; otherwise, return to continue training. After the model training is completed, the model can be used for online prediction. When performing fault prediction, first, the equipment status data in the production process (i.e., the target domain data in Figure 1 , and the data can be vibration data, current data, voltage data, and / or temperature data of equipment in the production process, etc.) are preprocessed according to a predetermined rule, and the preprocessed data are used to predict the equipment status data at future times through the trained data prediction model. During the online prediction process, the types of equipment status data in the production process also correspond one-to-one with the data prediction models. Then, the trained fault diagnosis model is used to perform fault diagnosis on the predicted equipment status data at future times to obtain the results of fault prediction.

[0059] The data preprocessing method in this embodiment adopts wavelet packet transform. A wavelet packet function is expressed as and its calculation formula is as follows:

[0060]

[0061] where j≥0 represents the transformation of the wavelet function in the frequency domain, k≥0 represents the translation of the function in the time domain, t represents the time parameter, and n represents the decomposition level where the wavelet packet is located. The first two wavelet packet functions are the scaling function and the wavelet function ψ(t), as follows:

[0062]

[0063]

[0064] When n = 2, 3, …, the recurrence relation of the wavelet packet function is as follows:

[0065]

[0066]

[0067] where h(k) is the high-pass filter and g(k) is the low-pass filter. The calculation formula for the wavelet packet coefficients of the signal f(t) is as follows:

[0068]

[0069] As shown in Figure 2 , the signal decomposition process based on wavelet packet transform is shown. The first level of wavelet packet transform is the original data set, and one level represents one step of wavelet packet transform. Each wavelet packet transform decomposes the signal into a high-frequency part and a low-frequency part. In the next transform, the high-frequency and low-frequency parts decomposed in the previous level are decomposed again, and so on in a recursive loop.

[0070] Wavelet packet transform can not only convert the signal into features with both time domain and frequency domain, but also capture high-frequency and low-frequency information simultaneously. Therefore, after the original sequence is transformed by wavelet packet transform to obtain the time domain and frequency domain features of the signal and then sent into the model for learning, the effect is better than directly sending the data into the model.

[0071] The data prediction model designed in the present invention is an MKCNN-LSTM (Multi-Scale Kernel Convolutional Neural Network-Long Short-Term Memory Network) model based on the attention mechanism. The attention mechanism is a special mechanism used to measure the importance of input data to the output. This mechanism can evaluate the importance by calculating the score of each input, so it is widely used in various tasks, from question answering, machine translation to image recognition, etc.

[0072] Assume that H is the output of the LSTM cell, which consists of the hidden vectors [h 1 , h 2 , …, h N . In the first stage of calculating the attention value, the input is multiplied by a weight matrix and then activated through a tanh function to create a score matrix M, and its calculation method is as follows:

[0073] M = tanh(W h H + b h )

[0074] where M ∈ R d×N , b h ∈ R d×N , R is the set of real numbers, W h ∈ R d×d , b h ∈ R d×N , b h is the bias, W h is the weight coefficient, N is the total number of LSTM cells, and each LSTM cell includes N LSTM outputs in total. d is the output of each LSTM.

[0075] Then, a calculation method similar to softmax is used to convert the scores obtained in the first stage. The conversion result is similar to a probability ranking, and the scores of the first stage after conversion are sorted. The conversion result makes the weight α of important elements more prominent, and the following formula is commonly used for this calculation.

[0076]

[0077] where α ∈ R N , V h ∈ R d , V h represents the trainable weight, is the transpose of V h .

[0078] The calculation result of the second stage is the value that the element is worthy of attention. In order to obtain the weighted representation r of the features learned by LSTM, it is necessary to further multiply H by α, as follows:

[0079] r = Hα T

[0080] where r ∈ R d ; α T is the transpose of α.

[0081] The output h * of the last attention layer is calculated as follows:

[0082] h * = relu(W p r + W x h N )

[0083] where h * ∈ R d , W p and W x are trainable weights, and h N is the last hidden vector. As Figure 3 is a schematic diagram of the attention mechanism.

[0084] Based on the attention mechanism, the present invention proposes an MKCNN-LSTM model based on the attention mechanism for data prediction. Figure 4is its network structure diagram. The MKCNN-LSTM model based on the attention mechanism mainly consists of two convolutional modules, an LSTM (Long Short-Term Memory Network) module based on the attention mechanism, and a fully connected layer. Among them, the two convolutional modules perform convolution on the input signal through convolutional kernels of different sizes, enabling the model to learn different information in the same sample with different receptive fields, rather than extracting only single information from the sample. Convolutional module 1 and convolutional module 2 have convolutional layers composed of convolutional kernels of different sizes. In one embodiment, the convolutional kernel sizes in convolutional module 1 are all 3×8. The first convolutional layer has 24 convolutional kernels, and the second convolutional layer has 12 convolutional kernels. The input edges are filled with 0. Therefore, the output size of convolutional layer conv_1 in convolutional module 1 is 10×24, and the output size of conv_2 is 10×12. The convolutional kernel sizes in convolutional module 2 are all 5×8. The first convolutional layer has 24 convolutional kernels, and the second convolutional layer has 12 convolutional kernels. The input edges are filled with 0. The output sizes of the two convolutional layers conv_3 and conv_4 in convolutional module 2 are 10×24 and 10×12 respectively. The filter size of the pooling layer is 2×1, and the input is divided into regions with a size of 2×1 to reduce the dimension of the input. The output of the attention layer is 16×1, and the model finally outputs a one-dimensional vector of 8×1, which is the final predicted value.

[0085] After the data is preprocessed, it is input into the model. The model performs multi-scale convolution on the input, then pools the convolution results, concatenates the obtained pooling outputs, and inputs them into the LSTM module based on the attention mechanism. Finally, the output of the attention layer is integrated through the fully connected layer to obtain the final prediction result.

[0086] In the fault diagnosis stage of the present invention, a hybrid convolutional model for data feature extraction is designed. As Figure 5 shown, the hybrid convolutional model is mainly divided into three modules: input, convolution, and output. Among them, the convolution includes two convolutional modules: standard convolution and dilated convolution. The standard convolutional module has two standard convolutional layers (the first standard convolutional layer and the second standard convolutional layer) and a max pooling layer. The dilated convolutional module includes two dilated convolutional layers (the first dilated convolutional layer and the second dilated convolutional layer) and a max pooling layer. Data feature interaction and sharing occur during the convolution process of the hybrid convolutional model.

[0087] Among them, the output of the standard convolution Z r and the dilated convolution Z q are respectively:

[0088] Z r = f(W r * X m + b r )

[0089]

[0090] The output after concatenating standard convolution and dilated convolution is as follows:

[0091]

[0092] where X m is the input of the standard convolution, and X m (q) is the input of the dilated convolution. W r and b r are the weights and biases of the standard convolution, and are the weights and biases of the dilated convolution, both of which are trainable parameters. The data is input into the convolution module, and the operation process in the convolution module is as follows:

[0093] (1) Use the first standard convolution layer and the first dilated convolution layer to perform standard convolution and dilated convolution on the input respectively, to obtain the outputs Z r and Z q ;

[0094] (2) Concatenate the output feature map Z r of the first standard convolution layer and the output feature map Z q of the first dilated convolution layer to obtain Z;

[0095] (3) Use the second dilated convolution layer to perform dilated convolution operation on the feature Z again, and at the same time use the second standard convolution layer to perform standard convolution on the output feature map Z r of the first standard convolution layer again;

[0096] (4) Expand the two convolution output feature maps obtained in step (3) after pooling with the corresponding max - pooling layer, and then concatenate and output to obtain the effective features learned from the original data;

[0097] (5) Input the learned effective features into the softmax classifier for classification to obtain the final classification result.

[0098] The fault diagnosis model designed in the present invention is the MFSCNN (Multi - Feature Fusion Stacked Convolutional Neural Network) model. The network structure of the MFSCNN model is as Figure 6a and Figure 6b shown. It mainly consists of three types of modules, namely, the hybrid convolution module, a standard convolution module, and a classification module, where the number of hybrid convolution modules corresponds to the number of data types. After different types of data are used by the hybrid convolution module to extract features, the features are concatenated, and then the concatenated data features are used as the input of the standard convolution module to extract the joint features of multiple data sources. Finally, the joint features are classified by the softmax classification module.

[0099] In one embodiment, the input size is 10×8. In the hybrid convolution module of the MFSCNN model, the convolution kernel sizes of the standard convolution layers conv, conv_2, conv_4, and conv_6 are all 3×8, and the stride is 1. The convolution kernel sizes of the dilated convolution layers conv_1, conv_3, conv_5, and conv_7 are also 3×8, the stride is 1, and the dilation rate is 2. No padding is performed on the edges of the input feature maps of the standard convolution and the dilated convolution, and each convolution layer has 12 convolution kernels. The pooling window size of the max pooling layer is 2×1, and the input is downsampled with a stride of 2. The convolution kernel sizes in the standard convolution module added at the end of the MFSCNN model are all 5×1, the stride is 1, and the number of convolution kernels in the convolution layer is 12. The pooling window size in the max pooling layer is 2×1, and the pooling window traverses and downsamples the input without overlapping.

[0100] The present invention combines a data prediction model and a fault diagnosis model to obtain a fault prediction model. As Figure 7 shown.

[0101] The present invention also discloses a fault prediction system based on cloud-edge collaboration, as Figure 8 shown. The fault prediction system is divided into three parts, namely: terminal devices, edge computing nodes, and central cloud nodes. The terminal devices are responsible for collecting production data and sending it to the edge computing nodes; the edge computing nodes are responsible for collecting the production data sent by the terminal devices, preprocessing the data according to predetermined rules, and performing data prediction model inference, and sending the inference results to the central cloud nodes; the central cloud nodes are responsible for training the fault prediction model, receiving the inference results of the edge computing nodes, and performing fault diagnosis based on the inference results of the edge computing nodes.

[0102] The offline training process of the fault prediction system based on cloud-edge collaboration is as follows:

[0103] (1) Train the data prediction model and the fault diagnosis model based on the historical data in the production process at the central cloud node;

[0104] (2) After the model training is completed, the central cloud distributes the data prediction model to the edge computing nodes.

[0105] The online inference stage of the fault prediction system based on cloud-edge collaboration is as follows:

[0106] (1) Data preprocessing. The edge computing node receives the terminal device data and preprocesses the data according to the predetermined rules;

[0107] (2) Data prediction. The edge computing node performs data prediction on the preprocessed data through the data prediction model, and then sends the predicted results to the central cloud node;

[0108] (3) Fault diagnosis. After the central cloud node receives the data predicted by the edge computing node, it uses this data for fault diagnosis.

[0109] The prediction process in the fault prediction system is as Figure 9 shown. The fault prediction model of the system is a combination of the MKCNN-LSTM model based on the attention mechanism and the MFSCNN model. The terminal device includes multiple sensors, which are respectively used to collect device status data during the production process, including vibration data, current data, voltage data, temperature data, etc.; the number of data prediction models corresponds to the number of data types. Figure 9 Taking two types of device status data as input as an example, including two sensors. In other embodiments, it can be three types, four types, etc. After the offline training is completed, the central cloud distributes the trained MKCNN-LSTM model based on the attention mechanism to the edge computing device. When the system performs online fault prediction on different monitoring data, it first preprocesses the original data according to a predetermined rule to extract the shallow features of the data; the preprocessed data is respectively predicted by the MKCNN-LSTM model based on the attention mechanism to obtain the data at future moments; then the MFSCNN model deployed on the central cloud is used to perform feature fusion classification on the predicted multi-type data to obtain the fault status at future moments.

Claims

1. A fault prediction method, characterized in that: it includes an offline training process of the fault prediction system and an online inference process of the fault prediction system; wherein, the offline training process of the fault prediction system is: training the data prediction model and the fault diagnosis model according to the historical data in the production process; the data prediction model is an MKCNN-LSTM model based on the attention mechanism, and the fault diagnosis model is an MFSCNN model; wherein, the online inference process of the fault prediction system is: first, preprocess the device status data in the production process according to a predetermined rule; then, perform data prediction on the preprocessed device status data in the production process through the trained data prediction model to obtain the device status data at a future moment; finally, perform fault diagnosis on the device status data at the future moment through the trained fault diagnosis model to obtain the result of fault prediction; The MKCNN-LSTM model based on the attention mechanism includes two convolutional modules, an LSTM module based on the attention mechanism, and a fully connected layer; wherein each convolutional module is used to perform convolution on the input signal through convolutional kernels of different sizes; the LSTM module based on the attention mechanism is used to predict the features learned by the convolutional module; the fully connected layer is used to integrate the information output by the LSTM module based on the attention mechanism to obtain the final prediction result; The MFSCNN model includes three types of modules, namely a hybrid convolutional module, a standard convolutional module, and a classification module; the number of hybrid convolutional modules corresponds to the number of data types; different types of data output by each MKCNN-LSTM model based on the attention mechanism are respectively extracted with features by the corresponding hybrid convolutional module and then the features are concatenated, and then the concatenated data features are used as the input of the standard convolutional module to extract the joint features of multiple data sources, and finally the joint features are classified by the classification module; Each hybrid convolutional module includes a softmax classifier and two types of convolutional modules, and the two types of convolutional modules are a standard convolutional module and a dilated convolutional module respectively; in the standard convolutional module, a first standard convolutional layer, a second standard convolutional layer, and a max pooling layer are sequentially arranged; in the dilated convolutional module, a first dilated convolutional layer, a second dilated convolutional layer, and a max pooling layer are sequentially arranged; the operation process is: (1) Respectively use the first standard convolutional layer and the first dilated convolutional layer to perform standard convolution and dilated convolution on the input to obtain an output and; (2) The output feature map of the first standard convolutional layer is concatenated with the output feature map of the first dilated convolutional layer to obtain; (3) Use the second dilated convolutional layer to perform a dilated convolution operation on the features again, and at the same time use the second standard convolutional layer to perform standard convolution on the output feature map of the first standard convolutional layer again; (4) Pool and expand the two convolutional output feature maps obtained in step (3) by using the corresponding max pooling layer, and then concatenate and output to obtain the effective features learned from the original data; (5) Input the learned effective features into the softmax classifier for classification to obtain the final classification result.

2. The fault prediction method according to claim 1, characterized in that: Both the historical data in the production process and the equipment status data in the production process include n types of data, namely vibration data, current data, voltage data, and / or temperature data, where n is an integer greater than or equal to 2; The number of data prediction models corresponds to the number of data types; During the offline training process of the fault prediction system, the corresponding data prediction models are trained according to different types of historical data in the production process; during the online inference process of the fault prediction system, the trained data prediction models are used to perform data prediction on the preprocessed equipment status data of the corresponding type in the production process.

3. The fault prediction method according to claim 2, characterized in that: In the preprocessing of the equipment status data in the production process according to a predetermined rule, the wavelet packet transform method is used to achieve the preprocessing.

4. The fault prediction method according to claim 3, characterized in that: The data prediction model and the fault diagnosis model are trained according to the historical data in the production process, specifically: The data prediction model and the fault diagnosis model are trained at the central cloud node according to the historical data in the production process; and the trained data prediction model is sent to the edge computing node.

5. The fault prediction method according to claim 4, characterized in that: The preprocessing of the production data according to a predetermined rule is specifically: The edge computing node receives the production data sent by the terminal device and preprocesses the production data according to a predetermined rule.

6. The fault prediction method according to claim 5, characterized in that: The preprocessed equipment status data in the production process is subjected to data prediction by the trained data prediction model to obtain the equipment status data at a future time, specifically: The edge computing node performs data prediction on the preprocessed data by the trained data prediction model to obtain the equipment status data at a future time; and then sends the prediction result to the central cloud node.

7. The fault prediction method according to claim 6, characterized in that: The equipment status data at a future time is subjected to fault diagnosis by the trained fault diagnosis model to obtain the result of fault prediction, specifically: After receiving the equipment status data at a future time sent by the edge computing node, the central cloud node uses the trained fault diagnosis model to perform fault diagnosis to obtain the result of fault prediction.

8. A fault prediction system for implementing the fault prediction method according to any one of claims 1-7, characterized in that: It includes a terminal device, a data preprocessing and prediction module, and a fault diagnosis module; The terminal device is used to collect the equipment status data in the production process and send it to the data preprocessing and prediction module; The data preprocessing and prediction module is used to collect the equipment status data in the production process sent by the terminal device, preprocess the data according to a predetermined rule, and is used to perform data prediction on the preprocessed data by the trained data prediction model to obtain the equipment status data at a future time, and send the equipment status data at a future time to the fault diagnosis module; The fault diagnosis module is used for training the data prediction model and the fault diagnosis model, receiving the device status data at future moments, and performing fault diagnosis on the device status data at future moments through the trained fault diagnosis model to obtain the results of fault prediction.

9. The fault prediction system according to claim 8, wherein: The terminal device includes n sensors, which are respectively used for collecting device status data during the production process, including vibration data, current data, voltage data or temperature data, where n is an integer greater than or equal to 2.

10. The fault prediction system according to claim 9, wherein: The data preprocessing and prediction module is located at the edge computing node, and the fault diagnosis module is located at the central cloud node.

Citation Information

Patent Citations

  • CNN-LSTM fan fault prediction method and CNN-LSTM fan fault prediction system based on attention mechanism

    CN112633317A

  • Intelligent vehicle fault real-time prediction method and system based on end-cloud cooperation

    CN113821875A

  • CNN-Bi LSTM-based motor fault prediction method and system

    CN114118586A

  • Large-scale experimental device power device fault diagnosis method based on deep convolutional neural network

    CN111505424A

  • Equipment failure mode prediction method based on double deep learning models

    CN112989976A