Escalator transmission system fault diagnosis method based on multi-modal fusion
Through the multimodal fusion fault diagnosis method, using vibration and noise sensor data, combined with BiLSTM and cross-attention mechanism, the accuracy problem of escalator transmission system fault diagnosis is solved, and the loose anchor bolts of the drive host are efficiently identified.
Patent Information
- Application Number
- CN202510848588.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-06-24
AI Technical Summary
Existing technologies make it difficult to accurately diagnose faults in escalator drive systems, especially loose anchor bolts in the drive unit. Traditional methods rely on single signal features and the cost of training complex neural networks is high, resulting in insignificant diagnostic results.
A multimodal fusion fault diagnosis method is adopted. Vibration and noise sensors are used to collect time domain and frequency domain data. The BiLSTM model and cross-attention mechanism are used for feature extraction and information fusion. The decision-level data fusion algorithm is combined to improve the diagnostic accuracy.
The accuracy and reliability of escalator drive system fault diagnosis are improved, and the degree of looseness of the drive host anchor bolts can be identified, avoiding the limitations and abnormalities of single sensor and single modal data.
Smart Images

Figure CN120354215B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a mechanical fault diagnosis method, in particular to an escalator transmission system fault diagnosis method based on multi-modal fusion. Background Art
[0002] With the development of society and the economy, my country's infrastructure has gradually improved. In recent years, escalators, as highly efficient personnel transportation equipment in public places, have been widely used in subway stations, large shopping malls, hospitals, and other public places. With the increasing use, the demand for escalator maintenance has also continued to grow. Due to the complex and diverse operating environments of escalators, the frequency of escalator failures has also continued to increase. The transmission system is the most important system for the safe operation of escalators. Failures in components such as the drive unit, handrails, and steps are highly likely to cause serious safety accidents. Therefore, fault diagnosis of escalator transmission systems is of great significance to protecting people's lives and property.
[0003] When an escalator is in operation, the anchor bolts of the drive unit are prone to loosening, which can lead to abnormal faults. For rotating machinery such as escalators, vibration signals, noise signals, and speed signals are the most common fault diagnosis signals, which contain the fault characteristics of the escalator. Traditional fault diagnosis methods mostly use time domain signal processing technology to extract features from the time domain of the signal, and then input them into the classifier for diagnosis; traditional methods rely too much on expert experience and the quality of the extracted features. However, the operating environment of escalators is complex, and most escalators are in full load operation for a long time. It is impossible to accurately judge the fault of the escalator by relying solely on a single type of signal feature.
[0004] In recent years, with the continuous development of deep learning technology, end-to-end neural networks for fault diagnosis have developed rapidly. Due to their lack of reliance on expert experience and manual feature extraction, they are widely used in the field of fault diagnosis. Due to the complex operating environment of escalators and the strong background noise, simple neural networks have limited effectiveness. However, complex neural networks require a large number of parameters to train, which significantly increases training costs. Therefore, there is an urgent need for an efficient and lightweight neural network model for fault diagnosis of escalator drive systems. Summary of the Invention
[0005] The purpose of the present invention is to overcome the shortcomings of the above-mentioned background technology and provide an escalator transmission system fault diagnosis method based on multimodal fusion. By processing the original time domain signal and then undergoing multiple rounds of neural network training, the final diagnosis result is obtained, thereby improving the fault diagnosis rate of the escalator transmission system.
[0006] The technical solution provided by the present invention is:
[0007] A fault diagnosis method for escalator transmission system based on multimodal fusion is carried out in the following steps:
[0008] Step (1) Data acquisition and data preprocessing to obtain time domain vibration data , time domain noise data , frequency domain vibration data and frequency domain noise data .
[0009] (1.1) The data collection method is as follows: according to the looseness of the anchor bolts of the driving host, the looseness fault is defined as slight looseness, moderate looseness, and severe looseness, which correspond to one circle, two circles, and three circles of loose anchor bolts respectively. The vibration data and noise data of the anchor bolts of the escalator driving host in normal state, slight looseness state, moderate looseness state, and severe looseness state are collected; the time domain data collected by the vibration sensor and the noise sensor are recorded as the two types of data respectively. and , and There are four types of fault data in each, namely normal state, slightly loose state, moderately loose state, and severely loose state.
[0010] (1.2) The data preprocessing method is: the collected vibration data and noise data are respectively subjected to the normalization method based on the envelope value to realize the noise suppression function and enhance the distinguishability of the signal; the obtained time domain data is enhanced and labeled according to the sliding window segmentation method, and each group of segmented data is subjected to fast Fourier transform to obtain the corresponding frequency domain data, and finally the time domain vibration data is obtained. , time domain noise data , frequency domain vibration data and frequency domain noise data The processed time domain data and frequency domain data are used to create a data set, which is divided into a training set, a validation set, and a test set according to a ratio of 7:2:1.
[0011] Step (2) inputs the training set data into the neural network model proposed in the present invention for training, and uses the validation set to verify the training effect.
[0012] (2.1) The time domain vibration data obtained after the first step of collection and processing is and frequency domain vibration data The data is input into three consecutive time domain feature extraction blocks and frequency domain feature extraction blocks for preliminary feature extraction. The results are then input into the BiLSTM model to capture the bidirectional dependencies of long time series. The time domain output results and frequency domain output results are fused across modalities through the cross-attention mechanism. By establishing connections between different modalities, the complementary information and heterogeneous information between different modalities can be better captured, thereby improving the performance of the model in complex tasks. The data obtained through the cross-attention mechanism is input into the global average pooling layer to reduce the channel dimension to 1, and then input into the fully connected layer and the Softmax activation function to obtain the probability corresponding to each fault category, which is recorded as the preliminary vibration diagnosis result. .
[0013] Time domain noise data obtained after noise sensor processing and frequency domain noise data Train according to the method in 2.1 and finally obtain the preliminary diagnosis results of noise .
[0014] (2.2) Combine the probability vectors of the preliminary diagnosis results obtained from the two types of data into one The probability matrix , is the number of classifiers, which is 2 in this invention, The number of fault categories outputted in the final output is 4 in this invention. 、 and Matrix, will The probability of each failure in the matrix is summed up to get Vector, calculation The fault category corresponding to the maximum probability in the vector is the fault category predicted by the neural network; if If there are two or more equal maximum probabilities in the vector, one of them is randomly selected as the diagnosis result. The role of the matrix is to obtain the classification ability weights of different classifiers and multiply the weights by the preliminary diagnosis results; The role of the matrix is to obtain the weight of the classification ability of the same classifier for each category and multiply the weight by the preliminary diagnosis result; Matrix Pass and The dot product of the matrix obtains the preliminary diagnosis results after weighting according to the capabilities of different classifiers and the weighting of the classification capabilities of the same classifier for different categories; Add the probabilities of the two classifiers corresponding to the categories to get the final diagnosis result. Figure 5 .
[0015] In step (2.2) The matrix calculation formula is: ,in, is a matrix No. Rank Elements of the column, is the number of classifiers. , is a matrix No. Rank Elements of a column.
[0016] In step (2.2) The matrix calculation formula is: ,in, is a matrix No. Rank Elements of the column, is the fault classification number, in the present invention , is a matrix No. Rank Elements of a column.
[0017] In step (2.2) The calculation formula of the matrix is: ,Right now Each element in the matrix is and The corresponding elements in the matrix are multiplied element by element.
[0018] In step (2.2) The vector calculation formula is: ,Right now The sum of the probabilities of each fault type in the matrix.
[0019] (2.3) Train the model until one round of training is complete, save the final trained model parameters, and calculate the classification accuracy of the training set and validation set in one round of training.
[0020] Step (3) Repeat the second step above to train the neural network. If the number of training rounds reaches 100, stop training.
[0021] Step (4) Save the network parameters obtained from model training, input the test set into the trained neural network, and calculate the accuracy.
[0022] The formula used in the present invention is as follows:
[0023] The normalization method based on envelope value used in step (1.2) is as follows: (1) Extract envelope value: Use Hilbert transform on the original signal to obtain the analytical signal of the signal. The modulus of the analytical signal is the instantaneous amplitude of the signal, that is, the envelope. Then fit the envelope curve and divide the original signal by the envelope curve to obtain the normalized signal. The overall formula is: ,in is the normalized data, is the original data, is the imaginary part of the analytical signal obtained after Hilbert transform.
[0024] The sliding window segmentation method in step (1.2) is to use a length of The window moves data, split the original data into The length is Time series data, the final insufficient data is discarded, if the original data packet contains data, then .
[0025] The fast Fourier transform in step (1.2) is a method for quickly calculating the discrete Fourier transform. Its essence is still the discrete Fourier transform. The formula of the discrete Fourier transform is: ,in, is the kth output of the discrete Fourier transform, is the nth sample of the input sequence, N is the length of the input sequence, and j is the imaginary unit. The fast Fourier transform reduces the computational complexity of the discrete Fourier transform from O( ) is reduced to O( ).
[0026] The time domain feature extraction block and the frequency domain feature extraction block in step (2.1) sequentially include a one-dimensional convolution layer, a ReLu layer, a one-dimensional convolution layer, a ReLu layer, a one-dimensional maximum pooling layer, a ReLu layer, and a one-dimensional batch normalization layer. The formula for the one-dimensional convolution layer is: ,in, Output sequence elements, is the value of the region corresponding to the convolution kernel in the input sequence, is the convolution kernel elements, is the length of the convolution kernel. The formula of the ReLu layer is: , is the input value. The formula of the one-dimensional maximum pooling layer is: given a one-dimensional vector , and the pooling window size , the one-dimensional maximum pooling layer will output a new vector ,in ,in =1,2,…,m. The formula of the one-dimensional batch normalization layer is: given a one-dimensional vector , the one-dimensional batch normalization layer will output a new vector , the batch normalization process includes the following steps: (1) Calculate the mean: (2)Calculate the variance: . (3) Normalization: For each element Normalize to get , ,in is a very small constant to prevent the denominator from being zero. (4) Scaling and offsetting: Scale and offset the normalized elements to obtain the output elements , ,in and are learnable parameters representing the scaling factor and offset respectively.
[0027] The cross attention mechanism in step (2.1) is an attention mechanism for processing the relationship between two different sequences. Its input generally comes from two different sequences. In the present invention, the two input sequences of the cross attention mechanism are the time domain sequence and frequency domain sequence of each sensor, which are the query vector Q, the key vector K and the value vector V (see Figure 4 ). The calculation process of cross attention is as follows: (1) Calculate the attention weight: , where Q and K are the query and key matrices, is the dimension of the key vector, which is used to scale the dot product result to avoid excessive values. (2) Weighted summation: , applying the attention weight to the value V, where Linear is a fully connected layer. The cross-attention mechanism allows the model to dynamically extract information related to one sequence from another, which is particularly suitable for multimodal or cross-modal tasks.
[0028] The global average pooling layer formula in step (2.1) is: , where for each and Where B is the number of batches, C is the number of channels, and L is the signal length.
[0029] The fully connected layer formula in step (2.1) is: ,in, is the output vector, is a weight matrix with dimension (n,m), where n is the number of neurons in the fully connected layer and m is the number of features in the input vector. is the input vector, As the bias vector, the fully connected layer can achieve the effect of data dimensionality reduction, and the data can be reduced to the number of faults that need to be classified for classification tasks.
[0030] In step (2.1) The formula for the activation function is: ,in, is the input vector, is the dimension of the input vector, The role of the activation function is to convert a vector or a set of real numbers into a probability distribution, that is, the value of each element is between 0 and 1, and the sum of all elements is 1. This makes the Softmax function very suitable for representing probabilities in classification problems.
[0031] The formula for calculating the classification accuracy of the training set in step (2.2) is: ,in The number of correct samples predicted in all training set samples, is the number of samples in all training sets. The formula for the classification accuracy of the validation set is: ,in is the number of samples predicted correctly in all validation set samples, is the number of samples in all validation sets.
[0032] The loss function used in step (2.2) is the cross entropy loss function, and the formula is: ,in, is the true label, To predict the label, is the loss value, is the number of categories. This measures the distance between two probability distributions, using the true label and the predicted label. One probability distribution is the true label distribution, and the other is the model's predicted probability distribution. A smaller cross-entropy loss function indicates a more accurate model prediction. Since the cross-entropy loss function is convex, it is minimized using gradient descent. The optimizer used in this invention is the Adam optimizer. This optimizer adjusts the learning rate of each parameter and updates the parameters according to the gradient descent direction to minimize the cross-entropy loss function.
[0033] The formula for calculating the classification accuracy of the training set in step (4) is: ,in The number of correct samples predicted in all training set samples, is the number of samples in all training sets.
[0034] The beneficial effects of the present invention are:
[0035] The present invention collects multi-sensor data to diagnose faults in the escalator transmission system, opening up a new research direction in the field of escalator fault diagnosis. The method proposed in the present invention has good performance in diagnosing faults in the escalator transmission system and can accurately identify loosening faults and the degree of faults in the reduction gearbox anchor bolts of the escalator drive main unit.
[0036] Currently, most existing technologies use single sensor data for fault diagnosis, and in the feature extraction process, only extract data features under a single mode, and the feature representation is single. At the same time, existing technologies often only use the output of one classifier as the final result, without considering the situation where the probability of the fault category is extremely close when a single classifier is used for fault classification, resulting in misclassification. In response to the above problems of the existing technologies, the present invention improves and combines the existing technologies from the following three points.
[0037] This paper employs a cross-attention mechanism and a multimodal BiLSTM neural network, using signal vibration and noise signals, along with corresponding time-domain and frequency-domain signals, to form multimodal data. This fully overcomes the limitations of single-sensor data. Using multi-sensor and multimodal data not only extracts complementary features between different types of data, but also avoids anomalies in single-sensor or single-modal data. The cross-attention mechanism fully extracts correlations between different sequences and fully integrates the complementarity of data from different modalities. Furthermore, the use of multi-source data and multimodality can avoid fault diagnosis errors caused by a single data anomaly.
[0038] After generating preliminary diagnostic results, the present invention uses a decision-level data fusion algorithm to perform decision-level data fusion. This overcomes the problems of misclassification caused by extremely close classification probabilities for fault categories in a single classifier, as well as the randomness of a single classifier. In summary, the present invention can significantly improve various indicators of escalator transmission system fault diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 This is a flow chart of the method adopted in an embodiment of the present invention.
[0040] Figure 2 This is a flowchart of the neural network used in an embodiment of the present invention.
[0041] Figure 3 for Figure 2 Flowchart of the time domain feature extraction block and frequency domain feature extraction block in .
[0042] Figure 4 Flowchart of the cross-attention mechanism in an embodiment of the present invention.
[0043] Figure 5 This is a flow chart of the decision-level fusion algorithm in an embodiment of the present invention. DETAILED DESCRIPTION
[0044] The present invention will be further described below with reference to the embodiments shown in the accompanying drawings.
[0045] like Figures 1 to 5 As shown, this embodiment performs fault diagnosis on an escalator drive system based on time-domain and frequency-domain signals, obtaining fault classification probabilities and final prediction labels. This approach uses multimodal data based on deep learning to solve the escalator drive system fault diagnosis problem. The method is divided into a data preprocessing process, a preliminary fault diagnosis process, and a decision-level data fusion process.
[0046] The escalator transmission system fault diagnosis method based on multimodal fusion adopted in the embodiment of the present invention is performed according to the following steps:
[0047] Step (1): Data acquisition and data preprocessing to obtain time domain vibration data , time domain noise data , frequency domain vibration data and frequency domain noise data .
[0048] Step (1.1) defines looseness faults as slight looseness, moderate looseness, and severe looseness according to the looseness of the anchor bolts of the drive main unit, which correspond to one, two, and three turns of loose anchor bolts, respectively. Vibration data and noise data of the anchor bolts of the escalator drive main unit in normal, slight looseness, moderate looseness, and severe looseness are collected. The time domain data collected by the vibration sensor and the noise sensor are recorded as and .
[0049] There are four types of fault data: normal state, slightly loose state, moderately loose state, and severely loose state; as described below:
[0050] [Normal:[0.05327525, 0.0560805, 0.07944125, 0.074767875, 0.0700945, 0.051407125, 0.048601875, 0.045796625, 0.053272188, 0.06074775, 0.0691635, …, 0.038318],
[0051] Mild: [0.141126125, 0.103739125, 0.118696375, 0.1074815, 0.113085875, 0.0943985, 0.0943985, 0.09252425, 0.055143375, 0.114960125, 0.10467625, …, 0.14206325],
[0052] Moderate: [0.113085875, 0.122432625, 0.11682825, 0.114960125, 0.131779375, 0.126175, 0.118696375, 0.126175, 0.1074815, 0.113085875, 0.118696375, …, 0.121501625],
[0053] Heavy : [0.113085875, 0.109349625, 0.09159325, 0.1074815, 0.0943985, 0.09813475, 0.100002875, 0.093461375, 0.086919875, 0.087857, 0.090656125, …, 0.10467625]].
[0054] The data format collected from the noise signal is consistent with that of the vibration signal (omitted).
[0055] In step (1.2), the collected vibration and noise data are subjected to envelope-based normalization to achieve noise suppression and enhance signal differentiation. The resulting time-domain data is augmented and labeled using a sliding window segmentation method. Each segmented data set is subjected to a fast Fourier transform to obtain the corresponding frequency-domain data. The processed time-domain and frequency-domain data are then used to construct a dataset, which is partitioned into training, validation, and test sets in a 7:2:1 ratio.
[0056] The above signal preprocessing takes vibration data as an example, and the preprocessing method of noise data is the same.
[0057] The steps of the envelope value-based normalization method are as follows: (1) Extract the envelope value: Use the Hilbert transform on the original signal to obtain the analytical signal of the signal. (2) Fit the envelope curve: The modulus of the analytical signal is the instantaneous amplitude of the signal, that is, the envelope, and then fit the envelope curve. (3) Signal normalization: Divide the original signal by the envelope curve to obtain the normalized signal. The overall formula is: ,in is the normalized data, is the original data, the imaginary part of the analytic signal obtained after Hilbert transform.
[0058] The specific implementation steps are:
[0059] (1.21) Extract envelope value: the formula of Hilbert transform is: wherein, is the time domain signal, is the Hilbert transform operator, which is used to convert a real number signal into the imaginary part of an analytic signal. The formula for calculating the analytic signal is: wherein is the imaginary unit, that is, the analytic signal is the original time domain signal as the real part, and the signal after the Hilbert transform as the imaginary part, to form a new analytic signal. Take the time domain signal in 1.1 as an example for demonstration, and the obtained analytic signal is:
[0060] Normal: [0.053275-0.008017j, 0.056081-0.000628j, 0.079441-0.008838j, 0.074768+0.019686j, 0.070095+0.018162j, 0.051407+0.027814j, 0.048602+0.007864j, 0.045797+0.008192j, 0.053272-0.007968j, 0.060748-0.002380j, 0.069164-0.003472j, …, 0.038318+0.032575j],
[0061] Mild: [0.141126+0.029766j, 0.103739+0.017909j, 0.118696+0.011501j, 0.107481+0.019085j, 0.113086+0.017383j, 0.094398+0.029167j, 0.094398+0.006957j, 0.092524+0.030716j, 0.055143-0.007076j, 0.114960-0.030877j, 0.104676+0.023956j, …, 0.142063-0.020732j],
[0062] Moderate: [0.113086-0.000172j, 0.122433-0.006380j, 0.116828+0.003627j, 0.114960-0.010890j, 0.131779-0.006523j, 0.126175+0.008513j, 0.118696+0.001778j, 0.126175+0.008643j, 0.107482+0.011338j, 0.113086-0.005713j, 0.118696+0.002944j,..., 0.121502+0.000191j],
[0063] Severe: [0.113086+0.000390j, 0.109350+0.018627j, 0.091593+0.007025j, 0.107482+0.004722j, 0.094399+0.013674j, 0.098135+0.001234j, 0.100003+0.010850j, 0.093461+0.010062j, 0.086920+0.007469j, 0.087857-0.002154j, 0.090656-0.004384j,..., 0.104676+0.001452j].
[0064] (1.22) Fitting the envelope curve: the modulus of the analytic signal is the instantaneous amplitude of the signal, i.e. the envelope , . Taking the analytic signal calculated in (1.21) as an example, the envelope value is obtained as:
[0065] Normal: [0.05387511, 0.056084012, 0.079931311, 0.077316012, 0.072409148, 0.058449265, 0.049233942, 0.046523626, 0.053864847, 0.060794351, 0.06925058,..., 0.050293033],
[0066] Mild: [0.144230982, 0.105273644, 0.119252275, 0.109162738, 0.114414124, 0.098801703, 0.09465452, 0.097489416, 0.055595569, 0.119034455, 0.10738254,..., 0.143568059],
[0067] Moderate: [0.113086005, 0.12259875, 0.116884528, 0.115474758, 0.131940713, 0.126461878, 0.118709685, 0.126470653, 0.108077867, 0.113230076, 0.118732869, …, 0.121501775],
[0068] Heavy : [0.113086548, 0.110924768, 0.091862278, 0.107585186, 0.095383661, 0.098142514, 0.100589761, 0.094001437, 0.087240231, 0.087883395, 0.09076208, …, 0.10468632]].
[0069] (1.23) Signal normalization: convert the original signal Divide by the envelope curve That is, the normalized signal is obtained The formula is: Using the above signal for demonstration, the normalized signal is:
[0070] [Normal:[0.487797396, 0.51348275, 0.727377814, 0.684587585, 0.641797356, 0.470692521, 0.445007167, 0.419321813, 0.48776936, 0.556216897, 0.633272959, …, 0.35084623],
[0071] Mild: [0.974663814, 0.716456795, 0.819756523, 0.742302877, 0.78100855, 0.651947341, 0.651947341, 0.63900315, 0.380838432, 0.793952741, 0.722928891, …, 0.98113591],
[0072] Moderate: [0.810561257, 0.877555596, 0.837385334, 0.823995246, 0.944549935, 0.904379672, 0.850775421, 0.904379672, 0.770390995, 0.810561257, 0.850775421, …, 0.870882503],
[0073] Heavy : [0.930376786, 0.89963802, 0.753553293, 0.884268637, 0.776632564, 0.80737133, 0.822740713, 0.768922676, 0.71510464, 0.722814527, 0.745843406, …, 0.861189366]].
[0074] Since it is difficult to collect escalator on-site fault data, there is a problem of insufficient data volume. The sliding window segmentation method uses a length of The window moves each time data, split the original data into The length is Time series data, the insufficient data will be discarded. Greater than , some of the data will be reused to expand the data, that is, data enhancement. Data enhancement can extract enough information from small sample data to improve the effect of model training. data, then .
[0075] Taking vibration data as an example, the vibration data contains four categories of data. Assuming that each category of data has 300,000, each category of data is segmented by a sliding window, and the insufficient data is discarded. , , we will eventually get 585 pieces of data for each category, and each piece of data contains 1024 data points.
[0076] After segmenting the vibration data into four categories, add a corresponding label to each piece of data. For example, if there are 585 pieces of normal state data, label each piece 0. Similarly, label the slightly loose, moderately loose, and severely loose data pieces 1, 2, and 3, respectively.
[0077] After labeling is completed, 585 pieces of vibration data of each category and corresponding labels will be obtained. The above data are all time domain data. The pre-processed time domain data are recorded as and
[0078] The above obtained and All of them are time domain data, and each of them needs to be fast Fourier transformed separately. The frequency domain vibration signal and frequency domain noise signal are obtained, which are recorded as and
[0079] Finally, after completing all signal acquisition and signal preprocessing, the final result is the time domain vibration data , time domain noise data , frequency domain vibration data and frequency domain noise data .
[0080] The four types of data described above are divided into a 7:2:1 ratio to create a training set, validation set, and test set, respectively. The training set is used to train the model. By learning the relationship between the input and output in the training set, the model's parameters are adjusted, thereby improving the model's classification ability. The validation set is used to adjust the model's hyperparameters and evaluate model performance. During the model training process, it is used to verify the model's generalization ability. The test set is used for the final evaluation of model performance. It is usually used after model development is completed to evaluate the model's generalization ability to unseen data.
[0081] Step (2): input the training set data into the neural network model proposed in the present invention, and use the validation set to verify the training effect;
[0082] Step (2.1) collects and processes the time domain vibration data obtained in the first step above. and frequency domain vibration data The data is input into three consecutive time domain feature extraction blocks and frequency domain feature extraction blocks for preliminary feature extraction. The results are then input into the BiLSTM model to capture the bidirectional dependencies of long time series. The time domain output results and frequency domain output results are fused through the cross-attention mechanism to achieve cross-modal information fusion. By establishing connections between different modes, the complementary information and heterogeneous information between different modes can be better captured, thus improving the performance of the model in complex tasks. The data obtained by the cross-attention mechanism is input into the fully connected layer and the Softmax activation function to obtain the probability corresponding to each fault category, which is recorded as the preliminary vibration diagnosis result. .
[0083] Step (2.2) The time domain noise data obtained after the noise sensor is processed and frequency domain noise data Follow the method in step (2.1) to train and finally get the preliminary diagnosis result of noise .
[0084] The time domain vibration data obtained in step (1) and frequency domain vibration data For example, the time domain data finally obtained in step (1) has four categories, each category has 585 data points, and each data point has 1024 data points. Therefore, there are 2340 data points in the time domain vibration data. Each data point is a one-dimensional vector. The data size is recorded as [2340, 1, 1024], where 2340 is the number of samples, 1 is the number of channels, and 1024 is the data length. Similarly, the frequency domain vibration data size is also [2340, 1, 1024].
[0085] In this embodiment, batch_size is set to 32, that is, 32 pieces of data are input into the neural network each time, and the size of each input is [32, 1, 1024].
[0086] Table 1
[0087] Time domain feature extraction block 1 Input size / Output size Frequency domain feature extraction block 1 Input size / Output size Convolutional layer 1 (Conv) KN=32,KS=3,S=1,P=1 [32,1,1024] / [32,32,1024] KN=16,KS=3,S=1 [32,1,1024] / [32,16,1024] ReLu layer default constant default constant Convolutional layer 2 (Conv) KN=32,KS=3,S=1,P=1 [32,32,1022] / [32,32,1024] KN=16,KS=3,S=1 [32,16,1024] / [32,16,1024] ReLu layer default constant default constant Max Pooling Layer (MAxPool) KS=2,S=2 [32,32,1024] / [32,32,512] KS=2,S=2 [32,16,1024] / [32,16,512] ReLu layer default constant default constant Batch Normalization layer (batchNormal) out_channels=32 constant out_channels=16 constant Time domain feature extraction block 2 Input size / Output size Frequency domain feature extraction block 2 Input size / Output size Convolutional layer 1 (Conv) KN=64,KS=3,S=1 [32,32,512] / [32,64,512] KN=32,KS=3,S=1 [32,16,512] / [32,32,512] ReLu layer default constant default constant Convolutional layer 2 (Conv) KN=64,KS=3,S=1,P=1 [32,64,512] / [32,64,512] KN=32,KS=3,S=1 [32,32,512] / [32,32,512] ReLu layer default constant default constant Max Pooling Layer (MAxPool) KS=2,S=2 [32,64,512] / [32,64,256] KS=2,S=2 [32,32,512] / [32,32,256] ReLu layer default constant default constant Batch Normalization layer (batchNormal) out_channels=64 constant out_channels=32 constant Time domain feature extraction block 3 Input size / Output size Frequency domain feature extraction block 3 Input size / Output size Convolutional layer 1 (Conv) KN=128,KS=3,S=1,P=1 [32,64,256] / [32,128,256] KN=64,KS=3,S=1 [32,32,256] / [32,64,256] ReLu layer default constant default constant Convolutional layer 2 (Conv) KN=128,KS=3,S=1,P=1 [32,128,256] / [32,128,256] KN=64,KS=3,S=1 [32,64,256] / [32,64,256] ReLu layer default constant default constant Max Pooling Layer (MAxPool) KS=2,S=2 [32,128,256] / [32,128,128] KS=2,S=2 [32,64,256] / [32,64,128] ReLu layer default constant default constant Batch Normalization layer (batchNormal) out_channels=128 constant out_channels=64 constant
[0088] Table 1 shows the parameter settings of each function in the time domain feature extraction block and the frequency domain feature extraction block (see Figure 3 ), where KN is the number of convolution kernels, KS is the kernel size, S is the stride, P is the padding, and out_channels is the number of output channels. According to the parameters set in the above table, the final result is a time domain feature with a size of [32, 128, 128] and a frequency domain feature with a size of [32, 64, 128].
[0089] Table 2
[0090] BiLSTM time-domain block Input size / Output size BiLSTM frequency domain block Input size / Output size BiLSTM layer [32,128,128] / [32,128,256] [32,64,128] / [32,128,256] ReLu layer default constant default constant Batch Normalization layer (batchNormal) out_channels=128 constant out_channels=128 constant
[0091] Table 2 shows the parameter settings of the BiLSTM time domain block and the BiLSTM frequency domain block. The time domain features and frequency domain features obtained after the time domain feature extraction block and the frequency domain feature extraction block are input into the BiLSTM time domain block and the BiLSTM frequency domain block respectively. The obtained time domain feature size is [32, 128, 256], and the frequency domain feature size is [32, 128, 256].
[0092] The time domain features and frequency domain features obtained by the BiLSTM time domain block and the BiLSTM frequency domain block are input into the cross-attention mechanism module, where Q is the time domain feature, K and V are the frequency domain features. The above table is the parameter setting of the cross-attention mechanism module, and the fusion feature with a size of [32, 128, 256] is obtained.
[0093] The fused features obtained by the cross-attention mechanism module are input into the global average pooling layer. The output channel of the global average pooling layer is set to 1, that is, global average pooling is performed on each feature channel, and the sequence length is compressed to 1, and the output size is [32, 1, 256]. Then, it is input into the linear layer, and the output dimension of the linear layer is set to 4, that is, the input feature dimension is mapped from 256 to 4, and the output size is [32, 4].
[0094] Finally, the output obtained after the linear layer is input into the Softmax layer, which normalizes the 4 output features of each sample into a probability distribution. The final output size is [32,4], where the 4 output features of each sample represent the probabilities of the 4 categories. The preliminary vibration diagnosis results are obtained according to the above method. .
[0095] According to the same method as above, the input data is noise data, and the preliminary diagnosis result of noise is finally obtained. .
[0096] Step (2.2) combines the two types of data into a single probability vector of the preliminary diagnosis results. The probability matrix , is the number of classifiers, which is 2 in this invention, The number of fault categories outputted in the final output is 4 in this invention. 、 and Matrix, will The probability of each failure in the matrix is summed up to get Vector, calculation The fault category corresponding to the maximum probability in the vector is the fault category predicted by the neural network. If there are two or more equal maximum probabilities in the vector, one of them is randomly selected as the diagnosis result. The role of the matrix is to obtain the classification ability weights of different classifiers and multiply the weights by the preliminary diagnosis results; The role of the matrix is to obtain the weight of the classification ability of the same classifier for each category and multiply the weight by the preliminary diagnosis result; Matrix Pass and The dot product of the matrix obtains the preliminary diagnosis results after weighting according to the capabilities of different classifiers and the weighting of the classification capabilities of the same classifier for different categories; The probabilities of the corresponding categories of the two classifiers are added together to obtain the final diagnosis result.
[0097] Taking a vibration and a noise sample as an example, assuming that the two samples pass through the Softmax layer and the final output probability is the vibration diagnosis result =[0.4,0.1,0.2,0.3], noise diagnosis results =[0.35,0.35,0.1,0.2], that is, the probability of the result obtained from the vibration data corresponding to the normal state is 0.4, the probability of corresponding to the slightly loose state is 0.1, the probability of corresponding to the moderately loose state is 0.2, and the probability of corresponding to the severely loose state is 0.3. The probability of the result obtained from the noise data corresponding to the normal state is 0.35, the probability of corresponding to the slightly loose state is 0.35, the probability of corresponding to the moderately loose state is 0.1, and the probability of corresponding to the severely loose state is 0.2.
[0098] because The probability of normal state is the largest, so the sample is classified as normal state, and The probabilities corresponding to the normal state and the slightly loose state are the same. In actual situations, there is a high possibility that there will be similar probabilities, which will affect the classification results.
[0099] The above and Input into the decision-level fusion algorithm and perform decision-level fusion. The specific steps are as follows:
[0100] Will and Splice into 2 4, that is, [[0.4, 0.1, 0.2, 0.3], [0.35, 0.35, 0.1, 0.2]], the calculated Q matrix is: [[0.53, 0.22, 0.67, 0.6], [0.47, 0.78, 0.33, 0.4]], the K matrix is: [[0.4, 0.1, 0.2, 0.3], [0.35, 0.35, 0.1, 0.2]], and the W matrix is :[[0.21,0.022,0.134,0.18],[0.1645,0.273,0.033,0.08]], the final diagnosis result P is: [0.3745,0.295,0.267,0.26]. It can be seen from the above P that the probability of the normal state is the largest, and the difference between the probability of the normal state and the probability of other states is large, avoiding the problem of similar probabilities of different states. Therefore, the sample is classified as normal, and its predicted label is recorded as 0. The predicted label and the true label are compared. If they are the same, the classification is correct, otherwise, the classification is wrong. The loss function of the sample is then calculated. In the present invention, the loss function uses the cross entropy loss function, and the formula is: ,in, is the true label, To predict the label, is the loss value, is the number of categories. Using the example above, whose true label is 0 (i.e., correctly classified), the loss calculated is 0.980. Following the above method, the losses for each example are summed, the gradients are calculated using backpropagation, and the Adam optimizer is used to update the model parameters. Training then proceeds to the next batch.
[0101] Follow the above steps until all training set samples are trained and save the final training parameters of the model. Completing the above steps once is counted as one round of neural network training. Calculate the classification accuracy of the training set and validation set in one round of training using the following formulas: ,in The number of correct samples predicted in all training set samples, is the number of samples in all training sets. ,in is the number of samples predicted correctly in all validation set samples, is the number of samples in all validation sets.
[0102] Step (3): Repeat the second step above to train the neural network. If the number of training rounds reaches 100, stop training.
[0103] Step (4): Save the network parameters obtained from model training, input the test set into the trained neural network, and calculate the accuracy.
[0104] Table 3
[0105] method Accuracy WDCNN 98.6% MSCNN 98.9% MA1DCNN 96.5% Method of the present invention 99.4%
[0106] Table 3 is a comparison table between the method of the present invention and other existing methods; wherein data acquisition, data preprocessing and decision-level fusion algorithm remain unchanged, only Figure 2 The network in
[15] is replaced with another network, and the other processes remain unchanged. As can be seen from the table, the accuracy of the present invention is significantly better than other existing methods.
[0107] In the present invention, the normalization method based on envelope value, sliding window segmentation method, fast Fourier transform, cross attention mechanism, global average pooling layer, fully connected layer, The activation function, cross-entropy loss function, and Adam optimizer are all existing technologies. In addition, one-dimensional convolutional layers, ReLu layers, one-dimensional max pooling layers, and one-dimensional batch normalization layers are also existing basic methods. Because neural networks often repeat these steps, a group of operations are combined and encapsulated into time-domain feature extraction blocks and frequency-domain feature extraction blocks. This is equivalent to a time-domain feature extraction block consisting of many basic methods. The time-domain feature extraction block and the frequency-domain feature extraction block have the same content, but different input data: one input is time-domain data, and the other input is frequency-domain data.
[0108] The application obtains a frequency domain signal by performing fast Fourier transform on an original time domain signal, inputs the time domain and frequency domain data into a multi-layer CNN convolutional neural network respectively to extract respective features, then inputs into a BiLSTM model respectively to capture a bidirectional dependency relationship of escalator operation data, the obtained result is subjected to feature-level data fusion through a cross-attention mechanism, and finally, a preliminary classification is performed through a Softmax function to obtain a preliminary diagnosis result; the preliminary diagnosis result obtained by a single sensor is input into a decision-level data fusion algorithm to realize fusion judgment of multi-sensor decision, and a final diagnosis result is obtained, thereby realizing fault diagnosis of the escalator transmission system.
[0109] The above embodiments are only preferred embodiments of the present application, and do not limit the protection scope of the present application, so: any equivalent changes made according to the structure, shape, principle of the present application should be covered within the protection scope of the present application.
Claims
1. A fault diagnosis method for an escalator transmission system based on multimodal fusion is carried out in the following steps: Step (1), data collection and data preprocessing; The data preprocessing method is: using the envelope value-based normalization method on the collected vibration data and noise data to achieve noise suppression function and enhance the distinguishability of the signal; The obtained time domain data is enhanced and labeled according to the sliding window segmentation method. Each group of segmented data is fast Fourier transformed to obtain the corresponding frequency domain data. The processed time domain data and frequency domain data are used to create a data set. Step (2): input the training set data into the neural network model for training, and use the validation set to verify the training effect; The training method is: (2.1) The time domain vibration data collected and processed by the vibration sensor and frequency domain vibration data , respectively input into three consecutive time domain feature extraction blocks and frequency domain feature extraction blocks for preliminary feature extraction, and then input the obtained results into the BiLSTM model to capture the bidirectional dependency of long time series. The obtained time domain output results and frequency domain output results are fused across modalities through the cross-attention mechanism, and connections are established between different modalities. Finally, the data obtained through the cross-attention mechanism is input into the global average pooling layer to reduce the channel dimension to 1, and then input into the fully connected layer and the Softmax activation function to obtain the probability corresponding to each fault category, which is recorded as the preliminary vibration diagnosis result. According to the above method, the time domain noise data collected and processed by the noise sensor is and frequency domain noise data Conduct training and finally obtain preliminary noise diagnosis results ; (2.2) Combine the two types of data into a single probability vector of the preliminary diagnosis results. The probability matrix ; is the number of classifiers, is the number of fault categories output in the end; calculated according to the formula 、 and Matrix, will The probability of each failure in the matrix is summed up to get Vector, calculation The fault category corresponding to the maximum probability in the vector is the fault category predicted by the neural network; like If there are two or more equal maximum probabilities in the vector, one of them is randomly selected as the diagnosis result; The role of the matrix is to obtain the classification ability weights of different classifiers and multiply the weights by the preliminary diagnosis results; The role of the matrix is to obtain the weight of the classification ability of the same classifier for each category and multiply the weight by the preliminary diagnosis result; Matrix Pass and The dot product of the matrix obtains the preliminary diagnosis results after weighting according to the capabilities of different classifiers and the weighting of the classification capabilities of the same classifier for different categories; The probabilities of the corresponding categories of the two classifiers are added together to obtain the final diagnosis result; step (3), repeat the above step (2) to train the neural network. If the number of training rounds has reached 100 rounds, stop training; step (4), save the network parameters obtained from model training, input the test set into the trained neural network, and calculate the accuracy.
2. The escalator transmission system fault diagnosis method based on multimodal fusion according to claim 1 is characterized in that: The data collection method in step (1) is as follows: according to the looseness degree of the anchor bolts of the driving main unit, the looseness fault is defined as slight looseness, moderate looseness, and severe looseness, which correspond to one circle, two circles, and three circles of looseness of the anchor bolts, respectively; and vibration data and noise data of the anchor bolts of the escalator driving main unit in normal state, slight loose state, moderate loose state, and severe loose state are collected through vibration sensors and noise sensors.
3. The escalator transmission system fault diagnosis method based on multimodal fusion according to claim 2 is characterized in that: The dataset prepared in step (1) is divided into a training set, a validation set, and a test set in a ratio of 7:2:
1.
4. The escalator transmission system fault diagnosis method based on multimodal fusion according to claim 3 is characterized in that: The method for verifying the training effect using the validation set in step (2) is: training the model until one round of model training is completed, and saving the parameters of the final training of the model; Calculate the classification accuracy of the training set and validation set in one round of training.
Citation Information
Patent Citations
Escalator footing loosening fault diagnosis method
CN113553898A
Bolt looseness detection method and device and medium
CN118277856A
Bearing fault diagnosis method and device, computer equipment and storage medium
CN120180090A