Mechanical Fault Diagnosis Method Based on Multi-Scale Component Analysis

Through the multi-scale component analysis model, multi-scale feature extraction and sparse constraints are performed on mechanical signals, combined with convolutional dictionary and pooling layers, the problem of extracting fault features of complex mechanical equipment and weakening noise interference in the prior art is solved, and mechanical fault diagnosis with high accuracy and reliability is achieved.

CN115391955BActive Publication Date: 2025-06-20XI AN JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211112504.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-13
Publication Date
2025-06-20
Estimated Expiration
2042-09-13

AI Technical Summary

Technical Problem

When existing mechanical fault diagnosis algorithms deal with large mechanical equipment with complex structures, it is difficult to effectively extract multi-scale fault characteristics and weaken noise interference, resulting in insufficient diagnostic accuracy and reliability.

Method used

Using a multi-scale component analysis method, multi-scale feature decoupling and extraction of mechanical signals through multi-scale component analysis model, sparse constraints are applied to weaken the noise influence, and the convolutional dictionary is used to reduce the amount of parameters, combining pooling layers and multi-layer perceptrons for classification prediction.

Benefits of technology

It realizes robust identification and classification of multiple fault modes, improves the accuracy and reliability of fault diagnosis, and provides transparent and trustworthy diagnostic results through the interpretability of the expanded network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115391955B_ABST
    Figure CN115391955B_ABST
Patent Text Reader

Abstract

The present invention discloses a mechanical fault diagnosis method based on multi-scale component analysis. The method includes the following steps: collecting mechanical vibration signals y of multiple fault types; establishing a multi-scale component analysis model to decouple and extract multi-scale features of the signals, and obtaining deep sparse codes γ of each scale. m,L ; using the iterative threshold shrinkage algorithm to solve the established multi-scale component analysis model, and expanding the optimization algorithm into a multi-scale component analysis network, concatenating the multi-scale codes as #imgabs0# and inputting it into the subsequent pooling layer and multi-layer perceptron as a classifier h. θ to complete the construction of the fault intelligent diagnosis network; using the training samples with fault labels to train the network model end-to-end, and using the backpropagation technique to learn the network model parameters; inputting the test signal into the network, and realizing fault diagnosis by outputting the predicted label #imgabs1#; visualizing the overall features of the reconstructed signal #imgabs2# of the input signal and the atomic features learned by the network to complete the post-event interpretability analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of mechanical fault diagnosis methods, and in particular to a mechanical fault diagnosis method based on multi-scale component analysis. Background Art

[0002] Since the faults of mechanical equipment can cause serious damage to the system and result in huge unplanned maintenance costs, it is crucial to conduct fault prediction and health management (PHM) for mechanical equipment. Large-scale mechanical equipment usually has the characteristics of complex structure, long signal transmission path, and serious environmental noise interference, which makes high requirements for the algorithms of its fault diagnosis. Traditional fault diagnosis algorithms such as time-frequency analysis and envelope spectrum analysis are somewhat inadequate and rely heavily on expert experience. With the development of the times and the breakthrough of information technology, intelligent algorithms such as deep neural networks (DNN) have brought major changes to image and signal processing technologies and provided great performance gains in PHM. However, the black-box characteristic of the neural network, that is, only the input and output can be seen, results in the lack of interpretability of the results and prevents it from being applied to situations that require high credibility, such as the fault diagnosis of important mechanical equipment with extremely high requirements for safety. Therefore, the research on interpretable artificial intelligence is crucial for providing more credible fault diagnosis results.

[0003] Generally speaking, the interpretability of machine learning models can be divided into two categories. One is pre-explainable modeling, and the other is post-explainable analysis. The former refers to using inherently explainable technologies to model specific problems, thus obtaining a model with an interpretable structure; the latter refers to developing specialized technologies to explain the already trained model, usually involving many visualization methods. Researchers have conducted a large number of studies on these two types of interpretability methods and their applications in interpretable fault diagnosis, such as wavelet kernel networks, attention mechanism-assisted localization, etc. They can all endow the neural network with interpretability from one aspect. However, most of the studies only focus on pre-explainable modeling or post-explainable analysis, lacking a more comprehensive exploration of interpretable models. The unfolded network formed by unfolding the optimization solution algorithm of the model along the iteration dimension corresponds to a specific mathematical and physical model and has inherent pre-explainability. Moreover, the self-reconfigurability of the model allows for the analysis of the features learned by the network, that is, post-explainable analysis. The unfolded network theory provides a new paradigm for establishing a more comprehensive interpretable network model. However, the unfolded network theory is still in its infancy. For mechanical fault diagnosis, the existing research does not fully consider the problems of serious noise interference and the multi-scale characteristics of fault features during modeling, resulting in the difficulty for the network to extract useful features for further fault diagnosis. Therefore, a method that focuses on mechanical fault diagnosis and is transparent and credible from multi-scale model construction to result analysis is needed to open up a new path for the application of high-performance intelligent diagnosis algorithms in practical engineering.

[0004] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present invention, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] Aiming at the problems existing in the prior art, the present invention proposes a mechanical fault diagnosis method based on multi-scale component analysis. This method uses a multi-scale component analysis model to decouple and extract multi-scale features of mechanical signals, applies sparse constraints during the hierarchical encoding process of multi-scale features to extract significant features and weaken the influence of environmental noise, and uses a convolutional dictionary to greatly reduce the number of parameters learned by the network. Then, the solution algorithm of the model is used as part of the forward process of the multi-scale component analysis network, and an average pooling layer, a max pooling layer, and a multi-layer perceptron are added at the back as a classifier to finally realize the label prediction of the input signal. Since the main body of the network is expanded by the optimization solution algorithm of a specific model, its structure is interpretable, that is, it has pre-explainability. And the visualization in the result analysis is based on the reconstructability of the morphological component analysis model, and the result is interpretable by reconstructing the input signal and viewing the atomic features learned by the network, that is, it has post-explainability.

[0006] The object of the present invention is achieved by the following technical solutions. A mechanical fault diagnosis method based on multi-scale component analysis, the method includes the following steps:

[0007] In the first step (S1), collect the mechanical vibration signals y of the mechanical equipment under various fault types and normal states, and divide them into non-overlapping training data sets Y and test data sets Y according to a fixed signal length, where the fault types include cracks and wear, and both the training data set Y and the test data set Y are labeled with fault labels; train and test data set Y test , where the fault types include cracks and wear, training data set Y train and test data set Y test both carry fault labels;

[0008] In the second step (S2), establish a multi-scale component analysis model, and decouple and extract the multi-scale features of the mechanical vibration signal y to obtain the deep sparse coding γ of each scale; m,L ;

[0009] In the third step (S3), based on the deep sparse coding γ, use the iterative threshold shrinkage algorithm to solve the multi-scale component analysis model, expand the optimization solution algorithm into a multi-scale component analysis network, splice the multi-scale coding, connect a pooling layer and a multi-layer perceptron after the coding network as a classifier, build a fault diagnosis network, and output the predicted label of the signal; m,L Connect a pooling layer and a multi-layer perceptron after the coding network as a classifier, build a fault diagnosis network, and output the predicted label of the signal; Connect a pooling layer and a multi-layer perceptron after the coding network as a classifier, build a fault diagnosis network, and output the predicted label of the signal;

[0010] In the fourth step (S4), by setting a fixed number of loop iterations, the training data set Y train is input into the fault diagnosis network for training, and the backpropagation technique is used to learn the parameters of the fault diagnosis network, reducing the gap between the predicted label and the fault label;

[0011] In the fifth step (S5), the test data set Y test is input into the trained fault diagnosis network, and fault diagnosis is performed by outputting the predicted label.

[0012] In the method described above, in the first step (S1), the mechanical vibration signal is decomposed into y = x + ε, where is the feature signal, is the noise interference, N is the signal length, the mechanical vibration signal is collected by an acceleration sensor, the training data set Y train includes time-domain training samples, and the test data set Y test includes time-domain test samples.

[0013] In the method described above, in the second step (S2), the multi-scale component analysis model is:

[0014]

[0015] where represents the deepest layer sparse coding of each scale, L is the number of coding layers, M is the dimension of the analyzed scale; y represents the input noisy signal; D m,l is the convolutional coding dictionary of the l-th layer of the m-th scale, D m,(l,L) = D m,l D m,l+ 1...D m,L represents the equivalent convolutional dictionary obtained by multiplying the convolutional coding dictionaries from the l-th layer to the deepest L-th layer of the m-th scale; γ m,l = D m,(l+1,L) γ m,L represents the sparse coding of the l-th layer of the m-th scale; the first term is the data fidelity term, ensuring that the input signal can be reconstructed from the coding of each scale; the remaining two terms are sparse regularization terms, making the coding of each level of each scale sparse; is the weight parameter for weighing the contributions of different terms. For simplicity, here let represents the square of the L2 norm; ||·||1 represents the L1 norm. In the multi-scale component analysis model, the first term is the data fidelity term, ensuring that the input signal can be reconstructed from the deep sparse coding of each scale through the coding dictionaries of each layer

[0016] features of y; the remaining terms are sparse regularization terms to ensure that the coding of each layer of each scale is sparse.

[0017] In the method described above, for each scale, the encoding dictionary of each layer satisfies the following relationship with the encoding and the encoding of the previous layer:

[0018]

[0019] Wherein, represents the sparse encoding of each layer of the m-th scale, L is the number of encoding layers; x m = γ m,0 represents the ideal feature signal component of the m-th scale, representing the sparse encoding of the 0-th layer; is the convolutional encoding dictionary of each layer of the m-th scale; is the sparsity constraint constant of each layer of the m-th scale.

[0020] In the method described above, the encoding dictionary is a convolutional dictionary, wherein, is a circulant matrix formed by shifting the j-th convolutional kernel of length n m,l along the column direction, j = 1, 2,..., m i , where the shifting stride is 1, and c l is the number of convolutional kernels.

[0021] In the method described above, in the third step (S3), the iterative threshold shrinkage solution algorithm of the multi-scale component analysis model is as follows:

[0022] Initialize the sparse encoding of each layer of each scale For

[0023] For each iteration k = 0, 1,..., K, for each scale m = 0, 1,..., M, calculate the outermost input of this scale:

[0024]

[0025] Then, solve the encoding values of each layer i = 0, 1,..., L in turn, and the solution formula is as follows:

[0026]

[0027] Wherein, I is the identity matrix with the same dimension as ; Soft T (·) is the soft threshold function, and its expression is Soft T (·) = max(|·| - T, 0) * sign(·), sign(·) is the sign function, T is the threshold, and in the l-th layer of the m-th scale, T = δ m,l = μm,l λ m,l ,μ m,l are the step sizes for each iteration of the proximal gradient mapping algorithm, and λ m,l is the trade-off parameter between the data fidelity term and the sparse regularization term.

[0028] Output the deepest layer sparse representation coefficients at each scale

[0029] In the method described above, in the multi-scale component analysis model, the encoding dictionary D m,l is multiplied by a vector and implemented using a transposed convolutional neural network layer, is multiplied by a vector and implemented using a convolutional neural network layer. The soft threshold function is equivalent to the activation function between layers.

[0030] In the method described above, in the third step (S3), after multi-scale component analysis sparse coding and before connecting the multi-layer perceptron, a global max pooling layer and a global average pooling layer are connected in parallel to extract the significant information of the feature map and the smooth information of the feature map respectively. Among them, both the max pooling layer and the average pooling layer reduce the original sparse coding to 4 dimensions. The multi-layer perceptron consists of a three-layer fully connected neural network, with the ReLU function as the activation function, encoding from the feature dimension to the dimension of the fault category.

[0031] In the method described above, in the fourth step (S4), at the end of the multi-layer perceptron, the Softmax activation function is used to output the probability that the signal belongs to each type of fault, calculate the cross-entropy loss, and update the network parameters using error backpropagation.

[0032] In the method described above, it also includes,

[0033] In the sixth step (S6), visualize the deep sparse coding at each scale output after the input signal is encoded by the multi-scale component analysis network and reconstruct the overall features of the input signal based on this

[0034]

[0035] and view the equivalent convolutional kernels in D m,(1,L) i.e., the atomic features, and analyze whether they are consistent with the signal fault features to complete the post hoc interpretability analysis of the network.

[0036] Compared with the prior art, the present invention has the following advantages:

[0037] By constructing a multi-scale component analysis model, using equivalent convolution dictionaries with different convolution kernel lengths as multi-scale encoding dictionaries therein, expanding them into a network, and using the backpropagation technology of the network to update the model parameters, the present invention can accurately decouple and extract multi-scale fault features, thereby robustly and effectively realizing the identification and classification of various fault modes, improving the accuracy and reliability of fault diagnosis. It has the characteristics of interpretability in both structure and results, which is conducive to the safe, reliable, and transparent fault prediction and health management of mechanical equipment. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] By reading the detailed descriptions in the preferred specific embodiments below, various other advantages and benefits of the present invention will become clear to those of ordinary skill in the art. The accompanying drawings in the specification are only for the purpose of showing the preferred embodiments and are not considered as limiting the present invention. Obviously, the drawings described below are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Moreover, throughout the drawings, the same reference numerals are used to represent the same components.

[0039] In the drawings:

[0040] Figure 1 is a step schematic diagram of a mechanical fault diagnosis method based on multi-scale component analysis according to an embodiment of the present invention;

[0041] Figure 2 is a schematic diagram of the Qianpeng gear test bench according to an embodiment of the present invention;

[0042] Figures 3(a) to 3(h) is a time-domain waveform diagram of the vibration signal of a helical gear according to an embodiment of the present invention. Figure 3(a) shows the signal in the normal state, Figure 3(b) shows the signal of a single tooth crack with a depth of 1 mm, Figure 3(c) shows the signal of a single tooth crack with a depth of 2 mm, Figure 3(d) shows the signal of a single tooth crack with a depth of 3 mm, Figure 3(e) shows the signal of a single tooth crack with a depth of 4 mm, Figure 3(f) shows the signal of a double tooth crack with a depth of 2 mm, Figure 3(g) shows the signal of uniform wear of a single tooth with a depth of 0.5 mm, and Figure 3(h) shows the signal of uniform wear of double teeth with a depth of 0.5 mm;

[0043] Figure 4 is a structural diagram of a mechanical fault interpretable intelligent diagnosis network expanded based on the multi-scale component analysis algorithm according to an embodiment of the present invention;

[0044] Figure 5 is a comparison of the performance (including accuracy and test time) of the mechanical fault diagnosis method based on multi-scale component analysis according to an embodiment of the present invention with other methods;

[0045] Figures 6(a) to 6(b)Performance comparison of the mechanical fault diagnosis method based on multi-scale component analysis in an embodiment of the present invention and other methods against noise attacks, Figure 6(a) Gaussian noise attack, Figure 6(b) Laplace noise attack;

[0046] Figure 7 Comparison of the test accuracy of the mechanical fault diagnosis method based on multi-scale component analysis in an embodiment of the present invention and other methods as the usage rate of the training data set changes;

[0047] Figure 8 (a) to Figure 8 (f) Comparison of the overall feature reconstruction ability of the mechanical fault diagnosis method based on multi-scale component analysis in an embodiment of the present invention and the ML-ISTA method, Figure 8 (a) Original signal of the helical gear, Figure 8 (b) Reconstructed signal by ML-ISTA, Figure 8 (c) Overall reconstructed signal by MCAN, Figure 8 (d) Small-scale reconstructed signal by MCAN, Figure 8 (e) Medium-scale reconstructed signal by MCAN, Figure 8 (f) Large-scale reconstructed signal by MCAN;

[0048] Figure 9 Correspondence between the atomic features learned by the mechanical fault diagnosis method based on multi-scale component analysis in an embodiment of the present invention and the fault mechanism.

[0049] The present invention will be further explained below with reference to the accompanying drawings and embodiments. Specific embodiments

[0050] The following will refer to the attached Figures 1 to 9 The specific embodiments of the present invention will be described in more detail. Although specific embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present invention can be completely conveyed to those skilled in the art.

[0051] It should be noted that certain terms are used in the specification and claims to refer to specific components. Those skilled in the art should understand that technicians may use different terms to refer to the same component. The specification and claims do not use the difference in terms as a way to distinguish components, but use the difference in the functions of components as the criterion for distinction. For example, the terms "comprising" or "including" mentioned throughout the specification and claims are open-ended terms and should be interpreted as "including but not limited to". The subsequent description in the specification is the preferred embodiment for implementing the present invention, but the description is for the purpose of the general principles of the specification and is not used to limit the scope of the present invention. The protection scope of the present invention shall be determined by the scope defined by the appended claims.

[0052] For the convenience of understanding the embodiments of the present invention, the following will further explain with specific examples in conjunction with the drawings, and each drawing does not constitute a limitation to the embodiments of the present invention.

[0053] For better understanding, Figure 1 is a schematic diagram of the steps of a mechanical fault diagnosis method based on multi-scale component analysis according to an embodiment of the present invention. As Figure 1 shown, a mechanical fault diagnosis method based on multi-scale component analysis includes the following steps:

[0054] In the first step (S1), collect the mechanical vibration signal y of the mechanical equipment under various fault types and normal conditions, and divide it into non-overlapping training data sets Y train and test data sets Y test in a fixed signal length according to a predetermined ratio, where the fault types include cracks and wear, and the training data set Y train and the test data set Y test both carry fault labels;

[0055] In the second step (S2), establish a multi-scale component analysis model, decouple and extract the multi-scale features of the mechanical vibration signal y to obtain the deep sparse coding γ m,L of each scale;

[0056] In the third step (S3), based on the deep sparse coding γ m,L use the iterative threshold shrinkage algorithm to solve the multi-scale component analysis model, expand the optimization algorithm into a multi-scale component analysis network, splice the multi-scale coding as connect a pooling layer and a multi-layer perceptron after the coding network as a classifier, build a fault diagnosis network, and output the predicted label of the signal;

[0057] In the fourth step (S4), by setting a fixed number of loops, use the training data set Y trainInput the fault diagnosis network for training, and use the backpropagation technique to learn the parameters of the fault diagnosis network, reducing the gap between the predicted label and the fault label;

[0058] In the fifth step (S5), input the test data set Y test into the trained fault diagnosis network, and perform fault diagnosis by outputting the predicted label.

[0059] In a preferred embodiment of the method, in the first step (S1), the mechanical vibration signal is decomposed into y = x + ε, where is the feature signal, is the noise interference, N is the signal length, the mechanical vibration signal is collected by an acceleration sensor, the training data set Y train includes time-domain training samples, and the test data set Y test includes time-domain test samples.

[0060] In a preferred embodiment of the method, in the second step (S2), the multi-scale component analysis model is:

[0061]

[0062] where represents the deepest layer sparse coding of each scale, L is the number of coding layers, M is the dimension of the analyzed scale; y represents the input noisy signal; D m,l is the convolutional coding dictionary of the l-th layer of the m-th scale, D m,(l,L) = D m,l D m,l+ 1...D m,L represents the equivalent convolutional dictionary obtained by multiplying the convolutional coding dictionaries from the l-th layer to the deepest L-th layer of the m-th scale; γ m,l = D m,(l+1,L) γ m,L represents the sparse coding of the l-th layer of the m-th scale; the first term is the data fidelity term, ensuring that the input signal can be reconstructed from the coding of each scale; the remaining two terms are sparse regularization terms, making the coding of each level of each scale sparse; is the weight parameter for weighing the contributions of different terms. For simplicity, here let represent the square of the L2 norm; ||·||1 represents the L1 norm. In the multi-scale component analysis model, the first term is the data fidelity term, ensuring that the features of the input signal y are reconstructed from the deep sparse coding of each scale through the coding dictionaries of each layer; the remaining terms are sparse regularization terms to ensure that the coding of each layer of each scale is sparse.

[0063] In a preferred embodiment of the method described above, for each scale, the following relationship is satisfied among the coding dictionaries of each layer, the coding, and the coding of the previous layer:

[0064]

[0065] Wherein, represents the sparse coding of each layer of the m-th scale, L is the number of coding layers; x m = γ m,0 represents the ideal characteristic signal component of the m-th scale, representing the sparse coding of the 0-th layer; is the convolutional coding dictionary of each layer of the m-th scale; is the sparsity constraint constant of each layer of the m-th scale.

[0066] In a preferred embodiment of the method described above, the coding dictionary is a convolutional dictionary, wherein, is a circulant matrix formed by shifting the j-th convolutional kernel of length n m,l along the column direction, j = 1, 2,..., m i , where the shift stride is 1, and c l is the number of convolutional kernels.

[0067] In a preferred embodiment of the method described above, in the third step (S3), the iterative threshold shrinkage solution algorithm of the multi-scale component analysis model is as follows:

[0068] Initialize the sparse coding of each layer of each scale For

[0069] For each iteration k = 0, 1,..., K, for each scale m = 0, 1,..., M, calculate the outermost input of this scale:

[0070]

[0071] After that, solve the coding values of each layer i = 0, 1,..., L in turn, and the solution formula is as follows:

[0072]

[0073] Wherein, I is the identity matrix of the same dimension as ; Soft T (·) is the soft threshold function, and its expression is Soft T (·) = max(|·| - T, 0) * sign(·), sign(·) is the sign function, T is the threshold, and in the l-th layer of the m-th scale, T = δ m,l = μm,l λ m,l ,μ m,l is the step size for each iteration of the proximal gradient mapping algorithm, and λ m,l is the trade-off parameter between the data fidelity term and the sparse regularization term.

[0074] Output the deepest layer sparse representation coefficients at each scale

[0075] In a preferred embodiment of the method described, in the multi-scale component analysis model, the encoding dictionary D m,l is implemented by a transposed convolutional neural network layer when multiplying with a vector, is implemented by a convolutional neural network layer when multiplying with a vector, and the soft threshold function is equivalent to the activation function between layers.

[0076] In a preferred embodiment of the method described, in the third step (S3), after multi-scale component analysis sparse coding and before connecting the multi-layer perceptron, a global max pooling layer and a global average pooling layer are connected in parallel to extract the significant information of the feature map and the smooth information of the feature map respectively. Among them, both the max pooling layer and the average pooling layer reduce the original sparse coding to 4 dimensions. The multi-layer perceptron consists of a three-layer fully connected neural network, with the ReLU function as the activation function, encoding from the feature dimension to the fault category dimension.

[0077] In a preferred embodiment of the method described, in the fourth step (S4), at the end of the multi-layer perceptron, the Softmax activation function is used to output the probability that the signal belongs to each type of fault, calculate the cross-entropy loss, and update the network parameters using error backpropagation.

[0078] In a preferred embodiment of the method described, it further includes,

[0079] In the sixth step (S6), visualize the deep sparse coding at each scale output after the input signal is encoded by the multi-scale component analysis network and reconstruct the overall features of the input signal based on this

[0080]

[0081] and view the equivalent convolutional kernels in D m,(1,L) i.e., the atomic features, and analyze whether they are consistent with the signal fault features to complete the post hoc interpretability analysis of the network.

[0082] To further understand the present invention, in one embodiment, Figure 1 is a schematic diagram of the steps of a mechanical fault diagnosis method based on multi-scale component analysis, including the following steps:

[0083] S1: Collect mechanical vibration signals y of multiple fault types (including normal states), and divide them into non-overlapping training data sets Y and test data sets Y according to a fixed signal length in a predetermined ratio. train and test data sets Y test ;

[0084] S2: Establish a multi-scale component analysis model to decouple and extract the multi-scale features of the mechanical vibration signal y, and obtain the deep sparse coding γ of each scale. m,L ;

[0085] S3: Use the iterative threshold shrinkage algorithm to solve the established multi-scale component analysis model, expand the optimization algorithm into a multi-scale component analysis network, splice the multi-scale coding, connect a pooling layer and a multi-layer perceptron after the coding network as a classifier, and complete the construction of the fault diagnosis network. In the coding network, connect a pooling layer and a multi-layer perceptron as a classifier to complete the construction of the fault diagnosis network.

[0086] S4: Use the training data set Y with fault labels to train the network model end-to-end, and use the backpropagation technique to learn the network model parameters. train Use the training data set Y with fault labels to train the network model end-to-end, and use the backpropagation technique to learn the network model parameters.

[0087] S5: Input the test data into the trained network, and implement fault diagnosis by outputting the predicted labels. Implement fault diagnosis.

[0088] S6: Visualize the overall features of the reconstructed signal after the input signal is encoded by the multi-scale component analysis network and the atomic features learned by the network, and perform post hoc interpretability analysis. Visualize the overall features of the reconstructed signal after the input signal is encoded by the multi-scale component analysis network and the atomic features learned by the network, and perform post hoc interpretability analysis.

[0089] The above embodiments constitute the complete technical solution of the present invention. Different from the prior art, the structure of the intelligent fault diagnosis network built in the above embodiments is expanded from the solution algorithm of the multi-scale component analysis model that combines morphological component analysis and multi-scale thinking. Through multi-layer multi-scale coding, the fault multi-scale features of complex signals are decoupled and extracted. The model can also ensure that the information learned by the model is effective fault features by reconstructing the signal and viewing the atomic features learned by the network, thus giving the model interpretability at both the structural and result levels and improving the accuracy and reliability of mechanical fault diagnosis.

[0090] Figure 2It is a schematic diagram of the Qianpeng gear test bench. The test system mainly consists of a variable-speed drive motor, a parallel-axis gearbox, shafts, a speed governor, etc., and can quickly simulate various faults through the organic combination of components. The test system includes an eddy current acceleration sensor, a laser speed sensor (key-phase signal), a Premax (Yiheng) data acquisition system, and a ThinkPad (Lenovo) notebook. Using it to conduct simulation experiments on helical gear wear and crack faults, where the parameters of the helical gear are: the number of teeth of the input shaft gear Z1 = 53, the number of teeth of the output shaft gear Z2 = 75, the module m = 2, the tooth width b = 20 mm, and the helix angle β is 10.0633°. A total of 1 normal state and 7 fault states (single-tooth cracks with depths of 1, 2, 3, 4 mm, multi-tooth cracks with a depth of 2 mm, single-tooth uniform wear with a depth of 0.5 mm, multi-tooth uniform wear with a depth of 0.5 mm) of data are collected.

[0091] In this embodiment, in step S1, the vibration signal is collected by an eddy current acceleration sensor. The rotational speed of the input shaft is 1300 r / min, the sampling frequency is 10240 Hz, and the sampling duration is 1.28 min. After downsampling the original signal by a factor of 2, a window with a length of 1024 points is used to intercept the signal with a sliding window ratio of 0.7 to obtain data samples of each fault category. The ratio of training samples to test samples is set to 2:1, and finally 4000 samples of 8 categories are obtained for training and 2000 samples for testing. Figures 3(a) to 3(h) The time-domain waveform diagram of the vibration signal in each state of the helical gear is shown.

[0092] In this embodiment, in step S2, the number of layers of sparse coding is set to 3, three-layer coding is performed, and the number of decomposed components is also 3 to extract small, medium, and large-scale features. The corresponding optimization problem is:

[0093]

[0094] Solving the above problem realizes the multi-scale feature decoupling coding of the original input signal and extracts deep multi-scale features that do not mix with each other.

[0095] In this embodiment, in step S3, by expanding the multi-scale component analysis model, the obtained multi-scale component analysis network can learn numerous parameters in the model end-to-end. The number of expansion layers is set to 4 layers, and the internal parameter settings of the model are shown in Table 1. The multi-scale component analysis coding part, as the feature extraction part, also needs to be connected to a max-pooling layer and an average-pooling layer to extract significant features and smooth features respectively, and finally connected to a multi-layer perceptron to complete the downstream task of classification. The final network structure is as Figure 4 shown.

[0096] Table 1 Hyperparameters of the multi-scale component analysis network

[0097]

[0098] In this embodiment, in step S4, the multi-scale component analysis network is implemented using PyTorch. The labeled training samples are input into the network, and the predicted labels are output to achieve end-to-end training. During the training process, the parameters for network training are set as follows: iterate 100 times, with a batch size of 64 samples used in each iteration process, the network learning rate is set to 0.001, the optimizer is Adam, and the learning rate is reduced to 0.3 times the original every 30 iterations for better convergence.

[0099] In this embodiment, in step S5, the test samples are input into the multi-scale component analysis network. Through multi-layer multi-scale component analysis encoding and decoupling, the multi-scale feature information of the signal is extracted, and finally the predicted labels are output by the multi-layer perceptron. Comparative experiments are conducted with other unfolding network methods (ML-ISTA) and popular neural network methods (CNN, AlexNet, ResNet18) to further illustrate the technical solution of the present invention.

[0100] Specifically, this comparative experiment includes the comparison of method performance, the resistance to noise interference, and the degree of dependence on data. When comparing method performance, the evaluation index used is the classification accuracy of the test samples, which is defined as: accuracy = the number of samples correctly predicted by the model / the total number of samples. The test accuracies of various methods are compared as Figure 5 shown. It can be seen that although the proposed method (MCAN) increases the test time to a certain extent, it can achieve the highest classification accuracy, superior to other unfolding network methods and popular deep neural network methods. To compare the resistance of the methods to noise interference, different intensities of noise can be added to the input signal, and the decrease in the classification accuracy of each method can be tested. From Figures 6(a) to 6(b) it can be seen that whether it is Gaussian noise attack or Laplace noise attack, as the noise intensity continuously increases, the classification accuracy of the proposed method decreases slowly, superior to other unfolding network methods, while the classification accuracy of the deep neural network method decreases rapidly, indicating that the proposed method has better feature extraction ability for the vibration dataset. To compare the degree of dependence of the methods on data, the relationship between the classification accuracy of each method and the size of the training dataset can be tested by reducing the size of the training dataset. From Figure 7 it can be seen that as the number of samples in the training dataset decreases, the accuracy of each method decreases to a certain extent, but the proposed method has the highest accuracy maintainability.

[0101] In this embodiment, in step S6, according to the reconstruction characteristics of sparse representation, the formula for reconstructing the input signal is:

[0102]

[0103] FromFigure 8 (a) to Figure 8 (f), it can be seen that ML-ISTA has a certain ability to reconstruct the original input signal, but the amplitude of the reconstructed signal is reduced to a certain extent, and some features are severely weakened. The reconstructed signal of MCAN, however, enhances the features of the signal. Further, according to the reconstructed signals at each scale, it can be seen that the impact features are mainly learned by the small-scale dictionary, the 2-times power frequency features in the normal state are mainly learned by the medium-scale dictionary, and the 6-times power frequency features in the worn state are mainly learned by the large-scale dictionary. This shows that the method can decouple and extract effective multi-scale features for fault diagnosis under noise interference. From Figure 9 the atomic features of the dictionaries at each scale can assist in explaining this point.

[0104] Although the embodiments of the present invention have been described above in conjunction with the accompanying drawings, the present invention is not limited to the above specific embodiments and application fields. The above specific embodiments are merely illustrative and guiding, rather than restrictive. Those of ordinary skill in the art can also make many forms under the inspiration of this specification and without departing from the scope protected by the claims of the present invention, and these all fall within the scope of protection of the present invention.

Claims

1. A mechanical fault diagnosis method based on multi-scale component analysis, the method comprising the following steps: In the first step, mechanical vibration signals of the mechanical equipment under various fault types and normal conditions are collected , and they are divided into non-overlapping training data sets and test data sets with a fixed signal length according to a predetermined ratio , where the fault types include cracks and wear, and both the training data set and the test data set are labeled with faults; ​ In the second step, a multi-scale component analysis model is established to decouple and extract the multi-scale features of the mechanical vibration signal to obtain the deep sparse coding of each scale ; In the third step, based on deep sparse coding Use the iterative thresholding shrinkage algorithm to solve the multi-scale component analysis model, and expand the iterative thresholding shrinkage algorithm into a multi-scale component analysis network, and splice the multi-scale coding into , connect a pooling layer and a multi-layer perceptron after the coding network as a classifier, build a fault diagnosis network, and output the predicted label of the signal; In the fourth step, by setting a fixed number of loop iterations, the training data set is input into the fault diagnosis network for training, and the backpropagation technique is used to learn the parameters of the fault diagnosis network, reducing the gap between the predicted label and the fault label; In the fifth step, the test data set is input into the trained fault diagnosis network, and fault diagnosis is performed by outputting predicted labels; Among them, in the second step, the multi-scale component analysis model is as follows: ; in, represents the deepest sparse coding at each scale, is the number of coding layers, is the dimensionality of the scale of analysis; For the The scale The convolutional encoding dictionary of the layer, Indicates Scale from Layer to the deepest The equivalent convolution dictionary obtained by multiplying the convolutional coding dictionary of the layer; Indicates The scale The first term is the data fidelity term, which ensures that the input signal can be reconstructed from the encoding at each scale; the other two terms are sparse regularization terms, which make the encoding at each level of each scale sparse; For the sake of simplicity, we assume that , represents the square of L2 norm; represents the L1 norm.

2. The method according to claim 1, wherein, In the first step, the mechanical vibration signal is decomposed into , where is the characteristic signal is the noise interference is the signal length, and the mechanical vibration signal is collected by an acceleration sensor. The training dataset includes time-domain training samples, and the test dataset includes time-domain test samples.

3. The method according to claim 1, wherein, In each scale, the following relationship is satisfied among the encoding dictionaries, encodings of each layer and the encoding of the previous layer: , Among them, represents the sparse coding of each layer at the scale, where is the number of coding layers; represents the ideal characteristic signal component at the scale, representing the sparse coding of the 0th layer; is the convolutional coding dictionary of each layer at the scale; is the sparsity constraint constant of each layer at the 4. The method according to claim 1, wherein, Coding dictionary is a convolution dictionary, , where, is a cyclic matrix formed by translating the -th convolution kernel with length along the column direction, where the translation stride is 1, and is the number of convolution kernels.

5. The method according to claim 3, wherein, In the third step, the iterative threshold shrinkage solution algorithm of the multi-scale component analysis model is as follows: Initialize the sparse coding for each layer at each scale ; For each iteration , for each scale , calculate the outermost input for that scale: , Then, the encoded values of each layer are solved in turn. The solution formula is as follows: , Among them, , is the identity matrix of the same dimension as ; is the soft threshold function, and its expression is , is the sign function, is the threshold. At the scale and the layer, , is the step size of each iteration of the proximal gradient mapping algorithm, is the trade-off parameter between the data fidelity term and the sparse regularization term; Output the deepest layer sparse representation coefficients at each scale 。 6. The method according to claim 5, wherein, In the multi-scale component analysis model, the coding dictionary The multiplication with the vector is implemented by a transposed convolutional neural network layer, The multiplication with the vector is implemented by a convolutional neural network layer, and the soft threshold function is equivalent to the activation function between layers.

7. The method according to claim 1, wherein, In the third step, before connecting the multi-layer perceptron after multi-scale component analysis sparse coding, a global max pooling layer and a global average pooling layer are connected in parallel to extract the significant information and smooth information of the feature map respectively. Among them, both the max pooling layer and the average pooling layer reduce the original sparse coding to 4 dimensions. The multi-layer perceptron consists of a three-layer fully connected neural network, with the ReLU function as the activation function, encoding from the feature dimension to the fault category dimension.

8. The method according to claim 1, wherein, In the fourth step, the Softmax activation function is used at the end of the multi-layer perceptron to output the probability that the signal belongs to each type of fault, calculate the cross-entropy loss, and update the network parameters using error backpropagation.

9. The method according to claim 3, wherein, It also includes In the sixth step, after the visual input signal is encoded by the multi-scale component analysis network, deep sparse codes of each scale are output, and the overall features of the input signal are reconstructed based on these. : , And view The equivalent convolutional kernel in it, that is, the atomic feature, is analyzed to see if it is consistent with the signal fault feature, and the post hoc interpretability analysis of the network is completed.

Citation Information

Patent Citations

  • NPC three-level inverter fault diagnosis method based on signal sparse representation

    CN110954761A

  • Multi-scale structure and feature fusion gearbox intelligent fault diagnosis method

    CN112116029A