A Gearbox Compound Fault Diagnosis Method Based on Capsule Spectrogram Wavelet Network

Through the method based on capsule spectral wavelet network, combined with the deep attention capsule network and the multi-layer spectral wavelet convolution network, the feature extraction and identification problems in gearbox composite fault diagnosis are solved, and high-precision and efficient fault diagnosis are achieved.

CN118606878BActive Publication Date: 2025-06-27ANHUI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410557305.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-07
Publication Date
2025-06-27
Estimated Expiration
2044-05-07

AI Technical Summary

Technical Problem

It is difficult for the prior art to accurately diagnose the composite failure of the gearbox, traditional methods and single-scale features are difficult to identify multiple fault components, and methods based on graph neural networks lack the ability to extract multi-scale features.

Method used

Using a method based on capsule spectral wavelet network, the deep attention capsule network learns the representation of fault vectors, and the multi-layer spectral wavelet convolutional network learns the relationship between different single-label faults, combining two networks to achieve composite fault diagnosis.

Benefits of technology

It improves the accuracy and efficiency of composite fault diagnosis, can effectively identify the prediction probability of each single fault component in composite fault, and has good stability and noise robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118606878B_ABST
    Figure CN118606878B_ABST
Patent Text Reader

Abstract

The present invention discloses a compound fault diagnosis method for a gearbox based on a capsule spectrogram wavelet network, including: Step 1: A multi-information fusion module performs signal preprocessing and weight adjustment on the gearbox fault signals acquired by multiple sensors, constructs a balanced data set, constructs an adjacency feature matrix, and generates a spatio-temporal graph of the adjacency feature matrix; Step 2: The balanced data set is input into a deep attention capsule network to obtain a vector feature matrix of single-label fault samples; Step 3: The spatio-temporal graph is input into a multi-layer spectrogram wavelet convolutional network to obtain a topological structure feature matrix between multiple single-label faults; Step 4: The vector feature matrix and the topological structure feature matrix are input into a multi-label classifier to obtain the prediction probabilities of each single-fault component in the compound fault. The present invention can effectively diagnose the compound faults of the gearbox, has good stability, strong robustness to noise, and has a certain generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of industrial machine fault diagnosis, and particularly relates to a gearbox compound fault diagnosis method based on a capsule spectrogram wavelet network. Background Technique

[0002] The gearbox is the most widely used component for transmitting speed and power in many industrial machines and is a key component for connecting and transmitting power in mechanical equipment. By using the gearbox, which is a key transmission component for changing the rotational speed and torque, mechanical equipment can achieve functions such as transmitting power, changing speed, and adjusting direction. The gearbox has strong bearing capacity, a compact structure, smooth and accurate transmission power, and high transmission efficiency. Therefore, the gearbox has been widely used in large and complex mechanical equipment such as wind turbines, helicopters, automobiles, agricultural machinery, and metallurgical machinery.

[0003] However, the gearbox is an integrated system. Affected by working conditions such as a harsh working environment, strong load, high speed, and long-term continuous operation, some typical components in the gearbox, such as gears and rolling bearings, are prone to various types of faults, which in turn affect the safety and reliability of the overall operation of the mechanical system. Gearbox faults may lead to unexpected shutdowns, causing huge economic losses and even major casualties. Therefore, more and more people are concerned about gearbox fault diagnosis to ensure the safety of the mechanical system.

[0004] Gearbox faults often do not occur independently. When a gearbox has a compound fault, there may be two or more fault components in the same vibration signal, which will increase the difficulty of fault feature extraction. In actual work, such compound faults are often misidentified as single faults, resulting in missed repairs. Traditional methods and single-scale features are difficult to accurately diagnose the compound fault types of gearboxes. With the development of deep learning (DL) methods, emerging graph neural networks (GNNs) have also been introduced into the field of fault diagnosis. However, these GNN-based fault diagnosis methods have some limitations. Due to the fixed receptive field of GNNs, they lack the ability to achieve multi-scale feature extraction.

[0005] Therefore, compound fault diagnosis is still challenging in terms of feature extraction and fault identification. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide a gearbox compound fault diagnosis method based on a capsule spectrogram wavelet network in view of the above-mentioned deficiencies of the prior art. First, the capsule attention network is used to learn the representation of the fault vector, then the multi-layer spectrogram wavelet convolution network is used to learn the relationship between different single-label faults, and finally the two networks are combined to achieve gearbox compound fault diagnosis, improving the accuracy and efficiency of compound fault diagnosis.

[0007] To achieve the above technical objectives, the technical solution adopted by the present invention is as follows:

[0008] A gearbox compound fault diagnosis method based on a capsule spectrogram wavelet network, characterized by comprising:

[0009] Step 1: The multi-information fusion module performs signal preprocessing and weight adjustment on the gearbox fault signals acquired by multiple sensors, constructs a balanced data set, constructs an adjacency feature matrix, and generates a spatio-temporal graph of the adjacency feature matrix;

[0010] Step 2: The balanced data set is input into the deep attention capsule network to obtain a vector feature matrix of single-label fault samples;

[0011] Step 3: The spatio-temporal graph is input into a multi-layer spectrogram wavelet convolutional network to obtain a topological structure feature matrix between multiple single-label faults;

[0012] Step 4: The vector feature matrix and the topological structure feature matrix are input into a multi-label classifier to obtain the prediction probabilities of each single-fault component in the compound fault.

[0013] To optimize the above technical solution, the specific measures taken also include:

[0014] The signal preprocessing in the above Step 1 includes: resampling the gearbox fault signals acquired by multiple sensors, obtaining multiple data volumes of the original signal by moving a fixed window, then performing FFT transformation to obtain a frequency spectrum, and using the frequency spectrum for normalization to accelerate the network convergence speed, and using the CapsNet module for the fusion of multiple sensor channels;

[0015] The strategy of the weight adjustment is: using data augmentation on the preprocessed data to generate a training set of abnormal single-label fault data, and adjusting the proportion of various single-label fault samples in the training set to construct a balanced data set.

[0016] The deep attention capsule network in the above Step 2 includes a frequency attention mechanism, a convolutional layer, an initial capsule layer, and a digital capsule layer;

[0017] The frequency attention mechanism introduces a global average pooling layer GAP to compress the global information of the balanced data set's mixed features LG ∈ R 2d ×L′ into the channel representation vector z ∈ R 2d×1 where the i-th element of z is expressed as:

[0018]

[0019] The convolutional layer obtains a non-linear representation v i through ReLU and sigmoid activation, specifically as follows:

[0020]

[0021]

[0022]

[0023] where Ker i and b i represent the weight and bias of the convolutional kernel respectively;

[0024] f and represent the non - linear activation function and convolutional calculation respectively;

[0025] The initial capsule layer and the digital capsule layer use the non - linear activation function squash to obtain the vector feature matrix of single - label fault samples, specifically as follows:

[0026]

[0027]

[0028]

[0029]

[0030]

[0031] where is the output vector of digital capsule j, which is the vector feature matrix of single - label fault samples output after n updates;

[0032] W i is the transformation matrix;

[0033] c ij is the coefficient updated by the protocol - based dynamic routing algorithm;

[0034] b ij is the logarithmic prior probability of capsule i coupling capsule j;

[0035] C represents the number of labels.

[0036] When the given input vector is u, the formula of the squash function is as follows:

[0037]

[0038] The multi-layer spectral wavelet convolution network described in step 3 above extracts multi-scale information from the spatiotemporal graph through the spectral wavelet convolution layer, and then learns the feature information hidden in the time domain and frequency domain through the convolution layer. Through layer-by-layer signal decomposition and feature learning, irrelevant information is gradually removed to obtain the topological structure feature matrix between multiple single-label faults.

[0039] The above-mentioned spectral wavelet convolution layer is based on the spectral wavelet transform. It realizes multi-scale feature extraction by performing wavelet decomposition of a low-pass filter and multiple scale band-pass filters on the graphic signal. The expression is as follows:

[0040]

[0041] in, is a learnable diagonal filter matrix in the wavelet domain, Θ is the weight parameter learned from the data, and W T represents the inverse SFWT operator;

[0042] SGWT operator W = [h(L), g(a1L), ..., g(a J L)] consists of a scale kernel function h(L) and J graph wavelet kernel functions g(a1L),...,g(a J L), h(L) corresponds to a low-pass filter, g(a1L),...,g(a J L) corresponds to J bandpass filters of different scales;

[0043] The scale kernel function and graph wavelet kernel function use Mexican hat wavelet and are approximately expressed by Chebyshev polynomials as follows:

[0044]

[0045]

[0046] in, is the kth coefficient of the truncated shift Chebyshev polynomial of the scaling kernel function;

[0047] It is the kth coefficient of the truncated shift Chebyshev polynomial of the wavelet kernel function.

[0048] The Mexican hat wavelet is selected for the above-mentioned scale kernel function and graph wavelet kernel function, as follows:

[0049]

[0050]

[0051] Among them, the scale value t j ∈{a|a1,a2,...,aJ}Determined by the largest Laplacian eigenvalue λ max and the hyperparameter Q;

[0052] The minimum and maximum scales are defined as and

[0053] The lower bound λ min of the Laplacian eigenvalue is

[0054] The intermediate scale where 1 < j < J;

[0055] The parameter γ is γ = e -1 .

[0056] The constraint conditions of the above scale kernel function and graph wavelet kernel function are:

[0057] g(0) = 0, lim λ→∞ g(λ) = 0 (22)

[0058] h(0) > 0, lim λ→∞ h(λ) = 0 (23)

[0059] where g(λ) is the scale kernel function;

[0060] h(λ) is the graph wavelet kernel function.

[0061] The multi-label classifier described in step 4 above learns the vector feature matrix and the topological structure feature matrix, and obtains the prediction probabilities of each single-fault component in the composite fault. The process is as follows:

[0062]

[0063]

[0064] where, is the topological structure feature matrix between multiple single-label faults obtained by the multi-layer spectral graph wavelet convolution network described in step 3; C is the number of categories, x is the vector feature matrix of the single-label fault samples obtained by the deep attention capsule network described in step 2, y is the true label, is the predicted multi-label learned by the classifier, that is, the prediction probabilities of each single-fault component in the composite fault.

[0065] The multi-label classifier described in step 4 above uses the following margin loss function to increase the feature discrimination of different types of faults and optimize the classification performance:

[0066] L c = K c max(0, b +-||v c ||) 2 +λK′max(0,||v c ||-b - ) 2 (28)

[0067] where K c ′ = 1 - K c , if category c exists in the object, then K C = 1;

[0068] The parameter b + = 0.9 and b - = 0.1, λ = 0.5.

[0069] The present invention has the following beneficial effects:

[0070] The present invention proposes a parallel processing method based on a capsule attention network and a multi-layer spectral graph wavelet convolution network to diagnose the compound faults of a gearbox, which can effectively diagnose the compound faults of the gearbox, has good stability, strong robustness to noise, and has a certain generalization ability.

[0071] The frequency spectrum of the fault signal can obtain the vector features of single-label fault samples through the deep attention capsule network, and the obtained vector features are more robust to pose changes and have stronger generalization ability than traditional CNNs;

[0072] And according to the graph structure and feature matrix obtained from the fault data through the multi-layer spectral graph wavelet convolution network, multi-scale feature extraction is realized and the over-smoothing problem is solved, and the topological structure between multiple single-label faults is obtained;

[0073] Finally, the dot product of the output feature matrices of the two sub-networks is used to output the prediction probabilities of each single-fault component in the compound fault. This algorithm has high recognition accuracy and fast processing speed, and can effectively diagnose the faults of the gearbox.

[0074] The attention capsule network module of the present invention designs a frequency attention mechanism, which can automatically adjust the network structure and parameters according to the characteristics of the input data, weight the importance of different frequency characteristics, and can extract the most useful features for fault diagnosis in a targeted manner, solving the problem of feature redundancy. At the same time, skip connections are used to solve the problem of gradient disappearance caused by network deepening.

[0075] The multi-layer spectrogram wavelet convolution network module of the present invention embeds the spectrogram wavelet convolution layer into the ordinary convolution layer. The input first passes through the spectrogram wavelet convolution layer to decompose and obtain multi-scale time-frequency information, and then passes through the ordinary convolution layer to obtain fault-related features from different frequency components, layer by layer learning the fault-related features hidden in the wavelet domain, and gradually removing interference information (such as noise), thereby constructing a deep frequency feature decomposition and feature learning architecture, improving the robustness to noise and the accuracy of fault diagnosis.

[0076] The present invention applies the Mexican hat wavelet in the spectrogram wavelet transform layer to extract multi-scale time-frequency information, decomposes the graph signal into scale function coefficients and multiple spectrogram wavelet coefficients, and by learning the information of different frequencies, improves the network's recognition ability for complex systems and details, and alleviates the over-smoothing problem caused by deepening the multi-layer spectrogram wavelet convolution network. Brief Description of the Drawings

[0077] Figure 1 It is the overall flowchart of the algorithm of the present invention;

[0078] Figure 2 It is the principle of multi-sensor data fusion of the present invention;

[0079] Figure 3 It is the frequency attention mechanism of the present invention;

[0080] Figure 4 It is the dynamic routing protocol of the present invention;

[0081] Figure 5 It is the detailed process of the SGWConv layer of the present invention. Detailed Embodiment

[0082] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0083] Although the steps in the present invention are arranged with reference numerals, they are not used to limit the order of the steps. Unless the order of the steps is clearly stated or the execution of a certain step requires other steps as a basis, the relative order of the steps can be adjusted. It can be understood that the term "and / or" used herein relates to and encompasses any and all possible combinations of one or more of the associated listed items.

[0084] The present invention provides a gearbox compound fault diagnosis method based on a capsule spectrogram wavelet network, which is mainly divided into three parts: The first part, the multi-information fusion module, uses multi-sensor data to obtain the frequency spectrum of the fault signal and constructs a label relationship graph of the adjacent feature matrix; The second part, the capsule attention network module, combines the relative relationship between objects through the capsule network to obtain representative features; The third part, the multi-layer spectrogram wavelet convolution network module, takes the graph convolutional neural network as the backbone network. First, it extracts multi-scale time-frequency information through the spectrogram wavelet convolution layer, and then learns features through the traditional convolution layer. After repeating multiple times to remove interference information such as noise, it obtains the topological structure between multiple single-label faults. Finally, it realizes the classification of compound faults through a multi-label classifier. The flowchart of the method is as Figure 1 , and the specific process is as follows:

[0085] Step 1: The multi-information fusion module performs signal preprocessing and weight adjustment on the gearbox fault signal obtained by the multi-sensor, constructs a balanced data set, constructs an adjacent feature matrix, and generates a spatio-temporal graph of the adjacent feature matrix;

[0086] Step 2: The balanced data set is input into the deep attention capsule network to obtain a vector feature matrix of single-label fault samples;

[0087] Step 3: The spatio-temporal graph is input into the multi-layer spectrogram wavelet convolution network to obtain a topological structure feature matrix between multiple single-label faults;

[0088] Step 4: The vector feature matrix and the topological structure feature matrix are input into the multi-label classifier to obtain the prediction probabilities of each single-fault component in the compound fault.

[0089] In the embodiment, the method of the present invention preprocesses the obtained acceleration signal of the gearbox, inputs the acceleration signals of 3 sensors into the multi-signal fusion module for multi-channel signal processing and data fusion and adopts a weight adjustment strategy to construct a balanced data set and obtain the input of the capsule network; uses the capsule attention network to learn the feature vectors of multiple single-label faults of the fault signal after data preprocessing; constructs an adjacent feature matrix and generates a spatio-temporal graph, and uses the feature matrix and the spatio-temporal graph as inputs to obtain the topological structure between multiple single-label faults through the multi-layer spectrogram wavelet convolution network; uses the generated classifier to perform a dot product on the feature matrices obtained by the two sub-networks and outputs the prediction probabilities of each single-fault component in the compound fault.

[0090] The specific introduction is as follows:

[0091] 1) Multi-signal fusion module:

[0092] This module is used for multi-channel signal processing and data fusion to obtain more health state information of mechanical equipment.

[0093] The signal preprocessing includes the following steps:

[0094] First, resample the acceleration vibration signals of the three sensors collected, and obtain 5 times the amount of data of the original signal by moving a fixed window. Secondly, perform FFT transformation on the original signal to obtain the spectrum. Thirdly, use the spectrum for normalization to accelerate the network convergence speed. Use the data fusion method to synthesize the preprocessed data collected from the three channels into the input unit in the CapsNet module, as Figure 2 shown. Specifically, each input network unit has three layers, and each layer contains data of more than one motion cycle of the sensor channel.

[0095] In addition, the weight adjustment strategy first uses data augmentation to generate abnormal single-label fault data, and then adjusts the proportion of various single-label fault samples in the training set to construct a balanced data set. It allows the model to fully learn the characteristics of various single-label fault components during the training phase.

[0096] 2) Capsule attention network module:

[0097] This module uses a deep attention capsule network to learn the representative features of the fault signals after data preprocessing. It includes the following three modules: frequency attention mechanism, convolutional layer, initial capsule layer and digital capsule layer.

[0098] As Figure 3 shown, the core idea of the frequency attention mechanism is to use the self-learning ability of CNN to filter out the signal components useful for the diagnosis task from the mixed features LG∈R 2d×L′ .

[0099] First, the frequency attention mechanism introduces a global average pooling layer (GAP) to compress the global information of the mixed features LG into the channel representation vector z∈R 2d×1 . The i-th element of z is expressed as:

[0100]

[0101] Then, the frequency attention mechanism uses a simple encoding and decoding mechanism to capture the importance of these channel signals. First, compress the channel representation vector into a hidden layer vector with a dimension of C×1, and then decode it back to the original dimension. Finally, output the channel weight vector z′∈R 2d×1 . The encoding and decoding operations of the frequency attention mechanism are completed by two convolutional layers respectively. The first convolutional layer uses the ReLU function to provide a non-linear transformation function. The second convolutional layer uses the Sigmoid function, mainly used to map the obtained feature vector to the range of 0 to 1, so as to generate the weight vector z′. The size of the elements of z′ represents the importance of the corresponding channel features. This can be expressed as:

[0102]

[0103]

[0104] Finally, the frequency attention mechanism embeds this weight information into the capsule network model using matrix multiplication to guide the feature learning of the capsule model. In addition, a residual connection is added in the frequency attention mechanism to optimize the gradient propagation of the network and prevent the feature response from being too small.

[0105] For a given input z′, the convolutional layer can obtain a non-linear representation v through ReLU or sigmoid activation i , as follows:

[0106]

[0107] where Ker i and b i represent the weights and biases of the convolutional kernel respectively. f and represent the non-linear activation function and the convolutional calculation respectively.

[0108] Then, a new non-linear activation function called "squash" is used to convert the output vector length into the interval of [0, 1] in the primary capsule, which can represent various categories. When the given input vector is u, the formula of the squash function is as follows:

[0109]

[0110] The output of the main layer of the initial capsule is named v p . u i is generated by multiplying v p with the transformation matrix W i , and d j is the weighted sum of all intermediate prediction vectors. They can be obtained as follows:

[0111]

[0112]

[0113]

[0114] where is the output vector of digital capsule j, which is the vector feature matrix of the single-label fault sample output after n updates, and c ij is the coefficient updated by the protocol-based dynamic routing algorithm, which can be expressed as follows:

[0115]

[0116] where b ij is the logarithmic prior probability that capsule i couples with capsule j, and C represents the number of labels. As Figure 4 shown, the dynamic routing protocol aims to construct a complex non-linear mapping between two consecutive capsule layers in a clustering manner, which is expressed as follows:

[0117]

[0118] By adding to the previous b i ' j to update b ij .

[0119] 3) Multi-layer spectral graph wavelet convolutional network module:

[0120] This module first maps the time domain space to the wavelet domain space through the spectral graph wavelet convolutional layer to extract multi-scale time-frequency information, and then the convolutional layer can well learn the feature information hidden in the time domain and frequency domain. Through layer-by-layer signal decomposition and feature learning, irrelevant information (such as noise) is gradually removed to obtain the topological structure between multiple single-label faults.

[0121] The spectral graph wavelet convolution (SGWConv) layer is based on the spectral graph wavelet transform. It can perform wavelet decomposition on the graph signal through a low-pass filter and multiple scale band-pass filters to achieve multi-scale feature extraction. With the help of SGWConv, SGWN can prevent the over-smoothing problem caused by long-range low-pass filtering by simultaneously extracting low-pass and band-pass features. In addition, to accelerate the calculation speed of SGWConv, the scale kernel function and graph wavelet kernel function in SGWConv are approximated as Chebyshev polynomials.

[0122] (a) Spectral graph wavelet convolution

[0123] The core of the spectral graph wavelet transform (SGWT) is the design of a low-pass filter and multi-scale discrete band-pass filters. This definition shows that SGWT can achieve multi-scale analysis by calculating the inner product between the graph signal and the band-pass filter. Therefore, by representing the SGWT operator as W, the SGWT of the graph signal x can be expressed as follows:

[0124]

[0125] where φ is the spectral scale function wavelet, ψ an represents the spectral graph wavelet at scale a and node n, U is the corresponding eigenvector matrix, Λ = dig(λ1, λ2,..., λ n ) is the eigenvalue λ NThe diagonal matrix, g(·) corresponds to the discrete scale a = (a1, a2,..., a J ), J represents the decomposition scale, and h(·) is the scale kernel function.

[0126] When performing signal fitting, or signal decomposition, in addition to introducing wavelet bases, scale bases are also introduced for low-frequency and constant information, as a supplement to wavelet functions in the frequency domain. And wavelet functions are in the shape of a band-pass filter in the frequency domain. Therefore, as a supplement to the ability to capture the blind area of low-frequency signals, the scale function is in the shape of a low-pass filter in the frequency domain.

[0127] Under the condition that the Laplacian matrix L allows eigenvalue decomposition L = UΛU T , Equation (11) can be rewritten as follows:

[0128] Wx = [h(L)x, g(a1L)x,..., g(a J L)x] (12)

[0129] The SGWT operator W = [h(L), g(a1L),..., g(a J L)] consists of a scale kernel function and J wavelet kernel functions of different scales. h(L) corresponds to the low-pass filter, and g(a1L),..., g(a J L) correspond to J band-pass filters of different scales. The SGWConv layer will decompose the input graph signal into SF coefficients and several scale-related wavelet function coefficients by using the designed filters, thus allowing multi-scale feature extraction.

[0130] Therefore, by using the SGW operator W to replace the Fourier basis U in the graph Fourier transform T , the expression of SGWConv can be obtained as follows:

[0131]

[0132] where, is the learnable diagonal filter matrix in the wavelet domain, Θ is the weight parameter learned from the data, and W T represents the inverse SFWT operator. The proposed SGWConv layer is specifically as Figure 5 shown.

[0133] (b) Polynomial approximation of the SGWConv layer

[0134] In the SGWConv layer, to obtain the SGWT operator W, it is necessary to calculate the scale kernel function h(L) and the graph wavelet kernel function g(a J L). Calculating g(a J L) corresponds to calculating g(a according to the properties of matrix functions.J λ l )。The light uses eigenvalue decomposition to obtain the eigenvalues λ of L l In this step, the QR decomposition algorithm will bring a computational complexity of O(N 3 ), and the computational complexity of the total process is even as high as O(N 3 (J + 1)). Therefore, in order to achieve better computational efficiency, the h(L) and g(a J L) operators are approximated by Chebyshev polynomials.

[0135] First, for the function y(t) to be fitted, according to the approximation theory, we can find its convergent Chebyshev series, which is expressed as follows:

[0136]

[0137] The original Chebyshev polynomial has the following form on y ∈ [-1, 1]: the first term T0(y) = 1, the second term T1(y) = y, and starting from the third term, all subsequent terms can be recursively calculated from the values of its first two terms, that is: T k (y) = 2yT k-1 (y) - T k-2 (y). Where k ≤ K is the order of the Chebyshev polynomial, and the expression of c k is as follows:

[0138]

[0139] Since y ∈ [-1, 1], so let where λ is max the largest eigenvalue, then it can make x ∈ [0, λ max , so the following expression can be obtained:

[0140]

[0141] Using the shifted Chebyshev polynomial and its coefficients, f(x) and the k-th coefficient of the shifted Chebyshev polynomial can be expressed as:

[0142]

[0143]

[0144]

[0145] Now, the approximate values of the scale and wavelet kernel function can be expressed as:

[0146]

[0147]

[0148] Among them, is the k-th coefficient of the truncated shifted Chebyshev polynomial of the scale kernel function;

[0149] is the k-th coefficient of the truncated shifted Chebyshev polynomial of the graph wavelet kernel function.

[0150] By approximating W with the truncated shifted Chebyshev polynomial, the total computational complexity becomes O(KN 2 (J + 1)). Generally, the order K of the Chebyshev polynomial is much smaller than N, i.e., K << N. Therefore, through this polynomial approximation, the computational complexity of SGWConv is significantly reduced.

[0151] (c) Core function design of the GWConv layer

[0152] The kernel function of the SGWConv layer should satisfy the following constraints:

[0153] g(0) = 0, lim λ→∞ g(λ) = 0 (22)

[0154] h(0) > 0, lim λ→∞ h(λ) = 0 (23)

[0155] Among them, g(λ) is the scale kernel function;

[0156] h(λ) is the graph wavelet kernel function.

[0157] This algorithm uses the Mexican hat wavelet as the kernel function of the SGWConv layer. The Mexican hat wavelet kernel function and its corresponding scale kernel function are as follows:

[0158]

[0159]

[0160] Among them, the scale value t j ∈{a|a1, a2,..., a J} is determined by the maximum Laplacian eigenvalue λ max and the designed hyperparameter Q, where the lower limit of the Laplacian eigenvalue is set to Therefore, the minimum and maximum scales are respectively defined as and After obtaining the maximum scale and the minimum scale, the intermediate scales can be set to decrease proportionally, such that where 1 < j < J. In addition, Q is taken to be 2, where the parameter γ of the scale kernel function is determined as γ = g max = e -1 .

[0161] 4) Classification and discrimination module:

[0162] The classification and discrimination module (multi-label classifier) performs a dot product on the feature matrices output by the two sub-network modules to obtain the output possibility results for each label. First, the data features of length N obtained by the capsule network module will be represented using vectorized neurons instead of scalar data. Second, the graph features from the adjacent feature matrix will also be obtained by the multi-scale spectral graph wavelet convolution module, which stacks 2 spectral graph wavelet convolution layers and activation functions. Then, a dot product will be performed on the two obtained output features to consider the contributions of the data samples and the graph to the class prediction.

[0163] In addition, is the set of correlations for each class learned by the multi-scale spectral graph wavelet convolution network . Finally, the calculation process of the sample multi-label prediction result is as follows:

[0164]

[0165]

[0166] where C is the number of classes, x is the feature vector learned by CapsNet, y is the true label, is the predicted multi-label learned by the classifier.

[0167] In the designed neural network, the margin loss function can be used not only to increase the feature distinguishability of different types of faults but also to optimize the classification performance of the network model. It adopts a new separate margin loss L c for each type of data c, and its calculation is as follows:

[0168] L c = K c max(0, b + - ||v c ||) 2 + λK′ max(0, ||v c || - b - ) 2 (28)

[0169] where, if class c exists in the object, then K C = 1, K′ c = 1 - K c . In addition, b + = 0.9 and b -= 0.1 respectively means ||v c the upper and lower boundaries of ||. The downward weighting parameter λ (λ = 0.5) is used to prevent the initial learning from shrinking the active vectors of all classes. Finally, the total margin loss calculated is the sum of the losses of all classes.

[0170] It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims concerned.

[0171] In addition, it should be understood that although this specification is described in terms of embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A gearbox composite fault diagnosis method based on capsule spectrum wavelet network, characterized in that: include: Step 1: The multi-information fusion module performs signal preprocessing and weight adjustment on the gearbox fault signals acquired by multiple sensors, builds a balanced data set, constructs an adjacency feature matrix, and generates a spatiotemporal graph of the adjacency feature matrix; Step 2: The balanced data set is input into the deep attention capsule network to obtain the vector feature matrix of the single-label fault sample; Step 3: The spatiotemporal graph is input into a multi-layer spectral wavelet convolutional network to obtain a topological structure feature matrix between multiple single-label faults; Step 4: Input the vector feature matrix and the topological structure feature matrix into a multi-label classifier to obtain the predicted probability of each single fault component in the composite fault; The deep attention capsule network described in step 2 includes a frequency attention mechanism, a convolutional layer, an initial capsule layer and a digital capsule layer; The frequency attention mechanism introduces a global average pooling layer GAP to balance the mixed features LG∈R 2d×L′ The global information is compressed into the channel representation vector z∈R 2d×1 In this example, the i-th element of z is represented as: The convolutional layer obtains nonlinear representation v through ReLU and sigmoid activation. i , as follows: Among them, Ker i and b i Respectively represent the weight and bias of the i-th convolution kernel; f and Represent nonlinear activation function and convolution calculation respectively; The initial capsule layer and the digital capsule layer use the nonlinear activation function squash to obtain the vector feature matrix of the single-label fault sample, as follows: in is the output vector of digital capsule j, which is the vector feature matrix of the single-label fault sample output after n updates; W i is the transformation matrix; c ij is the coefficient updated by the protocol-based dynamic routing algorithm; b ij is the logarithmic prior probability of capsule i coupling with capsule j; C represents the number of labels; The multi-layer spectral wavelet convolution network in step 3 extracts multi-scale information from the spatiotemporal graph through the spectral wavelet convolution layer, and then learns the feature information hidden in the time domain and frequency domain through the convolution layer. Through layer-by-layer signal decomposition and feature learning, irrelevant information is gradually removed to obtain the topological structure feature matrix between multiple single-label faults.

2. The gearbox composite fault diagnosis method based on capsule spectrum wavelet network according to claim 1 is characterized in that: The signal preprocessing in step 1 includes: resampling the gearbox fault signal acquired by multiple sensors, obtaining multiple data volumes of the original signal by moving the fixed window, and then performing FFT transformation to obtain the spectrum, and using the spectrum for normalization to speed up the network convergence speed, and using the CapsNet module to fuse the multi-sensor channels; The weight adjustment strategy is: use data enhancement on the preprocessed data to generate an abnormal single-label fault data training set, and adjust the proportion of various single-label fault samples in the training set to construct a balanced data set.

3. The gearbox composite fault diagnosis method based on capsule spectrum wavelet network according to claim 1 is characterized in that: When the given input vector is u, the formula of the squash function is as follows:

4. The gearbox composite fault diagnosis method based on capsule spectrum wavelet network according to claim 1 is characterized in that: The spectral wavelet convolution layer is based on the spectral wavelet transform, and realizes multi-scale feature extraction by performing wavelet decomposition of a low-pass filter and multiple scale band-pass filters on the graphic signal. The expression is as follows: in, is a learnable diagonal filter matrix in the wavelet domain, Θ is the weight parameter learned from the data, W T represents the inverse SFWT operator; SGWT operator W = [h(L), g(a1L), ..., g(a J L)] consists of a scale kernel function h(L) and J graph wavelet kernel functions g(a1L),...,g(a J L), h(L) corresponds to the low-pass filter g(a1L),...,g(a J L) corresponds to J scale bandpass filters; The scale kernel function and graph wavelet kernel function use Mexican hat wavelet and are approximately expressed by Chebyshev polynomials as follows: in, is the kth coefficient of the truncated shift Chebyshev polynomial of the scaling kernel function; It is the kth coefficient of the truncated shift Chebyshev polynomial of the wavelet kernel function.

5. The gearbox composite fault diagnosis method based on capsule spectrum wavelet network according to claim 4 is characterized in that: The Mexican hat wavelet selected by the scale kernel function and the graph wavelet kernel function is as follows: Among them, the scale value t j ∈{a|a1,a2,...,a J } by the maximum Laplace eigenvalue λ max and hyperparameter Q is determined; The minimum and maximum scales are defined as and The lower limit of the Laplace eigenvalue λ min for Intermediate scale 1 of them <j<J; The parameter γ is γ = e -1 .

6. The gearbox composite fault diagnosis method based on capsule spectrum wavelet network according to claim 5 is characterized in that: The constraints of the scale kernel function and the graph wavelet kernel function are: g(0)=0,lim λ→∞ g(λ)=0 (22) h(0)>0,lim λ→∞ h(λ)=0 (23) Among them, g(λ) is the scaling kernel function; h(λ) is the graph wavelet kernel function.

7. The gearbox composite fault diagnosis method based on capsule spectrum wavelet network according to claim 1 is characterized in that: In step 4, the multi-label classifier learns the vector feature matrix and the topological structure feature matrix to obtain the predicted probability of each single fault component in the composite fault. The process is as follows: in, is the topological structure feature matrix between multiple single-label faults obtained by the multi-layer spectral wavelet convolutional network in step 3; C is the number of categories, x is the vector feature matrix of single-label fault samples obtained by the deep attention capsule network in step 2, y is the true label, is the predicted multi-label learned by the classifier, that is, the predicted probability of each single fault component in the compound fault.

8. The gearbox composite fault diagnosis method based on capsule spectrum wavelet network according to claim 1 is characterized in that: The multi-label classifier described in step 4 uses the following margin loss function to increase the feature discrimination of different types of faults and optimize the classification performance: L c =K c max(0,b + -||v c ||) 2 +λK c ′max(0,||v c ||-b - ) 2 (28) Among them, K c ′=1-K c , if there is a type c in the object, then K c =1; Parameter b + = 0.9 and b - =0.1,λ=0.5.