Bearing fault diagnosis method based on dynamic multi-kernel jump graph convolutional network
Through the dynamic multi-core jump graph convolution network, the weighted graph is constructed using the adaptive robust perceptual graph module and the dynamic receptive field convolution module, which solves the problem of fault diagnosis accuracy of traditional methods in large-scale data sets and noise environments, and achieves efficient and robust bearing fault diagnosis.
Patent Information
- Application Number
- CN202510327724.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-04
AI Technical Summary
When processing large-scale data sets, traditional machine learning methods require professional knowledge and are not robust enough. Deep learning models are inaccurate in noisy environments. Traditional graph convolutional networks are difficult to adapt to complex and dynamic graph structures, resulting in a decrease in fault diagnosis accuracy.
A dynamic multi-core jump graph convolution network is adopted to construct weighted graphs through adaptive robust perceptual graph module and dynamic receptive field convolution module to reduce noise impact, dynamically adjust feature extraction range, combine jump connections to enhance information flow, and capture multi-scale features.
Improve fault diagnosis accuracy in noise environments, reduce dependence on expert experience, simplify the diagnosis process, show strong robustness and noise immunity, and adapt to complex fault scenarios.
Smart Images

Figure CN120257007A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of bearing fault diagnosis, and specifically discloses a bearing fault diagnosis method based on a dynamic multi-core jump graph convolutional network. Background Art
[0002] Mechanical systems are becoming increasingly complex, making them an indispensable part of modern industrial production. As a key component in mechanical equipment, rolling bearings ensure the safety and stability required for the equipment to operate efficiently and with high quality. Therefore, the health monitoring and fault diagnosis of rolling bearings have become a key research direction in the field of intelligent mechanical operation and maintenance.
[0003] In essence, bearing fault diagnosis aims to evaluate the health status of bearings. Machine learning-based methods have been proven to be effective in this diagnostic process. Although traditional machine learning methods perform well in dealing with small-scale data sets, they often encounter challenges when faced with large-scale data sets. In addition, traditional feature extraction techniques require a lot of professional knowledge, which may be a major obstacle for users.
[0004] With the progress of deep learning (DL) theory and computing resources, significant progress has been made in fault diagnosis technology in terms of feature extraction. By adopting neural network architectures in deep learning, various advanced techniques can be used to automatically extract features from raw signals. Among them, methods such as convolutional neural networks (CNNs), autoencoders (AEs), deep belief networks (DBNs), and long short-term memory networks (LSTMs) have attracted extensive attention in both academia and industry.
[0005] Deep learning strategies have shown significant value in automatically extracting complex patterns from raw data, especially in applications such as fault detection, signal processing, and pattern recognition. These methods are adept at handling massive amounts of information and revealing hidden structures, thus finding wide applications in both research and industrial fields. For example, Wen et al. successfully used Markov transfer fields to convert frequency signals into images to capture the dynamic characteristics of the signals, and employed a multi-branch residual convolutional neural network for efficient fault feature extraction. Yu et al. proposed a three-stage semi-supervised convolutional neural network for rolling bearing fault diagnosis, which combined data augmentation, k-means clustering, and feature distribution alignment. Shao et al. developed a deep wavelet autoencoder based on Gaussian wavelet functions for diagnosing faults in electric locomotive bearings. Shi et al. utilized a bidirectional convolutional long short-term memory network to extract features from time series data for early detection of rolling bearing faults. Although a large amount of research has focused on improving diagnostic accuracy, the impact of noise on the robustness of the model remains a key issue. Although some scholars have developed models specifically to enhance the noise resistance of the network, this problem has not been fully resolved, and the robustness and noise resistance of the network are still important areas that require in-depth research.
[0006] Deep neural networks (DNNs) are capable of learning the correlations in the input data, but in a noisy environment, the feature representations may not accurately capture the signal relationships that change with the health state of the device. Graph data represents these relationships through nodes and edges and performs well in handling noisy conditions. To meet the processing requirements of graph data, graph convolutional networks (GCNs) have been proposed. GCNs generate an embedded representation for the central node by integrating the feature information of neighboring nodes, enabling the central node to utilize the neighbor information for fault classification.
[0007] Although GCNs perform well in processing graph data, they still face several challenges. First, in a noisy environment, traditional graph construction methods often have difficulty effectively distinguishing meaningful signals from noise when calculating the edge weights between samples. This results in the generated graph structure being unable to accurately reflect the true relationships between sample points, thus affecting the extraction of key features. Second, the fixed receptive field design of traditional GCNs has certain limitations. On the one hand, the fixed receptive field restricts the network's ability to process features of different scales; on the other hand, in complex scenarios such as industrial fault diagnosis, the relationships between samples may be diverse and dynamic, and the fixed receptive field cannot adapt to this change and is difficult to dynamically adjust the feature extraction range according to different graph structures. In addition, as the number of network layers increases, the influence of neighboring nodes on the features of the central node gradually deepens. Although this helps to understand the overall structure, it may blur the key local features and weaken the model's ability to distinguish different fault types. Summary of the Invention
[0008] To solve the technical problems existing in the prior art, the present invention provides a bearing fault diagnosis method based on a dynamic multi-core jumping graph convolutional network. By using an adaptive anti-noise graph module to construct graph data, this method reduces the influence of noise on the connection weights between sample points. In addition, the dynamic multi-core jumping graph convolutional network adopts a dynamic allocation mechanism of multi-scale convolutional kernels, flexibly extracts features at different levels, and ensures the flow of information between multiple dynamic receptive field layers through skip connections, thereby alleviating the problem of feature homogenization in deep networks.
[0009] To achieve the above object, the technical solution adopted by the present invention is: a bearing fault diagnosis method based on a dynamic multi-core jumping graph convolutional network, and the specific steps are as follows:
[0010] Spectral graph convolution: For an undirected graph G=(V, ξ, A), V=n represents finite nodes, ξ represents the set of edges, A∈R n×n represents the adjacency matrix of graph G. For two nodes <v i , v j >, the value of A ij is expressed as:
[0011]
[0012] As a kind of data in the non-Euclidean domain, the graph is represented by the adjacency matrix. The Laplacian matrix of the graph is defined as:
[0013] L = D - A (2)
[0014] where L is the Laplacian matrix, D∈R n×n is the degree matrix;
[0015] Generally, it is necessary to perform symmetric normalization on the Laplacian matrix, and it can be defined as:
[0016] L sym = D -1 / 2 LD -1 / 2 = I N - D -1 / 2 AD -1 / 2 (3)
[0017] where L sym is the symmetric normalized Laplacian matrix, and I N is the identity matrix;
[0018] The eigenvalue decomposition of the symmetric normalized graph Laplacian operator is:
[0019]
[0020] where Λ = diag(λ1, λ2,..., λ n ) is composed of Lsym The diagonal matrix composed of the eigenvalues, and U is composed of the eigenvectors of L sym is an orthogonal matrix, for example, U -1 = U T ;
[0021] At this time, the spectral graph convolution of node v and node features can be defined as:
[0022] h = (x * G f) θ = U(U T xU T f) (5)
[0023] where x is the node feature, h represents the feature map after graph convolution, f is the eigenfunction of Λ, that is, f(Λ), θ is the learnable parameter, and * G represents graph convolution.
[0024] Considering f θ = U T f as a learnable graph convolution filter, the above formula is simplified to:
[0025] h = Uf θ U T x (6)
[0026] where U T x represents the graph Fourier transform of the node feature x.
[0027] Graph Convolutional Network: The graph convolution shown in Equation (6) is not spatially localized. In GCN, the filter f θ can be approximated by the truncated expansion of the Chebyshev polynomial proposed in the polynomial approximation θ , and the K-order approximation is:
[0028]
[0029] where K is the maximum order of the Chebyshev polynomial, is the rescaled eigenvalue, λ max represents the maximum eigenvalue of L, θ is the vector of polynomial coefficients, θ k ' ∈ R K is the vector of Chebyshev coefficients, is the Chebyshev polynomial of order k, which can be determined by the following recurrence relation, that is, T k (x) = 2xT k-1 (x) - T k-2 (x), where T0(x) = 1 and T1(x) = x.
[0030] After approximating the filter with the Chebyshev polynomial, the node feature X and the spectral domain filter fθ The graph convolution can be mathematically defined as follows:
[0031]
[0032] where is the rescaled Laplacian matrix.
[0033] To increase the non-linearity of the graph convolution, the graph convolution result is non-linearly activated by an activation function, expressed as:
[0034]
[0035] where σ is the activation function, and ReLU is usually used as the activation function, representing the activated feature representation.
[0036] The standard two-layer GCN model for node classification on a graph is expressed as:
[0037]
[0038] where X = {x1, x2,..., x n} is the set of node features, n is the number of nodes, Z is the convolutional signal matrix, and here, the softmax function defined as is applied row by row.
[0039] The Dynamic Multi-Kernel Skip Graph Convolutional Network (DMK-SGCN) integrates innovative modules, including the Adaptive Robustness-Aware Graph Module (ARAGM) and the Dynamic Receptive Field Convolution Module (DRFConv).
[0040] The Adaptive Robustness-Aware Graph Module (ARAGM) is an information-theory-based metric method for constructing a weighted graph by evaluating the difference between two probability distributions. By emphasizing the key regions in the spectrum and reducing the influence of noise, it ensures the robustness and noise resistance of the graph construction process. The specific steps are as follows:
[0041] Data preprocessing: For the original time series X with a signal length of L, data normalization is usually performed, and the normalized data can be expressed as:
[0042] X nol = normalize(X) (11)
[0043] where X nol is the normalized time series, and normalize(·) represents the normalization operation, and here, min-max normalization is adopted.
[0044] After data normalization, the normalized data can be divided into a series of subsamples with a predefined length or dimension d to construct a new dataset:
[0045]
[0046] where M is the constructed upper sample set, denotes the subsample, n is the total number of subsamples, floor represents rounding down, and d is the dimension of the subsample.
[0047] To reduce the impact of noise, perform a fast Fourier transform (FFT) on each subsample and use the resulting spectrum as a new sample. Through this method, the time-domain signal can be effectively converted into a frequency-domain signal, thus constructing a spectrum sample set with less noise impact. The entire process can be expressed as:
[0048]
[0049] where FFT(·) is an operator that transforms the subsample from the time domain to the frequency domain and extracts the first half of the result for further processing. Normalize each spectrum sample and convert it into a probability distribution.
[0050]
[0051] The numerator represents the amplitude of the frequency interval k, and the denominator is the sum of the amplitudes of all frequency intervals is the normalized spectrum sample.
[0052] By assigning labels to each sample, a labeled dataset can be obtained as:
[0053]
[0054] where y n is the label of the nth sample, and D is the labeled dataset;
[0055] The adaptive perception graph module can be constructed from the above labeled dataset. First, determine the number of nodes q. To find the neighbors of each node in the graph, use the ε-radius and obtain the neighbors of the node in the following way:
[0056]
[0057] where is the neighborhood of, ε is the selected radius, returns the set j ≤ q, Is the calculated node The bidirectional KL divergence between them.
[0058] The cumulative distribution function of the probability distribution is used to calculate the high-energy region, that is, the frequency range containing the main energy is concerned. The specific formula is as follows:
[0059]
[0060] Among them, F i (k) represents the cumulative probability of the i-th sample from frequency point 1 to k, represents the normalized energy ratio of the i-th sample at frequency point j, and d is the total number of frequency points in the spectrum.
[0061] Determine k * , such that F i (k * )≥τ, where (τ∈[0,1]), the formula is as follows:
[0062]
[0063] The high-energy region corresponds to the frequency range [1,k * ;
[0064] The bidirectional KL divergence is defined to measure the similarity between two subsequences. The specific formula is as follows:
[0065]
[0066] Among them, is a measure of the bidirectional divergence between the frequency spectrum distributions of two subsequences, is the value at point k, the value at point k, a and b represent the proportions of the high-energy region in the total energy, indicating the importance of these key regions.
[0067] While obtaining the neighborhood set of , the weights between each two nodes can be calculated through the threshold Gaussian kernel weight function:
[0068]
[0069] The factor can adjust the sensitivity of the model to the node similarity. By changing the value of the factor, the sensitivity of the model to the differences between nodes during weight calculation can be controlled.
[0070] Among them, w ij is the weight between nodes i and j, β is the bandwidth variance of the Gaussian function. When finding the node When the weights between neighbors and between their neighbors are obtained, a mathematically defined weighted graph can be obtained. Therefore, by repeating the steps shown in equations (16)-(22), a weighted graph dataset G can be generated. Generally, if the cosine similarity between two samples is greater than 0, there may be a positive correlation between them. During the construction of the affinity graph, the positive correlation between samples is mainly considered. Empirically, 0 is considered the threshold for the radius ε.
[0071] The network introduces an adaptive weight assignment mechanism through the AdaptiveMaxPool and Softmax layers. This mechanism adaptively assigns weights to different-scale convolutions according to the input data, highlighting important feature information during the feature fusion process, thereby enhancing the model's representation ability.
[0072]
[0073] Denotes the transformation of the Chebyshev polynomial to the matrix where, is the rescaled Laplacian matrix, where λ max denotes the largest eigenvalue of L sym X represents the feature matrix of the input data, [·] represents the concatenation operator, l represents the number of receptive fields, K1, K2, K l represent the receptive field sizes, k1, k2, k l represent the polynomial orders, are learnable parameters, and H is the fused feature matrix;
[0074] In most cases, the adjacency matrix and the node features of the graph are usually used as the inputs of GCNs. The weight parameters are denoted as W. Therefore, the above equation can be simplified as:
[0075]
[0076] where, Here, N is the number of samples, F is the number of sample features, and W0, W1, W l are the weight matrices for each receptive field.
[0077] By passing the fused feature H through one-dimensional convolution in two layers with the strategy of first compressing the number of channels and then increasing the number of channels, the feature learning ability of the model can be enhanced while ensuring computational efficiency. This method helps to capture multi-scale information, improve the non-linear expression ability of features, and improve the flow of gradients, thereby enhancing the performance of the entire network, which can be expressed as:
[0078] H expanded =Conv1D(H concat , in_channels=K l, out_channels = C') (25)
[0079] H compressed = Conv1D(H expanded , in_channels = C', out_channels = K l ) (26)
[0080] H expanded is the output feature matrix for increasing the number of channels, H compresses is the output feature matrix for increasing the number of channels, C' is the expanded number of channels, in_channels is the number of input channels, and out_channels is the number of output channels;
[0081] The obtained H is pooled through the adaptive pooling layer compressed to obtain K l one-dimensional vectors:
[0082] H pool = AdaptivePool(H compressed ) (27)
[0083] where H pool is the output feature matrix after the adaptive pooling layer, and AdaptiveAvgPool(·) is the adaptive average pooling operation;
[0084] Finally, through the Softmax layer, the adaptive weights between the vectors are obtained, and the weighted sum of different receptive fields is performed, which can be expressed as:
[0085] W = Softmax(H pool ) (28)
[0086]
[0087] where in formula (28) in formula (29), N is the total number of nodes, w i is the feature vector of the i-th node in W, and similarly h i is the feature vector of the i-th node in H.
[0088] The DRFConv model is combined with the JKN framework to enhance node features. The reason for choosing the DRFConv model is that it improves the feature representation in complex graph structures through multi-receptive fields and attention mechanisms, thereby more effectively capturing long-range dependencies.
[0089] For the k-layer aggregation model, the feature of the l-th layer node is expressed as:
[0090]
[0091] Among them, σ(·) represents the activation function, W l represents the weight matrix of the l-th layer, aggregate(·) represents the specific aggregation method, which represents the aggregation method of the dynamic receptive field convolution module here, and N(v) represents all adjacent nodes of node v.
[0092] After each layer of the dynamic receptive field convolution module, save the output features of this layer:
[0093] {h v (1) ,h v (2) ,...,h v (l)} (31)
[0094] Among them, h v (l) represents the node features extracted by the aggregation method of the dynamic receptive field convolution module after the l-th layer, and {·} is the concatenation operation;
[0095] By using the adaptive max pooling layer to aggregate the outputs of all layers, the final node representation is formed:
[0096] H final = AdaptiveMaxPool(h v (1) ,h v (2) ,...,h v (l) ) (32)
[0097] Among them, H final is the final node representation, and AdaptiveMaxPool(·) is the adaptive max pooling operation.
[0098] The present invention proposes a method for constructing a weighted graph based on the ARAGraph and DMK-SGCN models. Experimental results from two scenarios show that the combination of ARAGraph and DMK-SGCN achieves the highest classification accuracy.
[0099] The dynamic multi-kernel skip graph convolutional network can dynamically adjust the scope of information fusion in the graph convolutional network, enabling it to adaptively focus on the key regions in the graph. The introduction of skip connections promotes the unobstructed transmission of information in the network, reduces data loss caused by the increase in model depth, and strengthens the repeated use of features, thus more effectively capturing global information. The introduction of this mechanism not only improves the performance of the model but also demonstrates its potential in complex fault scenarios.
[0100] The present invention can construct a sample into a weighted graph, thereby accurately characterizing fault information in the case of strong noise, without relying on complex expert experience, avoiding the necessity of manual feature extraction, thus simplifying the diagnosis process and promoting efficient end-to-end intelligent diagnosis.
[0101] The method proposed by the present invention performs excellently in a noisy environment, demonstrating strong performance and robustness to noise. Compared with traditional machine learning models, deep learning models, and other graph neural network models, it can significantly reduce the impact of noise and maintain the best diagnostic performance in a challenging noisy environment. This method provides significant advantages in practical industrial applications, enabling it to effectively cope with the complexity and dynamics in the real environment. Brief Description of the Drawings
[0102] Figure 1 It is a schematic diagram of the structure of a dynamic multi-core jump graph convolutional network.
[0103] Figure 2 It is a schematic diagram of a four-layer jump knowledge network.
[0104] Figure 3 It is an assembly diagram of a dynamic simulator for a transmission system.
[0105] Figure 4 It is a schematic diagram of vector visualization of the dynamic multi-core jump graph convolutional network under different noise conditions. Figure 4 (a) is a schematic diagram of vector visualization at 0 dB. Figure 4 (b) is a schematic diagram of vector visualization at -2 dB. Figure 4 (c) is a schematic diagram of vector visualization at -4 dB. Figure 4 (d) is a schematic diagram of vector visualization at -6 dB. Figure 4 (e) is a schematic diagram of vector visualization at -8 dB.
[0106] Figure 5 It is a dot-line graph of ChebyNet and DMK-SGCN with different receptive fields under three graph construction methods. Figure 5 (a) is a dot-line graph of KNNGraph. Figure 5 (b) is a dot-line graph of PathGraph. Figure 5 (c) is a dot-line graph of ARAGraph.
[0107] Figure 6 It is a schematic diagram of the confusion matrix of DMK-SGCN under different graph construction methods and noise levels. Figure 6 (1) is a schematic diagram of the confusion matrix of the KNNGraph method. Figure 6 (2) is a schematic diagram of the confusion matrix of the PathGraph method. Figure 6(3) Schematic diagram of the confusion matrix of the ARAGraph method; the noise levels are 0 dB, -2 dB, -4 dB, -6 dB, and -8 dB from left to right.
[0108] Figure 7 Schematic diagram of the DMK-SGCN loss values under three graph construction methods.
[0109] Figure 8 Assembly diagram of the gearbox test bench.
[0110] Figure 9 Bar chart of the diagnostic results of each model under three graph construction methods at -15 dB noise.
[0111] Figure 10 Schematic diagram of the confusion matrix of DMK-SGCN under different graph construction methods and noise levels. Figure 10 (1) Schematic diagram of the confusion matrix of the KNNGraph method. Figure 10 (2) Schematic diagram of the confusion matrix of the PathGraph method. Figure 10 (3) Schematic diagram of the confusion matrix of the ARAGraph method, and the noise levels are -6 dB, -8 dB, -10 dB, and -15 dB from left to right.
[0112] Figure 11 Experimental effect diagram of the fault diagnosis of three models based on different graph convolutional layers in the KNN graph construction method at -12 dB noise. Specific implementation manner
[0113] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0114] A bearing fault diagnosis method based on a dynamic multi-core jump graph convolutional network, and the specific steps are as follows:
[0115] Spectral graph convolution: For an undirected graph G=(V,ξ,A), V=n represents a finite number of nodes, ξ represents the set of edges, A∈R n×n represents the adjacency matrix of graph G. For two nodes <v i ,v j > in the graph, the value of A ij is expressed as:
[0116]
[0117] As a kind of data in the non-Euclidean domain, the graph is represented by the adjacency matrix, and the Laplacian matrix of the graph is defined as:
[0118] L = D - A (2)
[0119] Among them, L is the Laplacian matrix, D ∈ R n×n is the degree matrix;
[0120] Under normal circumstances, it is necessary to perform symmetric normalization on the Laplacian matrix, and it can be defined as:
[0121] L sym = D -1 / 2 LD -1 / 2 = I N - D -1 / 2 AD -1 / 2 (3)
[0122] Among them, L sym is the symmetric normalized Laplacian matrix, I N is the identity matrix;
[0123] The eigenvalue decomposition of this symmetric normalized graph Laplacian operator is:
[0124]
[0125] Among them, Λ = diag(λ1, λ2,..., λ n ) is a diagonal matrix composed of the eigenvalues of L sym , U is an orthogonal matrix composed of the eigenvectors of L sym , for example U -1 = U T ;
[0126] At this time, the spectral graph convolution of node v and node features can be defined as:
[0127] h = (x * G f) θ = U(U T xU T f) (5)
[0128] Among them, x is the node feature, h represents the feature map after graph convolution, f is the eigenfunction of Λ, that is, f(Λ), θ is a learnable parameter, and * G represents graph convolution.
[0129] Considering f θ = U T f as a learnable graph convolution filter, the above formula is simplified to:
[0130] h = Uf θ U T x (6).
[0131] Among them, U Tx represents the graph Fourier transform of node feature x.
[0132] Graph Convolutional Network: The graph convolution shown in Equation (6) is not spatially localized. In GCN, the filter f θ can be approximated by the truncated expansion of the Chebyshev polynomial proposed in the polynomial approximation. θ , and the K - order approximation is:
[0133]
[0134] where K is the maximum order of the Chebyshev polynomial, is the rescaled eigenvalue, λ max represents the largest eigenvalue of L, θ is the vector of polynomial coefficients, θ k ′ ∈R K is the vector of Chebyshev coefficients, is the Chebyshev polynomial of order k, which can be determined by the following recurrence relation, i.e., T k (x) = 2xT k-1 (x) - T k-2 (x), where T0(x) = 1 and T1(x) = x.
[0135] After approximating the filter by the Chebyshev polynomial, the graph convolution of node feature X and the spectral - domain filter f θ can be mathematically defined as follows:
[0136]
[0137] where, is the rescaled Laplacian matrix.
[0138] To increase the non - linearity of the graph convolution, the graph convolution result is non - linearly activated by an activation function, denoted as:
[0139]
[0140] where σ is the activation function, and ReLU is usually used as the activation function, represents the activated feature representation.
[0141] The standard two - layer GCN model for node classification on a graph is represented as:
[0142] Z = softmax(h(ReLU(h(X, A)), A)) (10)
[0143] where X = {x1, x2,..., x n} is the set of node features, n is the number of nodes, Z is the convolutional signal matrix. Here, it is defined row - by - row as The softmax function.
[0144] The Dynamic Multi-Kernel Skip Graph Convolutional Network (DMK-SGCN) integrates innovative modules, including the Adaptive Robustness-Aware Graph Module (ARAGM) and the Dynamic Receptive Field Convolution Module (DRFConv).
[0145] The Adaptive Robustness-Aware Graph Module (ARAGM) is a metric method based on information theory for constructing a weighted graph by evaluating the difference between two probability distributions. By emphasizing the key regions in the spectrum, it reduces the influence of noise, ensuring the robustness and noise resistance of the graph construction process. The specific steps are as follows: Data preprocessing: For the original time series X with signal length L, data normalization is usually performed. The normalized data can be expressed as:
[0146] X nol = normalize(X) (11)
[0147] where X nol is the normalized time series, and normalize(·) represents the normalization operation. Here, the min-max normalization is adopted.
[0148] After data normalization, the normalized data can be divided into a series of subsamples with a predefined length or dimension d to construct a new dataset:
[0149]
[0150] where M is the constructed up-sampled set, represents the subsample, n is the total number of subsamples, floor represents rounding down, and d is the dimension of the subsample.
[0151] To reduce the influence of noise, the fast Fourier transform (FFT) is performed on each subsample, and the resulting spectrum is used as a new sample. By this method, the time-domain signal can be effectively converted into the frequency-domain signal, thus constructing a spectrum sample set with less noise influence. The whole process can be expressed as:
[0152]
[0153] where FFT(·) is an operator that transforms the subsample from the time domain to the frequency domain and extracts the first half of the result for further processing. Each spectrum sample is normalized and converted into a probability distribution.
[0154]
[0155] The numerator represents the amplitude of the frequency interval k, and the denominator It is the sum of the amplitudes in all frequency intervals is the normalized spectral sample.
[0156] By assigning labels to each sample, a labeled dataset can be obtained as follows:
[0157]
[0158] where y n is the label of the nth sample, and D is the labeled dataset;
[0159] The adaptive perception graph module can be constructed from the above labeled dataset. First, determine the number of nodes q. To find the neighbors of each node in the graph, use the ε-radius and obtain the neighbors of the node in the following way of neighbors:
[0160]
[0161] where is the neighborhood of, and ε is the selected radius, returns the set j ≤ q, is the calculated node between the two-way KL divergence.
[0162] The cumulative distribution function of the probability distribution is used to calculate the high-energy region, that is, to focus on the frequency range containing the main energy. The specific formula is as follows:
[0163]
[0164] where F i (k) represents the cumulative probability of the ith sample from frequency point 1 to k, represents the normalized energy ratio of the ith sample at frequency point j, and d is the total number of frequency points in the spectrum.
[0165] Determine k * , such that F i (k * ) ≥ τ, where (τ ∈ [0, 1]), the formula is as follows:
[0166] k * = min{k | F i (k) ≥ τ} (18)
[0167] The high-energy region corresponds to the frequency range [1, k * ;
[0168] The two-way KL divergence is defined to measure the similarity between two subsequences. The specific formula is as follows:
[0169]
[0170] Among them, is a measure of the two-way divergence between the frequency spectrum distributions of two subsequences, is the value at point k, the value at point k, where a and b represent the proportions of the high-energy regions in the total energy, indicating the importance of these key regions.
[0171] While obtaining the neighborhood set of, the weights between every two nodes can be calculated through a threshold Gaussian kernel weight function:
[0172]
[0173] The factor can adjust the sensitivity of the model to node similarity, and by changing the value of the factor, the sensitivity of the model to the differences between nodes during weight calculation can be controlled.
[0174] Among them, w ij is the weight between nodes i and j, β is the bandwidth variance of the Gaussian function. When finding the neighbors of node and the weights between their neighbors, a mathematically defined weighted graph can be obtained. Therefore, by repeating the steps shown in equations (16)-(22), a weighted graph dataset G can be generated. Generally, if the cosine similarity between two samples is greater than 0, there may be a positive correlation between them. During the construction of the affinity graph, the positive correlation between samples is mainly considered, and empirically, 0 is considered as the threshold with a radius of ε.
[0175] The network introduces an adaptive weight allocation mechanism through the AdaptiveMaxPool and Softmax layers. This mechanism adaptively allocates the weights of different-scale convolutions according to the input data, highlighting important feature information during the feature fusion process, thereby enhancing the representational ability of the model.
[0176]
[0177] represents the transformation of the Chebyshev polynomial for the matrix where is the rescaled Laplacian matrix, where λ max represents the largest eigenvalue of L sym X represents the feature matrix of the input data, [·] represents the concatenation operator, l represents the number of receptive fields, K1, K2, K l represent the receptive field sizes, k1, k2, k l represent the polynomial orders, θ' k1 、θ′ k2 、θ k ′ l are learnable parameters, and H is the fused feature matrix;
[0178] In most cases, the adjacency matrix and the node features of the graph are usually used as the inputs of GCNs, and the weight parameter is denoted as W. Therefore, the above formula can be simplified as:
[0179]
[0180] where here N is the number of samples, F is the number of sample features, and W0, W1, W l are the weight matrices for each receptive field.
[0181] By passing the fused feature H through a one-dimensional convolution in two layers with a strategy of first compressing the number of channels and then increasing the number of channels, it is possible to enhance the feature learning ability of the model while ensuring computational efficiency. This method helps to capture multi-scale information, improve the non-linear expression ability of features, and improve the flow of gradients, thereby enhancing the performance of the entire network, which can be expressed as:
[0182] H expanded = Conv1D(H concat , in_channels = K l , out_channels = C′) (25)
[0183] H compressed = Conv1D(H expanded , in_channels = C′, out_channels = K l ) (26)
[0184] H expanded is the output feature matrix for increasing the number of channels, H compresses is the output feature matrix for increasing the number of channels, C′ is the expanded number of channels, in_channels is the number of input channels, and out_channels is the number of output channels;
[0185] By passing the obtained H compressed features through an adaptive pooling layer to obtain K l one-dimensional vectors:
[0186] H pool = AdaptivePool(H compressed ) (27)
[0187] where, H poolis the output feature matrix after the adaptive pooling layer, and AdaptiveAvgPool(·) is the adaptive average pooling operation;
[0188] Finally, through the Softmax layer, the adaptive weights between vectors are obtained, and the weighted sum of different receptive fields is performed, which can be expressed as:
[0189] W = Softmax(H pool ) (28)
[0190]
[0191] Among them, in formula (28) In formula (29), N is the total number of nodes, w i is the feature vector of the i-th node in W, and similarly h i is the feature vector of the i-th node in H.
[0192] The DRFConv model is combined with the JKN framework to enhance node features. The reason for choosing the DRFConv model is that it improves the feature representation in complex graph structures through multi-receptive fields and attention mechanisms, thereby more effectively capturing long-range dependencies.
[0193] Skip Knowledge Network: Traditional graph convolutional networks (GCNs) aggregate information from neighboring nodes in each layer through graph convolution operations. However, as the network depth increases, this method may lead to problems such as vanishing gradients or exploding gradients. In addition, the features captured in different layers may have different importance in complex graph structures, and simple layer-by-layer aggregation may not fully utilize these features. The Skip Knowledge Network (JKN) solves the limitations of traditional fixed aggregation methods by proposing an adaptive layer aggregation framework for aggregating neighbor node information in heterogeneous graphs. Similar to traditional aggregation techniques, the hidden layer of each layer collects neighbor information from the previous layer; however, the output layer is processed differently. The information from each layer's hidden aggregation is passed to the operation layer, where various operations such as concatenation, max pooling, and layer attention are performed. After that, each node autonomously and adaptively integrates the aggregated features according to the structural details of its subgraph. This method ensures that the representation of each node is customized, thereby improving the model's feature extraction and learning capabilities. Figure 2 shows the structural flowchart of this method.
[0194] The DRFConv model is combined with the JKN framework, as Figure 1As shown by DMK-SGCN in [reference], to enhance node features, the DRFConv model is chosen because it improves feature representation in complex graph structures through multi-receptive fields and attention mechanisms, thus capturing long-range dependencies more effectively. Additionally, compared with the Graph Convolutional Network (GCN), it offers relatively greater flexibility.
[0195] For the k-layer aggregation model, the features of the nodes in the l-th layer are represented as:
[0196]
[0197] where σ(·) represents the activation function, W l represents the weight matrix of the l-th layer, aggregate(·) represents the specific aggregation method, which here represents the aggregation method of the dynamic receptive field convolution module, and N(v) represents all adjacent nodes of node v.
[0198] After each layer of the dynamic receptive field convolution module, the output features of this layer are saved:
[0199] {h v (1) ,h v (2) ,…,h v (l)} (31)
[0200] where h v (l) represents the node features extracted by the aggregation method of the dynamic receptive field convolution module after the l-th layer, and {·} is the concatenation operation;
[0201] By using an adaptive max pooling layer to aggregate the outputs of all layers, the final node representation is formed:
[0202] H final =AdaptiveMaxPool(h v (1) ,h v (2) ,...,h v (l) ) (32)
[0203] where H final is the final node representation, and AdaptiveMaxPool(·) is the adaptive max pooling operation.
[0204] The present invention proposes a method for constructing a weighted graph based on the ARAGraph and DMK-SGCN models. Experimental results from two scenarios show that the combination of ARAGraph and DMK-SGCN achieves the highest classification accuracy.
[0205] To verify the effectiveness of the proposed method, a case study was conducted in this section. The first case was the fault diagnosis of a gearbox at Southeast University, and the second case was the fault diagnosis of a gearbox at Xi'an Jiaotong University. ARAGM and DMK-SGCN were written in Python 3.7, and the experiments were carried out on the hardware of Windows 10 operating system, Pytorch 1.9 machine learning framework, Intel i7-8750H CPU, and GTX1050Ti GPU.
[0206] Gearbox dataset of Southeast University: The test bench consists of a motor, a motor controller, a planetary gearbox, a reduction gearbox, a brake, and a brake controller. The dataset was collected from the Figure 3 transmission system dynamic simulator (DDS) shown in. Two different working conditions of the rotational speed system load of 20HZ - 0V and 30HZ - 2V were studied. Each working condition contains 5 states of the bearing, namely bearing ball fault, inner ring fault, outer ring fault, inner and outer ring compound fault, and healthy state, and 5 states of the gear, namely tooth defect, broken tooth, tooth root wear, tooth surface wear, and healthy state. Different types of faults of the bearing and the gearbox are shown in Table 1. In this study, each healthy state was regarded as a fault type, so it became a 20-class classification task.
[0207] Table 1 Description of bearing and gearbox fault types
[0208]
[0209] In this experiment, the input data was processed by mean normalization. Subsequently, the original vibration signal was segmented using a sliding window of length 1024 to ensure that there was no overlap between samples during the segmentation process. Each fault type in this dataset contains 1,000 samples. To verify the effectiveness of the method in dealing with mixed faults, gear faults and bearing faults were combined to create a mixed dataset, which contains four gear fault types, four bearing fault types, and two healthy states under different working conditions. To create the training set and the test set, the dataset was divided according to the ratio of 80:20. The training set contains 16,000 samples, and the test set contains 4,000 samples. In addition, to verify the robustness of the model, random Gaussian noise was added to the samples in both the training set and the test set. These steps together constitute a comprehensive experimental dataset, laying a solid foundation for subsequent analysis and model evaluation.
[0210] In this experiment, 10 samples were taken to construct 1 graph, with the ε-radius set to 0 and the factor set to 100. Finally, 1600 graphs were constructed for training and 400 graphs for testing. It should be noted that the labels of the nodes in the associated graph are the same because the features of the nodes come from the same operating condition.
[0211] To demonstrate the effectiveness and advantages of the proposed method, it is necessary to conduct a comparative analysis with various deep learning techniques. The comparison methods include the standard GCN under different noise levels and three different graph construction methods (KNNGraph, PathGraph, and ARAGraph), and three different receptive field Chebyshev graph convolutional networks (ChebyNet), GraphSage, and simplified graph convolutional network (SGCN).
[0212] KNNGraph: In KNNGraph, the top-k nearest neighbors are found for each node, and the neighbors of node x i can be expressed as Equation (33):
[0213] Ne(x i ) = KNN(k', x i , Ψ) (33)
[0214] where KNN(·) returns the top-k nearest neighbors of node x i in the set ψ, k' is taken as 5; Ψ = [x i+1 , x i+2 ,..., x i+m , representing a subset of m samples, and Ne(x i ) represents the neighbors of node x i .
[0215] The edge weights between the nodes in KNNGraph can be estimated using the Gaussian kernel weight function, defined as:
[0216]
[0217] where e ij represents the edge weight value between node x i and node x j , and β represents the bandwidth of the Gaussian kernel.
[0218] PathGraph: In PathGraph, the samples in the subset ψ are sequentially connected, and there is an edge between every two adjacent samples. In addition, the edge weights can also be estimated by Equation (34).
[0219] According to the fault diagnosis accuracy results shown in Table 2, the proposed graph construction method ARAGraph and the network DMK-SGCN exhibit significant advantages under different noise levels, demonstrating their excellent performance. The accuracy of ARAGraph is always higher than that of all other models, including SGCN, GCN, GraphSage, and ChebyNet, especially prominent at high noise levels (such as -6dB and -8dB). As Figure 7 shown, the loss values of the three graph construction methods decrease with the increase in the number of iterations. Among them, the loss value of PathGraph is the highest, followed by KNNGraph, while ARAGraph maintains the lowest loss value. This indicates that ARAGraph can effectively capture the key features in the data, improve the accuracy of fault diagnosis, and exhibit excellent stability and robustness.
[0220] After combining DMK-SGCN with ARAGraph, it performs excellently under all noise levels, and the accuracy reaches or approaches 99% at 0dB, -2dB, and -4dB. Even at high noise levels (such as -6dB and -8dB), the accuracy of DMK-SGCN still remains above 97%. The clustering performance under different noise conditions is as Figure 4 and Figure 6 shown. It is obvious that the proposed method can always achieve strong clustering results, highlighting its ability to handle fault diagnosis tasks in complex environments. Compared with the traditional MLP model, the accuracy of various graph neural networks using ARAGraph is significantly improved under different noise levels, further confirming the potential and advantages of graph neural networks and advanced graph construction methods in fault diagnosis.
[0221] From Table 2 and Figure 5 it can be observed that as the order of the Chebyshev polynomial increases, the accuracy of ChebyNet improves, indicating that a larger receptive field size (K) corresponds to a higher accuracy. In addition, DMK-SGCN combines multiple receptive fields and assigns importance weights to each receptive field. This enables the aggregation of features extracted from neighbor nodes according to the given weights, thereby enhancing the information capture ability, which further clarifies the excellent performance of DMK-SGCN.
[0222] Table 2 Fault diagnosis results of the gearbox dataset of Southeast University with different graph construction methods and noises
[0223]
[0224] Gearbox dataset: As Figure 8As shown, the experimental setup includes a planetary gearbox, a parallel gearbox, a drive motor, a brake, and a controller. The motor used is a three-phase motor with a rated power of three horsepower, an operating voltage of 230V AC, and a frequency of 60 / 50HZ. To monitor vibration signals, two one-dimensional acceleration sensors are installed on the X-axis and Y-axis of the planetary gearbox, with a particular focus on capturing data in the Y direction.
[0225] In this setup, four different fault modes related to the bearings inside the gearbox are achieved: including gear wear, missing teeth, root cracks, and tooth fractures. In addition, several types of bearing faults are considered, specifically including faults involving spheres, inner rings, outer rings, and their combinations. Therefore, a total of nine different fault types are recorded, in addition to the baseline state, and the details are shown in Table 3. The motor runs at a speed of 1800 revolutions per minute (r / min), and the sampling frequency for data acquisition is 20,480HZ.
[0226] Table 3 Health status of the gearbox
[0227]
[0228] To further verify the superiority of ARAGraph and DMK-SGCN, the noise level is increased to test the performance of the models under more severe conditions. Subsequently, by testing the models at higher noise levels, their reliability in practical applications is comprehensively evaluated.
[0229] As shown in Table 4 and Figure 9 As shown, ARAGraph performs excellently among almost all models, especially in a high-noise environment. In all cases, whether it is SGCN, GraphSage, ChebyNet, or GCN, the accuracy of ARAGraph is better than the two graph structures considered, even at different noise levels. It is worth noting that in the DMK-SGCN model, even under the most severe noise conditions, the accuracy still remains above 93%, demonstrating its excellent anti-noise ability. In contrast, the performance of PathGraph and KNNGraph gradually decreases as the noise increases. In a -15dB noise environment, the accuracy of the ChebyNet family models increases as the receptive field expands. In particular, the improvement in accuracy of the ARAGraph graph structure is significantly higher than other combination methods. It is worth mentioning that the combination of ARAGraph and DMK-SGCN achieves an accuracy of 93.8341%, significantly higher than the ChebyNet family models. This indicates that ARAGraph can more effectively maintain the integrity of graph structure information in a high-noise environment. In addition, the dynamic receptive field mechanism of DMK-SGCN significantly enhances the expressive ability of feature extraction.
[0230] Figure 10 The confusion matrix in further demonstrates the superiority of ARAGraph. Regardless of the noise level, ARAGraph is always superior to other graph structures. Especially in the case of high noise, it can still distinguish various fault types with higher accuracy. These results indicate that the choice of graph structure in ARAGraph, the dynamic receptive field mechanism in DMK-SGCN, and the appropriate receptive field configuration all play important roles in enhancing the fault diagnosis ability in a noisy environment.
[0231] Table 4 Fault diagnosis results of the gearbox dataset of Xi'an Jiaotong University
[0232]
[0233]
[0234] Deep neural networks (DNNs) are superior to shallow neural networks (SNNs), mainly because they extract meaningful information from raw data by constructing approximate complex non-linear functions. However, when the number of network layers is too large, the graph convolution operation will fuse the features of neighbor nodes in each iteration. This process will extract node features containing information of a large number of nodes in the graph, resulting in a decrease in the identifiability of node features. This phenomenon is called "over-smoothing". To study the "over-smoothing" problem related to the GCN diagnosis model, the ChebyNet network was extended from a single layer to eight layers, and a fault diagnosis experiment was conducted using the aforementioned dataset. Figure 11 and Table 5 shows the experimental results conducted at a noise level of -12 dB.
[0235] Table 5 Diagnosis of different graph convolution layers of three models based on the KNN graph construction method under -12 dB noise
[0236]
[0237] In the evaluation of different numbers of layers, there are significant differences among ChebyNet_1 (receptive field k = 1), ChebyNet_2 (receptive field k = 2), and ChebyNet_3 (receptive field k = 3). The accuracy of ChebyNet_1 steadily increases as the number of layers increases, reaching a maximum value (0.8644) at six layers and then slightly decreasing. ChebyNet_2 remains stable in the first five layers, reaching the highest accuracy (0.9489) at the fifth layer, but the accuracy significantly decreases after the sixth layer, being 0.8633 at the eighth layer. ChebyNet_3 performs well in the first three layers, reaching a peak (0.9406) at the second layer, but the accuracy significantly decreases starting from the fourth layer, dropping to 0.5778 at the eighth layer. This indicates that ChebyNet_1 and ChebyNet_2 have smaller receptive fields and can maintain a high accuracy within a certain number of layer depths, but the accuracy will decrease as the number of layers increases. In contrast, ChebyNet_3 converges the features of too many neighbor nodes at deep layers, resulting in over-smoothing and a significant decrease in performance.
[0238] DMK-SGCN always performs excellently at different numbers of layers, maintaining a high accuracy. Especially at the seventh layer, the accuracy reaches the highest value of 0.9872. Although the accuracy slightly decreases to 0.9817 at the eighth layer, the model still demonstrates strong stability and robustness at higher numbers of layers. This stability is attributed to the multi-receptive field aggregation mechanism of DMK-SGCN, which effectively extracts and distributes feature information from different receptive fields, thus alleviating the common over-smoothing problem encountered by GNNs at deep layers.
[0239] The experimental results further show that the size of the receptive field (k) has a significant impact on the model performance. Comparing the performance of ChebyNet_1, ChebyNet_2, and ChebyNet_3, a smaller receptive field is more conducive to maintaining the distinguishability of features, and the model can maintain a high accuracy even when the number of layers increases. In contrast, although a larger receptive field (k = 3) can capture more extensive information, it is more prone to over-smoothing as the number of layers increases, resulting in a significant decrease in accuracy. This phenomenon highlights the importance of selecting an appropriate receptive field size and controlling the number of layers to avoid over-smoothing and enhance the fault diagnosis ability of the model. By adopting the multi-receptive field aggregation mechanism, DMK-SGCN successfully alleviates this problem, demonstrating the advantages of its design in improving performance.
[0240] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the scope of the present invention.
Claims
1. A bearing fault diagnosis method based on a dynamic multi-core jump graph convolutional network, characterized in that, The specific steps are as follows: Step 1: Construct a dynamic multi-core jumping graph convolutional network, and use an adaptive anti-noise graph module to construct graph data; Step 2: An adaptive robustness-aware graph module is integrated into the dynamic multi-core jumping graph convolutional network, which adaptively adjusts the connection weights between sample points according to the correlation between different frequency intervals, and focuses the attention on the region where the energy is concentrated in the spectrum; A dynamic receptive field convolutional module is integrated into the dynamic multi-core jumping graph convolutional network, which dynamically adjusts the size and weight of the convolutional kernel, so that each kernel extracts features of different scales according to the relationship between samples, thereby capturing global features and local features; The time series data is converted into anti-noise graph data through the adaptive robustness-aware graph module, the features are extracted and aggregated through the dynamic receptive field convolutional module, and finally classified through the fully connected layer; Step 3: Combine cross-layer jumping connections, retain and integrate the fine-grained features of the shallow layer and the abstract representations of the deep layer, and perform end-to-end fault diagnosis on the rolling bearing.
2. The bearing fault diagnosis method based on a dynamic multi-core jump graph convolutional network according to claim 1, wherein In Step 1, the specific steps are as follows: For an undirected graph G = (V, ξ, A), where V = n represents finite nodes, ξ represents the set of edges, and A ∈ R n×n represents the adjacency matrix of graph G. For two nodes <v i , v j > in the graph, the value of A ij is expressed as: As a kind of data in the non-Euclidean domain, a graph is represented by an adjacency matrix, and the Laplacian matrix of the graph is defined as: L = D - A (2) where \(L\) is the Laplacian matrix and \(D\in\mathbb{R}\) n×n is the degree matrix; The Laplacian matrix is symmetrically normalized and defined as: L sym = D -1 / 2 LD -1 / 2 = I N - D -1 / 2 AD -1 / 2 (3) Among them, L sym is the symmetric normalized Laplacian matrix, and I N is the identity matrix; The eigenvalue decomposition of the symmetrically normalized graph Laplacian matrix is: where, Λ = diag(λ1, λ2,..., λ n ) is a diagonal matrix composed of the eigenvalues of L sym , and U is an orthogonal matrix composed of the eigenvectors of L sym , for example, U -1 = U T ; At this time, the spectral graph convolution of node v and node features is defined as: h = (x * G f) θ = U(U T xU T f)(5) where x is the node feature, h represents the feature map after graph convolution, f is the eigenfunction of Λ, i.e., f(Λ), θ is the learnable parameter, * G represents graph convolution; Consider f θ = U T If f is a learnable graph convolutional filter, the above formula is simplified to: h = Uf θ U T x (6) Among them, U T x represents the graph Fourier transform of node feature x; Graph convolutional network: Filter f θ Approximate f by the truncated expansion of the Chebyshev polynomial proposed in the polynomial approximation θ , and the K-order approximation is as follows: where K is the maximum order of the Chebyshev polynomial, is the rescaled feature, λ max represents the largest eigenvalue of L, θ is the vector of polynomial coefficients, θ k ′∈R K is the vector of Chebyshev coefficients, is the Chebyshev polynomial of order k, determined by the following recurrence relation, i.e., T k (x) = 2xT k-1 (x) - T k-2 (x), T0(x) = 1 and T1(x) = x; After approximating the filter by Chebyshev polynomials, the graph convolution of node features X and the spectral domain filter f θ is mathematically defined as follows: Among them, is the rescaled Laplacian matrix; The graph convolution result is non-linearly activated by the activation function and expressed as: where σ is the activation function, and ReLU is commonly used as the activation function, representing the activated feature expression; The standard two-layer GCN model for node classification on the graph is expressed as: Z = softmax(h(ReLU(h(X, A)), A)) (10) Among them, X = {x1, x2,..., x n} is a set of node features, n is the number of nodes, and Z is a convolutional signal matrix, which is defined by applying softmax function of 3. The bearing fault diagnosis method based on a dynamic multi-core jump graph convolutional network according to claim 2, wherein In Step 2, the specific steps are as follows: The adaptive robustness-aware graph module is used to construct a weighted graph by evaluating the difference between two probability distributions. The specific steps are as follows: Data preprocessing: Normalize the original time series X with signal length L, and the normalized data is expressed as: X nol = normalize(X) (11) Among them, X nol is a standardized time series, and normalize(·) represents the normalization operation. Here, the maximum-minimum normalization is adopted; After data normalization, the normalized data is divided into a series of sub-samples with a predefined length or dimension d to construct a new data set: where M is the constructed sample set, represents a subsample, n is the total number of subsamples, floor represents rounding down, and d is the dimension of the subsample; Perform a fast Fourier transform on each sub-sample, and use the obtained spectrum after the transformation as a new sample to convert the time-domain signal into a frequency-domain signal and construct a spectrum sample set. The whole process is expressed as: Among them, FFT(·) is an operator; Normalize each spectral sample and convert it into a probability distribution; Molecule represents the amplitude of the frequency interval k, and the denominator is the sum of the amplitudes of all frequency intervals is the normalized spectral sample; By assigning labels to each sample, a labeled data set is obtained as: where y n is the label of the nth sample, and D is the labeled dataset; Construction of the adaptive perception graph module: Determine the number of nodes q and obtain the neighbors of the nodes using ε-radius and in the following manner: of: Among them, is 's neighborhood, ε is the selected radius, Return the set as the calculated node and the bidirectional KL divergence between; The cumulative distribution function of the probability distribution is used to calculate the high-energy region. The specific formula is as follows: Among them, F i (k) represents the cumulative probability of the i-th sample from frequency point 1 to k, represents the normalized energy ratio of the i-th sample at frequency point j, and d is the total number of frequency points in the spectrum; Determine k * , such that F i (k * ) ≥ τ, where (τ ∈ [0, 1]), and the formula is as follows: k * = min{k | F i (k) ≥ τ} (18) The high-energy region corresponds to the frequency range [1, k * ; Define the bidirectional KL divergence to measure the similarity between two subsequences. The formula is as follows: wherein, is a measure of the two-way divergence between the frequency spectrum distributions of two subsequences, is the value at point k, is the value at point k, and a and b represent the proportions of the high-energy region in the total energy; While obtaining the neighborhood set, calculate the weight between each two nodes through the threshold Gaussian kernel weight function: Among them, w ij is the weight value between nodes i and j, β is the bandwidth variance of the Gaussian function. When the neighbors of the node and the weights between their neighbors are found, a weighted graph defined mathematically is obtained. By repeating the steps shown in formulas (16)-(22), a weighted graph dataset G is generated; Dynamic receptive field convolutional module: Introduce an adaptive weight allocation mechanism, which adaptively allocates the weights of different-scale convolutions according to the input data: Among them, represents the transformation of the Chebyshev polynomial for the matrix , where is the rescaled Laplacian matrix, where λ max represents the largest eigenvalue of L sym , X represents the feature matrix of the input data, [·] represents the concatenation operator, l represents the number of receptive fields, K1, K2, K l represent the receptive field sizes, k1, k2, k l represent the polynomial orders, are learnable parameters, and H is the fused feature matrix; Use the adjacency matrix and the node features of the graph as the input of GCNs, and denote the weight parameter as W. Equation (23) is simplified to: Among them, where N is the number of samples, F is the number of sample features, W0, W1, W l is the weight matrix for each receptive field; The fused feature H is passed through one-dimensional convolution in two layers to first increase the number of channels and then compress the number of channels, and it is expressed as: H expanded = Conv1D(H concat , in_channels = K l , out_channels = C') (25) H compressed = Conv1D(H expanded , in_channels = C′, out_channels = K l ) (26) H expanded The output feature matrix for increasing the number of channels, H compresses The output feature matrix for increasing the number of channels, C' is the expanded number of channels, in_channels is the number of input channels, and out_channels is the number of output channels; The obtained H is pooled through an adaptive pooling layer to obtain K compressed one-dimensional vectors, which are denoted as: l H pool = AdaptiveAvgPool(H compressed ) (27) Among them, H pool is the output feature matrix after the adaptive pooling layer, and AdaptiveAvgPool(·) is the adaptive average pooling operation; Finally, an adaptive weight between vectors is obtained through the Softmax layer, and weighted summation is performed on different receptive fields, and it is expressed as: W = Softmax(H pool ) (28) Among them, in formula (28) in formula (29), N is the total number of nodes, and w i is the feature vector of the i-th node in W, and h i is the feature vector of the i-th node in H.
4. A bearing fault diagnosis method based on a dynamic multi-core jump graph convolutional network according to claim 3, characterized in that In step three, the specific steps are as follows: For the k-layer aggregation model, the features of the nodes in the l-th layer are represented as: where, σ(·) represents the activation function, W l represents the weight matrix of the l-th layer, aggregate(·) represents the specific aggregation method, and N(v) represents all adjacent nodes of node v; After each layer of the dynamic receptive field convolution module, save the output features of that layer: {h v (1) ,h v (2) ,...,h v (l)} (31) Among them, h v (l) represents the node features extracted by the aggregation method of the dynamic receptive field convolution module passing through the l-th layer, and {·} is the concatenation operation; Finally, aggregate the outputs of all layers by using an adaptive max pooling layer to form the final node representation: H final = AdaptiveMaxPool(h v (1) , h v (2) ,..., h v (l) ) (32) Among them, H final is the final node representation, and AdaptiveMaxPool(·) is the adaptive max pooling operation.