A Mechanical Fault Diagnosis Method Based on a Multi-Scale Graph Attention Fusion Network
The family graph and feature fusion layer are constructed through the multi-scale graph attention fusion network, which solves the problem of difficulty in capturing local and global feature information in existing methods, and realizes high-accuracy mechanical fault diagnosis under data imbalance and strong noise conditions.
Patent Information
- Application Number
- CN202310793938.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-06-30
AI Technical Summary
Existing mechanical fault diagnosis methods based on graph neural networks are difficult to effectively capture local and global feature information of data, and have poor ability to process limited and noisy fault vibration signals in industrial scenarios.
A multi-scale graph attention fusion network is adopted, and by constructing a family graph and a multi-scale feature fusion layer, the weights of adjacent nodes are automatically learned, the node feature representation is enhanced, and the model is optimized with the cross entropy loss function.
The accuracy of fault diagnosis is improved under data imbalance and strong noise conditions, and the robustness and generalization capabilities of the model are improved.
Smart Images

Figure CN116821762B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mechanical fault diagnosis, and particularly relates to a mechanical fault diagnosis method based on a multi-scale graph attention fusion network. Background Art
[0002] As a key component of modern manufacturing equipment, the safe operation of rotating machinery directly affects the stability of industrial systems. In rotating machinery, bearings will bear huge and changing loads, which will lead to failures. Statistics show that the failure rate of rotating machinery caused by bearings is as high as 40%. In order to reduce the losses of equipment during the production process and increase the stability of the equipment, bearing condition monitoring and fault diagnosis technologies have become the central issues of concern to scholars and engineers.
[0003] Currently, fault diagnosis methods can be divided into three categories: model-based diagnosis methods, signal processing-based diagnosis methods, and data-driven diagnosis methods. Model-based physical methods require the establishment of detailed mathematical models to describe the physical characteristics and fault modes of the system. Signal processing-based fault diagnosis methods mainly use time-frequency domain analysis, wavelet packet transform, envelope spectrum analysis, high-order statistic analysis and other technologies to analyze and process signals for fault diagnosis. These two methods rely too much on high-quality professional domain knowledge. In addition, for complex mechanical systems, it is very difficult to establish a mathematical model, which requires a large amount of computing and time costs and reduces the overall efficiency of fault diagnosis.
[0004] Fault diagnosis methods based on data-driven technologies do not require the establishment of complex mathematical models, but construct feature extraction models and classifiers, and use a large amount of historical data to optimize the models. This diagnosis method has good effects and can be applied to the fault diagnosis of complex mechanical equipment. In recent years, data-driven methods have become very popular in the field of fault diagnosis and have achieved many research results.
[0005] Among them, the method based on graph neural network has recently attracted wide attention because it can introduce the geometric relationship information between signals into the machine fault diagnosis model. The graph neural network can aggregate the features of the central node and adjacent nodes of the graph into enhanced node features, automatically learn different weights between nodes through the attention mechanism, and aggregate the information according to the learned weights, so as to expand the influence of useful information and ignore the influence of interference information, and realize node classification. Zhang et al. converted the collected acoustic signals into graphs with geometric structures and applied graph convolutional networks to achieve fault diagnosis of rolling bearings. Sun et al. used the proposed multi-channel residual network (MCRN) to extract weak features in the signals, generated signals and finite graph data of different scales through the autoencoder (AE) graph generation layer, and finally achieved fault diagnosis based on graph convolutional networks. Jiang et al. used the dynamic time warping method to convert the original vibration signals into graph data and first applied GAT to the field of fault diagnosis. Although the existing methods have achieved great success, the above-mentioned proposed methods still have great limitations: traditional fault diagnosis methods based on graph neural networks (GNNs) are difficult to effectively capture the local and global feature information of data, and most GNN models do not consider the inherent differences between adjacent nodes, and have poor ability to process limited and noisy fault vibration signals in industrial scenarios. Summary of the Invention
[0006] To solve the above technical problems, the present invention provides a mechanical fault diagnosis method based on a multi-scale graph attention fusion network to achieve the purpose of obtaining a higher accuracy fault diagnosis under data imbalance and strong noise conditions.
[0007] To achieve the above object, the technical solution of the present invention is as follows:
[0008] A mechanical fault diagnosis method based on a multi-scale graph attention fusion network, comprising the following steps:
[0009] Step 1, collect mechanical vibration data, construct a sample data set, and divide the data set into a training set and a test set;
[0010] Step 2, respectively construct the training set family graph and the test set family graph of the training set and the test set of the data set according to the cosine similarity between samples;
[0011] Step 3, input the training set family graph into the constructed multi-scale graph attention fusion network to train the network to obtain the best fault diagnosis model;
[0012] Step 4, input the test set family graph into the best fault diagnosis model for testing to evaluate the model performance;
[0013] Step 5: Input the mechanical vibration data to be diagnosed into the evaluated fault diagnosis model for fault diagnosis.
[0014] In the above solution, in Step 1, the collected sample signals are constructed into a sample data set together with the corresponding fault type labels.
[0015] In the above solution, the specific method of Step 2 is as follows:
[0016] Step1: Perform a normalization operation on the collected data;
[0017] Step2: Convert the normalized time series data into frequency domain data;
[0018] Step3: Select the n nodes with the highest similarity as the ancestor graph G in the family graph according to the similarity between samples. ancestor , secondly, randomly divide these n nodes into two groups, and use the cosine similarity method for each group to obtain their respective parent graphs G parent ; finally, randomly divide the n / 2 nodes of each parent graph G parent into two groups again, and use the cosine similarity method for each group to obtain their respective subgraphs G child ; G ancestor , G parent , G child together constitute the family graph G clan , and are used as the input of the model.
[0019] In the above solution, in Step 3, the multi-scale graph attention fusion network includes an input layer, a multi-scale feature fusion layer, and a node classification output layer. The multi-scale feature fusion layer includes three parallel graph attention modules with different attention scales and a feature fusion module. The node classification output layer includes two layers of fully connected layers, a LeakyReLU activation layer, and a Dropout layer.
[0020] In the above solution, the graph attention module includes a graph attention layer, a BatchNomlization layer, a Dropout layer, and a LeakyReLU activation layer.
[0021] In the above solution, the loss function of the multi-scale graph attention fusion network selects the cross-entropy loss function, and its formula is:
[0022] Loss(p,q) = -∑p(x)log q(x)
[0023] where p(x) is the label of the training set, and q(x) is the label value predicted by the network.
[0024] In the above solution, the output size of the input layer is 10×512×1, the output size of the multi-scale feature fusion layer is 10×1024×3, the output size of the node classification output layer is 10×512, and the Dropout layer ratio in the graph attention module and the node classification output layer is 0.6.
[0025] In the above solution, the output result of the multi-scale feature fusion layer is expressed as:
[0026]
[0027] Among them, [·] represents the concatenation operator; σ is the Sigmoid activation function; H1, H2, and H3 respectively represent that the graph attention layers of the three graph attention modules have H1, H2, and H3 independent attention mechanisms; represents the normalized attention coefficient calculated by node i and its neighbor node m under the h1, h2, and h3 attention mechanisms; respectively represent that the normalized attention coefficients are the corresponding linear transformation weighting matrices; represents the node feature of neighbor node m, and each (·) represents the feature representation from different scales; h1, h2, and h3 respectively represent the h1, h2, and h3 independent attention mechanisms of the graph attention layers of the three graph attention modules; represents the neighborhood of node i in the graph.
[0028] In the above solution, the output Y of the graph with node feature F in the multi-scale graph attention fusion network is expressed as: Y = FCL2(FCL1(Leaky_ReLU(MSFFL1(F)),
[0029] Leaky_ReLU(MSFFL2(F)), Leaky_ReLU(MSFFL3(F))))
[0030] Among them, FCL1 and FCL2 are the fully connected layers of the multi-scale graph attention fusion network, Leaky_ReLU is the activation function used in the multi-scale graph attention fusion network, and MSFFL1, MSFFL2, and MSFFL3 are the multi-scale feature fusion layers built in the multi-scale graph attention fusion network.
[0031] Through the above technical solution, a mechanical fault diagnosis method based on a multi-scale graph attention fusion network provided by the present invention has the following beneficial effects:
[0032] 1. Regarding the problem that existing data-driven graph neural network methods can only represent the geometric relationship information between signals and ignore the local and global feature information of graph-structured data, the present invention introduces a family graph. By converting data samples into family graphs with multiple information scales, the local and global information of graph-structured data can be effectively represented, breaking through the defect that the local and global feature information of graph-structured data cannot be expressed simultaneously, better characterizing the representative features of different fault types, and improving the feature extraction ability.
[0033] 2. Aiming at the problem that most existing GNN models do not consider the differences between adjacent nodes, the present invention designs a multi-scale graph attention fusion network (MSGAFN). By cooperating the designed multi-scale feature fusion layer with the family graph, the weights of adjacent nodes are automatically learned at multiple scales to represent their importance to the central node and reflect the differences between different adjacent nodes, making the learned node embedding representation more representative and distinguishable, maximizing the use of the connection between local information and the overall structure, fully mining the hidden connections in the network topology structure, enhancing the robustness of the model, and improving the fault diagnosis accuracy under data imbalance and strong noise conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art.
[0035] Figure 1 It is a schematic diagram of the overall process of a mechanical fault diagnosis method based on a multi-scale graph attention fusion network disclosed in the embodiments of the present invention;
[0036] Figure 2 It is the network structure diagram of MSGAFN;
[0037] Figure 3 It is the flowchart of family graph construction;
[0038] Figure 4 It is the structure diagram of the graph attention module;
[0039] Figure 5 It is the influence diagram of the number of attention heads in different scales on MSGAFN;
[0040] Figure 6 It is the influence of single-scale and multi-scale construction graphs on the model;
[0041] Figure 7 It is the result of the bearing data comparison test;
[0042] Figure 8 It is the result of the bearing and gear mixed data comparison test. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0044] The present invention provides a mechanical fault diagnosis method based on a multi-scale graph attention fusion network. This method is mainly used to determine whether there are faults in bearings and gears and the types of faults by detecting the collected vibration signals of bearings and gears. The overall flowchart of this method is as Figure 1 shown.
[0045] To verify the effect of the mechanical fault diagnosis method based on the MSGAFN network structure proposed by the present invention in the case of sample imbalance and strong noise scenarios for bearing and gear fault diagnosis, the data of two datasets will be used in this experiment, namely the gearbox dataset collected by Southeast University (SEU) and the 12KHz drive-end data collected by the Bearing Data Center of Case Western Reserve University (CWRU).
[0046] This method specifically includes the following steps:
[0047] S1: Select part of the data from the 12KHz drive-end dataset collected by the Bearing Data Center of Case Western Reserve University and the gearbox dataset collected by Southeast University.
[0048] Specifically, the collection of part of the 12KHz drive-end data collected by the Bearing Data Center of Case Western Reserve University used in Embodiment 1 includes the following steps:
[0049] Step1: Select the rolling bearing model SKF6205 at the motor drive end.
[0050] Step2: Use electric spark to process single-point damage on the rolling bearing at the motor drive end. The damage positions are the inner ring and the ball of the rolling bearing (the damage diameters are 0.007 inches, 0.014 inches, and 0.021 inches), and 12 o'clock on the outer ring of the rolling bearing.
[0051] It should be added that using electric spark single-point processing of rolling bearing fault damage can imitate the bearing damage in actual working conditions. Common bearing damage points are on the outer ring, inner ring, and ball of the rolling bearing, and there are usually differences in the damage diameters. Here, the damage diameters of the inner ring and the ball are set to 0.007 inches, 0.014 inches, and 0.021 inches. Thus, there are a total of 9 different types of rolling bearing faults. At the same time, there is also a type of normal rolling bearing without any loss. Therefore, there are a total of 10 types of rolling bearings.
[0052] Step 3: Place the vibration sensor above the bearing housing of the SKF6205 model at the motor drive end and connect it to a 16-channel data recorder to form a data acquisition system for collecting vibration signals of the rolling bearing.
[0053] Step 4: Apply the same single load to the same type of rolling bearings. The vibration signals of the rolling bearings at the motor drive end are collected by a 16-channel data recorder at a sampling frequency of 12 KHZ. In this embodiment, to verify that the present invention can be applied to different loads, multiple loads are set, namely 1HP and 2HP. Under each load, the vibration signals of 10 types of rolling bearings are collected.
[0054] Step 5: For the vibration signals of the 10 types of rolling bearings collected under each load, respectively divide the collected rolling bearing vibration signals into a series of sample signals according to the set sample signal length, and then the sample signals can be constructed into a sample data set together with the corresponding fault type labels.
[0055] It should be added that in this embodiment, the division is carried out according to each sample signal length of 1024 vibration points. After division and arrangement, in order to test in the unbalanced sample scenario, different numbers of different types of samples are set. Under each load, there are 80 sample signals for 9 different fault types of rolling bearings and 100 sample signals for normal rolling bearings. To simulate the actual industrial scenario, the sample signals under 2 loads are mixed, and the same fault type under different loads is regarded as one label to form the data set D bear For testing the model diagnosis effect in a strong noise environment, for D bear add Gaussian white noise with signal-to-noise ratios of -5, -6, -7, -8, -9, -10 dB to obtain the data set D bear-5 D bear-6 D bear-7 D bear-8 D bear-9 D bear-10 ; For the convenience of subsequent model training and testing, divide D bear-5 D bear-6 D bear-7 D bear-8 D bear-9 D bear-10 into a training set and a testing set according to the same ratio.
[0056] It should be added that for each load, 9 different fault types of rolling bearings obtain 720 sample signals, and normal rolling bearings obtain 100 sample signals. A total of 1640 rolling bearing vibration signals under two loads are mixed to obtain D bear, the specific division of the training set and test set of the dataset is shown in Table 1 as follows.
[0057] Table 1
[0058]
[0059] Some of the data collection of the gearbox dataset collected by Southeast University used in Example 2 includes the following steps:
[0060] Step1: Collect steady-state signals from the Dynamic Drive Simulator (DDS). Among them, bearing faults are caused by cracks in the inner ring, outer ring, rolling elements, etc. Gear faults are divided into four types: chipped teeth, missing teeth, root faults, and surface faults; chipped teeth and gear root cracks are both caused by cracks, but at different positions; missing tooth faults are caused by the absence of teeth; surface wear indicates wear on the gear surface.
[0061] The collected steady-state signals contain five health states of bearing data and gear data respectively: four fault states and one healthy state. Combining gear faults and bearing faults forms a 10-class mixed dataset including four gear faults, four bearing faults, and gears and bearings in the healthy state respectively.
[0062] Step2: Apply the same single load to the bearings and gears, and use the Dynamic Drive Simulator (DDS) for data collection. In this embodiment, in order to verify that the present invention can be applied to different loads, multiple loads are set in total, which are 20HZ - 0V and 30HZ - 2V respectively. Under each load, 10 types of mixed bearing and gear signals are collected.
[0063] Step3: For the 10 types of bearing and gear vibration signals collected under each load, respectively divide them into a series of sample signals according to the set sample signal length, and then the sample signals can be constructed into a sample dataset together with the corresponding fault type labels.
[0064] Specifically, in this embodiment, the division is carried out according to each sample signal length of 1024 vibration points. After segmentation and arrangement, in order to test in the unbalanced sample scenario, multiple quantities of different types of samples are set. Under each load, there are 80 sample signals for bearings and gears of 8 different fault types respectively, and 100 sample signals for bearings and gears of 2 normal types respectively. In order to simulate the actual industrial scenario, the sample signals under 2 loads are mixed, and the same fault types under different loads are regarded as one label, forming dataset D mix ; in order to test the model diagnosis effect in a strong noise environment, Gaussian white noise with signal-to-noise ratios of 0, -1, -2, -3, -4, -5dB is added to D mix to obtain dataset Dmix-0 , D mix-1 , D mix-2 , D mix-3 , D mix-4 , D mix-5 ; For the convenience of subsequent model training and testing, D mix-0 , D mix-1 , D mix-2 , D mix-3 , D mix-4 , D mix-5 is divided into a training set and a test set according to the same ratio.
[0065] It should be noted that 640 sample signals are obtained for bearings and gears of 8 different fault types for each load, and 200 sample signals are obtained for bearings and gears of 2 normal types. The vibration signals of a total of 1680 bearings and gears under the two loads are mixed to obtain D mix , and the specific dataset division of the training set and the test set is shown in Table 2.
[0066] Table 2
[0067]
[0068] S2: The training set and the test set of the dataset are respectively constructed into a family graph G clan-train , G clan-test ;
[0069] Specifically, the flow chart for constructing the family graph is as Figure 3 shown. The dataset for constructing the family graph G clan is used when training the neural network for diagnosing bearings D bear , and is used when training the neural network for diagnosing the mixture of gears and bearings D mix .
[0070] The method for constructing G clan includes the following steps:
[0071] Step1: Perform a normalization operation on the collected data;
[0072] It should be noted that the Min-Max normalization method is used when performing the normalization operation on the collected data.
[0073] Step2: Convert the collected time series data into frequency domain data;
[0074] It should be noted that the FFT (Fast Fourier Transform) method is used to convert the time series data into frequency domain data.
[0075] Step3: Select the 20 nodes with the highest similarity as the ancestor graph G in the family graph according to the similarity between samplesancestor , secondly, randomly divide these 20 nodes into two groups, and for each group, use the cosine similarity method to obtain their respective parent graphs G parent . Finally, randomly divide the parent graph into two groups again, and for each group, use the cosine similarity method to obtain their respective subgraphs G child . G ancestor , G parent , G child together constitute G clan , and are used as the input of the model.
[0076] It should be noted that the cosine similarity is used to estimate the distance between nodes when calculating the similarity between samples, and the formula is expressed as: where, is 's neighborhood, and ∈ is the selected radius length. Get the set and define the threshold as 0. If the similarity is greater than the threshold, there will be an edge e between two nodes i .
[0077] S3: Input G clan-train into the MSGAFN network model for training, and obtain an optimal rolling bearing fault diagnosis model after meeting the accuracy requirements.
[0078] Specifically, building the MSGAFN network model includes the following steps:
[0079] Step1: Design a multi-scale feature fusion layer, and use the constructed G clan as the input for multi-scale feature extraction and feature fusion. The multi-scale feature fusion layer includes three parallel graph attention modules with different attention scales and a feature fusion module. The multi-scale feature fusion layer performs multi-scale feature extraction on the input G clan , obtains the weights of the output features at three scales through the graph attention mechanism, multiplies the weights with their output features at each scale to obtain more representative features, denoted as Fea rep1 , Fea rep2 , Fea rep3 . Through the LeakyReLU activation function for Fea rep1 , Fea rep2 , Fea rep3 complete the non-linear feature transformation, and then the representative features are fused into an enhanced feature representation in the feature fusion module, denoted as Fea enhance , and use Fea enhance as the input of the subsequent network. The output result of node multi-scale feature fusion is expressed as:
[0080]
[0081] Among them, [·] represents a connection operator; σ is the Sigmoid activation function; H1, H2, and H3 respectively represent that the graph attention layers of the three graph attention modules have H1, H2, and H3 independent attention mechanisms; represents the normalized attention coefficient calculated by node i and its neighbor node m under the h1, h2, and h3 attention mechanisms; respectively represent that the normalized attention coefficient is the corresponding linearly transformed weighted matrix; represents the node feature of neighbor node m, and each (·) represents a feature representation from different scales; h1, h2, and h3 respectively represent the h1, h2, and h3 independent attention mechanisms of the graph attention layers of the three graph attention modules; represents the neighborhood of node i in the graph.
[0082] It should be supplemented that the three attention head numbers of the multi-scale feature fusion layer are set as: the values of H1, H2, and H3 are [4, 8, 16] respectively. The graph attention module consists of a graph attention layer, a BatchNomlization layer, a Dropout layer, and a LeakyReLU activation layer. As Figure 4 shown, where the graph attention layer calculates the attention weights between nodes to determine the importance of different nodes. The BatchNomlization layer accelerates training convergence, suppresses overfitting, and improves the performance of the model. The Dropout layer reduces overfitting and improves the robustness of the model by randomly discarding neurons. The LeakyReLU activation layer accelerates the convergence speed of the model and enhances the generalization ability of the model. The feature fusion module realizes the concatenation of the output results of the three attention modules to aggregate multi-scale features.
[0083] Step 2: Design the node classification output layer, which includes two fully connected layers, a LeakyReLU activation layer, and a Dropout layer. Set the loss function to complete the construction of the multi-scale graph attention fusion network.
[0084] It should be supplemented that the cross-entropy (Cross Entropy Loss) loss function is selected as the loss function, and its formula is:
[0085] Loss(p,q) = -∑p(x)log q(x)
[0086] Among them, p(x) is the label of the training set, q(x) is the label value predicted by the network, and the Dropout ratio of the Dropout layer is 0.6.
[0087] It should be noted that the MSGAFN network model proposed by the present invention can be used for fault diagnosis of rolling bearings in industrial scenarios where the fault vibration signals are limited and have strong noise, and can also be used for hybrid fault diagnosis of bearings and gears, with high robustness and generalization. The structure of this network will be described in detail below.
[0088] As Figure 2 shown, the MSGAFN network structure in the present invention is successively an input layer, a multi-scale feature fusion layer, and a node classification output layer. The input of the model input layer is G clan , and the output of the model output layer is the diagnostic result of the fault type.
[0089] In this MSGAFN network model, the specific network parameters of each layer can be determined through optimization. The parameters of each layer determined through optimization in the present invention are as follows:
[0090] The output size of the model input layer is 10×512×1, the output size of the multi-scale feature fusion layer is 10×1024×3, the output size of the node classification output layer is 10×512, and the Dropout layer ratio in the graph attention module and the node classification output layer is 0.6.
[0091] It should be noted that when using the sample data set to train the pre-constructed MSGAFN network model, it can be divided into a training set and a test set according to the conventional model training method. After the training set inputs the optimized parameters of the model, it is verified through the test set. After multiple trainings and optimizations, the optimal model parameters can be obtained.
[0092] Specifically, the model training specifically includes the following steps:
[0093] Step1: Adjust the network structure and parameters of the MSGAFN network model according to the training set results in S3;
[0094] Step2: Repeatedly perform the above process to obtain an optimal mechanical fault diagnosis model based on MSGAFN.
[0095] It should be noted that during the model training process, D bear and D mix train the model independently, and each forms an optimal fault diagnosis model based on MSGAFN.
[0096] Here, through repeated experiments, the specific network parameters of this layer are obtained as follows: the output size of the model input layer is 10×512×1, the output size of the multi-scale feature fusion layer is 10×1024×3, the output size of the node classification output layer is 10×512, the Dropout layer ratio in the graph attention module and the node classification output layer is 0.6, and the output Y of the graph with node feature F in the MSGAFN network structure can be expressed as:
[0097] Y = FCL2(FCL1(Leaky_ReLU(MSFFL1(F)),
[0098] Leaky_ReLU(MSFFL2(F)),Leaky_ReLU(MSFFL3(F))))
[0099] where FCL1 and FCL2 are the fully connected layers of the MSGAFN network model, Leaky_ReLU is the activation function used in the MSGAFN network model, and MSFFL1, MSFFL2, and MSFFL3 are the multi-scale feature fusion layers built in the MSGAFN network model.
[0100] S4: Input the previously divided test set into the best MSGAFN network model for testing and evaluate the model performance. This evaluation is divided into two forms, specifically as follows:
[0101] Step1: Use D bear-5 , D bear-6 , D bear-7 , D bear-8 , D bear-9 , D bear-10 Construct G clan-train according to S2 to train the MSGAFN network model and perform performance evaluation through the corresponding G clan-test ;
[0102] Step2: Use D mix-0 , D mix-1 [[ID=3S]], D mix-2 , D mix-3 , D mix-4 , D mix-5 Construct G clan-train according to S2 to train the MSGAFN network model and perform performance evaluation through the corresponding G clan-test ;
[0103] Step3: Compare the experimental results obtained from the processes described in Step1 - Step2 with other methods to prove that the method proposed in the present invention has relatively superior performance. The specific results are shown as follows:
[0104] (1) Network parameter selection
[0105] Different numbers of attention heads focus on different features. To study the impact of attention heads with different scales on the MSGAFN network model, we used the data of D mix-5 for experimental comparison. The accuracy rate and training time curve of the experimental results are as Figure 5 shown. As the scale of the number of attention heads increases, the accuracy of MSGAFN gradually increases. However, if the number of heads is too large, it will lead to excessive information fusion, which will instead affect the accuracy and training efficiency of the model. On the premise of ensuring the training efficiency, we finally choose 4-8-16 as the combination of the number of attention heads.
[0106] (2) Performance comparison
[0107] In actual industrial applications, the number of fault samples and normal samples of mechanical equipment is unbalanced, and there is a lot of noise interference. Therefore, constructing a family graph including multi-scale information can fully express the local and overall information of graph data, enhancing the generalization ability of the network model. The impact of single-scale and multi-scale constructed graphs on the accuracy rate of the MSGAFN network model is as Figure 6 shown.
[0108] The MSGAFN network model conducts experiments on bearing data and bearing-gear mixed data under different signal-to-noise ratio scenarios and compares them with GCN, ChebyNet, MRF_GCN, and MHGAT. The comparison test results of the bearing data in the CWRU dataset are shown in Table 3 and Figure 7 shown, and the comparison test results of the bearing-gear mixed data in the SEU dataset are shown in Table 4 and Figure 8 shown.
[0109] Table 3
[0110]
[0111] Table 4
[0112]
[0113]
[0114] The experimental results show that the MSGAFN network structure proposed by the present invention has a higher test accuracy rate than the comparison methods in both the CWRU dataset and the SEU dataset. The accuracy rate can still reach 84.83% in the strong noise scenario with a signal-to-noise ratio of -10dB.
[0115] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A mechanical fault diagnosis method based on a multi-scale graph attention fusion network, characterized in that It includes the following steps: Step 1: Collect mechanical vibration data, construct a sample data set, and divide the data set into a training set and a test set; Step 2: Respectively construct the training set family graph and the test set family graph from the training set and the test set of the data set according to the cosine similarity between samples; Step 3: Input the training set family graph into the constructed multi-scale graph attention fusion network to train the network and obtain the optimal fault diagnosis model; Step 4: Input the test set family graph into the optimal fault diagnosis model for testing to evaluate the model performance; Step 5: Input the mechanical vibration data to be diagnosed into the evaluated fault diagnosis model for fault diagnosis; The specific method of Step 2 is as follows: Step1: Perform normalization operation on the collected data; Step2: Convert the normalized time series data into frequency domain data; Step 3: Select the n nodes with the highest similarity as the ancestor graph G in the family graph according to the similarity between samples ancestor , secondly, randomly divide these n nodes into two groups, and for each group, use the cosine similarity method to obtain their respective parent graphs G parent ; finally, randomly divide the n / 2 nodes of each parent graph G parent into two groups again, and for each group, use the cosine similarity method to obtain their respective subgraphs G child ; G ancestor , G parent , G child together constitute the family graph G clan , and are used as the input of the model.
2. The mechanical fault diagnosis method based on a multi-scale graph attention fusion network according to claim 1, wherein, In Step 1, the collected sample signals and the corresponding fault type labels are constructed into a sample data set together.
3. A mechanical fault diagnosis method based on a multi-scale graph attention fusion network according to claim 1, characterized in that In Step 3, the multi-scale graph attention fusion network includes an input layer, a multi-scale feature fusion layer, and a node classification output layer. The multi-scale feature fusion layer includes three parallel graph attention modules with different attention scales and a feature fusion module. The node classification output layer includes two layers of fully connected layers, a LeakyReLU activation layer, and a Dropout layer.
4. A mechanical fault diagnosis method based on a multi-scale graph attention fusion network according to claim 3, wherein, The graph attention module includes a graph attention layer, a BatchNomlization layer, a Dropout layer, and a LeakyReLU activation layer.
5. A mechanical fault diagnosis method based on a multi-scale graph attention fusion network according to claim 1 or 3, characterized in that The loss function of the multi-scale graph attention fusion network selects the cross-entropy loss function, and its formula is: Loss(p,q)=-∑p(x)log q(x) where p(x) is the label of the training set, and q(x) is the label value predicted by the network.
6. The mechanical fault diagnosis method based on a multi-scale graph attention fusion network according to claim 4, wherein, The output size of the input layer is 10×512×1, the output size of the multi-scale feature fusion layer is 10×1024×3, the output size of the node classification output layer is 10×512, and the Dropout layer ratio in the graph attention module and the node classification output layer is 0.
6.
7. A mechanical fault diagnosis method based on a multi-scale graph attention fusion network according to claim 4, characterized in that The output result of the multi-scale feature fusion layer is expressed as: Among them, [·] represents the connection operator; σ is the Sigmoid activation function; H1, H2, and H3 respectively represent that the graph attention layers of the three graph attention modules have H1, H2, and H3 mutually independent attention mechanisms; represents the normalized attention coefficient calculated by node i and its neighborhood node m under the h1, h2, and h3 attention mechanisms; respectively represent that the normalized attention coefficient is the corresponding linear transformation weighted matrix; represents the node feature of neighborhood node m, and each (·) represents the feature representation from different scales; H1, H2, and H3 respectively represent the H1, H2, and h3 independent attention mechanisms of the graph attention layers of the three graph attention modules; represents the neighborhood of node i in the graph.
8. A mechanical fault diagnosis method based on a multi-scale graph attention fusion network according to claim 4, characterized in that, The output Y of the graph with node feature F in the multi-scale graph attention fusion network is expressed as: Y=FCL2(FCL1(Leaky_ReLU(MSFFL1(F)), Leaky_ReLU(MSFFL2(F)),Leaky_ReLU(MSFFL3(F)))) where FCL1 and FCL2 are the fully connected layers of the multi-scale graph attention fusion network, Leaky_ReLU is the activation function used in the multi-scale graph attention fusion network, and MSFFL1, MSFFL2, and MSFFL3 are the multi-scale feature fusion layers built by the multi-scale graph attention fusion network.
Citation Information
Patent Citations
Hydroelectric generating set fault diagnosis method and system based on multi-sensory domain GCN
CN115238739A