Mechanical fault detection method and device under variable speed working condition, medium and equipment

By adopting parallel multi-head attention mechanism layer and adaptive dynamic weight allocation in mechanical fault diagnosis under variable speed conditions, the limitations of single-scale feature extraction are solved, multi-scale feature capture and dynamic weight adjustment for mechanical faults are achieved, and the accuracy and robustness of fault detection are improved.

CN120067874AActive Publication Date: 2025-05-30HUNAN UNIV

Patent Information

Application Number
CN202510541441.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

In the prior art, in mechanical fault diagnosis under variable speed conditions, feature extraction is limited to a single scale, making it difficult to fully capture the complex fault characteristics of mechanical equipment, resulting in the need to improve detection accuracy.

Method used

The parallel single-head, double-head and triple-head attention mechanism layer architecture is adopted, combined with adaptive dynamic weight allocation, and multi-level graph structure information is extracted through multi-scale feature fusion and dynamic feature weighting fusion to enhance the ability to capture diverse fault features.

Benefits of technology

It realizes multi-particle size coverage for mechanical failures under variable speed conditions, improves the comprehensiveness and robustness of feature extraction, and enhances the accuracy of fault detection and the adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067874A_ABST
    Figure CN120067874A_ABST
Patent Text Reader

Abstract

The invention discloses a mechanical fault detection method and device under a variable speed working condition, a medium and equipment, and relates to the technical field of fault detection. The existing method is difficult to simultaneously and flexibly capture multi-scale features of data, and the detection effect is influenced. Therefore, the invention provides a mechanical fault detection method which comprises the following steps of: firstly, constructing a sparse threshold graph, matching variable speed data characteristics through flexible adjustment, and simultaneously, randomly deleting edges to enhance the robustness of a model; secondly, efficiently capturing graph structure information of different levels through an attention graph convolutional layer structure, and improving the expression ability of the model to complex graph data; and finally, dynamically adjusting weight distribution by calculating a loss ratio through an adaptive loss weight strategy, focusing fault types difficult to diagnose, and enhancing the fault detection performance of the model. The method can effectively adapt to speed change characteristics on a variable speed working condition data set, and shows higher detection accuracy and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fault detection, and particularly relates to a mechanical fault detection method, device, medium and equipment under variable speed conditions. Background Art

[0002] With the complication of the operating environment of industrial equipment, the data characteristics show significant localization and heterogeneity characteristics. Especially under variable speed conditions, different speed conditions will further exacerbate the diversity and complexity of fault types. This variable speed fault diagnosis plays an important role in timely discovering potential problems in equipment operation and improving the reliability and stability of industrial systems. Therefore, this has prompted many researchers to focus on mechanical fault diagnosis under variable speed condition data, providing potential directions for future research.

[0003] In recent research, GNN (Graph Neural Networks) has shown unique advantages in the field of mechanical fault diagnosis, especially in dealing with the monitoring data of equipment running at variable speeds. With its efficient graph construction ability for non-Euclidean data, this network architecture realizes the learning and fault detection of structured data, can capture the complex relationships and topological structure features between nodes, and provides new ideas and methods for fault diagnosis.

[0004] However, the feature extraction in GNN is often limited to a single scale. The single scale processes diverse fault patterns and variable speed conditions within a fixed range, which may cause the neglect of local details or global patterns of mechanical equipment, resulting in the need to improve the detection accuracy of the model for mechanical defects. Summary of the Invention

[0005] Based on this, in order to solve the technical problems in the prior art, the present invention provides a mechanical fault detection method, device, medium and equipment under variable speed conditions.

[0006] The present invention provides a mechanical fault detection method under variable speed conditions, including: Construct a neural network, the neural network includes a parallel single-head attention mechanism layer, double-head attention mechanism layer and triple-head attention mechanism layer, a single-head dynamic feature fusion layer, double-head dynamic feature fusion layer and triple-head dynamic feature fusion layer respectively connected to the output ends of the single-head, double-head and triple-head attention mechanism layers, and a fully connected layer connected to the output ends of the single-head, double-head and triple-head dynamic feature fusion layers at the same time; Obtain different types of variable speed fault graph data of mechanical equipment under variable speed conditions, and construct a data set based on the different types of variable speed fault graph data; use the data set to train the neural network to obtain a fault diagnosis model; Input the data of the variable-speed fault diagram to be detected into the fault diagnosis model. Through the single-attention head, double-attention heads, and triple-attention heads in the single-head attention mechanism layer, double-head attention mechanism layer, and triple-head attention mechanism layer respectively, extract features of different scales from the variable-speed fault diagram data to obtain a single-head fault feature map, a double-head fault feature map, and a triple-head fault feature map. Through the single-head dynamic feature fusion layer and double-head dynamic feature fusion layer respectively, perform adaptive feature weighted fusion operations on the single-head fault feature map, double-head fault feature map, and triple-head fault feature map to obtain a single-head fault weighted feature map, a double-head fault weighted feature map, and a triple-head fault weighted feature map. Through the fully connected layer, splice the single-head fault weighted feature map, double-head fault weighted feature map, and triple-head fault weighted feature map and then perform probability prediction to output the fault type label corresponding to the variable-speed fault diagram data to be detected.

[0007] Further, the extracting features of different scales from the variable-speed fault diagram data specifically includes: ; Among them, is the feature of node . is the linear transformation matrix of the k th attention head; is the attention weight of node k output by the th attention head; represents that node belongs to the set of neighbor nodes of node ; K is the number of attention heads, K = {1, 2, 3}; is the feature of node generated after being processed by the attention mechanism.

[0008] Further, the adaptive feature weighted fusion operation specifically includes: ; Among them, is the feature weight of the k th attention head, is the fault weighted feature map.

[0009] Further, the constructing a data set based on different types of variable-speed fault diagram data specifically includes: Calculate the Euclidean distance between each pair of nodes in the fault diagram data, set a threshold according to the distribution of the Euclidean distances between each pair of nodes, and retain the edges between the nodes whose Euclidean distance is less than or equal to the threshold to obtain the basic graph structure: ; ; ;

[0010] Among them, and are the eigenvalues of nodes and nodes in the m dimension; n is the dimension of the eigenvector; is the Euclidean distance between nodes and nodes ; Percentile (·) represents the quantile threshold function, D represents the set of Euclidean distances between all node pairs, which is the universal set; q represents the percentile; is the adjacency matrix in the basic graph structure; Set the deletion ratio p , and randomly delete some edges in the basic graph structure according to the deletion ratio p to obtain a sparsified graph structure: ; Among them, is the adjacency matrix of the sparsified graph structure; Add the corresponding fault type labels to each sparsified graph structure to obtain a dataset.

[0011] Furthermore, the training of the neural network using the dataset specifically includes: Use the cross-entropy loss function to calculate the initial loss between the predicted fault type label and the true fault type label output by the neural network; Dynamically calculate the loss ratio of different types of faults based on the ALW strategy, and introduce a guiding factor to adaptively adjust the weight allocation of each fault type to obtain the total loss: ; ; ; ;

[0012] Among them, C is the guiding factor; N is the total number of fault types; is the epoch at the t moment of the training stage; is the maximum epoch; is the fault type label, ={1, 2, 3..., N}; is the initial loss of the t th type of fault at time during the training phase; is the initial loss of the t- 1st type of fault at time during the training phase; is the training speed of the t th type of fault at time during the training phase; is the loss weight of the t th type of fault at time during the training phase; is the total loss at time t during the training phase: Taking the minimum total loss as the optimization objective of the optimizer, the optimizer is used to iteratively update the parameters in the neural network.

[0013] Further, the probability prediction after splicing the single - head fault weighted feature map, double - head fault weighted feature map and triple - head fault weighted feature map specifically includes: Splice the single - head fault weighted feature map, double - head fault weighted feature map and triple - head fault weighted feature map to form a spliced feature vector; After weighting the spliced feature vector, perform non - linear processing through the ReLU activation function to obtain a non - linear weighted spliced feature vector; Map the non - linear weighted spliced feature vector to a fault category probability distribution through the Softmax function, and take the fault category with the highest probability as the output.

[0014] Further, the obtaining of different types of variable - speed fault map data of mechanical equipment under variable - speed working conditions specifically includes: Collect the fault vibration sensor data of mechanical equipment under variable - speed working conditions; Change the fault vibration sensor data from the time domain to the frequency domain. In the frequency domain, taking each sampling point as a node, using the corresponding frequency - domain eigenvalue of each sampling point as the point feature, and constructing edges between adjacent sampling points to obtain the fault map data of mechanical equipment under variable - speed working conditions.

[0015] The present invention provides a mechanical fault detection device under variable - speed working conditions, including: A model construction module for constructing a neural network, where the neural network includes parallel single - head attention mechanism layers, double - head attention mechanism layers and triple - head attention mechanism layers, single - head dynamic feature fusion layers, double - head dynamic feature fusion layers and triple - head dynamic feature fusion layers that are respectively connected to the output ends of the single - head, double - head and triple - head attention mechanism layers, and a fully - connected layer that is simultaneously connected to the output ends of the single - head, double - head and triple - head dynamic feature fusion layers; A model training module, configured to obtain different types of variable-speed fault map data of mechanical equipment under variable-speed working conditions, construct a data set based on the different types of variable-speed fault map data; use the data set to train a neural network to obtain a fault diagnosis model. A fault detection module, configured to input the variable-speed fault map data to be detected into the fault diagnosis model, and respectively perform feature extraction of different scales on the variable-speed fault map data through single attention heads, double attention heads, and triple attention heads in the single-head attention mechanism layer, double-head attention mechanism layer, and triple-head attention mechanism layer to obtain a single-head fault feature map, a double-head fault feature map, and a triple-head fault feature map; respectively perform adaptive feature weighted fusion operations on the single-head fault feature map, double-head fault feature map, and triple-head fault feature map through a single-head dynamic feature fusion layer and a double-head dynamic feature fusion layer to obtain a single-head fault weighted feature map, a double-head fault weighted feature map, and a triple-head fault weighted feature map; splice the single-head fault weighted feature map, double-head fault weighted feature map, and triple-head fault weighted feature map through a fully connected layer and perform probability prediction to output a fault type label corresponding to the variable-speed fault map data to be detected.

[0016] The present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the mechanical fault detection method under the above variable-speed working conditions is implemented.

[0017] The present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the mechanical fault detection method under the above variable-speed working conditions is implemented.

[0018] At least one of the above technical solutions adopted by the present invention can achieve the following beneficial effects: In the mechanical fault detection method under variable-speed working conditions provided by the present invention, through the parallel single-head, double-head, and triple-head attention mechanism layer architecture and adaptive dynamic weight allocation, the model can effectively cope with the non-stationarity problem of signals under variable-speed working conditions: the multi-scale network architecture can extract information of different time or frequency scales, enhance the adaptability to complex working conditions, achieve multi-granularity coverage of fault features, and improve the comprehensiveness and robustness of feature extraction; while the adaptive weight allocation can dynamically adjust the importance of different-scale features, make key features more prominent, optimize feature selection, and thus improve the accuracy of fault detection. The combination of the two enables the model to dynamically adjust the importance of information at each scale according to the current working condition, which helps the model to more accurately identify fault features under variable-speed working conditions and reduce the probability of false detection and missed detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings described herein are used to provide a further understanding of the present invention, and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 It is a schematic flow diagram of a mechanical fault detection method under variable speed conditions provided by the present invention; Figure 2 It is a schematic flow diagram of the GATv2 weighted feature aggregation provided by the present invention; Figure 3 It is a schematic diagram of the MS-GATv2 model structure provided by the present invention; Figure 4 It is a schematic diagram of the data acquisition device provided by the present invention, Figure 4 in which (a) is a bearing fault simulation test bench, Figure 4 in which (b) is a gearbox fault simulation test bench; Figure 5 It is a schematic diagram of the visualization of multi-state vibration signals under variable speed conditions provided by the present invention, Figure 5 in which (a1) is a normal bearing waveform, Figure 5 in which (b1) is a normal gearbox waveform, Figure 5 in which (a2) is a bearing rolling element defect waveform, Figure 5 in which (b2) is a gear tooth root crack waveform, Figure 5 in which (a3) is a bearing compound defect waveform, Figure 5 in which (b3) is a gear tooth surface wear waveform, Figure 5 in which (a4) is a bearing inner ring defect waveform, Figure 5 in which (b4) is a gear tooth missing waveform, Figure 5 in which (a5) is a bearing outer ring defect waveform, Figure 5 in which (b5) is a gear tooth breakage waveform; Figure 6 It is a schematic flow diagram of the dynamic graph data construction provided by the present invention; Figure 7 It is a schematic flow diagram of the ALW strategy provided by the present invention; Figure 8 It is the construction method and its corresponding confusion matrix results based on different graph structures provided by the present invention, Figure 8 in which (a) is the STG method, Figure 8 in which (b) is the FCG method, Figure 8 in which (c) is the KCG method; Figure 9 It is a schematic diagram of the influence of different p values on the model accuracy provided by the present invention, Figure 9 in which (a) is the result of Case 1, Figure 9 in which (b) is the result of Case 2; Figure 10 The two-dimensional visualization results of the encoding features of different models provided by the present invention Figure 10 Among them, (a) is the two-dimensional visualization result of the encoding features of the SAMS-GAT model Figure 10 Among them, (b) is the two-dimensional visualization result of the encoding features of the Chebynet model Figure 10 Among them, (c) is the two-dimensional visualization result of the encoding features of the GCN model Figure 10 Among them, (d) is the two-dimensional visualization result of the encoding features of the SGCN model Figure 10 Among them, (e) is the two-dimensional visualization result of the encoding features of the GAT model Figure 10 Among them, (f) is the two-dimensional visualization result of the encoding features of the GIN model Figure 10 Among them, (g) is the two-dimensional visualization result of the encoding features of the GraphSage model Figure 10 Among them, (h) is the two-dimensional visualization result of the encoding features of the HO-GNN model Figure 11 The schematic diagram of the ALW strategy update process provided by the present invention Figure 11 Among them, (a) is the ALW strategy update process of case1 Figure 11 Among them, (b) is the ALW strategy update process of case2 Figure 12 The schematic diagram of the comparison of the test accuracies of the ALW, SWA, and RWU methods provided by the present invention Figure 12 Among them, (a) is the ALW strategy update process of case1 Figure 12 Among them, (b) is the ALW strategy update process of case2 Detailed implementation manners

[0020] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention

[0021] With the complication of the operating environment of industrial equipment, the data characteristics show significant localization and heterogeneity characteristics. Especially under variable-speed working conditions, different speed conditions will further exacerbate the diversity and complexity of fault types. This variable-speed fault diagnosis plays an important role in timely detecting potential problems in equipment operation and improving the reliability and stability of industrial systems. Therefore, this has prompted many researchers to focus on mechanical fault diagnosis under variable-speed working condition data, providing potential directions for future research

[0022] Since traditional methods usually build models based on fixed operating condition data, it is difficult to flexibly handle the dynamic characteristics and diverse fault types brought about by variable speed operating conditions. In recent research, GNN has received increasing attention in the field of variable speed operating condition fault diagnosis. With its efficient graph construction ability for non-Euclidean data, this architecture has achieved the learning and fault detection of structured data, can capture the complex relationships and topological structure features between nodes, and provides new ideas and methods for fault diagnosis. Yuan et al. integrated fault data with the physical structure and dynamic characteristic information of bearings into a unified graph model, and through combining node features, edge relationships and physical constraints, realized the comprehensive modeling and diagnosis of time-varying bearing faults. Yan et al. proposed a method combining label propagation strategy and dynamic graph attention network, effectively utilized the label dependence between samples, fully mined the limited label information, and dynamically extracted the feature information of different adjacent nodes under speed fluctuations. Liang et al. proposed an end-to-end semi-supervised subdomain adaptive graph convolutional network method for extracting transferable features, reducing the distribution difference between the source domain and the target domain, and overcoming the influence of speed fluctuations. Li et al. proposed a semi-supervised fault diagnosis method of multi-head dynamic graph attention network, and by introducing dynamic time warping similarity and pseudo-label mechanism, effectively solved problems such as time shift and insufficient utilization of neighborhood information.

[0023] In fault diagnosis tasks, there are often significant differences in the complexity and optimization difficulty of different fault types. Traditional loss functions usually adopt fixed weight allocation strategies and are difficult to dynamically adapt to these differences. In recent years, fault diagnosis methods based on dynamic loss function updates have gradually become a research hotspot. Zhang et al. proposed dynamic metric learning and an improved triplet loss function to more accurately simulate the complex relationships between vibration signals, generate an adjacency matrix that can reflect the similarity and correlation between nodes, thereby enhancing the robustness and expressive ability of graph structure construction. Yan et al. designed a dynamic weight optimization strategy, by adaptively adjusting the data reconstruction loss weights under different operating conditions, flexibly balancing the contributions of various types of data, aiming to guide the model to more effectively learn generalization features with wide applicability. Wang et al. proposed a dynamic spectrum loss to effectively break through the synthesis bottleneck of high-difficulty frequency components in data. Based on spectral distance, this method reduces the frequency weights of components that are easy to synthesize in the spectrum by introducing a dynamic weight matrix, and at the same time adaptively enhances the attention to frequency components that are difficult to synthesize. The loss function designed by Zhou et al. does not require the introduction of additional adjustable parameters. With its inherent dynamic mechanism, while maintaining the simplicity of the model, it effectively improves the prediction accuracy and the progressiveness of the training process, showing strong adaptability and robustness.

[0024] The fault diagnosis research under the above variable-speed working conditions faces three major problems, which significantly restrict the performance and robustness of the method under complex working conditions. First, the traditional method of fixed graph data is difficult to adapt to the changes in dynamic working conditions and is hard to fully capture the dynamic correlations and complex topological characteristics between nodes. Especially in scenarios with diverse working conditions and speed fluctuations, the fixed graph structure is prone to weakening the expression ability of system characteristics due to partial information loss, while dynamic graph data can effectively compensate for the impact caused by information loss through dynamic adjustment and reconstruction of the correlation structure. Second, feature extraction is often limited to a single scale and lacks the ability to comprehensively utilize multi-scale information. This limitation is particularly obvious when dealing with diverse fault modes or complex dynamic working conditions, which may lead to the neglect of local details or global patterns, further restricting the generalization performance and adaptability of the model. In addition, the single-scale strategy is also difficult to handle the diversity of fault characteristics under different working conditions and cannot provide comprehensive information support for the model. Finally, the traditional loss function calculates with equal attention to all fault types, ignoring the significant differences in the diagnostic difficulty of different fault types and making it difficult to dynamically adjust the attention to high-difficulty fault types or key features. In the case of complex working conditions and variable data distributions, this strategy not only weakens the learning ability of key features but also limits the processing effect of the model on unbalanced data, further affecting the accuracy and robustness of diagnosis. Therefore, in order to better cope with the diverse challenges under variable-speed working conditions, it is necessary to dynamically capture the correlation characteristics between nodes by introducing a dynamic graph structure, adopt a multi-scale feature fusion strategy to comprehensively extract key information, and at the same time combine a dynamic optimization mechanism of loss values to flexibly adjust the focus of the model, so as to comprehensively improve the diagnostic ability and robustness for complex working conditions and diverse faults.

[0025] Based on this, the present invention proposes a new fault detection method aimed at coping with the complexity of data under variable-speed working conditions and improving the detection accuracy and model robustness. The main innovations of this method can be summarized into the following three aspects: (1) A sparse threshold graph is constructed to flexibly capture the local geometric structure of variable-speed data and enhance the adaptability of the graph model to the dynamic environment; (2) A multi-scale attention graph convolutional layer is designed to effectively extract multi-level graph structure information, achieve multi-scale feature fusion, and improve the ability to capture diverse fault features; (3) An adaptive loss weight strategy is developed to dynamically adjust the loss weights to judge the diagnostic difficulty of different fault types and strengthen the model's attention and optimization ability for complex faults. Experimental results show that this method exhibits excellent performance on variable-speed working conditions and complex data sets, can effectively adapt to the data distribution characteristics under different speed changes, and demonstrates higher detection accuracy and robustness.

[0026] The following will detail the technical solutions provided by each embodiment of the present application in conjunction with the accompanying drawings.

[0027] Figure 1Schematic diagram of the process of a mechanical fault detection method under variable speed conditions in this embodiment, specifically including the following steps: S1: Construct a neural network, which includes a parallel single-head attention mechanism layer, a double-head attention mechanism layer, and a triple-head attention mechanism layer, a single-head dynamic feature fusion layer, a double-head dynamic feature fusion layer, and a triple-head dynamic feature fusion layer respectively connected to the output ends of the single-head, double-head, and triple-head attention mechanism layers, and a fully connected layer connected to the output ends of the single-head, double-head, and triple-head dynamic feature fusion layers simultaneously.

[0028] Traditional GNNs are difficult to capture local and global features simultaneously when dealing with complex graph data, and have limited expressive ability for multi-scale structures. In addition, its feature fusion method lacks flexibility, is difficult to adapt to dynamic working conditions and non-uniform data distributions, and has insufficient robustness to abnormal data. To solve these problems, the present invention uses MS-GATv2 (Multi-Scale GraphAttenTion convolution v2) to efficiently capture graph features of different scales, and uses the DWU (Dynamic Weight Update) method to improve the expressive ability and robustness of the model under complex working conditions, so as to better meet the needs of complex graph data processing.

[0029] Before introducing MS-GATv2, it is necessary to understand the GATv2 network model. GATv2 is an improved model of the traditional GAT (GraphAttenTion network). By introducing a dynamic attention mechanism and non-linear feature interaction, it significantly improves the expressive ability of node relationship modeling. The traditional GAT uses a static linear transformation to calculate attention weights, which limits the modeling ability of complex non-linear feature interactions. GATv2, on the other hand, dynamically calculates the attention weights between nodes through an MLP (Multi LayerPerceptron), so as to be able to capture more complex relationships between nodes. Specifically, GATv2 can not only flexibly adjust the weight allocation according to the dynamic interaction relationship of node features, but also more efficiently combine the features of neighbor nodes to construct high-quality node representations through weighted aggregation. The key calculation process of GATv2 includes the following steps:

[0030] (1) The attention score i between node j and node e ij is calculated by the following formula: ; where and are the features of node i and nodej Features; || represents feature splicing; MLP represents a multi-layer perceptron, which is used for dynamic modeling of the feature relationships between nodes.

[0031] (2), Use the Softmax function to normalize the attention scores of the neighbor nodes of node to obtain the attention weights : ; Among them, represents the set of neighbor nodes of node .

[0032] (3), Achieve feature aggregation through weighted summation, and combine the linear transformation matrix W and the activation function to complete the update of node features. The GATv2 weighted feature aggregation process is as Figure 2 shown, where i = 1, j = {1, 2, 3, 4, 5, 6}: ; Among them, represents the activation function; W is a learnable linear transformation matrix.

[0033] The MS-GATv2 network structure constructed by the present invention is as Figure 3 shown. This network is an innovative multi-scale graph convolution method. By combining the multi-head attention mechanisms with different numbers of heads (Head = 1, Head = 2, Head = 3), it effectively captures the feature information at different levels in the graph structure. For each number of heads within the same neighborhood range, the feature relationships between nodes are independently modeled through different attention mechanisms, and the results of each head are averaged when output. Head = 1 provides a single attention allocation to capture the basic feature relationships. Head = 2 adds an independent head to make the features more abundant. Head = 3 further enhances the diversity and robustness of feature learning through the multi-head mechanism. As Figure 3 shown, the model consists of MS-GATv2, DFF layer (Dynamic feature fusion layer), and FC layer (Fully Connected). The feature interaction relationship between nodes and their neighborhoods is dynamically adjusted through the multi-head attention mechanism. Each head independently calculates the attention weights and feature updates to ensure that the feature information at different scales is fully expressed. The GATv2 feature calculations for different numbers of heads are as follows: ; Among them, is a node feature; is the k linear transformation matrix of the is the k node output by the attention head; represents that the node belongs to the set of neighbor nodes of the node ; K is the number of attention heads, K ={1, 2, 3}; is the node feature generated after processing by the attention mechanism. Through the parallel single-head, double-head, and triple-head attention mechanism layer architectures, the model can integrate information from different attention branches to form a more comprehensive and multi-scale fault feature representation. This fused feature representation is more conducive to making accurate fault detection decisions. Specifically: The single-head attention mechanism captures local detailed information and can focus on the specific location and subtle features of the fault occurrence; the double-head attention mechanism balances local and slightly larger-scale features, helping to understand the propagation and impact of the fault within the local area; the triple-head attention mechanism captures information from a more global perspective, achieving multi-granularity coverage of the fault features; each attention mechanism branch has an independent adaptive weight allocation mechanism, which means that the model can dynamically adjust the importance of information at each scale according to the current working conditions, helping the model to more accurately identify fault features under variable speed working conditions.

[0034] After obtaining multi-head attention features with different numbers, the DWU strategy is used to fuse features at different scales. This strategy can adaptively adjust the contribution ratio of different features, enabling the model to flexibly allocate attention according to the task requirements. By dynamically learning weights, the model can highlight the most important feature scales while suppressing irrelevant or redundant information, thereby enhancing the effectiveness of feature representation. In addition, dynamic weight fusion can integrate multi-scale information on the basis of the multi-head attention mechanism, not only retaining the diversity of features but also optimizing the balance of feature fusion, greatly enhancing the robustness and generalization ability of the model in complex data structures. The calculation process is as follows: ; where, is the feature weight of the k attention head, is the fault-weighted feature map. After passing the result output by MS-GATv2 through the batch normalization BN operation and non-linearly activating it with the ReLU function, it is fused with the result output by the DWU module to obtain the output of the DFF layer.

[0035] Finally, the single - head fault weighted feature map, double - head fault weighted feature map, and triple - head fault weighted feature map are concatenated to form the final feature vector as the input. Then, the input features undergo a linear transformation, that is, weighted summation through a trainable weight matrix and bias term to extract deeper features. Subsequently, through a non - linear activation function (ReLU), non - linear capabilities are introduced to enhance the model's expressive ability. Finally, the output layer calculates the probability distribution of each category through Softmax to achieve the classification prediction of mechanical faults.

[0036] S2: Obtain different types of variable - speed fault map data of mechanical equipment under variable - speed working conditions, and construct a data set based on the different types of variable - speed fault map data. Use the data set to train a neural network to obtain a fault diagnosis model.

[0037] S21: Construct a data set.

[0038] S211: Obtain the original fault data set.

[0039] Case 1: Data set 1 comes from the mechanical bearing fault test bench of the University of Ottawa. As shown in Figure 4 (a), it includes a motor, an encoder, a healthy bearing, an accelerometer, a test bearing, and an AC drive. The data set contains a total of 5 bearing states, including normal state, inner - race defect, outer - race defect, ball defect, and a combined defect of the inner - race, outer - race, and ball. During the experiment, the vibration signals during two deceleration processes were collected respectively, and the acceleration under the speed fluctuation of 14.1 Hz - 23.8 Hz was sampled and analyzed, with a sampling frequency of 200 kHz.

[0040] Case2: Data set 2 comes from the fault simulation test bench of the drive system of the rail transit train bogie of Beijing Jiaotong University. As shown in Figure 4 (b), it includes a motor, a gearbox, and an accelerometer. The data set contains a total of 9 gearbox states, including normal, tooth - root crack, tooth - surface wear, tooth missing, tooth breakage, inner - race fault of the bearing, outer - race fault of the bearing, rolling - element fault of the bearing, and cage fault of the bearing. During the experiment, the vibration data at rotational speeds of 20 Hz, 40 Hz, and 60 Hz were collected for analysis, with a sampling frequency of 64 kHz.

[0041] For the above two datasets, the samples used in the present invention are shown in Table 1. In Case 1, the time-varying speed of the test bench increases from 14.1 Hz to 23.8 Hz. There are 50 training samples and 50 test samples for each state, and the length of each sample is 4096. In Case 2, the variable-speed samples for each state of the test bench are 20 Hz, 40 Hz, and 60 Hz respectively. There are 30 training samples and 30 test samples with different speeds for each state. The three speeds are spliced into variable-speed data. Therefore, the total number of training samples and test samples for each state is 90, and the length of each sample is also 4096. Figure 5 shows the differences in time-domain waveform features of two variable-speed datasets under five equipment operating states, including the normal state and multiple fault states. Figure 5 In (a1) of [Figure / Table] [X], it is the normal waveform of the bearing. Figure 5 In (b1) of [Figure / Table] [X], it is the normal waveform of the gearbox. Figure 5 In (a2) of [Figure / Table] [X], it is the waveform of the bearing rolling element defect. Figure 5 In (b2) of [Figure / Table] [X], it is the waveform of the gear tooth root crack. Figure 5 In (a3) of [Figure / Table] [X], it is the waveform of the bearing compound defect. Figure 5 In (b3) of [Figure / Table] [X], it is the waveform of the gear tooth surface wear. Figure 5 In (a4) of [Figure / Table] [X], it is the waveform of the bearing inner ring defect. Figure 5 In (b4) of [Figure / Table] [X], it is the waveform of the gear tooth missing. Figure 5 In (a5) of [Figure / Table] [X], it is the waveform of the bearing outer ring defect. Figure 5 In (b5) of [Figure / Table] [X], it is the waveform of the gear tooth breakage. It can be seen from the figure that with the change of the equipment state, the complexity of the waveform features increases significantly: the waveform is stable under the normal state, and the amplitude change is small; while in the fault state, the amplitude fluctuation increases, and the irregularity and suddenness of the signal increase significantly. Especially in complex fault states such as combined defects and tooth breakage, highly nonlinear and multi-band coupling characteristics are exhibited. These significant complexities increase the difficulty of fault diagnosis based on time-domain signals, and at the same time reflect the diversity and challenges of variable-speed equipment fault data.

[0042] Table 1 Usage of experimental samples

[0043] In the actual scenario, the graph data often has the problem of redundant connections due to noise interference. This not only increases the computational burden, but also may mask the key relationships between nodes and reduce the learning effect of the model. In addition, the real variable-speed data has a certain degree of volatility, and the traditional graph structure with fixed connections is difficult to effectively simulate this characteristic, which limits the adaptability of the model to complex dynamic environments. To address the above problems, the present invention proposes a method for constructing a sparse threshold graph, aiming to reduce redundant connections and improve the adaptability of the graph model to complex data. The process is as Figure 6 shown.

[0044] S212: Transform the original fault data of the mechanical equipment collected under variable speed conditions from the time domain to the frequency domain through FFT (Fast Fourier Transform); in the frequency domain, use "sampling points" as nodes, "adjacent sampling points" as edges, and "node eigenvalues" as node features.

[0045] A vibration sensor collects sensor data during the operation of the bearing within a preset time, then divides the collected data, obtains several nodes in the time-domain signal, and then transforms these nodes from the time domain to the frequency domain through FFT, and constructs edges by finding adjacent nodes in the frequency domain. Through the above processing, a graph model based on frequency-domain data can be constructed, where each node represents a frequency component, the eigenvalue of the node represents the amplitude of the frequency component, and the edge between nodes represents the relationship between adjacent frequency components.

[0046] S213: Retain the edges between nodes in the fault graph data whose Euclidean distance is less than or equal to the threshold to obtain the basic graph structure.

[0047] First, calculate the Euclidean distance between nodes and set the threshold according to the data distribution d th , determine whether the nodes are connected, and only retain the edges with a distance less than or equal to the threshold, thereby constructing a basic graph structure. This method can effectively filter out weakly related or irrelevant connections, retain high-quality node relationships, reduce noise interference and improve the representation ability of the graph. The calculation formula is as follows: ; ; ; Among them, and are the eigenvalues of node and node in the m dimension; n is the dimension of the eigenvector; is the Euclidean distance between node and node ; Percentile (·) represents the quantile threshold function, D represents the set of Euclidean distances between all node pairs, which is the universal set; q represents the percentile; is the adjacency matrix in the basic graph structure.

[0048] S214: Set the deletion ratiop , according to the deletion ratio p Randomly delete some edges in the basic graph structure to obtain a sparsified graph structure.

[0049] On the basis of the constructed basic graph, further introduce a random edge deletion strategy. By setting the deletion ratio p , randomly delete or retain the edges that meet the distance threshold in a randomized manner, and finally generate a sparsified graph structure. As Figure 6 shown, different deletion ratios p can randomly delete adjacent edges to reduce the density of the graph, optimize the calculation efficiency, and at the same time introduce appropriate perturbation to enhance the robustness of the model to dynamic working conditions and complex topological structures. The calculation process is as follows: ; Among them, is the adjacency matrix of the sparsified graph structure after random deletion; p is the adjacent edge deletion ratio in the graph, p which can be set to 0.1, 0.5 or 0.9.

[0050] S215: Add the corresponding fault type labels to each sparsified graph structure to obtain a dataset.

[0051] S22: Use the dataset to train the neural network.

[0052] S221: Set the model-related parameters.

[0053] The set model-related parameters are shown in Table 2.

[0054] Table 2 Model-related parameters

[0055] S222: Set the loss function.

[0056] Most existing fault diagnosis methods fail to fully consider the differences between different fault modes in loss design. This defect makes it difficult for the model to balance the diagnostic ability for all faults when facing complex working conditions or diverse fault types, especially insufficient performance for some difficult-to-diagnose fault modes, thus seriously affecting the overall diagnostic performance. In addition, some methods usually adopt a fixed loss allocation strategy, paying insufficient attention to the characteristics of some important fault types in complex working conditions, weakening the adaptability and generalization ability of the model. To solve this problem, this method proposes the ALW (Adaptive lossweighting) strategy.

[0057] The process of the ALW strategy is as Figure 7 shown. First, a guiding factor for dynamically adjusting the training stage is designed C. At the initial stage of training, the dynamic guiding factor in the formula C has a relatively large value. At this time, the guiding factor will amplify the influence of the loss ratio of different fault modes, ensuring that the loss weights of all fault modes are relatively close, so as to guide the model to pay more balanced attention to all fault modes globally. Then, as the training progresses, gradually approaches the maximum value , resulting in C gradually decreasing and approaching 1. At this time, the amplification effect of the guiding factor on the training rate weakens and is dominated by the loss ratio. Finally, for the fault modes that are difficult to diagnose (with a slow loss decrease), their training speed will be relatively large, so that its value in the weight will increase, guiding the model to focus more on diagnosing these difficulties. The process is as Figure 7 shown. Compared with the ALW-free strategy, this method can more accurately guide the update of the model parameters: ; ; ; ; where C is the guiding factor; N is the total number of fault types; is the epoch at the t moment in the training stage; is the maximum epoch; is the fault type label, ={1, 2, 3..., N}; is the initial loss of the t -th type of fault at the moment in the training stage; is the initial loss of the t- -th type of fault at the moment in the training stage; is the training speed of the t -th type of fault at the moment in the training stage; The loss weight of the t -th type of fault at the moment in the training stage; is the total loss at the t moment in the training stage.

[0058] S223: Taking the minimum total loss as the optimization objective of the optimizer, use the optimizer to iteratively update the parameters in the neural network. The weight update formula based on the loss function is: ; where, Represents the finally fused features; is the predicted label, and y is the true label; is the cross-entropy loss function; is the learning rate; t and t+1 is the number of iterations.

[0059] S23: Model effect verification.

[0060] To verify the superiority of the proposed method, 8 existing methods are used as comparison methods, including Chebynet, graph convolutional networks (GCN), simplifying graph convolutional networks (SGCN), grapg sample and aggregate (GraphSage), graph isomorphism network (GIN), higher-order graph neural network (HO-GNN), graph attention network (GAT). The calculation of the fault diagnosis evaluation metrics involved includes: Accuracy, Precision, Recall, AUC.

[0061] The experimental results of Case 1 and Case 2 are shown in Table 3 and Table 4 respectively. In Case 1, SAMS-GAT (Sparse Threshold Graph and Adaptive Loss Weighting Multi-Scale Attention Model) demonstrated excellent performance, with an accuracy of 99.20%, significantly exceeding other methods. Among them, the accuracy of GCN was 95.60%, and the gap between the two was close to 3.6%. At the same time, SAMS-GAT also led comprehensively in key metrics such as precision (98.03%), recall (98.00%), and AUC (99.99%). These data not only reflect the high accuracy of SAMS-GAT in classification tasks but also its stability and robustness in complex network environments. Among the comparison methods, the performance of traditional models (such as ChebNet and HO-GNN) was relatively low, while other advanced models (such as GCN and GAT), although performing well in some metrics, were still inferior to SAMS-GAT overall.

[0062] In Case 2, SAMS-GAT once again demonstrated its excellent performance, with an accuracy rate of 98.15%, significantly higher than other comparison methods. The accuracy rate of GAT was 90.13%, and the gap between the two was close to 8%. This result fully reflects the advantage of SAMS-GAT in terms of accuracy. At the same time, SAMS-GAT still performed outstandingly in other key indicators. For example, the precision rate was 98.24%, the recall rate was 97.22%, and the AUC was as high as 99.96%. These results not only verified the robustness of SAMS-GAT in classification tasks but also indicated its strong adaptability in diverse scenarios.

[0063] Table 3 Experimental Results of Case 1

[0064] Table 4 Experimental Results of Case 2

[0065] Through comparison with two other graph construction methods, the present invention evaluated the performance of the STG (Sparse Threshold Graph) module in classification tasks. The first method was FCG (Fully Connected Graph), which used all node data for global information fusion but had a relatively high information redundancy. The second method was K-Nearest Connected Graph (KCG), which constructed a sparse graph by selecting 4 neighbor nodes based on KNN and could capture local features. The STG method proposed in the present invention reduced noise interference while retaining key connections by randomly removing redundant edges. The three graph construction methods are as Figure 8 shown. It can be seen from the confusion matrix that the STG method achieved a relatively high classification accuracy rate for all categories. The number of correctly classified samples for categories 0 to 4 was 50, 48, 50, 50, and 50 respectively, and the misclassification phenomenon was extremely rare, reflecting the robustness and accuracy of this method in complex data environments.

[0066] In contrast, the FCG method shows relatively weak classification performance for Class 2 and Class 3. Among them, the number of samples misclassified from Class 1 to Class 0 is 11, the number of samples misclassified from Class 2 to Class 1 is 9, and the number of samples misclassified from Class 3 to Class 2 and Class 4 are 19 and 26 respectively, indicating that the redundant information in the fully connected graph may lead to a higher sensitivity of the model to noise. The KCG method has strong local feature extraction ability, but the number of correct classifications for Class 1 and Class 4 is only 46 and 41, respectively, showing obvious misclassification phenomena, especially the confusion between Class 2 and Class 3 is more significant. Generally speaking, the STG method effectively enhances the balance between global information and local association by optimizing the graph structure construction, and at the same time significantly reduces the impact of noise on the graph construction quality, thereby improving the classification accuracy and the robustness of the system.

[0067] In addition, the present invention also explores the influence of the p-value (the proportion of deleted adjacent edges in the graph) on the model, and the results are as Figure 9 shown. Figure 9 (a) in Figure 9 is the result of Case 1, and

[0068] (b) in Figure 10 is the result of Case 2. Generally speaking, as the p-value increases, the model accuracy in both cases shows a slightly decreasing trend. Among them, the accuracy in Case 2 is slightly higher and the fluctuation is smaller at high p-values, showing stronger adaptability. Both cases show that when the p-value is low, the model accuracy is above 97%. Even when there is partial information loss or incomplete structure in the actual network, the model can still maintain a high accuracy. As the p-value increases to 0.9, the accuracy decreases, but still remains above 95%, indicating that the model has a certain stability and adaptability in the face of network structure interruption and can better handle the problem of incomplete graph structure that may occur in actual work.

[0068] To verify the feature extraction ability of the MS-GATv2 module in the proposed method, the experiment applied SAMS-GAT and 7 comparison methods to Case 2 for comparative experimental analysis. By performing two-dimensional visualization on the encoded features of each method, as Figure 10 shown, Figure 10 (a) in Figure 10 is the two-dimensional visualization result of the encoded features of the SAMS-GAT model, Figure 10 (c) in Figure 10 is the two-dimensional visualization result of the encoded features of the Chebynet model, Figure 10 (d) in Figure 10 is the two-dimensional visualization result of the encoded features of the GCN model, Figure 10Among them, (g) is the two-dimensional visualization result of the encoded features of the GraphSage model. Figure 10 Among them, (h) is the two-dimensional visualization result of the encoded features of the HO-GNN model. The results show that the proposed method has significantly better discrimination between normal samples and faulty samples than other methods, demonstrating strong feature separation ability.

[0069] It can be observed from the figure that: Figure 10 For (a) of, the features extracted by SAMS-GAT have good clustering effects. Clear and separated clustering clusters can be formed for normal samples and various faulty samples, showing high feature discrimination ability. Figure 10 For (b) of Figure 10 and (c) of GCN also show a certain degree of sample separation, but the clustering boundaries are relatively fuzzy and there are overlaps in some categories. Figure 10 For (d) of SGCN, the feature extraction ability is relatively weak, the sample distribution is relatively dispersed, and the category discrimination is not obvious. Figure 10 For (e) of GAT and Figure 10 For (g) of GraphSage, the clustering structure has been improved, but there are still problems with fuzzy category boundaries compared to SAMS-GAT. Figure 10 For (f) of GIN and Figure 10 For (h) of HO-GNN, the sample distribution has no obvious pattern, and the discrimination ability between normal samples and faulty samples is the weakest. Therefore, the proposed SAMS-GAT method is superior to other comparison methods in terms of feature extraction and classification performance, and can effectively achieve accurate identification of different types of faults, further verifying the superiority of the MS-GATv2 module.

[0070] Figure 11 shows the change of weight update during the training process of the ALW strategy. Figure 11 For (a) of, it is the update process of the ALW strategy for case1. Figure 11 For (b) of, it is the update process of the ALW strategy for case2. In the initial stage of training, the weight changes of various faults are relatively small, mainly driven by the guiding factor C, and the model focuses on the overall loss distribution. At this stage, the weight differences of various faults are not significant, and the main goal is to balance the global loss. However, as the training progresses, the change of the loss ratio gradually dominates the weight update. At this time, the weights of different fault categories begin to show differences, and the model begins to focus on more difficult-to-diagnose fault categories. Therefore, the ALW strategy gradually amplifies the weights of the categories with larger losses in the later stage of training to more effectively diagnose difficult problems. In addition, to better observe the differences, the weights of different categories are amplified by the same proportion every 20 batches.

[0071] As can be seen from the figure, the weight changes are recorded every 20 training batches. The weights gradually diverge as the training progresses, reflecting the dynamic adjustment ability of the model during the optimization process. Taking Figure 9 's (a) as an example, the weights of different labels have formed significant differences at epoch 100. The weights of the difficult-to-diagnose categories gradually increase, while the weights of the easy-to-diagnose categories decrease, indicating that the model tends to focus on difficult samples. Similarly, Figure 9 's (b) shows a similar trend. The divergence of the weight curves further verifies that the ALW strategy can dynamically adjust the model's attention to different categories. By capturing the category difficulty information in real time, it adaptively improves the learning priority of difficult-to-diagnose categories. This dynamic adjustment method not only improves the optimization effect of the model on complex categories but also maintains the stability and balance of global optimization, showing significant potential in dealing with the scenarios of class imbalance and complex samples.

[0072] In addition, two methods, Static Weight Adjustment (SWA) and Random Weight Update (RWU), are selected for comparison with the proposed method (ALW). Figure 12 's (a) is the ALW strategy update process for case 1, Figure 12 's (b) is the ALW strategy update process for case 2. As can be seen from Figure 12 , the ALW method shows significant advantages in terms of test accuracy compared with the SWA and RWU methods. Specifically, the test accuracy of ALW reaches about 0.85 at the 20th epoch, significantly better than 0.75 of SWA and 0.65 of RWU; the final test accuracy even reaches 0.98, significantly higher than 0.88 of SWA and 0.82 of RWU. In addition, after the 80th epoch, the accuracy curve of ALW tends to be stable with less fluctuation, always remaining at a high level of 0.97 - 0.98, while SWA and RWU not only have lower accuracy but also show greater fluctuations. This indicates that the ALW method demonstrates significant superiority in terms of convergence speed, final performance, and stability.

[0073] S3: Input the data of the variable-speed fault diagram to be detected into the fault diagnosis model. Through the single attention head, double attention heads, and triple attention heads in the single-head attention mechanism layer, double-head attention mechanism layer, and triple-head attention mechanism layer respectively, extract features of different scales from the variable-speed fault diagram data to obtain the single-head fault feature map, double-head fault feature map, and triple-head fault feature map. Through the single-head dynamic feature fusion layer and double-head dynamic feature fusion layer, perform adaptive feature weighted fusion operations on the single-head fault feature map, double-head fault feature map, and triple-head fault feature map respectively to obtain the single-head fault weighted feature map, double-head fault weighted feature map, and triple-head fault weighted feature map. Through the fully connected layer, splice the single-head fault weighted feature map, double-head fault weighted feature map, and triple-head fault weighted feature map and perform probability prediction to output the fault type label corresponding to the variable-speed fault diagram data to be detected.

[0074] Through the above method, the relevant process of the overall steps of the proposed method is described as follows: (1) First, obtain the vibration data under speed fluctuation on the mechanical equipment fault simulation test bench, and preprocess the collected data, dividing it into a training set and a test set to ensure that the model can learn and be verified at different stages. Subsequently, perform frequency domain analysis on the divided samples through the FFT method, extract spectral features from the time domain signal, and further enhance the expression ability of fault features. Finally, generate a graph structure for the sample data based on the constructed distance threshold graph, and simulate the incompleteness and noise influence of the graph by randomly deleting some edges.

[0075] (2) Input the constructed graph structure into the designed MS-GATv2 model, and use its multi-scale attention mechanism to efficiently extract graph features at different levels, fully capturing local and global structural information. Subsequently, the extracted features are further feature mapped through the fully connected layer, and finally the prediction probability is output.

[0076] (3) Input the obtained prediction probability and the true label into the cross-entropy loss function to calculate the training loss. On this basis, develop and apply the ALW strategy, adaptively adjust the weight distribution of each fault type by dynamically calculating the loss ratio of each fault mode and introducing a guiding factor, so as to calculate the total loss. Finally, input the total loss into the optimizer for parameter update, continuously optimize the model performance, and improve the accuracy and robustness of fault diagnosis.

[0077] (4) Input the test set into the trained model, obtain the diagnosis result using the prediction output of the model, and evaluate its detection ability and accuracy for different fault modes. At the same time, comprehensively compare and analyze this method with other fault diagnosis methods of common GNNs, evaluate its superiority from multiple indicators, and verify the robustness and applicability of the proposed method in dealing with variable-speed working conditions and diverse fault types.

[0078] In view of the problems of redundant connections and computational overhead caused by fully connected graphs faced by graph neural network models when processing complex graph data, as well as the challenge that it is difficult for the model to balance the processing of different categories due to the differences in fault categories, the present invention proposes a multi-scale attention model based on dynamic perception and adaptive optimization. By introducing STG, the model can dynamically perceive and capture the local geometric structure of the data, enhancing the expressive ability and robustness of the characteristics of graph data. In addition, the present invention designs MS-GATv2, which combines the DWU strategy to efficiently capture graph structure information at different levels and effectively adapt to the characteristics of complex fault categories. Finally, the ALW strategy is developed, enabling the model to optimize the attention to difficult-to-classify categories in real time and further improving the overall detection performance.

[0079] The experimental results show that the proposed SAMS-GAT method performs excellently in multiple key performance indicators (Accuracy, Precision, Recall, AUC). In the experiments of Case 1 and Case 2, SAMS-GAT demonstrates excellent classification performance, with the accuracy reaching 99.20% and 98.15% respectively, and also leading in indicators such as precision, recall, and AUC. This proves that the method of the present invention can not only effectively solve the problem of fault diagnosis under variable speed conditions, but also dynamically adapt to the characteristics of graph data and enhance the detection ability for complex fault categories. Overall, the dynamic perception and multi-scale attention model proposed by the present invention provides an effective solution for complex graph data analysis and also lays a foundation for the research and practice in related fields. Future work will combine multi-modal data fusion technology to expand the actual application scenarios. Further optimizing the dynamic perception mechanism and improving the efficiency will contribute to promoting the wide application of the model in industrial intelligent diagnosis.

[0080] The above is the mechanical fault detection method under variable speed conditions provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding mechanical fault detection device under variable speed conditions, including: A model construction module, used to construct a neural network, where the neural network includes a parallel single-head attention mechanism layer, a double-head attention mechanism layer, and a triple-head attention mechanism layer, a single-head dynamic feature fusion layer, a double-head dynamic feature fusion layer, and a triple-head dynamic feature fusion layer connected to the output ends of the single-head, double-head, and triple-head attention mechanism layers respectively, and a fully connected layer connected to the output ends of the single-head, double-head, and triple-head dynamic feature fusion layers simultaneously.

[0081] A model training module, used to obtain different types of variable speed fault graph data of mechanical equipment under variable speed conditions, construct a data set based on the different types of variable speed fault graph data; use the data set to train the neural network to obtain a fault diagnosis model.

[0082] A fault detection module is configured to input the to-be-detected variable-speed fault map data into a fault diagnosis model. The single-attention head, double-attention head, and triple-attention head in the single-head attention mechanism layer, double-head attention mechanism layer, and triple-head attention mechanism layer respectively perform feature extraction on the variable-speed fault map data at different scales to obtain a single-head fault feature map, a double-head fault feature map, and a triple-head fault feature map. The single-head dynamic feature fusion layer and double-head dynamic feature fusion layer respectively perform adaptive feature weighted fusion operations on the single-head fault feature map, double-head fault feature map, and triple-head fault feature map to obtain a single-head fault weighted feature map, a double-head fault weighted feature map, and a triple-head fault weighted feature map. The single-head fault weighted feature map, double-head fault weighted feature map, and triple-head fault weighted feature map are concatenated through a fully connected layer and then probability prediction is performed to output the fault type label corresponding to the to-be-detected variable-speed fault map data.

[0083] For the specific limitations of the mechanical fault detection device under variable-speed conditions, reference can be made to the limitations of the mechanical fault detection method under variable-speed conditions in the above text, which will not be elaborated here. Each module in the above mechanical fault detection device under variable-speed conditions can be implemented in whole or in part by software, hardware, and their combinations. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above-mentioned modules.

[0084] The present invention also provides a computer-readable storage medium storing a computer program, which can be used to execute the above-mentioned Figure 1 mechanical fault detection method under variable-speed conditions.

[0085] The present invention also provides the structure of a computer device. At the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, other hardware required for other services may also be included. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above-mentioned Figure 1 mechanical fault detection method under variable-speed conditions.

[0086] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0087] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded by the present invention.

Claims

1. A method for detecting mechanical faults under speed change conditions, characterized in that: include: Constructing a neural network, the neural network comprising a single-head attention mechanism layer, a dual-head attention mechanism layer, and a triple-head attention mechanism layer in parallel, a single-head dynamic feature fusion layer, a dual-head dynamic feature fusion layer, and a triple-head dynamic feature fusion layer connected to the output ends of the single-head, dual-head, and triple-head attention mechanism layers in one-to-one correspondence, and a fully connected layer connected to the output ends of the single-head, dual-head, and triple-head dynamic feature fusion layers at the same time; Obtain different types of speed change fault diagram data of mechanical equipment under speed change conditions, and construct a data set based on the different types of speed change fault diagram data; use the data set to train a neural network to obtain a fault diagnosis model; The speed change fault map data to be detected is input into the fault diagnosis model, and the single attention head, double attention head and triple attention head in the single-head attention mechanism layer, double-head attention mechanism layer and triple-head attention mechanism layer are used to extract features of different scales on the speed change fault map data, so as to obtain the single-head fault feature map, double-head fault feature map and triple-head fault feature map; the single-head dynamic feature fusion layer and the double-head dynamic feature fusion layer are used to perform adaptive feature weighted fusion operations on the single-head fault feature map, double-head fault feature map and triple-head fault feature map, so as to obtain the single-head fault weighted feature map, double-head fault weighted feature map and triple-head fault weighted feature map; the single-head fault weighted feature map, double-head fault weighted feature map and triple-head fault weighted feature map are spliced ​​through the fully connected layer for probability prediction, and the fault type label corresponding to the speed change fault map data to be detected is output.

2. The mechanical fault detection method under the speed change condition according to claim 1, characterized in that: The feature extraction of different scales from the speed change fault diagram data specifically includes: ; in, Is a node Features; It is k The linear transformation matrix of the attention heads; It is k Nodes output by the attention head The attention weight of Representation Node Belongs to Node The set of neighbor nodes; K is the number of attention heads, K ={1, 2, 3}; It is the node generated after the attention mechanism is processed feature.

3. The mechanical fault detection method under the speed change condition as claimed in claim 2, characterized in that: The adaptive feature weighted fusion operation specifically includes: ; in, It is k The feature weights of the attention heads, is the fault weighted feature map.

4. The mechanical fault detection method under the speed change condition as claimed in claim 1, characterized in that: The data set is constructed based on different types of speed change fault diagram data, specifically including: Calculate the Euclidean distance between each node in the fault graph data, set a threshold value based on the distribution of the Euclidean distance between each node, and retain the Euclidean distance between nodes in the fault graph data that is less than or equal to the threshold value The edges between the nodes of the graph are obtained: ; ; ; in, and The nodes are With Node In the m The eigenvalue of the dimension; n is the dimension of the feature vector; Is a node With Node The Euclidean distance between Percentile (·) represents the quantile threshold function, D represents the set of Euclidean distances between all pairs of nodes, which is Complete works of; q represents percentile; is the adjacency matrix in the underlying graph structure; Set the deletion ratio p , according to the deletion ratio p Randomly delete some edges in the basic graph structure to obtain a sparse graph structure: ; in, is the adjacency matrix of the sparse graph structure; Add the corresponding fault type label to each sparse graph structure to obtain the data set.

5. The mechanical fault detection method under the speed change condition as claimed in claim 1, characterized in that: The use of the data set to train the neural network specifically includes: Use the cross entropy loss function to calculate the initial loss between the predicted fault type label output by the neural network and the true fault type label; Based on the ALW strategy, the loss ratio of different types of faults is dynamically calculated, and a guiding factor is introduced to adaptively adjust the weight distribution of each fault type to obtain the total loss: ; ; ; ; Where C is the guiding factor; N is the total number of fault types; It is the training phase t The epoch of the moment; is the maximum epoch; is the fault type label, ={1,2,3..., N }; It is the training phase t Moment Initial loss of similar fault types; It is the training phase t- 1st moment Initial loss of similar fault types; It is the training phase t Moment The training speed of the class fault type; Training phase t Moment Loss weights for class fault types; It is the training phase t Total loss at time: The optimization objective of the optimizer is to minimize the total loss, and the optimizer is used to iteratively update the parameters in the neural network.

6. The method for detecting mechanical failure under speed change conditions according to claim 1, characterized in that: The single-head fault weighted feature map, the double-head fault weighted feature map and the triple-head fault weighted feature map are spliced ​​together to perform probability prediction, specifically including: The single-head fault weighted feature map, the double-head fault weighted feature map and the triple-head fault weighted feature map are spliced ​​to form a spliced ​​feature vector; After weighting the concatenated feature vector, nonlinear processing is performed through the ReLU activation function to obtain a nonlinear weighted concatenated feature vector; The nonlinear weighted concatenated feature vector is mapped to the probability distribution of fault categories through the Softmax function, and the fault category with the highest probability is taken as the output.

7. The method for detecting mechanical failure under speed change conditions according to claim 1, characterized in that: The obtaining of different types of speed change fault diagram data of the mechanical equipment under speed change conditions specifically includes: Collect fault vibration sensor data of mechanical equipment under variable speed conditions; The fault vibration sensor data is transformed from the time domain to the frequency domain. In the frequency domain, each sampling point is used as a node, the frequency domain eigenvalue corresponding to each sampling point is used as a point feature, and edges are constructed between adjacent sampling points to obtain the fault graph data of the mechanical equipment under speed change conditions.

8. A mechanical fault detection device under speed change conditions, characterized in that: include: A model construction module, for constructing a neural network, wherein the constructed neural network includes a parallel single-head attention mechanism layer, a dual-head attention mechanism layer, and a triple-head attention mechanism layer, a single-head dynamic feature fusion layer, a dual-head dynamic feature fusion layer, and a triple-head dynamic feature fusion layer connected to the output ends of the single-head, dual-head, and triple-head attention mechanism layers in a one-to-one correspondence, and a fully connected layer connected to the output ends of the single-head, dual-head, and triple-head dynamic feature fusion layers at the same time; A model training module is used to obtain different types of speed change fault diagram data of mechanical equipment under speed change conditions, and to construct a data set based on different types of speed change fault diagram data; the data set is used to train a neural network to obtain a fault diagnosis model; A fault detection module is used to input the speed change fault map data to be detected into the fault diagnosis model, and extract features of different scales on the speed change fault map data through the single attention head, double attention head and triple attention head in the single attention mechanism layer, double attention mechanism layer and triple attention mechanism layer, respectively, to obtain the single-head fault feature map, double-head fault feature map and triple-head fault feature map; perform adaptive feature weighted fusion operation on the single-head fault feature map, double-head fault feature map and triple-head fault feature map through the single-head dynamic feature fusion layer and the double-head dynamic feature fusion layer, respectively, to obtain the single-head fault weighted feature map, double-head fault weighted feature map and triple-head fault weighted feature map; splice the single-head fault weighted feature map, double-head fault weighted feature map and triple-head fault weighted feature map through the fully connected layer, perform probability prediction, and output the fault type label corresponding to the speed change fault map data to be detected.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

10. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Graph neural network fault diagnosis method and device and storage medium

    CN116502175A

  • Rotary machinery variable working condition self-supervision domain adaptation fault diagnosis method and system

    CN118861792A

  • Traffic flow prediction method and system of adaptive graph learning space-time neural network

    CN119227861A

Cited By

  • Variable-working-condition cross-component causal coupling modeling and fault diagnosis method and device

    CN122262822A