Industrial equipment fault diagnosis method and system, computer equipment and storage medium

By fusing STFT images, weighted GAF ​​images and original time series through a multi-branch multimodal convolutional neural network model and attention mechanism, the problem that existing methods are difficult to fuse multimodal features under complex working conditions is solved, and industrial equipment fault diagnosis with higher accuracy and robustness is achieved.

CN120822100APending Publication Date: 2025-10-21BEIJING INFORMATION SCI & TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510950526.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing industrial fault diagnosis methods have difficulty in effectively integrating multimodal features when faced with complex and changeable working conditions. In particular, single-modal input may ignore frequency domain or time series change characteristics. Traditional methods also lack modeling of the importance differences between modes, making it difficult to adapt to the actual working conditions in industrial environments that are complex, changeable, and highly modally correlated.

Method used

A multi-branch multimodal convolutional neural network model (GTS-AttCNN) is adopted to convert the original vibration signal into STFT images, weighted GAF ​​images and original time series. The attention mechanism is combined to automatically weight and fuse the features of each modality, extract time domain, frequency domain and time series structure features, and enhance the model's attention to key fault features.

Benefits of technology

It achieves more comprehensive and robust feature extraction and expression, improves the classification accuracy and generalization ability of the model in complex tasks, significantly increases the focus on key information, and reduces the interference of redundant features on classification performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120822100A_ABST
    Figure CN120822100A_ABST
Patent Text Reader

Abstract

The invention provides an industrial equipment fault diagnosis method and system, computer equipment and a storage medium, belongs to the field of industrial fault detection, and provides a convolutional neural network model GTS-AttCNN based on multi-modal fusion and an attention mechanism. The model receives an original vibration signal acquired by a sensor, and constructs multi-modal feature input from three different angles: constructs a GAF image to extract a time sequence structure feature, constructs an STFT image to extract a frequency domain feature, and retains an original time sequence to retain the time domain feature. And the multi-modal features are input into the GTS-AttCNN for feature extraction. An attention mechanism is introduced into the model in a feature fusion stage, and weighted fusion is performed on each modal feature, so that the sensitivity and diagnosis precision of the model to key fault features are improved. Compared with an existing method, the method has higher recognition precision and higher generalization ability, and can be widely applied to fault diagnosis scenes of complex industrial equipment such as bearings, motors and gearboxes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of industrial fault detection, and in particular relates to an industrial equipment fault diagnosis method, system, computer equipment and storage medium. Background Art

[0002] Equipment may experience various faults during industrial manufacturing processes, and timely identification and effective diagnosis are crucial. Industrial fault diagnosis methods can currently be categorized into three main categories: analytical model-based, knowledge-based, and data-driven. Knowledge-driven methods rely on the experience or prior knowledge of domain experts and are only applicable when rules are clear, knowledge acquisition is simple, and experts are experienced. However, due to the difficulties in acquiring knowledge and processing complex data, they have not been widely adopted in the real world of industry. Model-based diagnostic methods rely heavily on precise physical or mathematical models, which enable effective fault diagnosis. However, this overreliance on models makes modeling difficult when faced with complex and changing operating conditions, and the computational cost of modeling increases dramatically with the amount of data. Recent advances in sensors and data acquisition methods have made data acquisition easier and more efficient. Consequently, data-driven fault diagnosis methods have gradually become the mainstream approach. Furthermore, with the development of neural networks, they have gradually replaced traditional machine learning methods due to their superior feature extraction capabilities and excellent learning and generalization capabilities.

[0003] Convolutional Neural Networks (CNNs) have been particularly successful in speech recognition, natural language processing, and especially image processing due to their local receptive fields and parameter sharing mechanisms. However, in the field of industrial process fault diagnosis, the collected fault data is typically one-dimensional time signals, making them unsuitable as neural network inputs. To fully leverage the advantages of CNNs for one-dimensional time signals, researchers have proposed a variety of methods: using methods such as the Gramian Angular Field (GAF) and Markov Transition Field (MTF) to convert one-dimensional sequences into two-dimensional images; using frequency domain methods such as the Continuous Wavelet Transform (CWT) and the Short-Time Fourier Transform (STFT) to extract frequency domain features from vibration signals, converting the original one-dimensional sequence into images with time-frequency distributions, which are then fed into CNN training. Some researchers, hoping to maximize the preservation of the original time series features, have proposed one-dimensional convolutional neural networks (1D CNNs). One-dimensional CNNs can directly extract local temporal features from raw one-dimensional signals, avoiding complex preprocessing operations, and have therefore been widely used in fault diagnosis. However, because they can only perform convolution along the time axis, it is difficult to fully capture global patterns and frequency domain features. In particular, their performance may be limited when facing multi-scale or non-stationary signals.

[0004] While these methods have achieved promising results in diagnosing bearing and gear faults, they also face challenges in comprehensively modeling complex fault characteristics. Single-modal inputs may overlook frequency-domain or time-series variations. Multimodal fusion methods are relatively simple, with mainstream approaches often employing feature-level concatenation (early fusion) or directly stacking convolutional branches to process different modalities. These approaches lack the ability to model the importance differences between modalities, making them difficult to adapt to the complex, variable, and highly modally correlated real-world operating conditions found in industrial environments. Summary of the Invention

[0005] To address the challenges of industrial fault detection using multimodal features, this paper provides a method for industrial equipment fault diagnosis. This method proposes a multi-branch multimodal convolutional neural network model (GTS-AttCNN) that fuses GAF images, STFT spectrograms, and raw time series. This model uses three different data representation methods to extract time domain, frequency domain, global nonlinear features, and temporal structure, respectively. An attention mechanism is also introduced to automatically weight the fusion of modal features, effectively enhancing the model's focus on key fault characteristics and improving classification accuracy and model generalization.

[0006] In order to achieve the above object, the present invention provides the following technical solutions: A method for diagnosing faults of industrial equipment, comprising: Obtain the original vibration signal of the industrial equipment, perform non-overlapping segmentation on the original vibration signal, and obtain multiple original time subsequences after segmentation; Convert part of the original time subsequences into STFT graphs, convert part of the original time subsequences into weighted GAF ​​graphs, and construct multimodal features through the remaining original time subsequences, STFT graphs, and weighted GAF ​​graphs; Extract local temporal pattern features, frequency domain texture information features, and temporal structure features from multimodal features, concatenate the extracted features, perform attention-weighted fusion on the concatenated features, learn the importance weights of each feature subspace, and obtain fused features; Fault diagnosis of industrial equipment is performed based on fusion features, and diagnostic results for multiple fault categories are obtained.

[0007] Preferably, the step of converting a portion of the original time subsequence into an STFT graph is as follows: Use short-time Fourier transform (STFT) to transform a part of the original time subsequence into an STFT image, resample the STFT output result through bilinear interpolation, and uniformly adjust the output result to N*N size; The conversion formula of the STFT graph is: ; Where, represents the original time subsequence, n represents a local time index, m Indicates the frame index or frame number, k represents the frequency index, H Indicates the frame shift or overlap step size, N Indicates the number of points used in Fourier transform. represents the window function, L represents the length of the window function, j Is an imaginary unit.

[0008] Preferably, converting a portion of the original time subsequences into a weighted GAF ​​graph specifically includes: A part of the original time subsequence is converted into GADF image and GASF image. According to the set weight, the weighted summation method is used to perform weighted summation on the GADF image and GASF image to obtain the fused weighted GAF ​​image. The specific conversion formula is: ; ; Where, represents the time series after normalization, for The angle in the polar coordinate system after encoding, Represents Consistent, represents the angle of the j-th point in the time series after encoding in the polar coordinate system, r represents the distance from the point in the polar coordinate system to the origin, t i express The corresponding timestamp, N Indicates the total length of the time series; represents the weighted GAF ​​graph, and represent the Gram angle sum field and Gram angle difference field matrices respectively, Represents weight.

[0009] Preferably, the multimodal features are input into the GTS-AttCNN model and the fused features are output; The attention mechanism of the GTS-AttCNN model consists of two fully connected layers. The first fully connected layer is used to compress the channel dimension and map the features to a low-dimensional space, while the second fully connected layer is used to restore the dimension and map the features back to a high-dimensional space. The output layer of the GTS-AttCNN model is a fully connected layer and a Softmax classifier.

[0010] Preferably, the concatenated features are subjected to attention-weighted fusion through an attention mechanism to learn the importance weights of each feature subspace and obtain fused features, specifically including: The concatenated features are input into the first fully connected layer to map the features into a low-dimensional space. After passing through the ReLU function, they are input into the second fully connected layer to map the sequence back to a high-dimensional space. Generate fusion weight vector through sigmoid function; Multiply the fusion weight vector and the original feature element by element to obtain the new fusion feature.

[0011] Preferably, before performing non-overlap segmentation on the original vibration signal, the method further includes performing data cleaning and normalization preprocessing on the original vibration signal; The preprocessed vibration signal is segmented without overlap using the sliding window method, and the original sequence is split into multiple original time subsequences.

[0012] The present invention also provides an industrial equipment fault diagnosis system, comprising: A data acquisition module is used to acquire the original vibration signal of the industrial equipment, perform non-overlapping segmentation on the original vibration signal, and obtain multiple original time subsequences after segmentation; The data conversion module is used to convert part of the original time subsequences into STFT graphs, and part of the original time subsequences into weighted GAF ​​graphs, and form multimodal features through the remaining original time subsequences, STFT graphs and weighted GAF ​​graphs; The feature fusion module is used to extract local temporal pattern features, frequency domain texture information features, and temporal structure features from multimodal features, and then concatenate the extracted features. The concatenated features are then subjected to attention-weighted fusion to learn the importance weights of each feature subspace and obtain fused features. The fault diagnosis module is used to diagnose faults of industrial equipment based on fusion features and obtain diagnostic results for multiple fault categories.

[0013] The present invention also provides a computer device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement any one of the steps in the industrial equipment fault diagnosis method.

[0014] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is loaded by a processor, it can execute any one of the steps in the industrial equipment fault diagnosis method.

[0015] The industrial equipment fault diagnosis method provided by the present invention has the following beneficial effects: The present invention first divides the original vibration signal of the industrial equipment into multiple time subsequences and converts the subsequences into STFT graphs and weighted GAF ​​graphs. It then constructs multimodal features by combining the time domain, frequency domain, and time series structure by using the time subsequences representing the time domain features, the STFT graph representing the frequency domain features, and the weighted GAF ​​graph representing the time series structure features. This method can more comprehensively capture the information in the vibration signal, comprehensively consider multiple dimensions, and improve the ability to describe complex fault characteristics, thereby achieving more comprehensive and robust feature extraction and expression, and significantly improving the model's expressiveness in complex tasks. At the same time, in the feature extraction stage, local time series pattern features, frequency domain texture information features, and time series structure features are extracted separately. This meticulous feature extraction method helps to capture fault characteristics more accurately. Compared with the existing simple splicing multimodal fusion method, the use of attention weighted fusion to process the spliced ​​features can automatically learn the importance weights of each feature subspace, thereby improving the collaborative expression ability between features, effectively enhancing the model's attention to key information, and reducing the interference of redundant features on classification performance. Compared with traditional splicing or fusion methods, it can more effectively highlight key features and improve diagnostic performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To more clearly illustrate the embodiments of the present invention and its design, the following briefly introduces the drawings required for this embodiment. The drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be derived from these drawings without inventive effort.

[0017] Figure 1 Schematic diagram of multimodal input; Figure 2 This is a structural diagram of the attention mechanism module of the present invention; Figure 3 This is a flow chart of the multi-branch GTS-AttCNN model training method proposed in Example 2 of the present invention; Figure 4 The results of multi-model comparison of conventional working condition test of bearing data; Figure 5 The confusion matrix of GTS-CNN for complex working condition test of bearing data; Figure 6 The confusion matrix of GTS-AttCNN for complex working condition test of bearing data; Figure 7 This is a comparison chart of the accuracy of each model during training; Figure 8 Comparison chart of the loss during training of each model; Figure 9 This is a flow chart of the industrial equipment fault diagnosis method proposed in Example 3 of the present invention. DETAILED DESCRIPTION

[0018] In order to enable those skilled in the art to better understand the technical solution of the present invention and to be able to implement it, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and are not intended to limit the scope of protection of the present invention.

[0019] Example 1 First, the present invention proposes a convolutional neural network model GTS-AttCNN based on multimodal fusion and attention mechanism, specifically a multi-branch GTS-AttCNN model.

[0020] First, the present invention is based on a convolutional neural network and introduces an attention mechanism into the network structure to achieve adaptive weighting of the importance of different modal features, thereby enhancing the model's perception of key features, improving the accuracy of fault identification and system robustness, and is suitable for multi-category fault diagnosis tasks under complex working conditions.

[0021] When constructing the input features of the GAF branch, a weighted fusion of GADF and GASF is used as the input of the GAF branch. As shown in Equation (1), due to the different construction methods of GASF and GADF, the transformed images each have different focuses. GASF emphasizes positive correlation information and global trends in time series, while GADF focuses on inverse correlation and local dynamic change trends. Therefore, the weighted fusion method involves performing element-by-element weighted summation of the GASF image and the GADF image according to preset weights to generate a fused GAF ​​image, enhancing the GAF branch's comprehensive expression of time series features. The weights can be manually set according to the requirements of the specific fault diagnosis task or optimized through a data-driven approach.

[0022] (1) When constructing the STFT branch, the bicubic interpolation method is used to resample the original STFT image to N*N size, where N is the length of the time subsequence, which is also the height and width of the GAF image. This is to adapt to the input of the convolutional neural network and facilitate subsequent feature fusion.

[0023] The attention mechanism consists of two fully connected layers. The first fully connected layer is used to compress the channel dimension and map the features to a low-dimensional space. The second fully connected layer is used to restore the dimension and map the features back to a high-dimensional space. The weight of each feature element is obtained through the Sigmoid function, as shown in Equations (2) and (3).

[0024] (2) (3) In formula (2), represents the weight of each element, represents the sigmoid function, represents the ReLU activation function, and Represents the learnable parameter matrices in the two fully connected layers. It represents the features of the three types of features after feature extraction and global average pooling, that is, the original features without the fusion of attention mechanism. In formula (3), represents the fusion feature, Represents element-wise multiplication.

[0025] The output layer of the model is a fully connected layer and a Softmax classifier, which is used to output prediction results for multiple fault categories. It is suitable for fault diagnosis tasks of various industrial equipment including bearings, gearboxes, motors, etc., and has significant application value and promotion prospects.

[0026] Example 2 The present invention also proposes a multi-branch GTS-AttCNN model training method, specifically as follows Figure 3 As shown, the following steps are included: Step 1: Data Acquisition and Preprocessing: Sensors are used to collect raw vibration signals from industrial equipment. Raw vibration signals are time series signals continuously acquired by sensors and recorded continuously at a fixed sampling frequency during equipment operation. They exhibit distinct temporal order and dynamic evolution characteristics. The collected raw vibration signals undergo preprocessing, including data cleaning and normalization. The preprocessed vibration signals are then segmented using a sliding window method, performing non-overlapping segmentation, splitting the original sequence into multiple original time subsequences.

[0027] Step 2: Multimodal feature construction. Convert multiple original time subsequences into time series structure feature graphs (weighted GAF), frequency domain feature graphs (STFT graphs), and retain the original time subsequences. The three types of modes respectively characterize the time series structure, frequency domain, and time domain information of the vibration signal, and together constitute multimodal input data for multi-channel input of subsequent models. Figure 1 As shown in the figure, the construction of each feature is as follows:

[0028] G-branch (GAF): In order to retain the ability of GAF images to express time series information, the GAF method is used to convert part of the original time subsequence into GADF images and GASF images. Finally, according to certain weights, the weighted summation method is used to obtain the fused GAF ​​image, and the features are then input into a lightweight CNN.

[0029] Specifically, the sequence is converted into a weighted GAF ​​image by using equations (4) and (1).

[0030] (4) (1) In the formula, in the formula, represents the time series after normalization, for The angle in the polar coordinate system after encoding, Represents Consistent, represents the angle of the j-th point in the time series after encoding in the polar coordinate system, r represents the distance from the point in the polar coordinate system to the origin, t i express The corresponding timestamp, N Indicates the total length of the time series; represents the weighted GAF ​​graph, and represent the Gram angle sum field and Gram angle difference field matrices respectively, Represents weight.

[0031] T branch (time): In order to preserve the temporal characteristics of the original signal, this branch selects to retain the original subsequence, that is, the time subsequence that characterizes the time domain features, so as to be subsequently input into the 1D CNN for feature extraction.

[0032] S-branch (STFT): To extract the frequency domain features of the signal, a portion of the original time subsequence is transformed into an STFT graph using the short-time Fourier transform. The STFT graph of the time series is extracted using Equation (5). To adapt to the CNN input and facilitate subsequent feature fusion, the present invention resamples the STFT output using bilinear interpolation and uniformly resizes it to an N*N size. Frequency domain features are then extracted using a lightweight CNN.

[0033] (5) Where, Represents the original time subsequence, which is a sequence obtained by sampling the vibration of industrial equipment or other physical quantities; n Represents a local time index, used to represent a time point within the current window; m Represents the frame index or frame number, indicating that the current analysis window is the mth one; k represents the frequency index, indicating the kth frequency component after transformation; H Indicates the frame shift or overlap step size (Hop size), that is, the offset between two adjacent windows on the time axis; N Indicates the number of points used in Fourier transform (i.e., spectrum length), usually taken as the window function length L or its extended value (such as an integer power of 2); Represents the window function, commonly used such as Hamming window, Hanning window, etc., which is used to weight the data in the current frame to reduce edge effects; L represents the length of the window function, j Is an imaginary unit.

[0034] The above three features, together with the retained subsequences and fault type labels, form a complete dataset. The dataset is randomly divided into training set, validation set and test set according to the proportion.

[0035] Step 3: Model training: Input the training set and validation set into the GTS-AttCNN model for training to establish a fault recognition model. Specifically: The fault type label, the time subseries representing the time domain features, the STFT graph representing the frequency domain features, and the weighted GAF ​​graph representing the temporal structure features constitute the multimodal feature input to the GTS-AttCNN model. After the features are extracted by CNN, they are input into the feature fusion module containing the attention mechanism for multimodal fusion. CNN is divided into a feature extraction (convolution) part and a classification part. The convolution part (including convolution layer, pooling layer, regularization layer, etc.) extracts the deep spatial features of each modal input, including local temporal patterns, frequency domain texture information, and abstract expressions of temporal structure features. The features extracted by the convolution part will be input into the classification part (including fully connected layers, softmax functions, etc.) to obtain the final classification results. The structure of the feature fusion module is as follows: Figure 2 As shown in Figure 1, the feature fusion module is located between the convolutional and classification components. After concatenating the multi-branch features, the module inputs them into the first fully connected layer, mapping the features to a low-dimensional space. After passing through the ReLU function, the module inputs them into the second fully connected layer, mapping the sequence back to a high-dimensional space. Finally, the sigmoid function is used to generate the fusion weight vector, as shown in Equation (6). Finally, the new fused features are multiplied element-wise with the original features, as shown in Equation (7).

[0036] (6) (7) In formula (6), represents the weight of each element, represents the sigmoid function, represents the ReLU activation function, and represents the learnable parameter matrix in fully connected layers 1 and 2, Represents the original features after simple concatenation without integrating the attention mechanism. In formula (7), Indicates fusion features. Represents element-wise multiplication.

[0037] After the above steps, weighted fusion of multiple modalities can be achieved, which can be used for subsequent classification tasks.

[0038] The improved model GTS-AttCNN proposed in the present invention, which fuses multimodal information and introduces an attention mechanism, is different from the traditional CNN model. This model comprehensively utilizes feature information of three different modalities, representing time domain features (time series), frequency domain features (STFT graph) and temporal structure features (weighted GAF ​​graph), thereby achieving more comprehensive and robust feature extraction and expression. By fusing these three types of information in the channel dimension, the present invention effectively overcomes the problems of limited information dimension and insufficient expression ability of single modal models, and significantly improves the expressiveness of the model in complex tasks. On the basis of multimodal fusion, the present invention further introduces an attention mechanism to achieve adaptive weighting of the importance of different modal features. Compared with the existing simple splicing multimodal fusion method, the attention mechanism can dynamically adjust the contribution of various features, thereby improving the collaborative expression ability between features, effectively enhancing the model's attention to key information, and reducing the interference of redundant features on classification performance.

[0039] Example 3 Based on the above-trained fault identification model, the present invention provides an industrial equipment fault diagnosis method for completing the task of equipment fault identification and classification through the vibration signal of industrial equipment. This method can be used to detect faults in components such as bearings, gears, gearboxes, and motors. Figure 9 As shown, the following steps are included:

[0040] S1. Sensors are used to collect raw vibration signals from industrial equipment and preprocess them, including data cleaning and normalization. The preprocessed vibration signals are then segmented using a sliding window method without overlapping, splitting the original sequence into multiple original time subsequences.

[0041] S2. Convert multiple original time subsequences into STFT graphs representing frequency domain features and weighted GAF ​​graphs representing temporal structure features, and construct multimodal features through time subsequences representing time domain features, STFT graphs representing frequency domain features, and weighted GAF ​​graphs representing temporal structure features.

[0042] S3. The multimodal features consisting of the time subsequence representing the time domain features, the STFT graph representing the frequency domain features, and the weighted GAF ​​graph representing the temporal structure features are input into the fault recognition model. After the features are extracted through CNN, they are input into the feature fusion module containing the attention mechanism for multimodal fusion, and the fused features are output.

[0043] S4. Input the fused features into the classifier and output the prediction results of multiple fault categories.

[0044] In order to verify the effectiveness and high-precision fault diagnosis capability of the GTS-AttCNN model, the present invention uses the rolling bearing dataset of Case Western Reserve University and the gearbox dataset of Southeast University to verify the effectiveness of the model in single-variable time series and multivariable time series fault diagnosis tasks. The comparison scheme is other single-branch models and dual-branch models. The verification results under normal working conditions of the bearing dataset are as follows: Figure 4 As shown. Figure 4 The proposed GTS-AttCNN model outperforms both single-branch and dual-branch models across all metrics. All four evaluation metrics reach above 99%, demonstrating that the model achieves high accuracy while also achieving high yield. It demonstrates significant advantages in multimodal fusion and attention mechanisms.

[0045] In order to further verify the discrimination ability and robustness of the proposed model under confusing working conditions, the complex working conditions of the bearing dataset were selected for comparison test again. At the same time, the GTS-CNN model without the introduction of the attention mechanism was added as a comparison experiment. The results are shown in Table 1: Table 1 Comparative experimental results of complex working conditions It can be seen that under the selected complex working conditions that are easy to confuse, the overall effect of the single-branch and dual-branch models is not good, and the classification accuracy is low. The two multi-branch models perform well, and the accuracy can reach more than 90%, indicating that they have better generalization ability and robustness when dealing with complex classification tasks that are difficult to distinguish. Further comparison shows that compared with the GTS-CNN model without the attention mechanism, the proposed GTS-AttCNN model has a higher accuracy, which confirms the rationality and effectiveness of the attention mechanism in feature fusion. The confusion matrices of the two multi-branch models are as follows: Figure 5 and Figure 6 As shown in the figure, the GTS-AttCNN model, which incorporates the attention mechanism, has fewer misclassified samples than the GTS-CNN model and performs better overall. The accuracy of each class is over 90%, and the overall classification accuracy reaches 97.2%. This further verifies the effectiveness of the attention mechanism in improving the model's feature perception and discrimination performance.

[0046] To further evaluate the generalization ability of the present invention on new datasets, we also conducted experiments using a gear dataset from Southeast University. This dataset is a multivariate time series dataset, which helps further verify the performance and robustness of the present invention in multidimensional time series fault diagnosis tasks. The experimental results are shown in Table 2:

[0047] Table 2 Comparative experimental results of gear dataset As shown in Table 2, on the gear dataset, both multi-branch models outperformed the single-branch and dual-branch models across all evaluation metrics. Compared to the traditional G-CNN model, which achieved only 63.35% accuracy, the multi-branch models achieved accuracy exceeding 90%. In particular, the GTS-AttCNN model, which incorporates the attention mechanism, achieved accuracy, precision, recall, and F1 score exceeding 99%, fully demonstrating the effectiveness of the multi-branch structure and attention mechanism in improving the model's discriminative ability and robustness.

[0048] The verification accuracy curve and verification loss curve of each model during training are as follows: Figure 7 and Figure 8 As shown in the figure, it can be clearly observed that the GTS-AttCNN model has a faster convergence speed in the early stage and higher accuracy and lower loss in the later stage compared with other models. This further demonstrates the model's good learning and generalization capabilities under complex working conditions.

[0049] The above experiments effectively demonstrate the superior performance of the proposed model in multimodal fusion and fault identification tasks. Compared with existing models, the proposed GTS-AttCNN achieves good classification performance on both datasets. Its performance in multimodal fusion and complex fault pattern recognition significantly outperforms traditional methods and existing deep learning models, demonstrating significant improvements in accuracy, robustness, and generalization, demonstrating its excellent engineering practicality and potential for widespread adoption.

[0050] Experimental results demonstrate that the proposed GTS-AttCNN model outperforms current mainstream deep learning models in terms of accuracy across multiple public datasets, achieving higher classification accuracy and stronger generalization capabilities. It demonstrates excellent adaptability and stability, particularly in industrial process fault diagnosis tasks. This method has significant application value and promising prospects for widespread adoption in areas such as rotating machinery monitoring, process control system diagnosis, and intelligent manufacturing.

[0051] Based on the same inventive concept, the present invention also proposes an industrial equipment fault diagnosis method, comprising: The data acquisition module is used to acquire the original vibration signal of the industrial equipment, perform non-overlapping segmentation on the original vibration signal, and obtain multiple original time subsequences after segmentation.

[0052] The data conversion module is used to convert part of the original time subsequences into STFT graphs, convert part of the original time subsequences into weighted GAF ​​graphs, and form multimodal features through the remaining original time subsequences, STFT graphs and weighted GAF ​​graphs.

[0053] The feature fusion module is used to extract local temporal pattern features, frequency domain texture information features and temporal structure features from multimodal features, and to splice the extracted features. The attention-weighted fusion of the spliced ​​features is performed to learn the importance weights of each feature subspace to obtain fused features.

[0054] The fault diagnosis module is used to diagnose faults of industrial equipment based on fusion features and obtain diagnostic results for multiple fault categories.

[0055] Each module in the aforementioned industrial equipment fault diagnosis system can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.

[0056] The present invention also provides a computer device comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the industrial equipment fault diagnosis method embodiment. The specific implementation method can be found in the method embodiment and will not be repeated here.

[0057] Furthermore, the present invention provides a non-transitory computer-readable storage medium containing instructions, wherein the storage medium stores a computer program. For example, this storage medium may contain instructions, and the instructions may be executed by a processor of a computer device to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device. When executed by the processor, this computer program can implement the steps of the industrial equipment fault diagnosis method embodiment. Specific implementation methods can be found in the method embodiment and will not be detailed here.

[0058] Those skilled in the art will appreciate that embodiments of the present invention may provide methods, systems, or computer program products. Accordingly, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.

[0059] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0060] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0061] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0062] It should be pointed out that the specific implementation methods described above can enable those skilled in the art to understand the invention more comprehensively, but do not limit the invention in any way. Therefore, although the present specification and examples have described the invention in detail, those skilled in the art should understand that the invention can still be modified or replaced by equivalents; and all technical solutions and improvements that do not deviate from the spirit and scope of the invention are included in the scope of protection of the patent for the invention. Any figure mark in the claims should not be regarded as limiting the claims involved. Any simple change or equivalent replacement of the technical solution that can be obviously obtained by any person familiar with the art within the technical scope disclosed in the present invention falls within the scope of protection of the present invention.

Claims

1. A method for diagnosing faults in industrial equipment, characterized in that: include: Obtain the original vibration signal of the industrial equipment, perform non-overlapping segmentation on the original vibration signal, and obtain multiple original time subsequences after segmentation; Convert part of the original time subsequences into STFT graphs, convert part of the original time subsequences into weighted GAF ​​graphs, and construct multimodal features through the remaining original time subsequences, STFT graphs, and weighted GAF ​​graphs; Extract local temporal pattern features, frequency domain texture information features, and temporal structure features from multimodal features, concatenate the extracted features, perform attention-weighted fusion on the concatenated features, learn the importance weights of each feature subspace, and obtain fused features; Fault diagnosis of industrial equipment is performed based on fusion features, and diagnostic results for multiple fault categories are obtained.

2. The industrial equipment fault diagnosis method according to claim 1, characterized in that: The conversion of a portion of the original time subsequence into an STFT graph is specifically as follows: Use Short-Time Fourier Transform (STFT) to transform a portion of the original time subsequence into an STFT graph, resample the STFT output through bilinear interpolation, and uniformly adjust the output to N*N size, where N is the length of the time subsequence; The conversion formula of the STFT graph is: ; Where, represents the original time subsequence, n represents a local time index, m Indicates the frame index or frame number, k represents the frequency index, H Indicates the frame shift or overlap step size, N Indicates the number of points used in Fourier transform. represents the window function, L represents the length of the window function, j Is an imaginary unit.

3. The industrial equipment fault diagnosis method according to claim 1, characterized in that: The step of converting a portion of the original time subsequences into a weighted GAF ​​graph specifically includes: A part of the original time subsequence is converted into GADF image and GASF image. According to the set weight, the weighted summation method is used to perform weighted summation on the GADF image and GASF image to obtain the fused weighted GAF ​​image. The specific conversion formula is: ; ; Where, represents the time series after normalization, for The angle in the polar coordinate system after encoding, Represents Consistent, represents the angle of the j-th point in the time series after encoding in the polar coordinate system, r represents the distance from the point in the polar coordinate system to the origin, t i express The corresponding timestamp, N Indicates the total length of the time series; represents the weighted GAF ​​graph, and represent the Gram angle sum field and Gram angle difference field matrices respectively, Represents weight.

4. The industrial equipment fault diagnosis method according to claim 1, characterized in that: Input multimodal features into the GTS-AttCNN model and output fused features; The attention mechanism of the GTS-AttCNN model consists of two fully connected layers. The first fully connected layer is used to compress the channel dimension and map the features to a low-dimensional space, while the second fully connected layer is used to restore the dimension and map the features back to a high-dimensional space. The output layer of the GTS-AttCNN model is a fully connected layer and a Softmax classifier.

5. The industrial equipment fault diagnosis method according to claim 4, characterized in that: The concatenated features are weightedly fused through the attention mechanism to learn the importance weights of each feature subspace and obtain fused features, including: The concatenated features are input into the first fully connected layer to map the features into a low-dimensional space. After passing through the ReLU function, they are input into the second fully connected layer to map the sequence back to a high-dimensional space. Generate fusion weight vector through sigmoid function; Multiply the fusion weight vector and the original feature element by element to obtain the new fusion feature.

6. The industrial equipment fault diagnosis method according to claim 1, characterized in that: Before performing non-overlap segmentation on the original vibration signal, the original vibration signal is also pre-processed by data cleaning and normalization; The preprocessed vibration signal is segmented without overlap using the sliding window method, and the original sequence is split into multiple original time subsequences.

7. An industrial equipment fault diagnosis system, characterized in that: include: A data acquisition module is used to acquire the original vibration signal of the industrial equipment, perform non-overlapping segmentation on the original vibration signal, and obtain multiple original time subsequences after segmentation; The data conversion module is used to convert part of the original time subsequences into STFT graphs, and part of the original time subsequences into weighted GAF ​​graphs, and form multimodal features through the remaining original time subsequences, STFT graphs and weighted GAF ​​graphs; The feature fusion module is used to extract local temporal pattern features, frequency domain texture information features, and temporal structure features from multimodal features, and then concatenate the extracted features. The concatenated features are then subjected to attention-weighted fusion to learn the importance weights of each feature subspace and obtain fused features. The fault diagnosis module is used to diagnose faults of industrial equipment based on fusion features and obtain diagnostic results for multiple fault categories.

8. A computer device comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is loaded into a processor, it can execute the steps of the method according to any one of claims 1 to 6.