Fault classification system and method, storage medium and electronic equipment

By converting the original vibration signal of the wind turbine into time-frequency images, and using feature extraction network, attention network and classifier for fault classification, the problem of low fault identification accuracy in the prior art is solved, and higher fault classification accuracy is achieved.

CN120030418APending Publication Date: 2025-05-23HUANENG RENEWABLES SHANGHAI POWER GENERATION CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510328736.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In the prior art, the method of determining the fault category of the target object through the residual network and the U-Net segmentation network has low recognition accuracy.

Method used

The continuous wavelet converter is used to convert the original vibration signal into a time-frequency image, and combine the feature extraction network, attention network and classifier to perform fault classification. The feature extraction network generates detailed information and feature maps of deep features. The attention network enhances the key features of the feature map through channel and spatial dimensions, suppresses irrelevant features, and the classifier finally determines the fault category.

Benefits of technology

It improves the accuracy of fault classification and can more accurately identify and classify complex fault patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030418A_ABST
    Figure CN120030418A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a fault classification system and method, a storage medium and electronic equipment, and the system comprises a continuous wavelet transformer which is used for converting an original vibration signal corresponding to a target object into a time-frequency image; the feature extraction network is connected with the continuous wavelet transformer and is used for performing feature extraction on the time-frequency image so as to generate a first feature graph used for indicating detail information and global information in the time-frequency image and a second feature graph used for indicating deep features of the time-frequency image; the attention network is used for performing splicing processing on the first feature map and the second feature map to obtain a multi-level feature map, and performing channel dimension and spatial dimension enhancement processing on the multi-level feature map to enhance key features of the multi-level feature map and suppress irrelevant features of the multi-level feature map; and the classifier is connected with the attention network and is used for carrying out fault classification on the multi-level feature map after enhancement processing so as to determine the fault category of the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of wind turbine fault diagnosis, and in particular to a fault classification system and method, a storage medium and an electronic device. Background Art

[0002] Xidian University proposed a satellite image segmentation method based on residual network and U-Net segmentation network in its patent application "Satellite image segmentation method based on residual network and U-Net segmentation network" (patent application number: 201910494013.1, application publication number: CN110211137A). The implementation of this method includes the following steps: First, construct a training sample set, ResNet34 and U-Net segmentation network. Subsequently, train ResNet34 and U-Net networks. Then, the satellite image to be segmented is input into the residual network ResNet34 for binary classification, and the ship target is judged. The U-Net segmentation network is used to perform binary segmentation on the positive samples in the classification results. Finally, for the negative samples in the classification results, a single-value mask map is directly output. This method uses ResNet34 to perform binary classification on the satellite image, and then uses U-Net embedded in the SE-ResNet module to re-segment the positive samples in the classification results. The two networks are serially connected. When ResNet34 misclassifies in the binary classification stage, the misclassified samples will directly output a single-value mask map without further segmentation processing by U-Net, resulting in the images that actually contain ship targets being ignored.

[0003] With regard to the problem of low recognition accuracy of methods for determining the fault category of a target object through a residual network and a U-Net segmentation network in related technologies, no effective solution has been proposed so far.

[0004] Therefore, it is necessary to improve the related technology to overcome the above-mentioned defects in the related technology. Summary of the invention

[0005] The embodiments of the present application provide a fault classification system and method, a storage medium and an electronic device to at least solve the problem of low recognition accuracy of the method of determining the fault category of a target object through a residual network and a U-Net segmentation network in the related art.

[0006] According to an embodiment of the present application, a fault classification system is provided, comprising: a continuous wavelet transformer, a feature extraction network connected to the continuous wavelet transformer, an attention network connected to the feature extraction network, and a classifier connected to the attention network, wherein: the continuous wavelet transformer is used to convert an original vibration signal corresponding to a target object into a time-frequency image; the feature extraction network is used to perform feature extraction on the time-frequency image to generate a first feature map and a second feature map, wherein the first feature map is used to indicate detail information and global information in the time-frequency image, and the second feature map is used to indicate deep features of the time-frequency image; the attention network is used to perform splicing processing on the first feature map and the second feature map to obtain a multi-level feature map, and perform channel dimension and spatial dimension enhancement processing on the multi-level feature map to enhance the key features of the multi-level feature map and suppress irrelevant features of the multi-level feature map; the classifier is used to perform fault classification on the enhanced multi-level feature map to determine the fault category of the target object.

[0007] In an exemplary embodiment, the continuous wavelet transformer is further used to: select a target wavelet basis function according to the frequency range of the original vibration signal, and determine multiple scale parameters and multiple translation parameters corresponding to the target wavelet basis function, wherein the scale parameter is used to adjust the time scale of the original vibration signal, and the translation parameter is used to adjust the translation position of the original vibration signal; determine multiple eigenvalues ​​of the original vibration signal at different time scales and different translation positions according to a first formula, and construct the time-frequency image according to the multiple eigenvalues, wherein the first formula is: x(t) is the original vibration signal, a is the scale parameter, b is the translation parameter, W(a,b) is the eigenvalue, t is the time variable corresponding to the converted time-frequency image, is the complex conjugate of the target wavelet basis function.

[0008] In an exemplary embodiment, the feature extraction network includes: a joint path network and a residual network, wherein the joint path network includes: an input layer, a downsampling path layer, a bottleneck layer, an upsampling path layer and an output layer, wherein: the input layer is used to receive the time-frequency image; the downsampling path layer includes a plurality of first convolutional layers, and a pooling layer connected to each first convolutional layer, wherein each first convolutional layer is used to identify frequency domain features and time domain features in the time-frequency image to generate each third feature map, each pooling layer is used to reduce the size of each third feature map, each first convolutional layer has a different number of output channels, and each third feature map after the size reduction is used to indicate the global features of the time-frequency image; the bottleneck layer is used to perform feature concentration processing on each third feature map after the size reduction. The upsampling path layer comprises a plurality of deconvolution layers and each second convolution layer connected to each deconvolution layer, wherein each deconvolution layer is used to perform deconvolution on each third feature map after feature concentration to increase the size of each third feature map after feature concentration, and the deconvolution third feature map is used to indicate the detail features of the time-frequency image, and each second convolution layer is used to fuse each third feature map after deconvolution with each third feature map after size reduction to generate each fused feature map, and the plurality of first convolution layers and the plurality of deconvolution layers correspond one to one; and the output layer is used to fuse the plurality of fused feature maps to generate the first feature map.

[0009] In an exemplary embodiment, the residual network is used to: convert the time-frequency image into a pseudo-RNB image; perform a convolution operation on the pseudo-RNB image to extract local features corresponding to the time-frequency image; perform a maximum pooling operation on the pseudo-RNB image after the local features are extracted to reduce the size of the pseudo-RNB image after the local features are extracted, wherein the pseudo-RNB image after the size reduction retains the local features; and learn the multi-level features of the pseudo-RNB image after the size reduction according to the residual block in the residual network to obtain the second feature map.

[0010] In an exemplary embodiment, the attention network is also used to: receive the first feature map and the second feature map, and concatenate the first feature map and the second feature map to obtain a multi-level feature map; input the multi-level feature map into the channel attention network in the attention network, so that the channel attention network performs global average pooling on the multi-level feature map in the channel dimension to obtain a first feature vector, and performs global maximum pooling on the multi-level feature map in the channel dimension to obtain a second feature vector; add and activate the first feature vector and the second feature vector through the fully connected layer in the attention network to generate a channel attention weight vector; multiply the multi-level feature map and the channel attention weight vector channel by channel to enhance the key features of the multi-level feature map to obtain a fourth feature map; enhance the fourth feature map in the spatial dimension to obtain the enhanced multi-level feature map.

[0011] In an exemplary embodiment, the attention network is also used to: input the fourth feature map into the spatial attention network in the attention network, so that the spatial attention module performs global average pooling on the fourth feature map in the spatial dimension to obtain a fifth feature map, and performs global maximum pooling on the fourth feature map in the spatial dimension to obtain a sixth feature map; splice the fifth feature map and the sixth feature map, and perform convolution on the spliced ​​feature map to generate a spatial attention weight vector; multiply the fourth feature map element by element by the spatial attention weight vector to obtain the enhanced multi-level feature map.

[0012] In an exemplary embodiment, the classifier is also used to: perform global average pooling on the enhanced multi-level feature map to obtain a third feature vector; input the third feature vector into the second fully connected layer in the classifier so that the second fully connected layer updates the dimension of the third feature vector, wherein the dimension of the updated third feature vector is equal to the number of fault categories corresponding to the multiple first fault categories corresponding to the target object; map the updated third feature vector to the fault category space according to a weight matrix and a bias to obtain multiple output values, wherein each first fault category corresponds to an output value; exponentially process the multiple output values ​​by a normalized exponential function, and weight the exponentially processed output values ​​to obtain the probability corresponding to each first fault category; determine the fault category with the highest probability among the multiple first fault categories as the fault category of the target object.

[0013] According to another embodiment of the present application, a fault classification method is provided, comprising: converting an original vibration signal corresponding to a target object into a time-frequency image through a continuous wavelet transformer; extracting features from the time-frequency image through a feature extraction network to generate a first feature map and a second feature map, wherein the first feature map is used to indicate detail information and global information in the time-frequency image, and the second feature map is used to indicate deep features of the time-frequency image; splicing the first feature map and the second feature map through an attention network to obtain a multi-level feature map, and enhancing the multi-level feature map in terms of channel dimension and spatial dimension to enhance key features of the multi-level feature map and suppress irrelevant features of the multi-level feature map; and performing fault classification on the enhanced multi-level feature map through a classifier to determine the fault category of the target object.

[0014] According to another embodiment of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the steps in the above method embodiment when running.

[0015] According to another embodiment of the present application, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in the above method embodiment.

[0016] According to another embodiment of the present application, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiment are implemented.

[0017] The fault classification system of the embodiment of the present application includes: a continuous wavelet transformer, a feature extraction network connected to the continuous wavelet transformer, an attention network connected to the feature extraction network, and a classifier connected to the attention network, wherein: the continuous wavelet transformer is used to convert the original vibration signal corresponding to the target object into a time-frequency image; the feature extraction network is used to extract features from the time-frequency image to generate a first feature map for indicating detail information and global information in the time-frequency image, and to generate a second feature map for indicating deep features of the time-frequency image; the attention network is used to splice the first feature map and the second feature map to obtain a multi-level feature map, and enhance the multi-level feature map in terms of channel dimension and spatial dimension to enhance the key features of the multi-level feature map and suppress irrelevant features of the multi-level feature map; the classifier is used to classify the multi-level feature map after the enhancement process to determine the fault category of the target object. That is to say, the fault classification system of the present application includes four parts: a continuous wavelet transformer, a feature extraction network, an attention network and a classifier. After the continuous wavelet transformer can convert the original vibration signal into a time-frequency image, the feature extraction network obtains the detail information, global information and deep features of the time-frequency image to generate a first feature map and a second feature map. Then, the attention network performs splicing processing on the first feature map and the second feature map, and enhances the key features of the multi-level feature map after the splicing processing, suppresses the irrelevant features of the multi-level feature map, and the classifier can determine the fault category of the target object according to the multi-level feature map after the enhanced processing. According to the embodiment of the present application, the problem of low recognition accuracy of the method of determining the fault category of the target object through the residual network and the U-Net segmentation network in the related art can be solved, and then the fault category of the target object can be more accurately identified through the four parts of the continuous wavelet transformer, the feature extraction network, the attention network and the classifier. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0020] Figure 1 is a framework diagram of a fault classification system according to an embodiment of the present application;

[0021] Figure 2 is an architecture diagram of a U-ResNet-CBAM model according to an optional embodiment of the present application;

[0022] Figure 3 is a schematic diagram of an original signal of a misalignment fault of a wind turbine generator set according to an optional embodiment of the present application;

[0023] Figure 4 is a schematic diagram of a time-frequency image obtained through continuous wavelet transform according to an optional embodiment of the present application;

[0024] Figure 5 is a structural block diagram of U-Net according to an optional embodiment of the present application;

[0025] Figure 6 is a structural block diagram of ResNet50 according to an optional embodiment of the present application;

[0026] Figure 7 is a structural block diagram of a CBAM according to an optional embodiment of the present application;

[0027] Figure 8 It is a hardware structure block diagram of a computer terminal device of a fault classification method according to an embodiment of the present application;

[0028] Fig. 9 It is a flowchart of a fault classification method according to an embodiment of the present application. DETAILED DESCRIPTION

[0029] Since the beginning of the 21st century, global energy security, ecological environment protection, climate change and other issues have become increasingly serious. With good economic efficiency, the domestic and foreign wind power industry has achieved a growth model of both land and sea. With the annual increase in installed capacity of wind turbines and long-term operation, the number of faulty units has increased significantly, and the accidents and losses caused have also continued to increase. Common faults of wind turbines mainly include: blade damage, gearbox wear, bearing wear, wind turbine failure, control system failure and transmission system misalignment failure. Among them, the misalignment failure of the wind turbine transmission system refers to the misalignment of the rotor shaft system composed of the speed increase gearbox and the generator connected by the coupling. In the actual operation of the wind turbine, due to the installation error of the unit, deformation after loading and uneven sinking of the installation foundation of the unit, it is easy to cause misalignment between the high-speed output end axis of the speed increase gearbox and the generator shaft axis when the unit is working. Misalignment failure will inevitably cause unit vibration and endanger the reliability of the transmission and power generation system. Therefore, the monitoring and diagnosis of transmission system misalignment failure has become an important research content in the development of wind power technology.

[0030] The misalignment fault diagnosis methods of transmission systems can be mainly divided into signal analysis methods, machine learning methods and deep learning methods.

[0031] Signal analysis methods mainly include time domain-based analysis methods, frequency domain-based analysis methods, and time-frequency domain-based analysis methods. However, signal analysis methods mainly rely on artificial prior knowledge to extract frequencies and judge the spectrum of signals collected on site, which places high demands on staff. With the development of computer technology, machine learning fault diagnosis methods are becoming more and more widely used. Although machine learning has achieved relatively remarkable results, it cannot achieve automatic feature extraction and cannot adapt well to the needs of fault diagnosis under big data. In recent years, deep learning has been widely used in fault diagnosis due to its advantages such as powerful feature learning, multimodal data fusion, and the ability to process large-scale data. In order to process complex monitoring data, a variety of different deep neural networks have been proposed, such as convolutional neural networks, long-term and short-term neural networks, autoencoders, residual neural networks, etc. These research methods can automatically learn deep discriminant features from raw data, which greatly improves the diagnostic accuracy compared with traditional methods.

[0032] Among them, deep convolutional neural networks have made significant progress in the fields of image processing and computer vision, such as image segmentation and image classification. Due to the powerful feature learning ability of convolutional neural networks, they have also been widely used in fault diagnosis. The U-Net network deployed in a fully convolutional manner has shown amazing performance in medical image segmentation tasks. This method uses skip connections to combine the shallow features of the encoder with the deep features of the decoder to ensure that the feature map finally recovered incorporates more low-level features. In order to extract more deep abstract features that are beneficial to classification, the network must be processed in depth. However, as the number of network layers increases, the parameters will increase significantly, and the gradient will disappear or explode as the neural network deepens, making the network difficult to train. In order to alleviate the gradient disappearance caused by the deepening of the network depth, ResNet based on convolutional neural networks was proposed. The most important element in ResNet is the residual block, which can be divided into two types: identity mapping and projection mapping. ResNet50 is formed by stacking multiple identity mapping blocks and projection mapping blocks after the first convolution layer. In order to emphasize the importance of features and improve the expressiveness of the model, enhance key feature areas, and suppress irrelevant features, the Convolutional Block Attention Module (CBAM) is proposed. When processing a given output feature map, the module uses two independent attention mechanisms: channel attention mechanism and spatial attention mechanism. The channel attention mechanism (CAM) realizes adaptive weighting of feature channel dimensions by quantitatively evaluating the importance of different channels. The spatial attention mechanism (SAM) emphasizes important areas and suppresses secondary or redundant information by weighting spatial positions. The two attention mechanisms generate corresponding weight maps respectively, which are fused with the input feature map in an element-by-element multiplication manner to realize adaptive refinement of features. The present invention proposes a wind turbine misalignment fault diagnosis method based on U-Net and ResNet50 network feature fusion, which aims to adaptively extract effective features from the vibration signals of complex faulty machinery to realize fault diagnosis.

[0033] For example, in the prior art, Minjiang College proposed a white blood cell segmentation method based on UNet and ResNet in its patent application "White blood cell segmentation method based on U-Net and ResNet" (patent application number: 202011513533.1, application publication number: CN112508931A). The implementation of this method includes three stages: feature encoding stage, feature refinement stage, and feature decoding stage. Among them, in the feature encoding stage, a context-aware feature encoder with residual convolution is used to extract multi-scale feature maps; in the feature refinement stage, parallel multi-scale hole convolution is used to capture multi-scale feature map information to obtain higher-level semantic information; in the feature decoding stage, a feature decoder with convolution and bilinear interpolation is used to adjust the size of the multi-scale feature map to achieve end-to-end white blood cell segmentation. This method uses the U-Net network framework, and uses ResNet to extract features in its downsampling stage (encoder) to obtain multi-scale feature maps, and uses a feature decoder with convolution and bilinear interpolation to reconstruct the feature map in the upsampling stage (decoder). This method does not enhance the feature map, and the segmentation capability of the model needs to be improved.

[0034] In order to solve the problem of low recognition accuracy of the method of determining the fault category of the target object through the residual network and the U-Net segmentation network in the above-mentioned related art. The optional embodiments of the present application combine the advantages of the Continuous Wavelet Transform (CWT), the U-Net (U-Net: Convolutional Networks for Biomedical Image Segmentation, U-Net for short), the 50-layer residual network (Deep Residual Learning for Image Recognition, ResNet50 for short), the Convolutional Block Attention Module (CBAM for short) and the normalized exponential function classifier (Softmax classifier), and propose a new model architecture - the U-ResNet-CBAM model (i.e., the fault classification system). The U-ResNet-CBAM model integrates the multi-scale analysis and high time and frequency resolution processing capabilities of CWT, the detail and global information reconstruction capabilities of U-Net, the deep feature learning capabilities of ResNet50, the ability of CBAM to focus on key feature areas and reduce interference from irrelevant features, and the highly accurate classification capabilities of the Softmax classifier in multi-classification tasks. This combined advantage enables the model to perform excellently in wind turbine fault diagnosis and classification tasks, and can more accurately identify and classify complex fault modes.

[0035] The embodiments of the present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0036] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0037] In this embodiment, a fault classification system is provided. Figure 1 is a framework diagram of a fault classification system according to an embodiment of the present application, such as Figure 1 As shown, the system includes: a continuous wavelet transformer 12, a feature extraction network 14 connected to the continuous wavelet transformer, an attention network 16 connected to the feature extraction network, and a classifier 18 connected to the attention network, wherein:

[0038] The continuous wavelet transformer 12 is used to convert the original vibration signal corresponding to the target object into a time-frequency image;

[0039] Among them, the above-mentioned target object may be a wind turbine.

[0040] The feature extraction network 14 is used to extract features from the time-frequency image to generate a first feature map and a second feature map, wherein the first feature map is used to indicate detail information and global information in the time-frequency image, and the second feature map is used to indicate deep features of the time-frequency image;

[0041] The attention network 16 is used to perform splicing processing on the first feature map and the second feature map to obtain a multi-level feature map, and perform channel dimension and spatial dimension enhancement processing on the multi-level feature map to enhance key features of the multi-level feature map and suppress irrelevant features of the multi-level feature map;

[0042] The classifier 18 is used to perform fault classification on the enhanced multi-level feature graph to determine the fault category of the target object.

[0043] That is to say, the optional embodiment of the present application combines the advantages of CWT, U-Net, ResNet50, CBAM and Softmax classifiers to propose a new system architecture - U-ResNet-CBAM model (i.e., fault classification system), such as Figure 2 As shown, Figure 2This is an architecture diagram of a U-ResNet-CBAM model according to an optional embodiment of the present application, wherein the above-mentioned U-Net and ResNet50 are feature extraction networks, CBAM is an attention network, and Softmax classifier is a classifier. The U-ResNet-CBAM model combines the detail and global information reconstruction capabilities of U-Net, the deep feature learning capabilities of ResNet50, the ability of CBAM to focus on key feature areas and reduce interference from irrelevant features, and the highly accurate classification capabilities of the Softmax classifier in multi-classification tasks.

[0044] Specifically: The vibration signal of the collected wind turbine is converted into a time-frequency image using CWT. The encoding and decoding structure of U-Net enables it to effectively capture and reconstruct the details and global information in the image. By finely processing the context information of the feature map, it provides a solid foundation for high-precision image analysis. At the same time, ResNet50 can extract richer deep features from the data through its deep residual learning structure. Its residual block design enables the network to avoid the gradient vanishing problem while maintaining efficient training, which helps to identify more complex patterns and features, which is especially important when processing high-dimensional time-frequency images. Finally, CBAM can dynamically adjust the weights of the feature map by introducing channel and spatial attention mechanisms, so that the model can focus more on important feature areas and reduce interference with irrelevant features, thereby improving classification accuracy. By combining these modules together, the U-ResNet-CBAM model can not only efficiently extract and reconstruct image features, but also strengthen the ability to focus on key features through deep learning and attention mechanisms. This combined advantage enables the model to have excellent performance in the task of wind turbine misalignment fault diagnosis and classification, and can more accurately identify and classify complex fault modes.

[0045] The system includes: a continuous wavelet transformer, a feature extraction network connected to the continuous wavelet transformer, an attention network connected to the feature extraction network, and a classifier connected to the attention network, wherein: the continuous wavelet transformer is used to convert the original vibration signal corresponding to the target object into a time-frequency image; the feature extraction network is used to extract features from the time-frequency image to generate a first feature map for indicating detail information and global information in the time-frequency image, and to generate a second feature map for indicating deep features of the time-frequency image; the attention network is used to splice the first feature map and the second feature map to obtain a multi-level feature map, and to enhance the multi-level feature map in terms of channel dimension and spatial dimension to enhance the key features of the multi-level feature map and suppress irrelevant features of the multi-level feature map; the classifier is used to classify the faults of the enhanced multi-level feature map to determine the fault category of the target object. That is to say, the fault classification system of the present application includes four parts: a continuous wavelet transformer, a feature extraction network, an attention network and a classifier. After the continuous wavelet transformer can convert the original vibration signal into a time-frequency image, the feature extraction network obtains the detail information, global information and deep features of the time-frequency image to generate a first feature map and a second feature map. Then, the attention network performs splicing processing on the first feature map and the second feature map, and enhances the key features of the multi-level feature map after the splicing processing, suppresses the irrelevant features of the multi-level feature map, and the classifier can determine the fault category of the target object according to the multi-level feature map after the enhanced processing. According to the embodiment of the present application, the problem of low recognition accuracy of the method of determining the fault category of the target object through the residual network and the U-Net segmentation network in the related art can be solved, and then the fault category of the target object can be more accurately identified through the four parts of the continuous wavelet transformer, the feature extraction network, the attention network and the classifier.

[0046] Optionally, the above-mentioned continuous wavelet transformer 12 is also used to select a target wavelet basis function according to the frequency range of the original vibration signal, and determine multiple scale parameters and multiple translation parameters corresponding to the target wavelet basis function, wherein the scale parameter is used to adjust the time scale of the original vibration signal, and the translation parameter is used to adjust the translation position of the original vibration signal; determine multiple eigenvalues ​​of the original vibration signal at different time scales and different translation positions according to a first formula, and construct the time-frequency image according to the multiple eigenvalues, wherein the first formula is: x(t) is the original vibration signal, a is the scale parameter, b is the translation parameter, W(a,b) is the eigenvalue, t is the time variable corresponding to the converted time-frequency image, is the complex conjugate of the target wavelet basis function.

[0047] It is understandable that Continuous Wavelet Transform (CWT) is used as a preprocessing step to extract time-frequency information from the original vibration signal of the wind turbine and generate a time-frequency image. CWT can automatically adjust the time resolution according to the different frequency characteristics of the signal, and is very suitable for analyzing non-stationary signals, such as the vibration signal of the wind turbine, which has characteristics that change over time.

[0048] First, you need to select the target wavelet basis function: the selection of the wavelet basis function is based on the frequency range of the original vibration signal, so that the characteristics of the signal at different frequencies can be effectively captured. Common wavelet basis functions include Morlet wavelet, Meyer wavelet, etc. Each wavelet function is suitable for different types of signal analysis.

[0049] Secondly, it is necessary to determine the scale parameter and translation parameter: the scale parameter is used to adjust the time scale of the wavelet basis function, which determines the degree of localization of the signal in the time domain. When a increases, the scale of the wavelet basis function also increases, which means that the resolution in the time domain decreases, but the resolution in the frequency domain increases, which is more suitable for capturing the characteristics of low-frequency signals. Conversely, when a decreases, the scale of the wavelet basis function decreases, the resolution in the time domain increases, and the resolution in the frequency domain decreases, which is more suitable for capturing the characteristics of high-frequency signals. The translation parameter is used to adjust the position of the wavelet basis function on the time axis, which ensures that the wavelet transform can capture the local characteristics of the signal at different time points. By changing the value of b, the entire signal can be scanned in the time domain to obtain the characteristic values ​​at different time points.

[0050] Finally, the first formula can be used to determine multiple eigenvalues ​​at different time scales and different translation positions, and then construct a time-frequency image: Through different scale parameters (a) and translation parameters (b), the continuous wavelet transformer will generate a series of eigenvalues ​​(W(a,b)), which constitute the time-frequency image. In the time-frequency image, the horizontal axis represents time, the vertical axis represents scale (indirectly reflects frequency), and the color or brightness represents the size of the eigenvalue, that is, the energy distribution of the signal.

[0051] In an alternative embodiment, Figure 3 is a schematic diagram of an original signal (i.e., an original vibration signal) of a misalignment fault of a wind turbine according to an optional embodiment of the present application, Figure 4 FIG. 1 is a schematic diagram of a time-frequency image obtained by continuous wavelet transform according to an optional embodiment of the present application. Figure 3 , Figure 4 As shown:

[0052] Sensors can be used to collect various operating conditions of wind turbines, and the obtained vibration signals are intercepted using a sliding window with a window size of 512 and a step size of 128 to obtain vibration data samples (i.e. Figure 3 The original vibration signal in ).

[0053] After constructing the vibration data sample, the continuous wavelet transformer (CWT) can be used to convert the vibration data sample into a time-frequency signal. Figure 4 Time-frequency images in .

[0054] CWT can automatically adjust the time resolution according to the frequency characteristics of the signal, provide fine detail features and extensive overview information, and can adapt to the analysis needs of various types of signals by flexibly adjusting the parameters of the wavelet basis function. It presents obvious ridge features in the joint time-frequency domain, namely the wavelet ridge line, making the time-frequency changes of the signal more intuitive and easy to find abnormal or fault features. For a given original vibration signal x(t), the definition of CWT is as follows: Among them, a is the scale parameter, b is the translation parameter, W(a,b) is the eigenvalue, t is the time variable corresponding to the converted time-frequency image, and is the complex conjugate of the target wavelet basis function. The scale parameter a changes the oscillation frequency by stretching or compressing the wavelet function, and the translation parameter b changes the position of the time window. The original vibration signal samples of the wind turbine are converted into time-frequency images as the input of the feature extraction network using CWT.

[0055] Optionally, the feature extraction network includes: a joint path network and a residual network, wherein the joint path network includes: an input layer, a downsampling path layer, a bottleneck layer, an upsampling path layer and an output layer, wherein: the input layer is used to receive the time-frequency image; the downsampling path layer includes multiple first convolutional layers and a pooling layer connected to each first convolutional layer, wherein each first convolutional layer is used to identify frequency domain features and time domain features in the time-frequency image to generate each third feature map, each pooling layer is used to reduce the size of each third feature map, each first convolutional layer corresponds to a different number of output channels, and each third feature map after the size reduction is used to indicate the global features of the time-frequency image; the bottleneck layer is used to perform feature concentration processing on each third feature map after the size reduction, and to generate each third feature map. The third feature map after each feature concentration process is transmitted to the upsampling path layer; the upsampling path layer includes multiple deconvolution layers and each second convolution layer connected to each deconvolution layer, wherein each deconvolution layer is used to perform deconvolution processing on the third feature map after each feature concentration process to increase the size of the third feature map after each feature concentration process, and the third feature map after deconvolution processing is used to indicate the detail features of the time-frequency image; each second convolution layer is used to fuse each third feature map after deconvolution processing with each third feature map after size reduction to generate each fused feature map, and the multiple first convolution layers and the multiple deconvolution layers correspond one to one; the output layer is used to fuse multiple fused feature maps to generate the first feature map.

[0056] It is understood that the joint path network may include: an input layer, a downsampling path layer, a bottleneck layer, an upsampling path layer and an output layer, wherein:

[0057] Input layer: responsible for receiving the preprocessed time-frequency image. The time-frequency image contains the energy distribution of the wind turbine vibration signal at each time point and frequency.

[0058] Downsampling path layer (encoder): includes multiple first convolutional layers: The downsampling path of U-Net consists of multiple first convolutional layers, each of which is responsible for identifying local features in the time-frequency image, such as frequency domain features (vibration patterns of different frequencies) and time domain features (temporal changes in vibration signals). The convolutional layer performs convolution operations on the input image through its weight parameters to generate multiple third feature maps. Different convolutional layers have different numbers of output channels, which means that they capture different feature dimensions and complexities. It also includes pooling layers connected to each first convolutional layer: The role of the pooling layer is to reduce the size of the feature map while maintaining the most important information of the feature map. In U-Net, the pooling layer usually uses maximum pooling or average pooling, which can reduce the amount of calculation and parameters, while making the model more robust to position changes. The reduced-size third feature map can indicate the global characteristics of the time-frequency image, that is, the overall vibration pattern and trend.

[0059] Bottleneck layer: The third feature maps obtained in the downsampling path are subjected to feature concentration, that is, the high-dimensional expressions of these feature maps are further extracted and fused to generate more abstract feature representations, thereby reducing feature redundancy while maintaining the key information of the signal.

[0060] Upsampling path layer (decoder): includes multiple deconvolution layers: The upsampling path is composed of multiple deconvolution layers (or upsampling layers), which are used to deconvolve the third feature map after feature concentration processing, that is, to increase the size of these feature maps to restore the detail information compressed by the downsampling path layer. The third feature map after deconvolution indicates the detail features of the time-frequency image. And each second convolution layer connected to each deconvolution layer: Each second convolution layer fuses the third feature map after deconvolution processing with the third feature map of the corresponding level in the downsampling path to generate each fused feature map. This fusion process is achieved through jump connections, so that the decoder can use the feature information of the encoder part to reconstruct the details of the time-frequency image.

[0061] Output layer: responsible for fusing multiple fused feature maps to generate the final first feature map. In U-Net, this usually involves a convolution operation to reduce the number of channels of the feature map, making the feature map more suitable for subsequent classification tasks. The first feature map is obtained after the encoder-decoder path and contains the details and global information of the time-frequency image.

[0062] In an optional embodiment, the structure of U-Net is as follows: input layer → downsampling path (encoder) → bottleneck layer → upsampling path (decoder) → output layer. The design feature of U-Net is its symmetrical encoder-decoder structure and the information is transmitted through jump connections. In the implementation of wind turbine misalignment diagnosis, the input size is compressed from 256×256×3 to 16×16×1024 and then decoded back to 256×256×1. The model structure is as follows Figure 5 The size change of the time-frequency image after U-Net downsampling is shown in Table 1, and the size change of the time-frequency image after U-Net upsampling is shown in Table 2. The parameters of each U-Net module include:

[0063] 1) Input layer: The input size is 256×256×3 (i.e., input time-frequency image).

[0064] 2) Downsampling path (encoder):

[0065] Layer 1: Two 3×3 convolutional layers (output channels: 64), each followed by a ReLU activation. One 2×2 max pooling layer (stride: 2) with an output size of 128×128×64.

[0066] Layer 2: Two 3×3 convolutional layers (output channels: 128), each followed by a ReLU activation. One 2×2 max pooling layer (stride: 2) with an output size of 64×64×128.

[0067] Layer 3: Two 3×3 convolutional layers (output channels: 256), each followed by a ReLU activation. One 2×2 max pooling layer (stride: 2) with an output size of 32×32×256.

[0068] Layer 4: Two 3×3 convolutional layers (output channels: 512), each followed by a ReLU activation. One 2×2 max pooling layer (stride: 2) with an output size of 16×16×512.

[0069] 3) Bottleneck layer: two 3×3 convolutional layers (output channels: 1024), each followed by a ReLU activation. The output size is 16×16×1024.

[0070] 4) Upsampling path (decoder):

[0071] Layer 4': One 2×2 deconvolution layer (output channels: 512) with an output size of 32×32×512. Skip connection: concatenate the output of layer 4 of the encoder part (32×32×512). Two 3×3 convolution layers (output channels: 512), each followed by a ReLU activation.

[0072] Layer 3': One 2×2 deconvolution layer (output channels: 256) with an output size of 64×64×256. Skip connection: concatenate the output of layer 3 of the encoder part (64×64×256). Two 3×3 convolutional layers (output channels: 256), each followed by a ReLU activation.

[0073] Layer 2': One 2×2 deconvolution layer (output channels: 128) with an output size of 128×128×128. Skip connection: concatenate the output of layer 2 of the encoder part (128×128×128). Two 3×3 convolution layers (output channels: 128), each followed by a ReLU activation.

[0074] Layer 1': One 2×2 deconvolution layer (output channels: 64) with an output size of 256×256×64. Skip connection: concatenate the output of the encoder layer 1 (256×256×64). Two 3×3 convolutional layers (output channels: 64), each followed by a ReLU activation.

[0075] 5) Output layer: A 1×1 convolutional layer (output channel: 1) is used to generate the final feature map with an output size of 256×256×1. A 32×32 convolutional layer (output channel: 2048) is used to generate the fused feature map with an output size of 8×8×2048.

[0076] Table 1. Changes in the size of time-frequency images after U-Net downsampling

[0077]

[0078] Table 2. Changes in the size of time-frequency images after U-Net upsampling

[0079]

[0080] Optionally, the above-mentioned residual network is used to: convert the time-frequency image into a pseudo RNB image; perform a convolution operation on the pseudo RNB image to extract local features corresponding to the time-frequency image; perform a maximum pooling operation on the pseudo RNB image after the local features are extracted to reduce the size of the pseudo RNB image after the local features are extracted, wherein the pseudo RNB image after the size reduction retains the local features; and learn the multi-level features of the pseudo RNB image after the size reduction according to the residual block in the residual network to obtain the second feature map.

[0081] It can be understood that the residual network can convert the time-frequency image into a pseudo RNB image: after the U-Net processes the time-frequency image, the generated first feature map is input into ResNet50. The pseudo RNB image actually refers to the time-frequency image obtained after CWT conversion, which is regarded as the input image by ResNet50 for subsequent feature learning and classification tasks. ResNet50 can learn deeper feature representations from time-frequency images through specific convolution operations.

[0082] Perform convolution operation on the pseudo RNB image: The convolution operation can identify local features in the image, such as the shape of the vibration mode, the pattern of frequency distribution, etc.

[0083] Max pooling operation to reduce size: After the convolution operation, ResNet50 performs a max pooling operation on the pseudo RNB image to further reduce the size of the image. The max pooling operation selects the maximum value from each local region of the feature map, thereby reducing the spatial dimension of the feature map while retaining key information and local features in the image.

[0084] Residual blocks learn multi-level features: The core of ResNet50 is the residual block. The design of these residual blocks can solve the gradient vanishing and gradient exploding problems in deep network training, allowing the network to learn more complex feature representations. Through the learning of residual blocks, ResNet50 can extract more levels of features from the reduced-size pseudo RNB image, including more abstract and higher-level pattern information. Each residual block consists of multiple convolutional layers, in which the skip connection allows the network to directly pass part of the feature map to the output of the block, thereby preserving the continuity of the information flow, allowing the network to maintain stable training while increasing depth.

[0085] Obtain the second feature map: The second feature map is the result of ResNet50 performing a series of convolutions, pooling, and residual block learning on the input time-frequency image. It contains the multi-level features of the wind turbine vibration signal.

[0086] In an optional embodiment, the structure of the 50-layer residual network ResNet50 is: input layer→feature extraction layer→max pooling layer→combination module→output layer;

[0087] Among them, the combination module is composed of four stage modules connected in sequence, each stage module contains several residual modules, and each residual module consists of three convolutional layers.

[0088] Among them, the output of the third convolutional layer is connected to the input of the first convolutional layer. The model structure is as follows Figure 6 shown.

[0089] The ResNet50 time-frequency image size transformation is shown in Table 3:

[0090] Table 3. ResNet50 time-frequency image size transformation table

[0091]

[0092] The parameters of each ResNet50 module used include:

[0093] Stage 0: The feature map of the feature extraction layer is set to 64, the convolution kernel size is set to 7×7 pixels, and the stride is set to 2 pixels.

[0094] The first phase module (i.e. Figure 6 The feature maps of each convolutional layer in the three residual modules in stage 1) are set to 64, 64, and 256 respectively, the convolution kernel sizes are set to 1×1, 3×3, and 1×1 pixels respectively, and the strides are all set to 1 pixel. The stride of the first convolutional layer of the first residual module needs to be set to 1 to match the input size.

[0095] The second phase module (i.e. Figure 6 The feature maps of each convolutional layer in the four residual modules in stage 2) are set to 128, 128, and 512 respectively, and the convolution kernel sizes are set to 1×1, 3×3, and 1×1 pixels respectively. The stride of the first convolutional layer of the first residual module of the second combined module is set to 2 pixels, and the strides of other convolutional layers are all set to 1 pixel.

[0096] The third phase module (i.e. Figure 6 The feature maps of each convolutional layer in the six residual modules in stage 3) are set to 256, 256, and 1024 respectively, and the convolution kernel sizes are set to 1×1, 3×3, and 1×1 pixels respectively. The stride of the first convolutional layer of the first residual module of the third combination module is set to 2 pixels, and the strides of other convolutional layers are all set to 1 pixel.

[0097] The fourth phase module (i.e. Figure 6 The feature maps of each convolutional layer in the three residual modules in stage 4) are set to 512, 512, and 2048 respectively, and the convolution kernel sizes are set to 1×1, 3×3, and 1×1 pixels respectively. The stride of the first convolutional layer of the first residual module of the fourth combination module is set to 2 pixels, and the strides of other convolutional layers are all set to 1 pixel.

[0098] Optionally, the attention network is also used to: input the fourth feature map into the spatial attention network in the attention network, so that the spatial attention module performs global average pooling on the fourth feature map in the spatial dimension to obtain a fifth feature map, and performs global maximum pooling on the fourth feature map in the spatial dimension to obtain a sixth feature map; splice the fifth feature map and the sixth feature map, and perform convolution on the spliced ​​feature map to generate a spatial attention weight vector; multiply the fourth feature map element by element by the spatial attention weight vector to obtain the multi-level feature map after the enhanced processing.

[0099] It is understandable that the introduction of the attention network, especially the convolutional block attention module (CBAM), is to enhance the model's ability to focus on key feature areas while reducing the interference of irrelevant or minor features, thereby improving the accuracy of fault diagnosis. Specifically: Input the fourth feature map: The fourth feature map is the result of the preliminary fusion of the feature maps generated by U-Net and ResNet50. It contains multi-level information of the wind turbine misalignment fault, including local details and global patterns.

[0100] Global average pooling and global maximum pooling: In order to generate spatial attention weights, the fourth feature map needs to be processed by global average pooling and global maximum pooling in the spatial dimension. Global average pooling calculates the average value of each channel in the feature map to obtain the fifth feature map; global maximum pooling finds the maximum value of each channel to generate the sixth feature map.

[0101] Concatenate and convolve: Concatenate the fifth and sixth feature maps to form a new feature map, which contains the information after global average pooling and global maximum pooling. Then, the spatial attention weight vector is generated through convolution operation. The convolution operation here usually uses a smaller convolution kernel (1x1 or 7x7) in order to generate a weight map of the same size as the input feature map.

[0102] Generate spatial attention weight vector: The spatial attention weight vector is an evaluation of the importance of each spatial position in the feature map. The weight map generated by the above convolution process can intuitively show which areas are important and which areas can be suppressed. Each value in the weight map represents the weight of the corresponding spatial position. The larger the value, the more important the feature at that position.

[0103] Element-by-element multiplication to obtain an enhanced feature map: Finally, the fourth feature map is element-by-element multiplication with the spatial attention weight vector, i.e., element-level multiplication. This step adjusts the feature values ​​in the fourth feature map according to the spatial attention weights at each position. In this way, the model can adaptively enhance the features of key areas while suppressing information in irrelevant or secondary areas, thereby generating an enhanced multi-level feature map.

[0104] In an optional embodiment, in order to emphasize the importance of features and improve the expressiveness of the model, enhance key feature areas, and suppress irrelevant features, convolutional CBAM is introduced in the feature maps generated by U-Net and ResNet50. When processing a given output feature map, CBAM uses two independent attention mechanisms: channel attention mechanism and spatial attention mechanism. The channel attention mechanism (CAM) quantitatively evaluates the importance of different channels to achieve adaptive weighting of feature channel dimensions. The spatial attention mechanism (SAM) emphasizes important areas and suppresses secondary or redundant information by assigning weights to spatial positions. The two attention mechanisms generate corresponding weight maps respectively, which are merged with the input feature map in an element-by-element multiplication manner. This process is called adaptive refinement of features. The model structure is as follows: Figure 7 The specific steps include:

[0105] Feature map fusion: The two input feature maps are concatenated along the channel dimension. The size of the concatenated feature map is 8×8×4096.

[0106] Channel attention module: The size of the input feature map is 8×8×4096. First, global average pooling and global maximum pooling are performed to obtain two 1×1×4096 feature vectors. Subsequently, these two feature vectors are added and activated after passing through a shared fully connected layer to generate a channel attention weight vector. Finally, the input feature map is multiplied by the channel attention weight channel by channel to form an enhanced feature map.

[0107] Spatial attention module: First, the enhanced feature map formed after the channel attention module is used to perform global average pooling and global maximum pooling in the spatial dimension to form two 8×8×1 feature maps. Subsequently, the two feature maps are concatenated along the channel dimension into an 8×8×2 feature map. A convolution layer with a convolution kernel size of 7×7 is used to convolve the feature map to generate a spatial attention weight map. Then, the spatial attention weight is generated by the Sigmoid activation function. Finally, the input feature map is element-wise multiplied with the spatial attention weight to form the final output feature map.

[0108] Output feature map (i.e. multi-level feature map): The size of the feature map processed by CBAM is 8×8×4096.

[0109] Optionally, the classifier is also used to: perform global average pooling processing on the multi-level feature map after enhancement processing to obtain a third feature vector; input the third feature vector into the second fully connected layer in the classifier, so that the second fully connected layer updates the dimension of the third feature vector, wherein the dimension of the third feature vector after the updated dimension is equal to the number of fault categories corresponding to the multiple first fault categories corresponding to the target object; map the third feature vector after the updated dimension to the fault category space according to the weight matrix and the bias to obtain multiple output values, wherein each first fault category corresponds to an output value; index the multiple output values ​​respectively by a normalized exponential function, and weight the indexed output values ​​to obtain the probability corresponding to each first fault category; determine the fault category with the highest probability among the multiple first fault categories as the fault category of the target object.

[0110] It can be understood that the classifier can perform the following steps: Global average pooling: After feature extraction and attention mechanism processing, the model generates a multi-level feature map after enhancement. In order to convert it into a classifiable vector form, it is first necessary to perform global average pooling on this feature map. Global average pooling is to calculate the average value of the feature map in the spatial dimension to generate a fixed-size feature vector, namely the third feature vector.

[0111] Update the dimension of the third feature vector: The third feature vector is input to the second fully connected layer in the classifier, which is responsible for updating the dimension of the vector. The fully connected layer can be understood as a linear transformation, which processes the input feature vector by weighted summation to change its dimension. The dimension of the updated third feature vector is equal to the number of misalignment fault categories of the target object (i.e., wind turbine), which means that each fault category will correspond to a feature value.

[0112] Mapping to fault category space: The third eigenvector after the updated dimension is further linearly mapped through the weight matrix and bias term of the classifier, converting the numerical value of the eigenvector to the fault category space, generating multiple output values, each of which corresponds to a fault category. This process is essentially learning a function that maps the obtained eigenvector to a probability space, where each value represents the probability of the corresponding fault category.

[0113] Softmax function for indexation and weighting: The Softmax function is used to index the multiple output values ​​after mapping, and then normalize them to generate a probability distribution. The function of the Softmax function is to convert the output value into a probability so that the sum of the probabilities of all possible fault categories is 1.

[0114] Determine the fault category: Compare the probabilities of each fault category and select the fault category with the highest probability as the final diagnosis result.

[0115] In an optional embodiment, the steps of constructing a Softmax classifier may be: first, the feature map of size 8×8×4096 formed after the CBAM module is average pooled to obtain a feature map of size 1×1×4096. Then, the feature map of size 1×1×4096 is flattened and input into a fully connected classifier. The number of nodes in the first layer of the classifier is 4096, the number of nodes in the second layer is 1024, and the number of nodes in the third layer is 4.

[0116] The time-frequency images are input into the U-Net and ResNet networks in parallel for deep feature extraction. The obtained multi-level feature maps are fused and input into the convolutional block attention module for channel and spatial dimension enhancement. The enhanced feature maps are compressed and flattened and input into the Softmax classifier for training. Cross entropy can be used as the loss function: Among them, M is the number of categories, yic is the compliance function, and pic is the observed sample.

[0117] Then, the time-frequency image test set data is input into the U-ResNet-CBAM model and classified and labeled. The evaluation indicators are accuracy, precision, recall and F1-Score. The diagnosis process includes:

[0118]

[0119] Among them, TP, FP, FN, and TN represent the number of true positive, false positive, false negative, and true negative results, respectively.

[0120] In summary, the embodiment of the present application inputs the time-frequency image obtained by CWT into the U-Net and ResNet50 networks to extract detail features and capture deep abstract features. Subsequently, the feature maps generated by the two networks are fused and input into the convolutional block attention module. The convolutional block attention module can intelligently identify and enhance key feature areas through its unique channel and spatial attention mechanism, while effectively suppressing irrelevant features, thereby significantly improving the classification and detection accuracy of the model. Finally, the fused feature map is compressed and flattened and input into the Sofxmax classifier for fault diagnosis. That is, the optional embodiment of the present application proposes a new model architecture for the vibration signal of the wind turbine generator - the U-ResNet-CBAM model; that is, first use CWT to convert the vibration signal into time-frequency to obtain a time-frequency image, and then combine the advantages of U-Net in detail and global information reconstruction with the deep feature learning ability of ResNet50 to obtain an enhanced fused feature map (this method has a significant improvement in feature extraction capabilities compared to a single model), and then use CBAM to focus on key feature areas and suppress irrelevant or secondary feature information. This refined feature processing greatly improves the accuracy of the model in classification tasks.

[0121] The method embodiments provided in the embodiments of the present application can be executed in a computer terminal device or a similar computing device. Taking running on a computer terminal device as an example, Figure 8 1 is a hardware structure block diagram of a computer terminal device of a fault classification method according to an embodiment of the present application. Figure 1 As shown, the computer terminal device may include one or more ( Figure 8 Only one is shown in the figure) a processor 802 (the processor 802 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 804 for storing data, wherein the above-mentioned computer terminal device may also include a transmission device 806 and an input and output device 808 for communication functions. It can be understood by those skilled in the art that Figure 8 The structure shown is only for illustration and does not limit the structure of the above-mentioned computer terminal device. Figure 8 More or fewer components as shown, or with Figure 8 Different configurations are shown.

[0122] The memory 804 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the fault classification method in the embodiment of the present application. The processor 802 executes various functional applications and data processing by running the computer program stored in the memory 804, that is, to implement the above method. The memory 804 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 804 may further include a memory remotely arranged relative to the processor 802, and these remote memories can be connected to the computer terminal device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0123] The transmission device 806 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of a computer terminal device. In one example, the transmission device 806 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 806 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0124] This embodiment also provides a fault classification method. Fig. 9 is a flowchart of a fault classification method according to an embodiment of the present application, such as Fig. 9 As shown, the process includes the following steps:

[0125] Step S202, converting the original vibration signal corresponding to the target object into a time-frequency image through a continuous wavelet transformer;

[0126] Step S204, extracting features from the time-frequency image through a feature extraction network to generate a first feature map and a second feature map, wherein the first feature map is used to indicate detail information and global information in the time-frequency image, and the second feature map is used to indicate deep features of the time-frequency image;

[0127] Step S206, concatenating the first feature map and the second feature map through an attention network to obtain a multi-level feature map, and performing channel dimension and spatial dimension enhancement processing on the multi-level feature map to enhance key features of the multi-level feature map and suppress irrelevant features of the multi-level feature map;

[0128] Step S208: performing fault classification on the enhanced multi-level feature graph through a classifier to determine the fault category of the target object.

[0129] Through the above method, the original vibration signal corresponding to the target object is converted into a time-frequency image through a continuous wavelet transformer; the time-frequency image is feature extracted through a feature extraction network to generate a first feature map for indicating detail information and global information in the time-frequency image, and a second feature map for indicating deep features of the time-frequency image is generated; the first feature map and the second feature map are spliced ​​through an attention network to obtain a multi-level feature map, and the multi-level feature map is enhanced in terms of channel dimension and spatial dimension to enhance the key features of the multi-level feature map and suppress irrelevant features of the multi-level feature map; the enhanced multi-level feature map is fault classified through a classifier to determine the fault category of the target object. That is to say, after the continuous wavelet transformer of the fault classification method of the present application converts the original vibration signal into a time-frequency image, the feature extraction network obtains the detail information, global information and deep features of the time-frequency image to generate a first feature map and a second feature map, and then the attention network performs splicing processing on the first feature map and the second feature map, and enhances the key features of the multi-level feature map after the splicing processing, suppresses the irrelevant features of the multi-level feature map, and the classifier can determine the fault category of the target object according to the multi-level feature map after the enhanced processing. According to the embodiment of the present application, the problem of low recognition accuracy of the method of determining the fault category of the target object through the residual network and the U-Net segmentation network in the related art can be solved, and then the fault category of the target object can be more accurately identified through the four parts of the continuous wavelet transformer, the feature extraction network, the attention network and the classifier.

[0130] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0131] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps in the above method embodiment when running.

[0132] Optionally, in this embodiment, the storage medium may be configured to store program codes for executing the following steps:

[0133] S1, converting the original vibration signal corresponding to the target object into a time-frequency image through a continuous wavelet transformer;

[0134] S2, extracting features from the time-frequency image through a feature extraction network to generate a first feature map and a second feature map, wherein the first feature map is used to indicate detail information and global information in the time-frequency image, and the second feature map is used to indicate deep features of the time-frequency image;

[0135] S3, concatenating the first feature map and the second feature map through an attention network to obtain a multi-level feature map, and enhancing the multi-level feature map in terms of channel dimension and spatial dimension to enhance key features of the multi-level feature map and suppress irrelevant features of the multi-level feature map;

[0136] S4, performing fault classification on the enhanced multi-level feature graph through a classifier to determine the fault category of the target object.

[0137] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0138] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in the above method embodiment.

[0139] In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0140] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:

[0141] S1, converting the original vibration signal corresponding to the target object into a time-frequency image through a continuous wavelet transformer;

[0142] S2, extracting features from the time-frequency image through a feature extraction network to generate a first feature map and a second feature map, wherein the first feature map is used to indicate detail information and global information in the time-frequency image, and the second feature map is used to indicate deep features of the time-frequency image;

[0143] S3, concatenating the first feature map and the second feature map through an attention network to obtain a multi-level feature map, and enhancing the multi-level feature map in terms of channel dimension and spatial dimension to enhance key features of the multi-level feature map and suppress irrelevant features of the multi-level feature map;

[0144] S4, performing fault classification on the enhanced multi-level feature graph through a classifier to determine the fault category of the target object.

[0145] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in the above method embodiment are implemented.

[0146] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.

[0147] An embodiment of the present application also provides a computer program, which includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above method embodiment.

[0148] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:

[0149] S1, converting the original vibration signal corresponding to the target object into a time-frequency image through a continuous wavelet transformer;

[0150] S2, extracting features from the time-frequency image through a feature extraction network to generate a first feature map and a second feature map, wherein the first feature map is used to indicate detail information and global information in the time-frequency image, and the second feature map is used to indicate deep features of the time-frequency image;

[0151] S3, concatenating the first feature map and the second feature map through an attention network to obtain a multi-level feature map, and enhancing the multi-level feature map in terms of channel dimension and spatial dimension to enhance key features of the multi-level feature map and suppress irrelevant features of the multi-level feature map;

[0152] S4, performing fault classification on the enhanced multi-level feature graph through a classifier to determine the fault category of the target object.

[0153] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail herein.

[0154] Obviously, those skilled in the art should understand that the above modules or steps of the present application can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order from that herein, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.

[0155] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the principles of the present application shall be included in the protection scope of the present application.

Claims

1. A fault classification system, characterized in that: include: A continuous wavelet transformer, a feature extraction network connected to the continuous wavelet transformer, an attention network connected to the feature extraction network, and a classifier connected to the attention network, wherein: The continuous wavelet transformer is used to convert the original vibration signal corresponding to the target object into a time-frequency image; The feature extraction network is used to extract features from the time-frequency image to generate a first feature map and a second feature map, wherein the first feature map is used to indicate detail information and global information in the time-frequency image, and the second feature map is used to indicate deep features of the time-frequency image; The attention network is used to perform splicing processing on the first feature map and the second feature map to obtain a multi-level feature map, and perform channel dimension and spatial dimension enhancement processing on the multi-level feature map to enhance key features of the multi-level feature map and suppress irrelevant features of the multi-level feature map; The classifier is used to perform fault classification on the enhanced multi-level feature graph to determine the fault category of the target object.

2. The system according to claim 1, characterized in that The continuous wavelet transformer is further used for: Selecting a target wavelet basis function according to the frequency range of the original vibration signal, and determining a plurality of scale parameters and a plurality of translation parameters corresponding to the target wavelet basis function, wherein the scale parameter is used to adjust the time scale of the original vibration signal, and the translation parameter is used to adjust the translation position of the original vibration signal; A plurality of eigenvalues ​​of the original vibration signal at different time scales and different translation positions are determined according to a first formula, and the time-frequency image is constructed according to the plurality of eigenvalues, wherein the first formula is: x(t) is the original vibration signal, a is the scale parameter, b is the translation parameter, W(a,b) is the eigenvalue, t is the time variable corresponding to the converted time-frequency image, is the complex conjugate of the target wavelet basis function.

3. The system according to claim 1, characterized in that The feature extraction network comprises: A joint path network and a residual network, wherein the joint path network comprises: an input layer, a downsampling path layer, a bottleneck layer, an upsampling path layer and an output layer, wherein: The input layer is used to receive the time-frequency image; The downsampling path layer includes a plurality of first convolutional layers and a pooling layer connected to each first convolutional layer, wherein each first convolutional layer is used to identify frequency domain features and time domain features in the time-frequency image to generate each third feature map, and each pooling layer is used to reduce the size of each third feature map, each first convolutional layer has a different number of output channels, and each reduced-size third feature map is used to indicate a global feature of the time-frequency image; The bottleneck layer is used to perform feature concentration processing on each of the third feature maps after the size reduction, and transmit each of the third feature maps after the feature concentration processing to the upsampling path layer; The upsampling path layer includes a plurality of deconvolution layers and each second convolution layer connected to each deconvolution layer, wherein each deconvolution layer is used to perform deconvolution processing on each third feature map after feature concentration processing to increase the size of each third feature map after feature concentration processing, and the third feature map after deconvolution processing is used to indicate the detail features of the time-frequency image, and each second convolution layer is used to fuse each third feature map after deconvolution processing with each third feature map after size reduction to generate each fused feature map, and the plurality of first convolution layers and the plurality of deconvolution layers correspond to each other one by one; The output layer is used to fuse multiple fused feature maps to generate the first feature map.

4. The system according to claim 3, characterized in that The residual network is used to: Converting the time-frequency image into a pseudo RNB image; Performing a convolution operation on the pseudo RNB image to extract local features corresponding to the time-frequency image; Performing a maximum pooling operation on the pseudo RNB image after the local features are extracted to reduce the size of the pseudo RNB image after the local features are extracted, wherein the pseudo RNB image after the size reduction retains the local features; The multi-level features of the pseudo RNB image after the size reduction are learned according to the residual block in the residual network to obtain the second feature map.

5. The system according to claim 1, characterized in that The attention network is further used to: receive the first feature map and the second feature map, and concatenate the first feature map and the second feature map to obtain a multi-level feature map; Inputting the multi-level feature map into the channel attention network in the attention network, so that the channel attention network performs global average pooling processing on the multi-level feature map in the channel dimension to obtain a first feature vector, and performs global maximum pooling processing on the multi-level feature map in the channel dimension to obtain a second feature vector; Adding and activating the first feature vector and the second feature vector through a fully connected layer in the attention network to generate a channel attention weight vector; Multiplying the multi-level feature map and the channel attention weight vector channel by channel to enhance key features of the multi-level feature map to obtain a fourth feature map; The fourth feature map is enhanced in terms of spatial dimension to obtain the enhanced multi-level feature map.

6. The system according to claim 5, characterized in that The attention network is further used to: input the fourth feature map into the spatial attention network in the attention network, so that the spatial attention module performs global average pooling processing on the fourth feature map in the spatial dimension to obtain a fifth feature map, and performs global maximum pooling processing on the fourth feature map in the spatial dimension to obtain a sixth feature map; Splicing the fifth feature map and the sixth feature map, and performing convolution processing on the spliced ​​feature map to generate a spatial attention weight vector; The fourth feature map is element-wise multiplied by the spatial attention weight vector to obtain the enhanced multi-level feature map.

7. The system according to claim 1, characterized in that The classifier is also used for: Performing global average pooling processing on the enhanced multi-level feature map to obtain a third feature vector; Inputting the third feature vector into a second fully connected layer in the classifier so that the second fully connected layer updates the dimension of the third feature vector, wherein the dimension of the updated third feature vector is equal to the number of fault categories corresponding to the multiple first fault categories corresponding to the target object; Mapping the third eigenvector after the updated dimension to the fault category space according to the weight matrix and the bias to obtain a plurality of output values, wherein each first fault category corresponds to an output value; Exponentially processing the multiple output values ​​respectively by using a normalized exponential function, and weighting the indexed output values ​​to obtain a probability corresponding to each first fault category; A fault category with the highest probability among the plurality of first fault categories is determined as the fault category of the target object.

8. A fault classification method, characterized in that: include: The original vibration signal corresponding to the target object is converted into a time-frequency image through a continuous wavelet transformer; Extracting features from the time-frequency image through a feature extraction network to generate a first feature map and a second feature map, wherein the first feature map is used to indicate detail information and global information in the time-frequency image, and the second feature map is used to indicate deep features of the time-frequency image; The first feature map and the second feature map are concatenated by an attention network to obtain a multi-level feature map, and the multi-level feature map is enhanced in terms of channel dimension and spatial dimension to enhance key features of the multi-level feature map and suppress irrelevant features of the multi-level feature map; The enhanced multi-level feature graph is subjected to fault classification by a classifier to determine the fault category of the target object.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the method described in claim 8 is executed when the program is executed.

10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method of claim 8 through the computer program.

Citation Information

Patent Citations

  • Satellite image segmentation method based on residual network and U-Net segmentation network

    CN110211137A

  • Satellite image segmentation method based on residual network and U-Net segmentation network

    CN110211137B

  • Leukocyte segmentation method based on U-Net and ResNet

    CN112508931A