Transformer defect voiceprint recognition method and device based on improved LSTM-CNN, medium and product

Through the improved LSTM-CNN network, combined with convolutional neural network, long and short-term memory network and CBMA module, defect soundprint recognition is performed on the transformer, which solves the problem of low defect soundprint recognition accuracy of transformer, and achieves higher recognition accuracy and better generalization capabilities.

CN120164490APending Publication Date: 2025-06-17BEIJING GUOWANG FUDA SCI & TECH DEV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510197143.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The accuracy of the transformer defect soundprint recognition is low, especially when special acoustic sensors have not been widely deployed. The lack of sufficient abnormal sound samples leads to low model recognition accuracy.

Method used

The improved LSTM-CNN network is used, combined with the convolutional neural network, long and short-term memory network and CBMA module to identify the voiceprint of transformer defects. The network preprocesses the audio signal through feature extraction and wavelet noise reduction, and improves the generalization ability of the model through data augmentation methods.

Benefits of technology

The accuracy of transformer defect soundprint recognition is improved, and the model's ability to recognize defect soundprints is enhanced, especially when there are fewer data samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164490A_ABST
    Figure CN120164490A_ABST
Patent Text Reader

Abstract

The invention discloses a transformer defect voiceprint recognition method and device based on an improved LSTM-CNN, a medium and a product, and relates to the technical field of defect detection, and the method comprises the steps: obtaining an original audio signal of a to-be-recognized transformer; preprocessing the original audio signal of the to-be-identified transformer to obtain a preprocessed audio signal of the to-be-identified transformer; the preprocessing comprises feature extraction and wavelet denoising; inputting the preprocessed audio signal of the to-be-recognized transformer into the transformer defect voiceprint recognition model to obtain a predicted value of a defect recognition result of the to-be-recognized transformer; the transformer defect voiceprint recognition model is obtained by training an improved LSTM-CNN network, the improved LSTM-CNN network comprises a convolutional neural network, a long short-term memory network and a CBMA module, and the defect recognition result is that a defect exists or does not exist. The transformer defect voiceprint recognition precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of defect detection, and particularly to a method, device, medium and product for transformer defect voiceprint recognition based on an improved LSTM-CNN. Background Art

[0002] In recent years, with the expansion of the power grid scale and the increase in load, the number of abnormal defects in electrical equipment has also increased significantly, which poses higher requirements and challenges for the operation and maintenance of electrical equipment. The intelligent inspection means for the main substation equipment mainly rely on the inspection systems based on vision recognition technology, such as the popular application of high-definition camera image recognition and infrared camera temperature measurement. However, when some abnormal defects occur in the main power grid equipment, there may not be visible visual features or temperature changes, but the operating sound of the equipment changes. During the traditional substation inspection process, for major electrical equipment such as transformers, reactors, current transformers, and voltage transformers, experienced operation and maintenance personnel can judge whether there are defects in the equipment by listening for abnormal sounds, that is, the abnormal state of the main power grid equipment can be effectively detected through hearing. With the development of voiceprint recognition technology, the voiceprint monitoring of substation equipment has been gradually applied to substation operation and maintenance to replace the traditional manual listening and recognition, monitor the abnormal sounds in the audible and ultrasonic frequency bands of the equipment, and achieve all-weather abnormal sound and partial discharge monitoring and alarm of substation equipment.

[0003] Currently, the pattern recognition methods for substation equipment voiceprint recognition mainly adopt deep learning solutions, and the model network architectures of deep learning are generally selected as follows:

[0004] One is the deep network for supervised learning tasks. Such deep networks are also known as discriminative networks (Discriminative Networks), and generally include convolutional neural networks (Convolutional Neural Network, CNN), recurrent neural networks (Recurrent Neural Networks, RNN), deep neural networks (Deep Neural Network, DNN), etc. These models have achieved classification performance far better than shallow models in tasks such as image recognition and speech recognition. Due to the progress made in parameter initialization methods, training data volume, and training methods, these models can obtain better classification models in such classification tasks without the need to use pre-training methods.

[0005] Second, it is a deep network for unsupervised learning tasks. Such deep networks are also known as Generative Networks, and are often used to extract high-order correlations in observed data, including Restricted Boltzmann Machine (RBM), Deep Belief Network (DBN), and Denoising Autoencoders (DAE). When these models are used for unsupervised learning tasks, they are usually used to learn features that are more robust and effective than the input features.

[0006] However, in the scenario of on-line monitoring of the operation of substation equipment, since abnormal sound samples of substation equipment must be collected using special acoustic sensors, in the case where special sensors have not been widely deployed, abnormal sound samples of substation equipment are severely lacking. At this time, there will be a problem that the above-mentioned models have a low recognition accuracy for the acoustic fingerprints of transformer defects. Summary of the Invention

[0007] The purpose of this application is to provide a method, device, medium and product for identifying acoustic fingerprints of transformer defects based on an improved LSTM-CNN to solve the problem of low recognition accuracy of acoustic fingerprints of transformer defects.

[0008] To achieve the above purpose, this application provides the following solutions:

[0009] In the first aspect, this application provides a method for identifying acoustic fingerprints of transformer defects based on an improved LSTM-CNN, including:

[0010] Obtain the original audio signal of the transformer to be identified;

[0011] Preprocess the original audio signal of the transformer to be identified to obtain the preprocessed audio signal of the transformer to be identified; the preprocessing includes: feature extraction and wavelet denoising;

[0012] Input the preprocessed audio signal of the transformer to be identified into the transformer defect acoustic fingerprint recognition model to obtain the predicted value of the defect recognition result of the transformer to be identified; the transformer defect acoustic fingerprint recognition model is obtained by training an improved LSTM-CNN network, and the improved LSTM-CNN network includes: a convolutional neural network, a long short-term memory network, and a CBMA module, and the defect recognition result is that there is a defect or there is no defect.

[0013] Optionally, preprocessing the original audio signal of the transformer to be identified to obtain the preprocessed audio signal of the transformer to be identified includes:

[0014] Feature extraction is performed on the original audio signal of the transformer to be recognized to obtain the Mel frequency cepstral coefficients of the transformer to be recognized;

[0015] Wavelet denoising is performed on the Mel frequency cepstral coefficients of the transformer to be recognized to obtain the preprocessed audio signal of the transformer to be recognized.

[0016] Optionally, the training process of the transformer defect voiceprint recognition model includes:

[0017] Obtain a training set; the training set includes: the data-augmented audio signals of multiple training transformers and the true values of the corresponding defect recognition results;

[0018] Based on the convolutional neural network, the long short-term memory network, and the CBMA module, construct the improved LSTM-CNN network;

[0019] Use the training set to train the improved LSTM-CNN network to obtain the transformer defect voiceprint recognition model.

[0020] Optionally, obtaining a training set includes:

[0021] Obtain an original data set; the original data includes: the preprocessed audio signals of multiple training transformers and the true values of the corresponding defect recognition results;

[0022] Use the data augmentation method to perform data augmentation on the preprocessed audio signals of each training transformer respectively to obtain the corresponding data-augmented audio signals; the data augmentation method includes: frequency masking and time masking.

[0023] Optionally, the improved LSTM-CNN network includes: 1 convolutional neural network, 2 long short-term memory networks, and 1 CBMA module;

[0024] The convolutional neural network is a ResNet18 network; the ResNet18 network includes: a first convolutional layer, a first residual block, a second residual block, a third residual block, and a fourth residual block;

[0025] The CBMA module includes: a channel attention module and a spatial attention module; the channel attention module includes: an average pooling layer, a maximum pooling layer, and a multi-layer perceptron, and the spatial attention module includes: an average pooling layer, a maximum pooling layer, a splicing layer, a sixth convolutional layer, and a sigmoid function.

[0026] Optionally, the first residual block includes 2 second convolutional layers, the convolutional kernel of the second convolutional layer is 3*3, and the output channel is 64;

[0027] The second residual block includes two third convolutional layers. The convolutional kernel of the third convolutional layer is 3*3, and the output channels are 128;

[0028] The third residual block includes two fourth convolutional layers. The convolutional kernel of the fourth convolutional layer is 3*3, and the output channels are 256;

[0029] The fourth residual block includes two fifth convolutional layers. The convolutional kernel of the fifth convolutional layer is 3*3, and the output channels are 512.

[0030] Optionally, training the improved LSTM-CNN network using the training set to obtain the transformer defect voiceprint recognition model, including:

[0031] Taking the audio signals after data augmentation of each training transformer as the input and the true values of the corresponding defect recognition results as the output, training the improved LSTM-CNN network to obtain the transformer defect voiceprint recognition model.

[0032] In a second aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the transformer defect voiceprint recognition method based on the improved LSTM-CNN as described in any one of the above.

[0033] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the transformer defect voiceprint recognition method based on the improved LSTM-CNN as described in any one of the above.

[0034] In a fourth aspect, the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the transformer defect voiceprint recognition method based on the improved LSTM-CNN as described in any one of the above.

[0035] According to the specific embodiments provided by the present application, the following technical effects are disclosed in the present application:

[0036] The present application discloses a method, device, medium and product for transformer defect voiceprint recognition based on an improved LSTM-CNN. First, the original audio signal of the transformer to be recognized is obtained; then, the original audio signal of the transformer to be recognized is preprocessed to obtain the preprocessed audio signal of the transformer to be recognized; the preprocessing includes: feature extraction and wavelet denoising; finally, the preprocessed audio signal of the transformer to be recognized is input into the transformer defect voiceprint recognition model to obtain the predicted value of the defect recognition result of the transformer to be recognized; the transformer defect voiceprint recognition model is obtained by training an improved LSTM-CNN network, and the improved LSTM-CNN network includes: a convolutional neural network, a long short-term memory network and a CBMA module, and the defect recognition result is that there is a defect or there is no defect. The present application inputs the preprocessed audio signal of the transformer to be recognized into the transformer defect voiceprint recognition model constructed based on the convolutional neural network, the long short-term memory network and the CBMA module to obtain the predicted value of the defect recognition result of the transformer to be recognized, improving the accuracy of transformer defect voiceprint recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0038] Figure 1 Schematic flowchart of the method for transformer defect voiceprint recognition based on an improved LSTM-CNN provided by an embodiment of the present application;

[0039] Figure 2 Schematic diagram of the architecture for transformer defect voiceprint recognition based on an improved LSTM-CNN;

[0040] Figure 3 Schematic diagram of the structure of the improved LSTM-CNN network;

[0041] Figure 4 Schematic diagram of the structure of the ResNet18 network;

[0042] Figure 5 Schematic diagram of the structure of the residual block;

[0043] Figure 6 Schematic diagram of the structure of the CBMA module;

[0044] Figure 7 Schematic diagram of the structure of the channel attention module;

[0045] Figure 8 Schematic diagram of the structure of the spatial attention module;

[0046] Figure 9 This is a schematic structural diagram of a computer device provided in an embodiment of the present application. Specific implementation manners

[0047] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0048] The purpose of the present application is to provide a method, device, medium and product for transformer defect voiceprint recognition based on an improved LSTM-CNN, aiming to improve the accuracy of transformer defect voiceprint recognition.

[0049] To make the above objects, features and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0050] In an exemplary embodiment, as Figure 1 and Figure 2 shown, a method for transformer defect voiceprint recognition based on an improved LSTM-CNN is provided, including:

[0051] Step 1: Obtain the original audio signal of the transformer to be recognized.

[0052] Step 2: Preprocess the original audio signal of the transformer to be recognized to obtain the preprocessed audio signal of the transformer to be recognized; the preprocessing includes: feature extraction and wavelet denoising.

[0053] As an optional implementation manner, step 2 includes:

[0054] Step 21: Extract features from the original audio signal of the transformer to be recognized to obtain the Mel frequency cepstral coefficients of the transformer to be recognized.

[0055] Specifically, Mel-Frequency Cepstral Coefficients (MFCC) is an important feature in audio signal processing, used to capture the spectral characteristics of audio signals. The specific implementation process is as follows: decompose an original audio signal into multiple frames; pre-emphasize the audio signals within each frame respectively, through a high-pass filter, to obtain the corresponding high-pass filtered audio signals; perform Fourier transform on each of the high-pass filtered audio signals respectively, transform the signals into the frequency domain, to obtain the corresponding Fourier transformed audio signals; pass the Fourier transformed audio signals obtained for each frame through a Mel filter (triangular overlapping window) to obtain the corresponding Mel scale; extract the logarithmic energy on each Mel scale and perform inverse discrete Fourier transform, transform it into the cepstral domain, to obtain a cepstral map; determine the amplitude of the cepstral map as the Mel-Frequency Cepstral Coefficient.

[0056] Step 22: Perform wavelet denoising on the Mel-Frequency Cepstral Coefficients of the transformer to be identified to obtain the preprocessed audio signal of the transformer to be identified.

[0057] Specifically, wavelet denoising is implemented based on wavelet transform. Wavelet transform is a multi-resolution analysis method that can decompose a signal into different frequency sub-bands. Wavelet transform has good localization characteristics in both the time domain and the frequency domain, and is suitable for processing non-stationary signals. After the sound signal is transformed into the wavelet domain, noise is removed through spatial domain filtering. The denoised wavelet coefficients are subjected to inverse wavelet transform to reconstruct the voiceprint signal. The specific steps of wavelet denoising are as follows: perform wavelet transform on the Mel-Frequency Cepstral Coefficients using the Daubechies wavelet basis function to obtain the wavelet transformed audio signal; perform denoising on the wavelet transformed audio signal using spatial domain filtering to obtain the denoised audio signal; perform signal reconstruction on the denoised audio signal using inverse wavelet transform to obtain the preprocessed audio signal.

[0058] Step 3: Input the preprocessed audio signal of the transformer to be identified into the transformer defect voiceprint recognition model to obtain the predicted value of the defect recognition result of the transformer to be identified.

[0059] Among them, the transformer defect voiceprint recognition model is obtained by training an improved LSTM-CNN network. The improved LSTM-CNN network includes: a convolutional neural network, a long short-term memory network, and a CBMA module. The defect recognition result is either there is a defect or there is no defect.

[0060] As an optional implementation manner, in Step 3, the training process of the transformer defect voiceprint recognition model includes:

[0061] Step 31: Obtain a training set; the training set includes: the data-augmented audio signals of multiple training transformers and the true values of the corresponding defect recognition results.

[0062] As an alternative implementation, step 31 includes:

[0063] Step 311: Obtain an original data set; the original data includes: the preprocessed audio signals of multiple training transformers and the true values of the corresponding defect recognition results.

[0064] Step 312: Use a data augmentation method to perform data augmentation on the preprocessed audio signals of each training transformer respectively to obtain the corresponding data-augmented audio signals; the data augmentation method includes: frequency masking and time masking.

[0065] Specifically, the specific operation content of frequency masking includes: the frequency dimension of the preprocessed audio signal is F, set the frequency masking interval length to f, where f is an adjustable parameter. Then randomly select the start position f0 of the masking interval from the interval [0, F - f], and finally perform a masking operation on the [f0, f0 + f] interval of the voiceprint feature S (preprocessed audio signal), that is, set the values in the [f0, f0 + f] interval to 0, and this operation can be repeated multiple times.

[0066] The specific operation content of time masking: the time dimension of the preprocessed audio signal is T, set the frequency masking interval length to t, where t is an adjustable parameter. Then randomly select the start position t0 of the masking interval from the interval [0, T - t], and finally perform a masking operation on the [t0, t0 + t] interval of the voiceprint feature S (preprocessed audio signal), that is, set the values in the [t0, t0 + t] interval to 0, and this operation can be repeated multiple times.

[0067] In this application, the frequency masking parameter is set to 10, the number of repetitions is 1 time, the time masking parameter is 15, and the number of repetitions is 2 times, that is, the parameters F, Nf, T, and Nt are respectively set to 10, 1, 15, and 2.

[0068] Step 32: Based on a convolutional neural network, a long short-term memory network, and a CBMA module, construct an improved LSTM-CNN network.

[0069] As an alternative implementation, as Figure 3 shown, the improved LSTM-CNN network includes: 1 convolutional neural network, 2 long short-term memory networks, and 1 CBMA module.

[0070] The convolutional neural network is a ResNet18 network; the ResNet18 network includes: a first convolutional layer, a first residual block, a second residual block, a third residual block, and a fourth residual block. The ResNet18 network structure is asFigure 4 as shown

[0071] As an alternative embodiment, the first residual block includes two second convolutional layers, the convolutional kernel of the second convolutional layer is 3×3, and the output channels are 64.

[0072] The second residual block includes two third convolutional layers, the convolutional kernel of the third convolutional layer is 3×3, and the output channels are 128.

[0073] The third residual block includes two fourth convolutional layers, the convolutional kernel of the fourth convolutional layer is 3×3, and the output channels are 256.

[0074] The fourth residual block includes two fifth convolutional layers, the convolutional kernel of the fifth convolutional layer is 3×3, and the output channels are 512.

[0075] Specifically, the structure of any residual block (Residual block) is as Figure 5 shown, where x is the input, F(x) is the output after x passes through the convolutional layer, weight layer is the weight layer of the convolutional layer, and relu is the relu function.

[0076] The CBMA module includes: a channel attention module and a spatial attention module; the channel attention module includes: an average pooling layer, a max pooling layer, and a multi-layer perceptron, and the spatial attention module includes: an average pooling layer, a max pooling layer, a concatenation layer, a sixth convolutional layer, and a sigmoid function.

[0077] The CBMA module (Convolutional Block Attention Module) is a lightweight attention module in the channel and spatial dimensions. As Figure 6 shown, the CBAM module contains two independent modules, the channel attention module (Channel Attention Module, CAM) and the spatial attention module (Spartial Attention Module, SAM), which perform attention on the channel and space respectively. This not only saves parameters and computing power, but also ensures that it can be integrated into the existing network architecture as a plug-and-play module.

[0078] The structure of the channel attention module is as Figure 7As shown, the channel attention module compresses the feature map (i.e., the feature map extracted by the ResNet18 network) in the spatial dimension to obtain a one-dimensional vector and then performs operations. When compressing in the spatial dimension, both average pooling and max pooling are considered. Average pooling and max pooling can be used to aggregate the spatial information of the feature map and send it to a shared multi-layer perceptron network (SharedMLP) to compress the spatial dimension of the input feature map, and sum and merge element-wise to generate a channel attention map. The channel attention mechanism of the channel attention module can be expressed as:

[0079]

[0080] where M c (·) represents the weight of the channel attention module; F represents the feature map input to the channel attention module; σ(·) represents the activation function Relu; MLP(·) represents the multi-layer perceptron network; AvgPool(·) represents average pooling; MaxPool(·) represents max pooling; W1 and W0 represent weights; represents the feature map after average pooling in the channel attention module; represents the feature map after max pooling in the channel attention module; F' represents the feature map output by the channel attention module; represents element-wise multiplication.

[0081] The structure of the spatial attention module is as Figure 8 shown. The spatial attention module takes the feature map output by the channel attention module as the feature map input to the spatial attention module. First, a channel-based max pooling layer and an average pooling layer are performed, and then the results of the max pooling layer and the average pooling layer are concatenated based on the channel. Then, through a convolutional operation, it is reduced to 1 channel. After passing through the sigmoid function, spatial attention features are generated. Finally, the spatial attention features and the spatial attention features are multiplied to obtain the finally generated features, that is, the defect recognition results. The weight of the spatial attention mechanism of the spatial attention module can be expressed as:

[0082]

[0083] where M s (·) represents the weight of the spatial attention module; f 7×7 (·) represents the sixth convolutional layer; represents the feature map after average pooling in the spatial attention module; represents the feature map after max pooling in the spatial attention module.

[0084] Step 33: Train the improved LSTM-CNN network using the training set to obtain a transformer defect voiceprint recognition model.

[0085] As an alternative implementation, step 33 includes:

[0086] Use the data-augmented audio signals of each training transformer as the input and the true values of the corresponding defect recognition results as the output to train the improved LSTM-CNN network to obtain a transformer defect voiceprint recognition model.

[0087] In summary, to improve the accuracy and robustness of transformer voiceprint fault recognition, this application designs a system based on an improved LSTM-CNN architecture. At the same time, the trained deep learning model is converted into the ONNX (Open Neural Network Exchange) format, which makes the model easier to be deployed and optimized on different frameworks and platforms. First, in the data processing stage, the original audio signal is preprocessed and feature extracted. Specifically, Mel-frequency cepstral coefficients are used as the main feature representation. For the substation environment with more noise, a wavelet denoising step is also introduced. The Daubechies wavelet basis function is selected to perform multi-resolution analysis on the audio signal, and the signal is decomposed into different frequency sub-bands. In the wavelet domain, a spatial filter is used to remove the influence of background noise, and then the clean voiceprint signal is reconstructed through inverse wavelet transform. This method not only retains useful audio information but also effectively suppresses interference noise, improving the data quality for subsequent model training. To address the problem of fewer data samples, two data augmentation methods, frequency masking and time masking, are introduced. Second, ResNet18 is selected as the basic structure of the convolutional neural network, which has good feature extraction ability. ResNet18 consists of multiple residual blocks, and each residual block contains two 3*3 convolutional layers, which can effectively prevent the problem of gradient disappearance. To capture the dependencies in the time series, two layers of long short-term memory networks (LSTM) are added after the convolutional neural network. This combination enables the model to process information in both spatial and temporal dimensions and is very suitable for processing data with spatio-temporal characteristics such as voiceprints. In addition, to further enhance the model's expressiveness, the CBMA attention mechanism is embedded in the CNN-LSTM architecture. The steps of data preprocessing in the model training stage include feature extraction, wavelet denoising, and data augmentation; in the inference stage, only feature extraction and wavelet denoising are adopted.

[0088] In an exemplary embodiment, a computer device is provided, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the transformer defect voiceprint recognition method based on the improved LSTM-CNN.

[0089] In an exemplary embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements a transformer defect voiceprint recognition method based on an improved LSTM-CNN.

[0090] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, it implements a transformer defect voiceprint recognition method based on an improved LSTM-CNN.

[0091] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 9 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a transformer defect voiceprint recognition method based on an improved LSTM-CNN.

[0092] Those skilled in the art can understand that Figure 9 the structure shown in

[0093] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0094] The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.

[0095] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0096] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0097] In this text, specific examples are used to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A transformer defect voiceprint recognition method based on improved LSTM-CNN, characterized in that: The transformer defect voiceprint recognition method based on the improved LSTM-CNN includes: Obtaining the original audio signal of the transformer to be identified; Preprocessing the original audio signal of the transformer to be identified to obtain the preprocessed audio signal of the transformer to be identified; the preprocessing includes: feature extraction and wavelet noise reduction; The preprocessed audio signal of the transformer to be identified is input into the transformer defect voiceprint recognition model to obtain a predicted value of the defect recognition result of the transformer to be identified; the transformer defect voiceprint recognition model is obtained by training an improved LSTM-CNN network, and the improved LSTM-CNN network includes: a convolutional neural network, a long short-term memory network and a CBMA module, and the defect recognition result is the presence of a defect or the absence of a defect.

2. According to the improved LSTM-CNN-based transformer defect voiceprint recognition method of claim 1, it is characterized in that: Preprocessing the original audio signal of the transformer to be identified to obtain the preprocessed audio signal of the transformer to be identified, including: Extract features of the original audio signal of the transformer to be identified, and obtain the Mel-frequency cepstrum coefficients of the transformer to be identified; The Mel-frequency cepstrum coefficients of the transformer to be identified are subjected to wavelet denoising to obtain a preprocessed audio signal of the transformer to be identified.

3. According to claim 2, the transformer defect voiceprint recognition method based on improved LSTM-CNN is characterized in that: The training process of the transformer defect voiceprint recognition model includes: Obtaining a training set; the training set includes: audio signals after data enhancement of multiple training transformers and true values ​​of corresponding defect recognition results; Based on the convolutional neural network, the long short-term memory network and the CBMA module, construct the improved LSTM-CNN network; The improved LSTM-CNN network is trained using the training set to obtain the transformer defect voiceprint recognition model.

4. The transformer defect voiceprint recognition method based on improved LSTM-CNN according to claim 3 is characterized in that: Get the training set, including: Acquire an original data set; the original data includes: preprocessed audio signals of multiple training transformers and true values ​​of corresponding defect recognition results; The data enhancement method is used to perform data enhancement on the preprocessed audio signals of each training transformer to obtain corresponding data-enhanced audio signals; the data enhancement method includes: frequency mask and time mask.

5. The transformer defect voiceprint recognition method based on improved LSTM-CNN according to claim 3 is characterized in that: The improved LSTM-CNN network includes: 1 convolutional neural network, 2 long short-term memory networks and 1 CBMA module; The convolutional neural network is a ResNet18 network; the ResNet18 network includes: a first convolutional layer, a first residual block, a second residual block, a third residual block and a fourth residual block; The CBMA module includes: a channel attention module and a spatial attention module; the channel attention module includes: an average pooling layer, a maximum pooling layer and a multi-layer perceptron, and the spatial attention module includes: an average pooling layer, a maximum pooling layer, a splicing layer, a sixth convolutional layer and a sigmoid function.

6. The transformer defect voiceprint recognition method based on improved LSTM-CNN according to claim 5 is characterized in that: The first residual block includes two second convolutional layers, the convolution kernel of the second convolutional layer is 3*3, and the output channel is 64; The second residual block includes two third convolutional layers, the convolution kernel of the third convolutional layer is 3*3, and the output channel is 128; The third residual block includes two fourth convolutional layers, the convolution kernel of the fourth convolutional layer is 3*3, and the output channel is 256; The fourth residual block includes two fifth convolutional layers, the convolution kernel of the fifth convolutional layer is 3*3, and the output channel is 512.

7. The transformer defect voiceprint recognition method based on improved LSTM-CNN according to claim 3 is characterized in that: The improved LSTM-CNN network is trained using the training set to obtain the transformer defect voiceprint recognition model, including: The audio signal after data enhancement of each training transformer is used as input, and the true value of the corresponding defect recognition result is used as output, and the improved LSTM-CNN network is trained to obtain the transformer defect voiceprint recognition model.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the transformer defect voiceprint recognition method based on the improved LSTM-CNN as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the transformer defect voiceprint recognition method based on the improved LSTM-CNN described in any one of claims 1 to 7 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the transformer defect voiceprint recognition method based on the improved LSTM-CNN described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Acoustic detection method and system for bolt axial force of wind driven generator

    CN121453263A

  • A method and system for acoustic measurement of bolt axial force in wind turbine generators

    CN121453263B