Power transmission line flash explosion sound detection method and system, medium and processor

Through the improved MFCC feature extraction and CNN-LSTM model, remote real-time monitoring of transmission line flashover faults is achieved, solving the problem of manual inspection being labor-intensive and inefficient, improving the accuracy and efficiency of monitoring, and ensuring the safe operation of transmission lines.

CN120669064APending Publication Date: 2025-09-19ELECTRIC POWER RES INST OF GUANGXI POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510561027.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In the existing technology, the detection and troubleshooting of transmission line flashover faults mainly rely on manual inspections, which is manpower-consuming and not timely, resulting in a lag in fault troubleshooting, affecting the safety of transmission lines and operators.

Method used

An improved MFCC feature extraction method combined with a CNN-LSTM model is used to collect flash burst sound signals through a microphone. After preprocessing, Mel coefficient extraction and differential processing are used to identify the status of the transmission line and achieve remote real-time monitoring.

Benefits of technology

It improves the monitoring accuracy and efficiency of transmission line flashover faults, can quickly identify faults and issue alarms in a timely manner, improves the safety of transmission line operation and ensures the safety of operators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120669064A_ABST
    Figure CN120669064A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of power system monitoring and early warning, and discloses a power transmission line flash explosion sound detection method and system based on an improved MFCC feature and a CNN-LSTM model, and the method comprises the steps: collecting a sound signal during the flash explosion of a power transmission line through a microphone, and carrying out the preprocessing of the collected flash explosion sound; an improved MFCC feature is obtained through a Mel coefficient extraction method and differential processing; and a CNN-LSTM network model is used to classify and identify the improved MFCC features, and the state of the power transmission line is monitored in real time, so that the flash explosion fault of the power transmission line is effectively identified. According to the method, the accuracy of the MFCC feature extraction method on the flash explosion sound data of the power transmission line and the rapidity of the CNN-LSTM model in the aspect of feature recognition are comprehensively improved, and the detection effect on the flash explosion fault of the power transmission line is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power system monitoring and early warning, and in particular to a method, system, medium and processor for detecting flashover sound of a transmission line. Background Art

[0002] Transmission lines are the most extensive and longest transmission system in the power system, carrying the crucial responsibility of transporting electricity from power generation to users. Ensuring their safe and stable operation is essential for ensuring power supply. According to the 2022 China Electric Power Reliability Annual Report, the top three causes of transmission line outages are natural factors, climate factors, and flashovers caused by external forces. These factors account for a significantly higher proportion than other causes. Transmission line flashovers caused nearly 2,500 hours of outages. With the development of society, both residents and industry are placing new demands on stable power supply. Transmission line flashovers have a significant impact on the power system and human society. They can cause line protection tripping, power outages, disrupting normal power transmission, and even leading to the collapse of the entire power system, causing inconvenience to daily life, industrial production, and healthcare. Flashover explosions can damage transmission equipment, increase maintenance costs, affect system stability, and potentially trigger cascading failures. They pose a serious threat to personnel safety, potentially causing injuries such as electric shock and burns, and potentially triggering secondary disasters such as fires. Pollutants produced during flashover explosions can also adversely affect the ecological environment. The monitoring of flashover explosions on transmission lines is a critical issue and requires significant attention.

[0003] Nowadays, the detection and troubleshooting of flashover faults in transmission lines are usually carried out through manual inspections, which is manpower-consuming and not very timely. There is a certain lag in the fault troubleshooting and alarm, which reduces the safety of transmission line operation and cannot guarantee the safety of operators in this process.

[0004] In view of this, a method, system, medium and processor for detecting flashover sound on a transmission line are needed. Summary of the Invention

[0005] In response to the existing problem that the detection and troubleshooting of transmission line flashover faults are usually manual inspections, which are labor-intensive and time-consuming, the present invention provides a transmission line flashover sound detection method, system, medium, and processor. By using an improved MFCC feature extraction method combined with a CNN-LSTM model, it can more accurately identify transmission line flashover accidents, achieve remote real-time monitoring, and improve the accuracy and efficiency of monitoring. The specific technical solution is as follows:

[0006] A method for detecting transmission line flashover sound comprises: collecting sound signals when transmission line flashover occurs using a microphone and preprocessing the collected flashover sound; obtaining improved MFCC features through a Mel coefficient extraction method and differential processing; and classifying and identifying the improved MFCC features using a CNN-LSTM network model to monitor the status of the transmission line in real time.

[0007] Preferably, collecting the sound signal from a power line flashover includes using a microphone to separately collect the fault sound from the power line flashover and a sound signal similar to the flashover sound from the surrounding environment of the power line operation. To ensure that the sound signal can be restored to the original sound without distortion after sampling, the sampling frequency of the microphone must meet a condition: it must be greater than or equal to twice the highest frequency in the sound signal. In this way, the sound signal data obtained through sampling can contain all information of the original sound signal, thereby accurately restoring the original sound.

[0008] Preferably, the sound signal preprocessing includes pre-emphasis, framing and windowing of the sound signal.

[0009] Pre-emphasis is based on the principle of high-pass filtering. By designing a suitable high-pass filter to process the original audio signal, the high-frequency portion of the signal is enhanced while the low-frequency portion is relatively weakened. The resulting signal has higher resolution and richer information content in the high-frequency portion.

[0010] The pre-emphasis formula is as follows:

[0011] y[n]=x[n]-a*x[n-1]

[0012] Where y[n] represents the pre-emphasized signal, x[n] represents the original speech signal, and a is the pre-emphasis coefficient;

[0013] The sound signal is framed. Using framing technology, a long, non-stationary sound signal is divided into several short, stationary signals. Each segmented sound signal is called a frame, and the duration of each frame is generally selected to be between 10ms and 30ms. To ensure a smooth transition of the characteristic information of the segmented sound and the continuity of two adjacent frames, an overlapping segmentation method is generally adopted, inserting frames between two frames to create an overlap between them and avoid loss of characteristic information caused by sudden changes during frame switching. The overlapping portion between two adjacent frames is called a frame shift, and the frame shift is selected to be between one-third and one-half of the frame length.

[0014] The sound signal is windowed and Fast Fourier Transform (FFT) is performed on the sound signal. However, the framed sound signal is a non-periodic signal, and spectrum leakage occurs after FFT transformation. In order to reduce the leakage effect, windowing processing is required to smooth the signal boundaries.

[0015] The window function expression is as follows:

[0016]

[0017] Where w(n) represents the value of the window function at position n. n is the index of the sampling point in the window. N is the length of the window, that is, the total number of sampling points.

[0018] Preferably, the Mel coefficient extraction method includes performing fast Fourier transform operations, Mel filtering and logarithm taking, and discrete cosine transform (DCT) on the preprocessed sound data to obtain Mel frequency cepstral coefficients (MFCC).

[0019] Perform a fast Fourier transform on the preprocessed sound data, and then stack the FFT transform results of each frame to obtain the time-frequency signal after the short-time Fourier transform (STFT) of the sound signal. Take the square of the amplitude of the sound signal and then calculate the logarithmic power to obtain the power spectrum of the sound signal.

[0020] The logarithmic power calculation expression is as follows:

[0021]

[0022] Wherein, N represents the Nth frame of sound signal.

[0023] The human ear has different sensitivities to sounds of different frequencies. Based on experimental simulations of human auditory characteristics, the Mel Scale was proposed. Before Mel filtering, audio signals with a frequency unit of Hz need to be converted to signals with a Mel unit.

[0024] Among them, the conversion formula of Mel frequency is as follows:

[0025]

[0026] Where Mel(f) represents the perceived frequency; f represents the actual frequency;

[0027] The power spectrum of the sound signal of each frame passes through the Mel filter bank and the energy of the power spectrum in the filter bank is calculated. The calculation formula is as follows:

[0028]

[0029] Among them, S(i,m) represents the energy of the filter; m represents the mth Mel filter; H m (k) represents a Mel filter.

[0030] The Mel filter bank is constructed based on the Mel frequency characteristics. The Mel filter bank is actually composed of multiple triangular filters, usually containing 20 to 40 triangular filters. The center frequency of each filter is set according to the Mel frequency characteristics. It is linearly distributed when the sound signal is at low frequency and logarithmically distributed at high frequency. The expression of the Mel filter response function Hm(k) is as follows:

[0031]

[0032] Where, f(m) represents the center frequency;

[0033] The filtered sound signal is processed by discrete cosine transform. Discrete cosine transform has a good characteristic of concentrating frequency domain energy, which can reduce the correlation between information in different dimensions and map it to a space with lower dimensions. The expression of MFCC feature coefficient is as follows:

[0034]

[0035] Where n represents the spectrum line after DCT, and m represents the mth Mel filter.

[0036] The strobe sound signal undergoes signal preprocessing, including pre-emphasis, framing, and windowing. The Mel-frequency cepstral coefficients are obtained through fast Fourier transform (FFT), Mel filtering, logarithmization, and discrete cosine transform (DCT).

[0037] Preferably, the differential processing includes performing differential processing using Mel frequency cepstral coefficients to obtain first-order differentials and second-order differentials.

[0038] Among them, the expression of the first-order MFCC feature is as follows:

[0039]

[0040] Among them, C(n+i) is the sound MFCC feature parameter, D(n) is the first-order difference of the MFCC coefficient, and the k value is 2.

[0041] The second-order difference MFCC feature parameter expression is as follows:

[0042]

[0043] Wherein, D(n+i) is the sound MFCC feature parameter.

[0044] The improved MFCC features can be obtained by concatenating the Mel cepstral coefficients with their first-order difference coefficients and second-order difference coefficients.

[0045] Preferably, the method of using the CNN-LSTM network model to classify and identify the improved MFCC features includes preparing improved MFCC features of power transmission line flashover and other environmental sound signals as a training data set;

[0046] Preferably, the use of the CNN-LSTM network model to classify and identify the improved MFCC features also includes constructing a CNN-LSTM network for feature recognition and detection, using the improved MFCC features obtained after feature extraction as the input of the model, performing spatial feature extraction of the data through the convolution layer and the pooling layer, and then inputting the output features of the CNN into the LSTM structure for temporally associated information feature extraction, and then passing the feature extraction results through the Dropout layer to perform a random neuron deletion operation, and then integrating through the flattening layer and the fully connected layer, and finally entering the Softmax layer for normalization for label recognition and classification, and the output layer represents the classification result;

[0047] Another object of the present invention is to provide a system for detecting power line flashover sound based on improved MFCC features and a CNN-LSTM model, which improves the level of automated monitoring of power systems and is of great significance to the safe and stable operation of transmission lines.

[0048] As a preferred solution of the system for the power transmission line flashover sound detection method based on the improved MFCC features and CNN-LSTM model described in the present invention, it includes: a microphone sound acquisition module, a preprocessing module, an improved MFCC feature extraction module, a CNN-LSTM network training module, and a flashover recognition module;

[0049] The microphone sound collection module collects sound signals of the power transmission line during operation through a microphone;

[0050] The pre-processing module performs pre-emphasis, framing and windowing processing on the collected original sound signal;

[0051] The improved MFCC feature extraction module extracts improved MFCC features from the preprocessed sound signal;

[0052] The CNN-LSTM network training module uses the CNN-LSTM network to train the extracted features to identify and detect the operating status of the transmission line;

[0053] The flash burst recognition module preprocesses the real-time audio data collected by the microphone and extracts improved MFCC features, and inputs the obtained features into a trained neural network model to perform real-time flash burst recognition.

[0054] A computer-readable storage medium includes a stored program, wherein when the program is executed, the device containing the computer-readable storage medium is controlled to execute any one of the above methods for detecting flashover sound on a transmission line.

[0055] A processor is used to run a program, wherein the program, when running, executes the above-mentioned method for detecting flashover sound on a power transmission line.

[0056] Compared with the prior art, the present invention has the following beneficial effects:

[0057] This method, by using an improved MFCC feature extraction method combined with a CNN-LSTM model, can more accurately identify transmission line flashover incidents, enabling remote real-time monitoring and improving monitoring accuracy and efficiency. This method outperforms traditional manual inspections and can quickly identify flashover faults and promptly alert personnel, significantly improving the safety of transmission line operations and ensuring the safety of operators. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly describes the drawings required for the specific embodiments or the description of the prior art. Similar elements or parts are generally identified by similar reference numerals throughout the drawings. Elements or parts in the drawings are not necessarily drawn to scale.

[0059] Figure 1 A schematic flow chart of a method for detecting flashover sound on a transmission line according to an embodiment of the present invention;

[0060] Figure 2 A structural diagram of a CNN-LSTM model provided for one embodiment of the present invention;

[0061] Figure 3 A schematic diagram of the working modules of a transmission line flashover sound detection system provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0062] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0063] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0064] It should also be understood that the terms used in the present specification are only for the purpose of describing particular embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0065] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0066] Example 1

[0067] Reference Figures 1-2 , which is the first embodiment of the present invention, provides a method for detecting power line flashover sound based on improved MFCC features and CNN-LSTM model, such as Figure 1 Shown, including:

[0068] S1: The microphone is used to collect the sound signal when the transmission line flashover occurs, and the collected flashover sound is pre-processed.

[0069] Microphones are used to collect the sound of a power line flashover, as well as similar sound signals from the surrounding environment. To ensure that the sampled sound signals can be restored to their original form without distortion, the microphone's sampling frequency must meet a specific requirement: it must be greater than or equal to twice the highest frequency in the sound signal. This ensures that the sampled sound signal data contains all the information from the original signal, accurately reproducing the original sound.

[0070] It should be further explained that the sound signal preprocessing includes, and the sound signal preprocessing operations include pre-emphasis, framing and windowing.

[0071] Pre-emphasis is based on the principle of high-pass filtering. By designing a suitable high-pass filter to process the original audio signal, the high-frequency portion of the signal is enhanced while the low-frequency portion is relatively weakened. The resulting signal has higher resolution and richer information content in the high-frequency portion.

[0072] The pre-emphasis formula is as follows:

[0073] y[n]=x[n]-a*x[n-1]

[0074] Where y[n] represents the pre-emphasized signal, x[n] represents the original speech signal, and a is the pre-emphasis coefficient;

[0075] The sound signal is framed. Using framing technology, a long, non-stationary sound signal is divided into several short, stationary signals. Each segmented sound signal is called a frame, and the duration of each frame is generally selected to be between 10ms and 30ms. To ensure a smooth transition of the characteristic information of the segmented sound and the continuity of two adjacent frames, an overlapping segmentation method is generally adopted, inserting frames between two frames to create an overlap between them and avoid loss of characteristic information caused by sudden changes during frame switching. The overlapping portion between two adjacent frames is called a frame shift, and the frame shift is selected to be between one-third and one-half of the frame length.

[0076] The sound signal is windowed and Fast Fourier Transform (FFT) is performed on the sound signal. However, the framed sound signal is a non-periodic signal, and spectrum leakage occurs after FFT transformation. In order to reduce the leakage effect, windowing processing is required to smooth the signal boundaries.

[0077] The window function expression is as follows:

[0078]

[0079] Where w(n) represents the value of the window function at position n. n is the index of the sampling point in the window. N is the length of the window, that is, the total number of sampling points.

[0080] S2: The improved MFCC features are obtained through Mel coefficient extraction method and differential processing;

[0081] The Mel coefficient extraction method includes performing fast Fourier transform operations, Mel filtering and logarithm taking, and discrete cosine transform (DCT) on the pre-processed sound data to obtain Mel frequency cepstral coefficients (MFCC).

[0082] Perform a fast Fourier transform on the pre-processed sound data, and then stack the FFT transform results of each frame to obtain the time-frequency signal after the short-time Fourier transform (STFT) of the sound signal. Take the square of the amplitude of the sound signal, and then calculate the logarithmic power to obtain the power spectrum of the sound signal. The logarithmic power calculation expression is as follows:

[0083]

[0084] Wherein, N represents the Nth frame of sound signal.

[0085] The human ear has different sensitivities to sounds of different frequencies. Based on experimental simulations of human auditory characteristics, the Mel Scale was proposed. Before Mel filtering, audio signals with a frequency unit of Hz need to be converted to Mel units. The conversion formula for Mel frequency is as follows:

[0086]

[0087] Where Mel(f) represents the perceived frequency; f represents the actual frequency;

[0088] The power spectrum of the sound signal of each frame passes through the Mel filter bank and the energy of the power spectrum in the filter bank is calculated. The calculation formula is as follows:

[0089]

[0090] Among them, S(i,m) represents the energy of the filter; m represents the mth Mel filter; H m (k) represents the Mel filter. The Mel filter response function Hm(k) is expressed as follows:

[0091]

[0092] Where, f(m) represents the center frequency of the mth Mel filter;

[0093] The filtered sound signal is processed by discrete cosine transform. Discrete cosine transform has a good characteristic of concentrating frequency domain energy, which can reduce the correlation between information in different dimensions and map it to a space with lower dimensions. The expression of MFCC feature coefficient is as follows:

[0094]

[0095] Where n represents the spectrum line after DCT, m represents the mth Mel filter, and M represents the number of filters in the Mel filter bank.

[0096] The strobe sound signal undergoes signal preprocessing, including pre-emphasis, framing, and windowing. The Mel-frequency cepstral coefficients are obtained through fast Fourier transform (FFT), Mel filtering, logarithmization, and discrete cosine transform (DCT).

[0097] It should be further explained that the differential processing includes performing differential processing using Mel-frequency cepstral coefficients to obtain first-order differences and second-order differences.

[0098] Among them, the expression of the first-order MFCC feature is as follows:

[0099]

[0100] Among them, C(n+i) is the sound MFCC feature parameter, D(n) is the first-order difference of the MFCC coefficient, and the k value is 2.

[0101] The second-order difference MFCC feature parameter expression is as follows:

[0102]

[0103] Wherein, D(n+i) is the sound MFCC feature parameter.

[0104] The improved MFCC features can be obtained by concatenating the Mel cepstral coefficients with their first-order difference coefficients and second-order difference coefficients.

[0105] S3: Use the CNN-LSTM network model to classify and identify the improved MFCC features and monitor the status of the transmission lines in real time.

[0106] The use of the CNN-LSTM network model to classify and identify the improved MFCC features includes preparing the improved MFCC features of transmission line flashovers and other environmental sound signals as a training dataset;

[0107] Reference Figure 2 It should also be noted that the use of the CNN-LSTM network model to classify and identify the improved MFCC features also includes constructing a CNN-LSTM network for feature recognition and detection, using the improved MFCC features obtained after feature extraction as the input of the model, performing spatial feature extraction of the data through the convolution layer and the pooling layer, and then inputting the output features of the CNN into the LSTM structure for temporally correlated information feature extraction, and then passing the feature extraction results through the Dropout layer to randomly delete neurons, and then integrating through the flattening layer and the fully connected layer, and finally entering the Softmax layer for normalization for label recognition and classification, and the output layer represents the classification result.

[0108] Embodiment 2, the second embodiment of the present invention, is different from the previous embodiment in that:

[0109] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0110] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0111] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.

[0112] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0113] Example 3, reference Figure 3 , which is one embodiment of the present invention, provides a power transmission line flashover sound detection system based on improved MFCC features and a CNN-LSTM model, characterized by: comprising a microphone sound acquisition module, a preprocessing module, an improved MFCC feature extraction module, a CNN-LSTM network training module, and a flashover recognition module;

[0114] A microphone sound collection module, which collects sound signals during the operation of the transmission line through a microphone;

[0115] The pre-processing module performs pre-emphasis, framing and windowing on the collected original sound signal;

[0116] Improve the MFCC feature extraction module to extract improved MFCC features from the preprocessed sound signal;

[0117] CNN-LSTM network training module, which uses the CNN-LSTM network to train the extracted features to identify and detect the operating status of the transmission line;

[0118] The flash burst recognition module preprocesses the real-time audio data collected by the microphone and extracts improved MFCC features, then inputs the obtained features into a trained neural network model for real-time flash burst recognition.

[0119] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A method for detecting flashover sound in a power transmission line, characterized in that: The following steps are involved: The sound signal of the flashover on the transmission line is collected through a microphone, and the collected flashover sound is pre-processed; The improved MFCC features are obtained through Mel coefficient extraction method and differential processing; The CNN-LSTM network model is used to classify and identify the improved MFCC features and monitor the status of the transmission lines in real time.

2. A method for detecting flashover sound on a power transmission line according to claim 1, characterized in that: The collecting of the sound signal when the transmission line flash occurs includes: using a microphone to collect the fault sound when the transmission line flash occurs and the sound signal similar to the flash sound around the transmission line operating environment, and the sampling frequency of the microphone is greater than or equal to twice the highest frequency in the sound signal.

3. A method for detecting flashover sound on a power transmission line according to claim 1, characterized in that: Sound signal preprocessing includes pre-emphasis, framing and windowing of sound signals.

4. A method for detecting flashover sound on a power transmission line according to claim 1, characterized in that: The specific method for extracting Mel coefficients is as follows: Perform a fast Fourier transform on the preprocessed sound data, and then stack the FFT transform results of each frame to obtain the time-frequency signal after the short-time Fourier transform of the sound signal. Take the square of the amplitude of the sound signal, and then calculate the logarithmic power to finally obtain the power spectrum of the sound signal.

5. A method for detecting flashover sound on a power transmission line according to claim 1, characterized in that: The differential processing is specifically to use the Mel frequency cepstral coefficient to perform differential processing to obtain its first-order difference and second-order difference, where: The expression of the first-order MFCC feature is as follows: Among them, C(n+i) is the sound MFCC feature parameter, D(n) is the first-order difference of the MFCC coefficient, and the k value is 2. The second-order difference MFCC feature parameter expression is as follows: Wherein, D(n+i) is the sound MFCC feature parameter.

6. A method for detecting flashover sound on a power transmission line according to claim 1, characterized in that: Using the CNN-LSTM network model to classify and identify the improved MFCC features includes: Prepare improved MFCC features of transmission line flashover and other environmental sound signals as training datasets; Construct a CNN-LSTM network for feature recognition and detection. The improved MFCC features obtained after feature extraction are used as the input of the model. The spatial features of the data are extracted through convolutional layers and pooling layers. Input the output features of CNN into the LSTM structure to extract temporally related information features; The result of feature extraction is passed through the Dropout layer to randomly delete neurons; After being integrated through the flattening layer and the fully connected layer, it finally enters the Softmax layer for normalization for label recognition and classification, and the output layer represents the classification result.

7. A power transmission line flashover sound detection system, characterized in that: The method applied to any one of claims 1 to 6, comprising a microphone sound acquisition module, a preprocessing module, an improved MFCC feature extraction module, a CNN-LSTM network training module, and a flash burst recognition module; The microphone sound collection module collects sound signals of the power transmission line during operation through a microphone; The pre-processing module performs pre-emphasis, framing and windowing processing on the collected original sound signal; The improved MFCC feature extraction module extracts improved MFCC features from the preprocessed sound signal; The CNN-LSTM network training module uses the CNN-LSTM network to train the extracted features to identify and detect the operating status of the transmission line; The flash burst recognition module preprocesses the real-time audio data collected by the microphone and extracts improved MFCC features, and inputs the obtained features into a trained neural network model to perform real-time flash burst recognition.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the power transmission line flashover sound detection method according to any one of claims 1 to 6.

9. A processor, characterized in that: The processor is configured to run a program, wherein the program, when running, executes the method for detecting flashover sound on a transmission line according to any one of claims 1 to 6.