A thunder recognition method based on combined filtering

Through combined filtering and deep learning technology, the lightning monitoring problem of the lightning monitoring system in complex terrain and close range is solved, and the high accuracy of thunder recognition and accurate judgment of thunder arrival time is achieved, which is suitable for lightning positioning.

CN116665681BActive Publication Date: 2025-08-08NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211472891.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-23
Publication Date
2025-08-08
Estimated Expiration
2042-11-23

AI Technical Summary

Technical Problem

The existing lightning monitoring system is difficult to meet various terrain requirements, and cannot achieve lightning monitoring and early warning within a close range. The existing methods have shortcomings in filtering effects and thunder recognition accuracy, resulting in more false signals and unstable recognition effects, which makes it impossible to achieve real-time judgment of the arrival time of thunder.

Method used

Combined filtering methods are adopted, including Wiener filtering, spectral subtraction filtering and low-pass filtering. Combined with deep learning technology, the signal-to-noise ratio is improved through Wiener filtering, spectral subtraction filtering further enhances the thunder signal, low-pass filtering removes high-frequency harmonics, and combined with the frequency domain BARK subband variance endpoint detection, to achieve accurate judgment of the arrival time of the thunder.

Benefits of technology

It improves the accuracy and stability of thunder recognition, can achieve accurate judgment of thunder arrival time, is suitable for lightning positioning, and reduces the workload of manual judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116665681B_ABST
    Figure CN116665681B_ABST
Patent Text Reader

Abstract

The present invention discloses a thunder recognition method based on combined filtering. The method comprises the following steps: first, performing a combined filtering process consisting of data preprocessing, Wiener filtering, spectral subtraction filtering, and low-pass filtering on the data to be recognized and the training data; then, extracting the spectral features of the data; and training the eigenvectors of the training data using a deep convolutional neural network to obtain a thunder recognition model. Furthermore, the recognition result is obtained by combining the eigenvectors of the data to be recognized; and finally, the arrival time of the thunder is obtained by endpoint detection of the frequency domain BARK subband variance of the thunder audio. The present invention utilizes a rational combination of three filtering techniques and deep learning to significantly improve the accuracy and stability of thunder recognition and meet the requirements of lightning location for the arrival time of the lightning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of lightning signal monitoring and filtering, and in particular to a thunder recognition method based on combined filtering. Background Art

[0002] Real-time monitoring of lightning signals is fundamental to lightning location and early warning, a crucial step in achieving lightning disaster prevention and mitigation. While my country has established lightning detection systems in many provinces, current multi-station systems struggle to adapt to diverse terrains (e.g., ocean and mountainous terrain). Furthermore, limited by their monitoring range, they are unable to meet the needs of close-range (20 km) lightning monitoring and early warning systems in locations such as chemical plants, oil depots, and signal base stations. In 2012, Zhang Han et al. designed a single-station location system based on thunder signals, but unfortunately, their solution failed to address the problem of thunder recognition.

[0003] Single-station lightning location determines the source of thunder using a time difference algorithm based on the arrival time of thunder and electromagnetic signals. Since the collected audio contains multiple sounds, including thunder, thunder identification is necessary to ensure efficient and accurate data acquisition. While deep learning and other methods have been used for thunder identification, existing single-station lightning location methods lack sufficient filtering of background noise, resulting in a high number of false signals, low recognition accuracy, and unstable results. Furthermore, existing methods cannot accurately determine the arrival time of thunder, making them unsuitable for lightning location. Therefore, designing a thunder identification method with excellent filtering performance, stable recognition results, and the ability to determine the arrival time of thunder is crucial for lightning monitoring and early warning, providing strong support and a crucial industry need. Summary of the Invention

[0004] In view of the impact of filtering effect on the accuracy of thunder recognition and the role of the double threshold method in judging the time endpoint, the present invention proposes a thunder recognition method based on combined filtering. The purpose is to utilize the high accuracy and efficiency of deep learning for sound recognition and the role of endpoint detection in judging the time of thunder occurrence. At the same time, it can achieve a better filtering effect based on the distribution characteristics of thunder energy by using a combined filtering method of Wiener filtering, spectral subtraction filtering, and low-pass filtering. In combination with diversified training data, it meets the requirements of sound recognition accuracy, stability and the required thunder arrival time when thunder is located.

[0005] The technical solution involved in this scheme is as follows: a thunder recognition method based on combined filtering, the method comprising the following steps:

[0006] (1) Combined filtering is performed on the thunder data to be identified and the training data. Based on the characteristics of thunder energy mainly concentrated in the low-frequency part, low signal-to-noise ratio, and more noise in the natural background, Wiener filtering is first performed to improve the sound signal-to-noise ratio, filter out high-frequency signals above 400 Hz, and enhance the thunder signal; then spectral subtraction filtering is performed to further enhance the thunder signal and better realize the noise processing of the earlier signals in the time series; finally, low-pass filtering is performed to filter out the high-frequency harmonic part above 200 Hz to make up for the shortcomings of Wiener filtering and spectral subtraction filtering;

[0007] (2) Extracting spectral features from the filtered data;

[0008] (3) The thunder data and non-thunder data in the training data are labeled, and the spectral feature vectors and corresponding labels of the training data are input into the neural network for training to obtain a thunder recognition model. The spectral features of the thunder data to be identified are then used as feature vectors and combined with the trained neural network to obtain a thunder recognition model to determine whether the thunder data to be identified is thunder;

[0009] (4) Performing endpoint detection of the frequency domain BARK sub-band variance on the data identified as thunder to determine the time point of the thunder segment in the thunder audio;

[0010] (5) Output thunder arrival time and recognition results.

[0011] Preferably, the training data in step (1) are all audio data collected from natural environment backgrounds in different regions, including rain, thunder, road noise, car horns, waves, chickens and dogs barking, and other interfering sound samples, to ensure the diversity of the training data, so that the neural network can maintain a high recognition accuracy for various audio data.

[0012] Preferably, the specific process of designing the combined filter in step (1) is as follows:

[0013] First, the data is normalized, superimposed with Gaussian white noise, pre-emphasized, framed and windowed, and processed with fast Fourier transform. Then, the energy entropy ratio of the data is calculated, and the double threshold method is used to detect the energy entropy ratio endpoint. Then, the power spectrum estimate of the noisy signal is calculated based on the detection result to avoid the data containing thunder at the beginning of the Wiener filter. Then, the gain function of the Wiener filter is calculated and the amplitude is processed. The data after Wiener filtering is re-synthesized into speech, and normalized, superimposed with Gaussian white noise, framed and windowed, and processed with fast Fourier transform again. Then, the multi-window spectrum method is used to calculate the power spectrum estimate and smooth it. Finally, the gain factor of the spectral subtraction filter is calculated. The optimal subtraction factor is 2.8 (which can effectively remove the noise signal and ensure that the thunder signal is not distorted), the gain compensation factor is 0.001, the amplitude is processed, and the speech is synthesized again; finally, the time series of the signal after spectral subtraction filtering is input into the Chebyshev II low-pass filter (the optimal cutoff frequency is 200Hz and the stopband frequency is 250Hz). Since the frequency band of thunder is mainly concentrated in the low-frequency part, the high-frequency signals above 200Hz are filtered out, and finally the time series of the low-pass filtered signal is output to complete the combined filtering.

[0014] Preferably, the energy entropy ratio endpoint detection process is as follows:

[0015]

[0016]

[0017] Where p(i,n) represents the probability density corresponding to the i-th frequency component of the n-th frame, Y(i,n) represents the amplitude corresponding to the i-th frequency component of the n-th frame after fast Fourier transform, P represents the number of frames after framing, and E(n) represents the energy value of the i-th frame. The short-time spectral entropy of each frame is obtained, and the spectral entropy value function is defined as follows:

[0018]

[0019] Where H(n) represents the short-time spectrum entropy value of the nth frame. The energy value E(n) and the short-time spectrum entropy value H(n) of each frame obtained by the above process are brought into the energy entropy ratio function calculation, and the energy entropy ratio is normalized. The energy entropy ratio function is defined as follows:

[0020]

[0021] Where, E f(n) represents the energy entropy ratio of each frame, and the speech frames with normalized energy entropy ratio greater than T1 (T1 is preferably 0.1 in this example) are screened, and the adjacent frames are recorded as a voiced segment, and the information of each frame and each voiced segment is recorded. Among the voiced segments with energy entropy ratio greater than T1 obtained by screening, the voiced segments with frame length less than minl (minl is preferably 5 in this example) are eliminated, and the frame positions of the first frame and the last frame in the screened voiced segment in the original signal are recorded as l1 and l2 respectively, and an empty array SF of length P is set, and the value of index l1~l2 is set to 1, and whether each frame has sound; then the initial noise power spectrum variance is calculated. The specific calculation formula of L is as follows:

[0022]

[0023] Where L(i) represents the initial noise power spectrum variance of the i-th frequency component, and T represents the transpose; then according to the number of silent segments P n Update the array SF, for index less than or equal to P n The value of is reassigned to 0; finally, each frame is updated according to the value of the array SF (0 or 1). Taking the nth frame as an example, if SF(n) is 1, there is no need to update the power spectrum estimation, that is, n=1,2,…,P n ; Initialize L u (i,0)=L(i). If SF(n) is 0, the power spectrum estimate needs to be updated:

[0024]

[0025] Where, L u (i,n) represents the frequency spectrum estimation value of the i-th frequency component of the n-th frame.

[0026] Preferably, the Wiener filter gain function calculation formula is as follows:

[0027]

[0028]

[0029]

[0030] S(i,n)=G(i,n) 2 ×Y(i,n) 2 ,i=1,2,…,[(N / 2)+1],n=1,2,…,P (8)

[0031] Where G(i,n) represents the spectral gain function of the i-th frequency component signal of the n-th frame, SNR l(i,n) is the posterior signal-to-noise ratio of the i-th frequency component of the n-th frame, L u (i,n) represents the noise power spectrum estimation of the i-th frequency component of the n-th frame; the present invention selects α as 0.99, SNR p (i,n) is the prior signal-to-noise ratio of the i-th frequency component in the n-th frame, and S(i,n) is the power spectrum estimate of the i-th frequency component in the n-th frame of the pure sound signal;

[0032] The calculation formula of the gain factor of the spectral subtraction filter is as follows:

[0033]

[0034] Wherein, α is the over-reduction factor, and the preferred over-reduction factor of the present invention is 2.8; β is the gain compensation factor, and the preferred gain compensation factor is 0.001; L n (i) represents the short average power spectrum of the leading silent signal of the i-th frequency component, L p (i,n) is the power spectrum estimate of the smoothed i-th frequency component of the n-th frame, g(i,n) and g′(i,n) are two calculation formulas for the gain factor of the i-th frequency component of the n-th frame. When g(i,n) of the i-th frequency component of the n-th frame is greater than 0, the gain factor G′(i,n) is equal to g(i,n); conversely, when g(i,n) of the i-th frequency component of the n-th frame is less than 0, the gain factor G′(i,n) is equal to g′(i,n).

[0035] Preferably, the Chebyshev II low-pass filter has a cutoff frequency of 200 Hz, a stopband frequency of 250 Hz, a ripple passband of 1 dB, and a stopband attenuation of 80 dB.

[0036] Preferably, the calculation process of the sound spectrum feature is as follows:

[0037] First, for the combined filtered training data and the data to be identified with a duration of 5 seconds, only the time series of the signal with a duration of 2.97 seconds is extracted, and frame windowing and fast Fourier transform processing are performed on it. The shift selection is 1024, the frame length N is 2048, and the total number of frames is 127. Then, the filter endpoints are calculated. In the present invention, 40 Mel filters are used, the upper frequency is limited to 22.05KHz, and the lower frequency is limited to 300Hz. The upper and lower frequencies are converted to logarithmic frequencies of 3923.33 and 401.97 respectively. The conversion formula is as follows:

[0038]

[0039] Where f′ is the converted logarithmic frequency, and f is the Hz frequency. Since this example uses 40 filters, 42 points need to be divided between 3923.33 and 401.97, m = (401.97, 487.85, …, 3923.33). Convert m(i) back to the Hz frequency h(i) from equation (28) and calculate the FFT bin as follows:

[0040]

[0041] In the formula, f(i) corresponds to each adjacent frequency point, and floor means rounding down. Finally, the sound spectrum energy characteristics are calculated: the calculation formula is as follows:

[0042]

[0043]

[0044] Where H m (i) is the weight of each i-th frequency component output by the m-th filter, and T(n,m) represents the m-th eigenvalue of the n-th frame. Each frame of data is then normalized by subtracting the average value of each frame from the eigenvalues of that frame to ensure that the average value of each frame is zero. Finally, a 128×128×1 feature matrix is output, and missing values are padded with 0.

[0045] As a preference, the neural network architecture in step (3) uses 11 convolution layers with 3×3 convolution kernels and 4 maximum pooling layers with 2×2 pooling kernels for feature extraction, and uses multiple convolution layers with smaller convolution kernels (3×3) to replace a convolution layer with a larger convolution kernel. This can reduce parameters and have more nonlinear mappings, thereby increasing the fitting expression ability of the network; and uses 3 fully connected layers, 1 input layer and 1 output layer; the specific architecture is as follows Figure 4 shown.

[0046] Preferably, the steps of detecting the endpoints of the BARK sub-band variance in the frequency domain in step (4) are as follows:

[0047] First, the BARK subband is set, and the data sampling rate used in the present invention is f s =44.1KHz, There are 25 BARK subbands in the image; the data is normalized, superimposed with Gaussian white noise, framed and windowed, and fast Fourier transformed in sequence; then the data is interpolated by resampling, and the BARK subbands that are greater than the sampling rate at this time are removed again according to the resampling sampling rate; then the variance of the BARK subband in each frame is calculated, and then the variance of each frame is smoothed by multiple median filtering methods, and 15 times and 30 times the average value of the smoothed variance are used as the lower threshold T1 and the upper threshold T2 of the double threshold method respectively. Finally, the double threshold method is used to judge the variance of each frame to obtain the time segment of thunder, and the sound parameter is initialized to 1 and the silent parameter is 0; the judgment process starts with the first judgment module:

[0048] (2) When the variance of the nth frame exceeds T2, the frame is considered to be in the speech segment, the voice parameter is increased by 1, and the second judgment module is entered to directly judge whether the variance of the next frame exceeds T1:

[0049] (1.3) If the variance of the frame is greater than T1, the frame is considered to be in the speech segment, the voiced parameter is increased by 1, and the next frame variance is determined to be greater than T1 until it is less than T1;

[0050] (1.4) If the variance of the frame is less than T1, the silent parameter is increased by 1, and the process enters the third judgment module: first, determine whether the silent parameter is less than 8. If so, increase the voice parameter by 1, and return to the second judgment module to determine whether the variance of the next frame is greater than T1, until the silent parameter is greater than or equal to 8; if the silent parameter is greater than or equal to 8 and the voice parameter is less than 5, the speech segment is considered too short and is noise, and all parameters are reset to zero; if the silent parameter is greater than or equal to 8 and the voice parameter is greater than or equal to 5, the voice segment is considered to have ended, and all frames of the voice segment are recorded, and then all parameters are reset to zero to prepare for the next speech;

[0051] (2) When the variance of the nth frame is greater than T1 and less than T2, it is considered that the frame may be in the speech segment, the voice parameter is increased by 1, and the first judgment module is continued to judge the next frame;

[0052] (3) When the variance of the nth frame is less than T1, it is considered to be silent and the first judgment module continues to judge the next frame;

[0053] Among them, (1), (2), and (3) are the first judgment module, (1.1) is the second judgment module, and (1.2) includes the second judgment module and the third judgment module; the arrival time of thunder is obtained based on the information of each frame of the recorded sound segment.

[0054] Compared with the prior art, the present invention has the following beneficial effects:

[0055] 1. The combined filtering designed by the present invention based on the distribution characteristics of thunder energy in the frequency band can better filter out interference signals and improve the accuracy and stability of neural network recognition.

[0056] 2. The present invention uses a deep convolutional neural network to analyze and identify the sound spectrum characteristics of filtered thunder, and can achieve an accuracy rate of over 93% for identifying various types of thunder.

[0057] 3. The present invention can determine the time period of thunder arrival while identifying thunder, and is suitable for application in thunder location technology, reducing the workload of manual judgment. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is a flow chart of the method of the present invention;

[0059] Figure 2 This is a combined filtering flow chart of the present invention;

[0060] Figure 3 This is a flow chart of endpoint detection of frequency domain BARK sub-band variance of the present invention;

[0061] Figure 4 Schematic diagram of the convolutional neural network architecture;

[0062] Figure 5 This is a comparison diagram of the original audio time domain waveform and the combined filtered waveform;

[0063] Figure 6 This is a comparison chart of neural network recognition accuracy;

[0064] Figure 7 This is an example of the endpoint detection results of the BARK sub-band variance in the frequency domain. DETAILED DESCRIPTION

[0065] As attached Figure 1-4 A thunder recognition method based on combined filtering is shown, the method comprising the following steps:

[0066] (1) Combined filtering is performed on the thunder data to be identified and the training data. Based on the characteristics of thunder energy mainly concentrated in the low-frequency part, low signal-to-noise ratio, and more noise in the natural background, Wiener filtering is first performed to improve the sound signal-to-noise ratio, filter out high-frequency signals above 400 Hz, and enhance the thunder signal; then spectral subtraction filtering is performed to further enhance the thunder signal and better realize the noise processing of the earlier signals in the time series; finally, low-pass filtering is performed to filter out the high-frequency harmonic part above 200 Hz to make up for the shortcomings of Wiener filtering and spectral subtraction filtering;

[0067] (2) Extracting spectral features from the filtered data;

[0068] (3) The thunder data and non-thunder data in the training data are labeled, and the spectral feature vectors and corresponding labels of the training data are input into the neural network for training to obtain a thunder recognition model. The spectral features of the thunder data to be identified are then used as feature vectors and combined with the trained neural network to obtain a thunder recognition model to determine whether the thunder data to be identified is thunder;

[0069] (4) Performing endpoint detection of the frequency domain BARK sub-band variance on the data identified as thunder to determine the time point of the thunder segment in the thunder audio;

[0070] (5) Output thunder arrival time and recognition results.

[0071] The training data comes from audio data collected in natural environments in multiple regions. It is 5 seconds long and has a sampling rate of 44.1KHz. It includes thunder data and non-thunder data (non-thunder data contains various interfering sound samples such as rain, road noise, car horns, waves, chickens and dogs barking, etc. These sound samples are centrally classified into the non-thunder sample data set).

[0072] The combined filtering process is as follows: first, the data is normalized, superimposed with Gaussian white noise, pre-emphasized, framed and windowed, and fast Fourier transform processed in sequence; then the energy entropy ratio is calculated for the data, and the energy entropy ratio endpoint detection is performed using the double threshold method, and then the power spectrum estimation value of the noisy signal is calculated based on the detection result to avoid the data containing thunder at the beginning of the Wiener filter; then the gain function of the Wiener filter is calculated and the amplitude is processed; the data after Wiener filtering is re-synthesized into speech, and normalized, superimposed with Gaussian white noise, framed and windowed, and fast Fourier transform processed again; then the multi-window spectrum method is used to calculate the power spectrum estimation and After smoothing, the gain factor of the spectral subtraction filter is calculated. The optimal over-subtraction factor is 2.8 (which can effectively remove noise signals while ensuring that the thunder signal is not distorted), and the gain compensation factor is 0.001. The amplitude is processed and the speech is synthesized again. Finally, the time series of the signal after spectral subtraction filtering is input into a Chebyshev II low-pass filter (with an optimal cutoff frequency of 200 Hz and a stopband frequency of 250 Hz). Since the frequency band of thunder is mainly concentrated in the low-frequency part, the high-frequency signals above 200 Hz are filtered out. Finally, the time series of the low-pass filtered signal is output to complete the combined filtering.

[0073] The thunder recognition method based on combined filtering designed by the present invention is developed based on convolutional neural networks, multiple filtering methods and endpoint detection methods; this method sequentially pre-processes the thunder data to be recognized and the training data, performs Wiener filtering, spectral subtraction filtering and low-pass filtering to improve the signal-to-noise ratio of the sound and obtain the main frequency part of the thunder; then the spectral features of the filtered audio signal are extracted, and the spectral feature vectors of the training data are input into the neural network together with the corresponding data classification labels for training to obtain a thunder recognition model; then the spectral feature vectors of the audio to be recognized are input into the model for model matching to obtain the recognition result; finally, the frequency domain BARK sub-band variance of the audio identified as thunder is subjected to endpoint detection to output the time when the thunder occurred. Figure 1 The specific implementation steps are as follows:

[0074] S1, combined filtering: This process first normalizes the audio signal, superimposes Gaussian white noise, performs pre-emphasis, and performs frame windowing and other data pre-processing before filtering; then performs energy entropy ratio endpoint detection to avoid the signal from the beginning and containing thunder, in preparation for Wiener filtering; then calculates the Wiener filter gain function and the spectral subtraction filter gain factor in sequence to perform Wiener filtering and spectral subtraction filtering, and finally completes the filtering through the designed Chebyshev II low-pass filter to output the time series of the signal; Figure 2 As shown, the specific process is as follows:

[0075] S11, data preprocessing: First read the time series x(n1) of the audio signal and the sampling rate f of the audio s (where n1 is equal to the audio duration multiplied by the audio sampling rate f s ), in this example, the frame length N is set to 0.025×f s (i.e., the frame length is 25ms), and the frame shift M is 0.01×f s (i.e., the frame shift is 10ms), in this example, f s The frequency is 44.1KHz, and the leading silence length is 0.25s;

[0076] The first step is to eliminate the DC component of the audio signal and normalize the amplitude. The formula involved is as follows:

[0077]

[0078]

[0079] In the formula, max(|x(j)|) represents the value with the largest absolute value in the array x;

[0080] The second step is to superimpose Gaussian white noise on the audio signal x to generate a Gaussian white noise matrix n of the same size as the audio signal x. x, calculate the average energy x_power and n_power of the audio signal and Gaussian white noise respectively. The average energy calculation formula is as follows:

[0081]

[0082]

[0083] Then calculate the corresponding white noise n according to the average energy of the noise w In this example, the signal-to-noise ratio (SNR) is set to 5dB, and the signal after superimposing the noise is y(i)=x(i)+n w (i), where i represents the i-th sampling point, n w (i) is in the following form:

[0084]

[0085] The third step is pre-emphasis processing. The audio signal after superimposing Gaussian white noise has skin effect and dielectric loss during signal transmission. Pre-emphasis processing compensates for the loss of the high-frequency part to a certain extent. A one-dimensional filter is used to process the noisy signal. In this example, the filter processing function is: ay′(i)=)(1(y(b)+b(2)uyi-1). The parameters a, b(1) and b(2) selected in this example are 1, 1 and -0.97 respectively. In the formula, y′(i) is the signal of the i-th point after pre-emphasis processing.

[0086] The fourth step is to set the window function and frame division, and divide the noisy signal y′ into frames according to the frame length N and the frame shift SP. The data of the i-th sampling point of the n-th frame after framing is represented as y n (i), where i = 1, 2, ..., N, and the total number of frames P is (n1-N+M) / M. Using the Hamming window w(n) as the window function, the data of each frame is multiplied by the Hamming window. Taking the data of the i-th frequency component of the n-th frame as an example, the windowed form is, y′ n (i) = y n (i)×w(i), w(i) is in the following form:

[0087]

[0088] The fifth step is fast Fourier transform, which performs fast Fourier transform on the signal after frame division and windowing in the previous step to obtain the spectrum Y of each frame. n =[Y n (1),Y n (2),…,Y n (N)] T :

[0089]

[0090] Where Y n(i) Represents the spectrum value of the ith frequency component of the nth frame, and finally after the transformation, a matrix of size N×P can be obtained; due to the symmetry of the fast Fourier transform result, only the first The spectrum value of each frequency component of each frame is obtained by fast Fourier transform. and phase angle In the formula, the [] symbol represents the integer part of the number.

[0091] The sixth step is to calculate the number of frames P of the silent segment n , the formula is as follows:

[0092]

[0093] Where, 0.25 is the length of the silent segment, f s , N and M are sampling rate, frame length and frame shift respectively.

[0094] S12, energy entropy ratio endpoint detection: The Wiener filtering method requires setting a leading thunder-free segment length. In this example, the value is set to 0.25s. To prevent the audio signal from containing thunder signals starting from time 0, the sound signal needs to be subjected to energy entropy ratio endpoint detection before Wiener filtering. The judgment is made using a double threshold method. If the energy entropy ratio is greater than a certain threshold and lasts for a fixed period of time, the signal is considered to contain thunder signals. Then, the power spectrum estimate of the noisy signal is calculated based on the detection results. The specific steps are as follows:

[0095] The normalized spectrum probability density function and energy value of each frame are calculated based on the positive frequency spectrum values obtained in S11. The normalized spectrum probability density function and energy value calculation formulas are as follows:

[0096]

[0097]

[0098] Where p(i,n) represents the probability density corresponding to the i-th frequency component of the n-th frame, and E(n) represents the energy value of the i-th frame, thereby obtaining the short-time spectral entropy of each frame. The spectral entropy function is defined as follows:

[0099]

[0100] Where H(n) represents the short-time spectrum entropy value of the nth frame. The energy value E(n) and the short-time spectrum entropy value H(n) of each frame obtained from the above process are brought into the energy entropy ratio function calculation, and the energy entropy ratio is normalized. The energy entropy ratio function is defined as follows:

[0101]

[0102] Where, E f (n) represents the energy entropy ratio of each frame. The speech frames with normalized energy entropy ratio greater than T1 (T1 is preferably 0.1 in this example) are screened, and the adjacent frames are recorded as a voiced segment, and the information of each frame and each voiced segment is recorded. Among the voiced segments with energy entropy ratio greater than T1 obtained by screening, the voiced segments with frame length less than minl (minl is preferably 5 in this example) are eliminated, and the frame positions of the first and last frames in the screened voiced segments in the original signal are recorded as l1 and l2 respectively. An empty array SF of length P is set, and the value of index l1~l2 is set to 1. It is recorded whether each frame has sound. The array SF will be used in the next step to determine whether to update the power spectrum estimation of each frame of the noisy signal.

[0103] S13, update the power spectrum estimation of the noisy signal: First, calculate the initial noise power spectrum variance The specific calculation formula of L is as follows:

[0104]

[0105] Where L(i) represents the initial noise power spectrum variance of the i-th frequency component, T represents the transpose; then according to the number of silent segments P n Update the array SF, for index less than or equal to P n The value of is reassigned to 0; finally, each frame is updated according to the value of the array SF (0 or 1). Taking the nth frame as an example, if SF(n) is 1, there is no need to update the power spectrum estimation, that is, n=1,2,…,P n , initialize L u (i,0)=L(i); if SF(n) is 0, the power spectrum estimate needs to be updated:

[0106]

[0107] Where, l u (i,n) represents the frequency spectrum estimation value of the i-th frequency component of the n-th frame.

[0108] S14, calculate the Wiener filter spectral gain function and filter: The Wiener filter spectral gain function is as follows:

[0109]

[0110]

[0111]

[0112] S(i,n)=G(i,n) 2×Y(i,n) 2 ,i=1,2,…,[(N / 2)+1],n=1,2,…,P (20)

[0113] Where G(i,n) represents the spectral gain function of the i-th frequency component signal of the n-th frame, SNR l (i,n) is the posterior signal-to-noise ratio of the ith frequency component of the nth frame, and the square of the positive frequency spectrum value Y(i,n) obtained in step S11, that is, the power spectrum of the ith frequency component of the noisy signal in the nth frame and the noise power spectrum estimate L of the ith frequency component in the nth frame u (i,n) is calculated; in this example, the preferred α is 0.99, SNR p (i,n) is the prior signal-to-noise ratio of the ith frequency component of the nth frame, S(i,n) is the power spectrum estimate of the ith frequency component of the nth frame of the pure sound signal, which can be calculated by the spectral gain function G(i,n) of the ith frequency component signal of the nth frame and the positive frequency spectrum value Y(i,n) of the ith frequency component of the nth frame. Substituting Equation (19) into Equation (17), SNR p The calculation formula of (i,n) can be simplified as follows:

[0114]

[0115] Therefore, the spectral gain function and positive frequency spectrum value of each frequency component in the initial 0th frame are 1. The spectral gain function of each frequency component in each frame can be calculated in sequence according to the number of frames; the spectral gain function is multiplied by the corresponding amplitude to obtain the filtered amplitude:

[0116] X1(i,n)=G(i,n)×Y(i,n) (22)

[0117] Where X1(i,n) is the amplitude of the i-th frequency component of the n-th frame after Wiener filtering; Finally, speech synthesis is performed using the spectrum value X1 and phase angle Y a , frame length N and frame shift M are used to resynthesize the speech to obtain the signal time series x′1; and the pre-emphasis processing in S11 is eliminated. In contrast to this process, b(1)x1(i)=ax′1(i)-b(2)x1(i-1), where x1(i) is the signal at point i after the pre-emphasis processing is eliminated. Signal x1 is normalized and finally output as the normalized signal x1.

[0118] S15, data preprocessing: This step is the same as step S11 except that no pre-emphasis processing is performed and the frame length and frame shift are respectively 400 and 160 sampling points (to ensure the short-term stability of the sound signal). The main purpose is to re-superimpose Gaussian white noise on the signal; the input signal x1 is sequentially subjected to the elimination of DC components, superimposition of Gaussian white noise, and frame division and windowing to obtain each frame data x′1, and then the positive frequency spectrum value X′1 and the phase angle Y are obtained after fast Fourier transformation. a .

[0119] S16, calculate the power spectrum estimate, the average power spectrum of the leading silent segment, and perform smoothing: First, smooth the positive frequency amplitude obtained in S24 for three adjacent frames, with the first and last frames not processed. The formula is as follows:

[0120] Y'(i,n)=0.25X′1(i,n-1)+0.5X′1(i,n)+0.25X′1(i,n+1),n=2,3,…P-1 (23)

[0121] Where Y'(i,n) is the amplitude of the ith frequency component of the nth frame after smoothing. Then, the multi-window spectrum method is used to take each frame data x'1 after framing and windowing, multiply it with multiple slepian windows, and perform fast Fourier transform on it. The same frame data after fast Fourier transform is added and averaged, and then the square of the amplitude modulus is calculated to obtain the power spectrum estimate L p , L p (i,n) is the power spectrum estimation value of the i-th frequency component of the n-th frame. Then the power spectrum estimation is smoothed in the same way as in the previous step to obtain the smoothed power spectrum estimation value, which is also set as L p ; Finally, calculate the average power spectrum of the leading silent short period, the formula is as follows:

[0122]

[0123] Where, l n (i) represents the short average power spectrum of the leading silent signal of the i-th frequency component.

[0124] S17, calculate the spectral subtraction filter calculation factor and filter, the spectral subtraction gain factor calculation formula is as follows:

[0125]

[0126] Where α is the subtraction factor. In this example, the preferred subtraction factor is 2.8, which can effectively remove noise signals while ensuring that the thunder signal is not distorted. β is the gain compensation factor. In this example, the preferred gain compensation factor is 0.001. g(i,n) and g′(i,n) are two calculation formulas for the gain factor of the i-th frequency component in the n-th frame. When g(i,n) of the i-th frequency component in the n-th frame is greater than 0, the gain factor G′(i,n) is equal to g(i,n); conversely, when g(i,n) of the i-th frequency component in the n-th frame is less than 0, the gain factor G′(i,n) is equal to g′(i,n). The gain factor is multiplied by the corresponding amplitude to obtain the amplitude after spectral subtraction filtering:

[0127]

[0128] S18, speech synthesis: using spectrum value X2, phase angle Y' a , frame length N and frame shift M to resynthesize the speech to obtain the audio signal x1, and normalize the signal x2, and finally output the normalized signal time series x2.

[0129] S19, Chebyshev II low-pass filter: After the audio signal passes through Wiener filtering and spectral subtraction, the high-frequency portion of the signal is not effectively processed. Therefore, a Chebyshev II low-pass filter is designed using the fdesign.lowpass function in Matlab to further filter the high-frequency portion. Since the thunder energy in this example is primarily concentrated in the low-frequency portion, the cutoff frequency is set to 200 Hz, the stopband frequency to 250 Hz, the ripple passband to 1 dB, and the stopband attenuation to 80 dB. The time series x2 after spectral subtraction filtering is input into the low-pass filter for filtering, and the final output signal is the time series x3.

[0130] S2, Spectral Feature Extraction: This example extracts spectral energy features and inputs them into a neural network for training and recognition. The specific method for extracting spectral energy features of training data and data to be recognized is as follows:

[0131] S21, read in the audio signal: for the combined filtered training data and the data to be recognized with a duration of 5 seconds, only the time series x3 of the audio data with a duration of 2.97 seconds and the sampling rate f are extracted. s , in this example, f s It is 44.1KHz.

[0132] S22, set the window function, frame processing and fast Fourier transform: in this example, the frame shift is selected as 1024, the frame length N is 2048, and the total number of frames is 127. The specific steps are the same as S11 and will not be repeated. Finally, the positive frequency spectrum value X′3 after the fast Fourier transform is obtained. Due to the symmetry of the fast Fourier transform result, only the spectrum values of the first 1023 frequency components are used, which are recorded as N′.

[0133] S23, filter endpoint calculation: This example uses 40 Mel filters, with an upper frequency limit of 22.05 kHz and a lower frequency limit of 300 Hz. The upper and lower frequencies are converted to logarithmic frequencies of 3923.33 and 401.97, respectively. The conversion formula is as follows:

[0134]

[0135] Where f′ is the converted logarithmic frequency, and f is the Hz frequency. Since this example uses 40 filters, 42 points need to be divided between 3923.33 and 401.97, m = (401.97, 487.85, …, 3923.33). Convert m(i) back to the Hz frequency h(i) from equation (28) and calculate the FFT bin as follows:

[0136]

[0137] Where f(i) corresponds to each adjacent frequency point, and floor means rounding down.

[0138] S24, calculate the sound spectrum energy characteristics: the calculation formula is as follows:

[0139]

[0140]

[0141] Where H m (i) is the weight of each i-th frequency component output by the m-th filter, and T(n,m) represents the m-th eigenvalue of the n-th frame. Each frame of data is then normalized by subtracting the average value of each frame from the eigenvalues of that frame to ensure that the average value of each frame is zero. Finally, a 128×128×1 feature matrix is output, and missing values are padded with 0.

[0142] S3, model training and model matching: The convolutional neural network model structure in this example is as follows Figure 3 As shown in the figure, 11 convolution layers with 3×3 convolution kernels and 4 maximum pooling layers with 2×2 pooling kernels are selected for feature extraction. Multiple convolution layers with smaller convolution kernels (3×3) are used instead of one convolution layer with a larger convolution kernel. This can reduce parameters and have more nonlinear mappings, thereby increasing the fitting expression ability of the network. Three fully connected layers, one input layer, and one output layer are used. Before training, the thunder data and non-thunder data in the training data need to be marked separately as [1,0] and [0,1] respectively.

[0143] After the feature vectors of the training data and the data to be identified are input into the neural network, they are extracted by 11 convolutional layers and 4 pooling layers to obtain a multi-dimensional feature vector of size 1×1×512, and the multi-dimensional feature vector is converted into a one-dimensional feature vector. The parameter configuration of each convolutional layer and pooling layer is as follows: Figure 5 As shown; then the one-dimensional feature vector is input into the connection layer for weighted evaluation, where the number of neurons in the first fully connected layer is 1024, and the number of neurons in the second fully connected layer is 256. In the first and second connection layers, the dropout function is used to randomly stop 50% of the neurons from working to prevent overfitting and further improve the learning ability of the neural network. The number of neurons in the third fully connected layer is 2; finally, the softmax classifier is used to obtain the category probability represented by the one-dimensional vector of size 1×2. When the thunder data to be identified is input for recognition, the classification result is finally output according to the probability.

[0144] S4, endpoint detection of frequency domain BARK subband variance: The purpose of endpoint detection on the filtered audio data identified as thunder is to obtain the start and end time of thunder and realize the time positioning of lightning, such as Figure 4 The specific steps are as follows:

[0145] S41, set BARK subband: In this example, the sampling rate is 44.1KHz, and there are 25 BARK subbands F in the range of 0 to 22.05KHz. k , F k (i,2) represents the low-frequency critical frequency of the i-th BARK subband, F k (i,3) represents the high-frequency critical frequency of the i-th BARK subband.

[0146] S42, data preprocessing: The data processing method is exactly the same as that in S11 except that no pre-emphasis is performed, the signal-to-noise ratio of the superimposed noise is 10dB, and the frame length and frame shift are 400 and 160 sampling points respectively. The final positive frequency amplitude after rapid change is X4, the number of frames is P', and the number of frames of the silent segment is P' n , there are 201 sampling points in each frame.

[0147] S43, spectral line interpolation: In this example, 400 sampling points are used as the frame length and 160 sampling points are used as the frame shift, resulting in 201 positive frequency amplitude spectral lines with a frequency resolution of 55.125 Hz. According to the frequency group table, the first BARK subband is 20 Hz to 100 Hz, and the corresponding frequencies of 1 to 3 of the 201 amplitude spectral lines are 0 Hz, 55.125 Hz, and 110.25 Hz. Therefore, only two spectral lines can be selected within the first BARK subband. Using two spectral lines to calculate the variance will result in a large error. This example requires widening the spectral lines through resampling to more accurately calculate the variance value in the BARK subband: the target frequency after resampling is 22.05 kHz, resulting in a new positive frequency spectral value sequence X′4.

[0148] S44, determine the number of BARK sub-bands: after resampling, filter the BARK sub-bands within 0 to 22.05 kHz according to the sampling rate, and remove the high-frequency critical frequency F k (i,3) is greater than the first sub-band of 22.05KHz, and the number of sub-bands is Q.

[0149] S45, calculate the variance value in the BARK subband: the calculation formula is as follows:

[0150]

[0151]

[0152]

[0153] Where E(k,n) represents the average value of the spectrum in the nth frame in the kth BARK subband, Denotes the mean of the BARK subband of the nth frame, and D(n) denotes the variance of the BARK subband of the nth frame.

[0154] S46, calculate the threshold: use the method of multiple median filtering to smooth the variance and reduce the impact of variance mutation on endpoint detection; take the variance sequence as the input value, in this example, the parameter k of the median filter is 5, then the variance after smoothing is In the formula Indicates that the array index is arrive The median of the values of , the default value for the non-existent index is 0. In this example, the median filtering process is repeated 10 times; let TH be the first P ′ n The mean of the frame variance is , then the lower threshold T1 is 15 times of TH, and the upper threshold T2 is 30 times of TH.

[0155] S47, dual threshold endpoint detection: initialize the voice parameter to 1 and the silent parameter to 0; the first judgment module:

[0156] (1) When the variance of the nth frame exceeds T2, the frame is considered to be in the speech segment, the voice parameter is increased by 1, and the second judgment module is entered to directly judge whether the variance of the next frame exceeds T1:

[0157] (1.1) If the variance of the frame is greater than T1, the frame is considered to be in the speech segment, the voiced parameter is increased by 1, and the next frame variance is determined to be greater than T1 until it is less than T1;

[0158] (1.2) If the variance of the frame is less than T1, the silent parameter is increased by 1, and the process enters the third judgment module: first, determine whether the silent parameter is less than 8. If so, increase the voice parameter by 1, and return to the second judgment module to determine whether the variance of the next frame is greater than T1, until the silent parameter is greater than or equal to 8; if the silent parameter is greater than or equal to 8 and the voice parameter is less than 5, the speech segment is considered too short and is noise, and all parameters are reset to zero; if the silent parameter is greater than or equal to 8 and the voice parameter is greater than or equal to 5, the voice segment is considered to have ended, and all frames of the voice segment are recorded, and then all parameters are reset to zero to prepare for the next speech;

[0159] (2) When the variance of the nth frame is greater than T1 and less than T2, it is considered that the frame may be in the speech segment, the voice parameter is increased by 1, and the first judgment module is continued to judge the next frame;

[0160] (3) When the variance of the nth frame is less than T1, it is considered to be silent and the first judgment module is continued to judge the next frame.

[0161] Among them, (1), (2), and (3) are the first judgment module, (1.1) is the second judgment module, and (1.2) includes the second judgment module and the third judgment module.

[0162] S5. Finally, based on the information of each frame of the recorded sound segment, the arrival time of the thunder is obtained and output.

[0163] Application effect examples

[0164] The following example demonstrates the thunder recognition method based on combined filtering. A training set of 200 thunder samples, including thunder number 1, and 200 non-thunder samples were used, while a test set of over 200 thunder samples and 300 non-thunder samples was used. Both the training and test sets contain a variety of sound samples, including rain, thunder, road noise, car horns, ocean waves, and the sounds of chickens and dogs barking.

[0165] Taking the audio of thunder number 1 as an example, the original time domain waveform of the audio and the combined filtered waveform are compared. Figure 6As shown in the figure, after combined filtering, the noise amplitude is 0.0035, the thunder amplitude is 0.9392, and the signal-to-noise ratio is 24.2707, which is 14.9828dB higher than the signal-to-noise ratio before filtering. Compared with single filtering, the filtering effect is also improved.

[0166] After the audio is processed by combined filtering, Wiener filtering, spectral subtraction filtering, low-pass filtering, LMS filtering and no processing, the neural network recognition accuracy is as follows: Figure 7 As shown in the figure, the recognition accuracy after combined filtering reaches 93.18%; the recognition accuracy after Wiener filtering reaches 89.77%; the recognition accuracy after spectral subtraction filtering reaches 88.64%; the recognition accuracy after low-pass filtering reaches 81.52%; the recognition accuracy after LMS filtering reaches 78.55%; the recognition accuracy of unprocessed raw data reaches 80.23%; it can be seen that the recognition effect of the neural network after combined filtering is better than the commonly used filtering method; the sound recognition accuracy in the convolutional neural network environment reaches 93.18%.

[0167] Taking the audio of thunder number 1 as an example, the results after frequency domain BARK sub-band variance endpoint detection are as follows Figure 7 As shown, the judgment of the thunder segment is consistent with the time of thunder playback, which is consistent with the time domain waveform of the sound signal (such as Figure 5 ) also matches.

[0168] In summary, the thunder recognition method based on combined filtering provided by the present invention processes sound signals in nature based on multiple filtering technologies, and uses a deep convolutional neural network to analyze and identify the spectral characteristics of thunder. It can recognize various types of thunder with an accuracy rate of over 93%, and realizes the judgment of thunder arrival time through endpoint detection, which reduces the workload of manual judgment to a certain extent and provides a good technical foundation for lightning location technology.

Claims

1. A thunder recognition method based on combined filtering, characterized by: The method comprises the following steps: (1) Combined filtering is performed on the thunder data to be identified and the training data. Since thunder energy is mainly concentrated in the low-frequency part, the signal-to-noise ratio is low, and there are many noise characteristics in the natural background, Wiener filtering is first performed to improve the sound signal-to-noise ratio, filter out high-frequency signals above 400 Hz, and enhance the thunder signal; then spectral subtraction filtering is performed to further enhance the thunder signal and better realize the noise processing of the earlier signals in the time series; finally, low-pass filtering is performed to filter out the high-frequency harmonic part above 200 Hz to make up for the shortcomings of Wiener filtering and spectral subtraction filtering; (2) Extracting spectral features from the filtered data; (3) The thunder data and non-thunder data in the training data are labeled, and the spectral feature vectors and corresponding labels of the training data are input into the neural network for training to obtain a thunder recognition model. The spectral features of the thunder data to be identified are then used as feature vectors and combined with the trained neural network to obtain a thunder recognition model to determine whether the thunder data to be identified is thunder; (4) Performing endpoint detection of the frequency domain BARK sub-band variance on the data identified as thunder to determine the time point of the thunder segment in the thunder audio; (5) Output thunder arrival time and identification results; The combined filtering process is as follows: first, the data is normalized, superimposed with Gaussian white noise, pre-emphasized, framed and windowed, and fast Fourier transformed; then, the energy entropy ratio of the data is calculated, and the energy entropy ratio endpoint detection is performed using the double threshold method, and then the power spectrum estimation value of the noisy signal is calculated based on the detection result to avoid the data containing thunder at the beginning of the Wiener filter; then the gain function of the Wiener filter is calculated and the amplitude is processed; the data after Wiener filtering is re-synthesized into speech, and normalized, superimposed with Gaussian white noise, framed and windowed, and fast Fourier transformed again; and then the power spectrum is calculated using the multi-window spectrum method. After estimation and smoothing, the gain factor of the spectral subtraction filter is calculated. The over-subtraction factor is 2.8, which can effectively remove the noise signal while ensuring that the thunder signal is not distorted. The gain compensation factor is 0.

001. The amplitude is processed and the speech is synthesized again. Finally, the time series of the signal after spectral subtraction filtering is input into a Chebyshev II low-pass filter with an optimal cutoff frequency of 200Hz and a stopband frequency of 250Hz. Since the frequency band of thunder is mainly concentrated in the low-frequency part, the high-frequency signals above 200Hz are filtered out. Finally, the time series of the low-pass filtered signal is output to complete the combined filtering.

2. The thunder recognition method based on combined filtering according to claim 1, characterized in that: The training data comes from audio data collected in the natural environment of multiple regions. It is 5 seconds long and has a sampling rate of 44.1KHz. It includes thunder data and non-thunder data. The non-thunder data contains a variety of interference sound samples such as rain, road noise, car horns, waves, and chicken and dog barking. These sound samples are centrally classified into a non-thunder sample data set.

3. The thunder recognition method based on combined filtering according to claim 2, characterized in that: The energy entropy ratio endpoint detection process is as follows: Where p(i,n) represents the probability density corresponding to the i-th frequency component of the n-th frame, Y(i,n) represents the amplitude corresponding to the i-th frequency component of the n-th frame after fast Fourier transform, P represents the number of frames after framing, and E(n) represents the energy value of the i-th frame. The short-time spectral entropy of each frame is obtained, and the spectral entropy value function is defined as follows: Where H(n) represents the short-time spectrum entropy value of the nth frame. The energy value E(n) and the short-time spectrum entropy value H(n) of each frame obtained by the above process are brought into the energy entropy ratio function calculation, and the energy entropy ratio is normalized. The energy entropy ratio function is defined as follows: Where, E f (n) represents the energy entropy ratio of each frame, and the speech frames with energy entropy ratio greater than T1 after normalization and T1 being 0.1 are screened, and the adjacent frames are recorded as a voiced segment, and the information of each frame and each voiced segment is recorded. Among the voiced segments with energy entropy ratio greater than T1 obtained by screening, the voiced segments with frame number less than minl and minl being 5 are eliminated, and the frame number positions of the first frame and the last frame in the screened voiced segment in the original signal are recorded as l1 and l2 respectively, and an empty array SF of length P is set, and the value of index l1~l2 is set to 1, and whether each frame has sound; then the initial noise power spectrum variance is calculated. The specific calculation formula of L is as follows: Where L(i) represents the initial noise power spectrum variance of the i-th frequency component, T represents the transpose; then according to the number of silent segments P n Update the array SF, for index less than or equal to P n The value of is reassigned to 0; finally, each frame is updated according to the value of the array SF 0 or 1. Taking the nth frame as an example, if SF(n) is 1, there is no need to update the power spectrum estimation, that is, Initialize L u (i,0)=L(i). If SF(n) is 0, the power spectrum estimate needs to be updated: Where, L u (i,n) represents the frequency spectrum estimation value of the i-th frequency component of the n-th frame.

4. The thunder recognition method based on combined filtering according to claim 3, characterized in that: The Wiener filter gain function calculation formula is as follows: S(i,n)=G(i,n) 2 ×Y(i,n) 2 ,i=1,2,…,[(N / 2)+1],n=1,2,…,P (8) Where G(i,n) represents the spectral gain function of the i-th frequency component signal of the n-th frame, SNR l (i,n) is the posterior signal-to-noise ratio of the i-th frequency component of the n-th frame, L u (i,n) represents the noise power spectrum estimate of the i-th frequency component of the n-th frame; α is 0.99, SNR p (i,n) is the prior signal-to-noise ratio of the i-th frequency component in the n-th frame, and S(i,n) is the power spectrum estimate of the i-th frequency component in the n-th frame of the pure sound signal; The calculation formula of the gain factor of the spectral subtraction filter is as follows: Where, α is the over-reduction factor, which is 2.8; β is the gain compensation factor, which is 0.001; L n (i) represents the short average power spectrum of the leading silent signal of the i-th frequency component, L p (i,n) is the power spectrum estimate of the smoothed i-th frequency component of the n-th frame, g(i,n) and g'(i,n) are two calculation formulas for the gain factor of the i-th frequency component of the n-th frame. When g(i,n) of the i-th frequency component of the n-th frame is greater than 0, the gain factor G'(i,n) is equal to g(i,n); conversely, when g(i,n) of the i-th frequency component of the n-th frame is less than 0, the gain factor G'(i,n) is equal to g'(i,n).

5. The thunder recognition method based on combined filtering according to claim 4, characterized in that: The Chebyshev II low-pass filter has a cutoff frequency of 200 Hz, a stopband frequency of 250 Hz, a ripple passband of 1 dB, and a stopband attenuation of 80 dB.

6. The thunder recognition method based on combined filtering according to claim 5, characterized in that: The calculation process of the sound spectrum feature is as follows: First, for the combined filtered training data and the data to be identified with a duration of 5 seconds m = [401.97, 487.85 ... 3923.33], only the time series of the signal with a duration of 2.97 seconds is extracted, and frame windowing and fast Fourier transform processing are performed on it. The shift selection is 1024, the frame length N is 2048, and the total number of frames is 127. Then, the filter endpoints are calculated. In the present invention, 40 Mel filters are used, the upper frequency is limited to 22.05KHz, and the lower frequency is limited to 300Hz. The upper and lower frequencies are converted to logarithmic frequencies of 3923.33 and 401.97, respectively. The conversion formula is as follows: Where f' is the converted logarithmic frequency and f is the Hz frequency. Since this example uses 40 filters, 42 points need to be divided between 3923.33 and 401.

97. Convert m(i) back to the Hz frequency h(i) from equation (28) and calculate the FFT bin as follows: In the formula, f(i) corresponds to each adjacent frequency point, and floor means rounding down. Finally, the sound spectrum energy characteristics are calculated: The calculation formula is as follows: Where H m (i) is the weight of each i-th frequency component output by the m-th filter, and T(n,m) represents the m-th eigenvalue of the n-th frame. Each frame of data is then normalized by subtracting the average value of each frame from the eigenvalues of that frame to ensure that the average value of each frame is zero. Finally, a 128×128×1 feature matrix is output, and missing values are padded with 0.

7. The thunder recognition method based on combined filtering according to claim 6, characterized in that: The neural network architecture is composed of 11 convolution layers with a convolution kernel of 3×3 and 4 maximum pooling layers with a pooling kernel of 2×2 for feature extraction, 3 fully connected layers, 1 input layer and 1 output layer.

8. The thunder recognition method based on combined filtering according to claim 7, characterized in that: The endpoint detection steps of the frequency domain BARK sub-band variance are as follows: First, set the BARK subband and use the data sampling rate of f s =44.1KHz, There are 25 BARK subbands in the image; the data is normalized, superimposed with Gaussian white noise, framed and windowed, and fast Fourier transformed in sequence; then the data is interpolated by resampling, and the BARK subbands that are greater than the sampling rate at this time are removed again according to the resampling sampling rate; then the variance of the BARK subband in each frame is calculated, and then the variance of each frame is smoothed by multiple median filtering methods, and 15 times and 30 times the average value of the smoothed variance are used as the lower threshold T1 and the upper threshold T2 of the double threshold method respectively. Finally, the double threshold method is used to judge the variance of each frame to obtain the time segment of thunder, and the sound parameter is initialized to 1 and the silent parameter is 0; the judgment process starts with the first judgment module: (1) When the variance of the nth frame exceeds T2, the frame is considered to be in the speech segment, the voice parameter is increased by 1, and the second judgment module is entered to directly judge whether the variance of the next frame exceeds T1: (1.1) If the variance of the frame is greater than T1, the frame is considered to be in the speech segment, the voiced parameter is increased by 1, and the next frame variance is determined to be greater than T1 until it is less than T1; (1.2) If the variance of the frame is less than T1, the silent parameter is increased by 1, and the process enters the third judgment module: first, determine whether the silent parameter is less than 8. If so, increase the voice parameter by 1, and return to the second judgment module to determine whether the variance of the next frame is greater than T1, until the silent parameter is greater than or equal to 8; if the silent parameter is greater than or equal to 8 and the voice parameter is less than 5, the speech segment is considered too short and is noise, and all parameters are reset to zero; if the silent parameter is greater than or equal to 8 and the voice parameter is greater than or equal to 5, the voice segment is considered to have ended, and all frames of the voice segment are recorded, and then all parameters are reset to zero to prepare for the next speech; (2) When the variance of the nth frame is greater than T1 and less than T2, it is considered that the frame may be in the speech segment, the voice parameter is increased by 1, and the first judgment module is continued to judge the next frame; (3) When the variance of the nth frame is less than T1, it is considered to be silent and the first judgment module continues to judge the next frame; Among them, (1), (2), and (3) are the first judgment module, (1.1) is the second judgment module, and (1.2) includes the second judgment module and the third judgment module; the arrival time of thunder is obtained based on the information of each frame of the recorded sound segment.

Citation Information

Patent Citations

  • Urban noise identification method based on hypercomplex random neural network

    CN111540373A

  • Thunder signal recognition system and method based on convolutional neural network

    CN113011302A