Wind noise suppression method, device, equipment and computer-readable storage medium
By performing wind noise analysis on microphone signals and using deep neural network and non-neural network algorithms to process low-frequency and high-frequency signals, combined with fusion technology, the problem of stroke noise suppression in the existing technology increases hardware cost, achieving efficient wind noise suppression effect.
Patent Information
- Application Number
- CN202310180377.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-23
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-02-23
AI Technical Summary
The prior art increases the hardware cost and design difficulty of the equipment when dealing with wind noise, and cannot effectively suppress wind noise in outdoor sound pickup systems.
By performing wind noise analysis on microphone signals, low-frequency signals are processed using deep neural networks and high-frequency signals are processed using non-neural network algorithms, and wind noise suppression is achieved in combination with fusion processing technology.
Without increasing the hardware cost and design difficulty of the equipment, the wind noise suppression effect is significantly improved and the sound pickup quality is improved.
Smart Images

Figure CN116386654B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of noise reduction, and particularly to a method, device, equipment and computer-readable storage medium for suppressing wind noise. Background Art
[0002] Wind noise is one of the most common types of noise in the process of outdoor sound pickup, seriously affecting the sound pickup quality of outdoor sound pickup systems such as mobile phones and hearing aids. Wind noise is generated by the interaction between air flow and obstacles (such as buildings, human bodies, microphone cavities, etc.), so the characteristics of wind noise caused by different obstacles will also vary. In many cases, the sound pressure level of wind noise can reach 80 dB SPL, which can completely cover the voice signal, greatly reducing the voice intelligibility and causing auditory discomfort.
[0003] Traditional methods for dealing with wind noise include windshields, which are commonly found in handheld microphones and professional shotgun microphones, and are made of various materials such as sponge, artificial fur, and iron mesh. The principle is mainly to reduce the air flow velocity near the microphone diaphragm and disperse the air flow to reduce the generation of turbulence. The bone conduction sensor VPU (Voice Pick Up) designed for voice pickup can pick up voice by collecting the vibration signal of the human mandible. Since wind noise only exists in air-conducted sound and bone-conducted sound is not affected, the bone conduction sensor can directly avoid the wind noise problem when picking up voice. However, the wind noise suppression schemes based on windshields and VPU will increase the cost of the equipment and the difficulty of structural design. Summary of the Invention
[0004] The main object of the present invention is to provide a method, device, equipment and computer-readable storage medium for suppressing wind noise, aiming to provide a wind noise suppression scheme that can improve the wind noise suppression effect of the equipment without increasing the hardware cost and design difficulty of the equipment.
[0005] To achieve the above object, the present invention provides a method for suppressing wind noise, the method includes the following steps:
[0006] Obtain a microphone signal, perform wind noise analysis on the microphone signal to obtain a wind noise analysis result;
[0007] When it is determined according to the wind noise analysis result that there is wind noise in the microphone signal, perform noise cancellation processing on the low-frequency signal in the microphone signal by using a preset deep neural network to obtain a first processed signal, and perform noise cancellation processing on the high-frequency signal in the microphone signal by using a preset non-neural network algorithm to obtain a second processed signal, where the high-frequency signal is a signal with a frequency greater than a first preset frequency, and the low-frequency signal is a signal with a frequency less than or equal to the first preset frequency;
[0008] Fuse the first processed signal and the second processed signal to obtain a wind noise suppression result.
[0009] Optionally, the step of performing wind noise analysis on the microphone signal to obtain a wind noise analysis result includes:
[0010] When there are two or more microphone signals, calculate the target correlation between each pair of the microphone signals;
[0011] Match the wind noise analysis result according to the target correlation and the preset corresponding relationship between the correlation and the wind speed;
[0012] Or, calculate the target low-frequency energy of the signal with a frequency less than the second preset frequency in any one of the microphone signals;
[0013] Match the wind noise analysis result according to the target low-frequency energy and the preset corresponding relationship between the low-frequency energy and the wind speed.
[0014] Optionally, the step of calculating the target correlation between two microphone signals includes:
[0015] Calculate the number of sampling points with negative signals in each of the two microphone signals respectively;
[0016] Calculate the target correlation between the two microphone signals according to the number of sampling points.
[0017] Optionally, the deep neural network includes an encoder, a recurrent neural network module, a decoder, and a fully connected layer. The step of performing noise cancellation processing on the low-frequency signals in the microphone signals by using a preset deep neural network to obtain a first processed signal includes:
[0018] Input the low-frequency signals in each frame of the microphone signals into the encoder for processing respectively to obtain first signal processing results corresponding to each frame of the microphone signals;
[0019] Input the first signal processing results of each frame into the recurrent neural network module for processing respectively to obtain second signal processing results corresponding to the first signal processing results of each frame. Among them, when processing the target signal processing result through the recurrent neural network module, use the result obtained by processing the previous frame of the first signal processing result of the target signal processing result, and the target signal processing result is any one of the first signal processing results of each frame;
[0020] Input the second signal processing results of each frame into the decoder for processing respectively to obtain third signal processing results corresponding to the second signal processing results of each frame;
[0021] Input the third signal processing results of each frame into the fully connected layer for processing respectively to obtain the first processing signals corresponding to the microphone signals of each frame.
[0022] Optionally, the recurrent neural network module includes at least one recurrent neural network layer connected in series. The recurrent neural network layer includes a reset gate and a new memory gate. The step of inputting the target signal processing result into the recurrent neural network module for processing to obtain the second signal processing result corresponding to the target signal processing result includes:
[0023] Input the target signal processing result into the recurrent neural network module, and obtain the second signal processing result corresponding to the target signal processing result after the series processing of each recurrent neural network layer;
[0024] Wherein, the target recurrent neural network layer is any one of the recurrent neural network layers. In the process of serially processing the target signal processing result through each recurrent neural network layer, the step of inputting the target input data corresponding to the target signal processing result in the target recurrent neural network layer into the target recurrent neural network layer for processing to obtain the target output data corresponding to the target signal processing result in the target recurrent neural network layer includes:
[0025] Input the target input data, and the output data corresponding to the first signal processing result of the previous frame of the target signal processing result in the target recurrent neural network layer, into the reset gate of the target recurrent neural network layer to obtain the reset gate processing result corresponding to the target input data;
[0026] Input the target input data, the reset gate processing result corresponding to the target input data, and the output data corresponding to the first signal processing result of the previous frame of the target signal processing result in the target recurrent neural network layer, into the new memory gate of the target recurrent neural network layer to obtain the new memory gate processing result corresponding to the target input data;
[0027] Calculate and obtain the target output data according to the new memory gate processing result corresponding to the target input data, the reset gate processing result, and the output data corresponding to the first signal processing result of the previous frame of the target signal processing result in the target recurrent neural network layer.
[0028] Optionally, when there are two or more of the microphone signals, the steps of performing noise cancellation processing on the low-frequency signals in the microphone signals using a preset deep neural network to obtain a first processed signal, and performing noise cancellation processing on the high-frequency signals in the microphone signals using a preset non-neural network algorithm to obtain a second processed signal include:
[0029] Performing echo cancellation on each of the microphone signals using a far-end signal to obtain an echo-cancelled signal;
[0030] Performing beamforming on each of the echo-cancelled signals, and suppressing noise in a preset direction on each of the echo-cancelled signals based on the result of beamforming to obtain a directional noise-suppressed signal;
[0031] Performing noise cancellation processing on the low-frequency signals in the directional noise-suppressed signal using a preset deep neural network to obtain a first processed signal, and performing noise cancellation processing on the high-frequency signals in the directional noise-suppressed signal using a preset non-neural network algorithm to obtain a second processed signal.
[0032] Optionally, after the steps of obtaining a microphone signal and performing wind noise analysis on the microphone signal to obtain a wind noise analysis result, the method further includes:
[0033] When it is determined according to the wind noise analysis result that there is no wind noise in the microphone signal, performing noise cancellation processing on the microphone signal using the non-neural network algorithm to obtain a noise suppression result.
[0034] To achieve the above object, the present invention further provides a wind noise suppression device, the device includes:
[0035] A wind noise analysis module, configured to obtain a microphone signal and perform wind noise analysis on the microphone signal to obtain a wind noise analysis result;
[0036] A noise cancellation module, configured to, when it is determined according to the wind noise analysis result that there is wind noise in the microphone signal, perform noise cancellation processing on the low-frequency signals in the microphone signal using a preset deep neural network to obtain a first processed signal, and perform noise cancellation processing on the high-frequency signals in the microphone signal using a preset non-neural network algorithm to obtain a second processed signal, where the high-frequency signal is a signal with a frequency greater than a first preset frequency, and the low-frequency signal is a signal with a frequency less than or equal to the first preset frequency;
[0037] A fusion module, configured to fuse the first processed signal and the second processed signal to obtain a wind noise suppression result.
[0038] To achieve the above object, the present invention also provides a wind noise suppression device, which includes: a memory, a processor, and a wind noise suppression program stored on the memory and executable on the processor. When the wind noise suppression program is executed by the processor, the steps of the wind noise suppression method described above are implemented.
[0039] In addition, to achieve the above object, the present invention also proposes a computer-readable storage medium, on which a wind noise suppression program is stored. When the wind noise suppression program is executed by a processor, the steps of the wind noise suppression method described above are implemented.
[0040] In the present invention, by acquiring a microphone signal, performing wind noise analysis on the microphone signal to obtain a wind noise analysis result; when it is determined according to the wind noise analysis result that there is wind noise in the microphone signal, performing noise cancellation processing on the low-frequency signal in the microphone signal using a preset deep neural network to obtain a first processed signal, and performing noise cancellation processing on the high-frequency signal in the microphone signal using a preset non-neural network algorithm to obtain a second processed signal, where the high-frequency signal is a signal with a frequency greater than a first preset frequency, and the low-frequency signal is a signal with a frequency less than or equal to the first preset frequency; fusing the first processed signal and the second processed signal to obtain a wind noise suppression result. The present invention realizes a wind noise suppression scheme, which improves the wind noise suppression effect of the device without increasing the hardware cost and design difficulty of the device. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a schematic flowchart of an embodiment of the wind noise suppression method of the present invention;
[0042] Figure 2 It is a structural diagram of a deep neural network related to an embodiment of the present invention;
[0043] Figure 3 It is a structural diagram of a recurrent neural network layer related to an embodiment of the present invention;
[0044] Figure 4 It is a schematic flowchart of a wind noise suppression process related to an embodiment of the present invention;
[0045] Figure 5 It is a schematic structural diagram of the hardware operating environment related to the embodiment scheme of the present invention.
[0046] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0048] Reference Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the method for suppressing wind noise according to the present invention.
[0049] The embodiments of the present invention provide an embodiment of the method for suppressing wind noise. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from that here. In this embodiment, the execution subject of the method for suppressing wind noise may be devices such as headphones, smart phones, personal computers, servers, etc., which are not limited in this embodiment. In this embodiment, for the convenience of description, the noise reduction device is used as the execution subject to elaborate on each embodiment. In this embodiment, the method for suppressing wind noise includes the following steps:
[0050] Step S10, obtain a microphone signal, perform wind noise analysis on the microphone signal, and obtain a wind noise analysis result;
[0051] The noise reduction device can obtain one or more microphone signals and perform wind noise analysis on the obtained microphone signals. The purpose of wind noise analysis is to determine whether there is wind noise in the microphone signal or to determine the intensity of wind noise in the microphone signal. Correspondingly, the wind noise analysis result obtained through wind noise analysis can be a result indicating whether there is wind noise in the microphone signal or a result indicating the intensity of wind noise in the microphone signal. The specific implementation manner of wind noise analysis is not limited in this embodiment.
[0052] It should be noted that in the specific implementation manner, the noise reduction device can perform wind noise suppression processing on the obtained offline microphone signal, or can perform wind noise suppression processing on the microphone signal collected in real time by the microphone. For example, in a feasible implementation manner, the noise reduction device can be a headphone device, and at least one microphone is provided in the headphone device. The microphone signal is collected through the microphone. The noise reduction device obtains the microphone signal collected in real time by the microphone, performs wind noise suppression processing on the microphone signal, and then outputs the processing result through the speaker in the headphone device or sends it to other devices.
[0053] In the specific implementation manner, the noise reduction device can perform frame division processing on the microphone signal, and perform wind noise suppression processing on each frame of the microphone signal in sequence according to the frame sequence.
[0054] Step S20, when it is determined according to the wind noise analysis result that there is wind noise in the microphone signal, perform noise cancellation processing on the low-frequency signal in the microphone signal by using a preset deep neural network to obtain a first processed signal, and perform noise cancellation processing on the high-frequency signal in the microphone signal by using a preset non-neural network algorithm to obtain a second processed signal, where the high-frequency signal is a signal with a frequency greater than a first preset frequency, and the low-frequency signal is a signal with a frequency less than or equal to the first preset frequency;
[0055] The wind noise analysis result indicates whether there is wind noise in the microphone signal, or it can be determined whether there is wind noise in the microphone signal according to the wind noise analysis result. For example, when the wind noise analysis result is the intensity of wind noise in the microphone signal, the noise reduction device can determine that there is wind noise in the microphone signal when the intensity is greater than a certain level.
[0056] The deep neural network and the non-neural network algorithm can be set in advance according to needs and are not limited in this embodiment. Since extracting useful speech signals (that is, removing noise signals) from the noisy microphone signal is essentially a classification problem, compared with non-neural network algorithms (that is, traditional noise reduction algorithms), the neural network simulates the human brain recognition model and has certain advantages in solving this problem, thereby being able to improve the wind noise suppression effect.
[0057] The first preset frequency can be set according to needs and is not limited in this embodiment. By analyzing the spectrogram of the audio signal with wind noise, it is found that wind noise mainly affects the low-frequency band. In this embodiment, by performing noise cancellation processing on the low-frequency signal in the microphone signal by using a deep neural network, the wind noise suppression effect can be improved, and by performing noise cancellation processing on the high-frequency signal by using a non-neural network algorithm, the advantage of low computational complexity of the non-neural network algorithm is utilized to reduce the overall computational complexity when the noise reduction device performs wind noise suppression processing, thereby reducing the requirement for the hardware computing power of the noise reduction device.
[0058] Step S30, fuse the first processed signal and the second processed signal to obtain a wind noise suppression result.
[0059] After separately processing the microphone signal to obtain the first processed signal and the second processed signal, the noise reduction device can fuse the first processed signal and the second processed signal to obtain a wind noise suppression result. The fusion can specifically adopt a superposition or weighted fusion method, and the weight of the weighted fusion can be set according to needs and is not limited in this embodiment.
[0060] In a specific embodiment, the noise reduction device can fuse the first processed signal and the second processed signal in the time domain to obtain a fused signal in the time domain, which is the signal after suppressing wind noise from the microphone signal. The noise reduction device uses this signal as the wind noise suppression result. The noise reduction device can output the signal after suppressing wind noise in the time domain, or output it after further processing the signal. For example, it can be output after performing dynamic range control (DRC) on the signal.
[0061] In a specific embodiment, when there are multiple paths of microphone signals, the noise reduction device can process the multiple paths of microphone signals into one signal, and then perform noise cancellation processing on this signal. For example, the noise reduction device can perform beamforming processing on the multiple paths of microphone signals, suppress noise in a preset direction for each of the microphone signals based on the result of beamforming to obtain a directional noise suppression signal, and then perform noise cancellation processing on this directional noise suppression signal.
[0062] In a specific embodiment, the noise reduction device can copy one signal (one microphone signal or one signal obtained after processing multiple microphone signals) into two signals, hereinafter referred to as Signal 1 and Signal 2. In a feasible embodiment, the noise reduction device can perform full-band noise cancellation processing on Signal 1 using a deep neural network, and then perform low-pass filtering on the result of the noise cancellation processing. The upper cut-off frequency of the low-pass filtering is the first preset frequency, and the filtered signal is used as the first processed signal; perform noise cancellation processing on Signal 2 using a non-neural network algorithm, and then perform high-pass filtering on the result of the noise cancellation processing. The lower cut-off frequency of the high-pass filtering is the first preset frequency, and the filtered signal is used as the second processed signal. In another feasible embodiment, the noise reduction device can perform low-pass filtering on Signal 1. The upper cut-off frequency of the low-pass filtering is the first preset frequency, perform full-band noise cancellation processing on the filtered signal using a deep neural network, and use the processed signal as the first processed signal; perform high-pass filtering on Signal 2. The lower cut-off frequency of the high-pass filtering is the first preset frequency, perform full-band noise cancellation processing on the filtered signal using a non-neural network algorithm, and use the processed signal as the second processed signal. In a specific embodiment, the high-pass filtering and the low-pass filtering can be respectively implemented by a high-pass filter and a low-pass filter composed of 5 biquads (biquadratic filters) connected in series.
[0063] Further, in a feasible embodiment, after step S10, it further includes:
[0064] Step S40, when it is determined according to the wind noise analysis result that there is no wind noise in the microphone signal, perform noise cancellation processing on the microphone signal using the non-neural network algorithm to obtain a noise suppression result.
[0065] In the case where there is no wind noise in the microphone signal, a non-neural network algorithm is used to process the microphone signal for noise cancellation, which can reduce the computational load of the noise cancellation device, thereby reducing the requirement for the hardware computing power of the noise cancellation device.
[0066] Further, in a feasible implementation, when there are two or more of the microphone signals, the step S20 includes:
[0067] Step S211, performing echo cancellation on each of the microphone signals using a far-end signal to obtain an echo cancellation signal;
[0068] For example, there are two microphone signals respectively denoted as microphone signal 1 and microphone signal 2. The noise cancellation device uses the far-end signal to perform echo cancellation on microphone signal 1 to obtain an echo cancellation signal 1, and uses the far-end signal to perform echo cancellation on microphone signal 2 to obtain an echo cancellation signal 2.
[0069] Step S212, performing beamforming on each of the echo cancellation signals, and suppressing noise in a preset direction on each of the echo cancellation signals based on the result of the beamforming to obtain a directional noise suppression signal;
[0070] Step S213, performing noise cancellation processing on the low-frequency signals in the directional noise suppression signal using a preset deep neural network to obtain a first processed signal, and performing noise cancellation processing on the high-frequency signals in the directional noise suppression signal using a preset non-neural network algorithm to obtain a second processed signal.
[0071] The noise cancellation device can use mature algorithms to perform echo cancellation and beamforming on the microphone signal, and there is no limitation in this implementation.
[0072] In this embodiment, by acquiring the microphone signal, performing wind noise analysis on the microphone signal to obtain a wind noise analysis result; when it is determined according to the wind noise analysis result that there is wind noise in the microphone signal, performing noise cancellation processing on the low-frequency signals in the microphone signal using a preset deep neural network to obtain a first processed signal, and performing noise cancellation processing on the high-frequency signals in the microphone signal using a preset non-neural network algorithm to obtain a second processed signal, where the high-frequency signal is a signal with a frequency greater than a first preset frequency, and the low-frequency signal is a signal with a frequency less than or equal to the first preset frequency; fusing the first processed signal and the second processed signal to obtain a wind noise suppression result. A wind noise suppression scheme is provided in this embodiment, which improves the wind noise suppression effect of the device without increasing the hardware cost and design difficulty of the device.
[0073] Further, based on the above first embodiment, a second embodiment of the wind noise suppression method of the present invention is proposed. In this embodiment, a feasible implementation manner of wind noise analysis is proposed. The step S10 includes:
[0074] Step S101, when there are two or more of the microphone signals, calculate the target correlation between each of the microphone signals;
[0075] When there are two or more microphone signals, the noise reduction device can use the correlation between each microphone signal for wind noise analysis. In a specific implementation, when there are two microphone signals, the noise reduction device can directly calculate the correlation between the two microphone signals and use this correlation as the target correlation. When there are more than two microphone signals, the noise reduction device can calculate the correlation between each pair of microphone signals, calculate the average of each correlation (other fusion methods can also be used, such as addition) to obtain the target correlation, or directly use each correlation as the target correlation.
[0076] There are many ways to calculate the correlation between two microphone signals, which are not limited in this embodiment. In a feasible implementation, the Fourier transform can be performed on the two microphone signals in the time domain. For example, after the Fourier transform calculation, the 8khz bandwidth is divided into 128 sub-bands, and Y1(K) and Y2(K) respectively represent the Fourier transforms of microphone signal 1 and microphone signal 2. Calculate the coherence coefficient within the specified bandwidth, and the formula is defined as follows:
[0077]
[0078] Use this coherence coefficient as the correlation between microphone signal 1 and microphone signal 2.
[0079] Step S102, according to the target correlation and the preset corresponding relationship between the correlation and the wind speed, match to obtain the wind noise analysis result;
[0080] In advance, according to the experimental test results, the corresponding relationship between the correlation between microphone signals and the wind speed (which can represent the wind noise intensity) can be set in the noise reduction device. This corresponding relationship shows that when there is wind noise in the microphone signal or the wind speed is greater, the correlation between each microphone signal is smaller. After the noise reduction device calculates the target correlation, it can match the wind noise analysis result according to the preset corresponding relationship between the correlation and the wind speed. For example, when the wind noise analysis result is a result indicating whether there is wind noise in the microphone signal, the noise reduction device can match the wind speed corresponding to the target correlation according to the corresponding relationship. When the wind speed is greater than a certain wind speed, it is obtained that there is wind noise in the microphone signal.
[0081] In a feasible implementation, when there are more than two microphone signals and there are multiple target correlation degrees, the noise reduction device can also respectively match the wind speeds corresponding to each target correlation degree, then calculate the average of each wind speed, and then obtain the wind noise analysis result according to the calculation result.
[0082] In this embodiment, another feasible implementation of wind noise analysis is proposed. The step S10 includes:
[0083] Step S111, calculate the target low-frequency energy of the signal with a frequency less than the second preset frequency in any one of the microphone signals;
[0084] In this implementation, when there is one microphone signal, the noise reduction device performs wind noise analysis based on this microphone signal. When there are two or more microphone signals, the noise reduction device can select any one of the microphone signals from each path for wind noise analysis.
[0085] For one microphone signal, the noise reduction device calculates the low-frequency energy of the signal with a frequency less than the second preset frequency in this microphone signal (hereinafter referred to as the target low-frequency energy for distinction). Among them, the second preset frequency can be set in advance according to needs, for example, set to 1500HZ. There are many ways to calculate the target low-frequency energy, which is not limited in this implementation. For example, in a feasible implementation, the noise reduction device can first perform low-pass filtering on this microphone signal, and the upper cut-off frequency of the low-pass filtering is the second preset frequency. The low-pass filtering can be implemented by, but not limited to, an IIR (Infinite Impulse Response) filter. Let a frame of the filtered signal be x1 LP , the target low-frequency energy P low can be calculated in the following way:
[0086]
[0087] where k represents the frame length of a frame of microphone signal.
[0088] Step S112, according to the target low-frequency energy and the preset corresponding relationship between the low-frequency energy and the wind speed, match to obtain the wind noise analysis result.
[0089] In advance, according to the experimental test results, the corresponding relationship between the low-frequency energy and the wind speed (which can represent the wind noise intensity) can be set in the noise reduction device. This corresponding relationship shows that when there is wind noise in the microphone signal or the wind speed is greater, the low-frequency energy of the signal with a frequency less than the second preset frequency in the microphone signal is greater. After calculating the target low-frequency energy, the noise reduction device can match the wind noise analysis result according to the preset corresponding relationship between the low-frequency energy and the wind speed. For example, when the wind noise analysis result is a result representing whether there is wind noise in the microphone signal, the noise reduction device can match the wind speed corresponding to the target low-frequency energy according to the corresponding relationship. When the wind speed is greater than a certain wind speed, it is obtained that there is wind noise in the microphone signal.
[0090] Further, in a feasible implementation manner, the step of calculating the target correlation degree between the two microphone signals in step S101 includes:
[0091] Step S1011, respectively calculate the number of sampling points with negative signals in the two microphone signals;
[0092] For the two microphone signals for which the correlation degree needs to be calculated, the noise reduction device can respectively calculate the number of sampling points with negative signals in the two microphone signals. The specific calculation method is not limited in this implementation manner.
[0093] Step S1012, calculate the target correlation degree between the two microphone signals according to the number of sampling points.
[0094] The noise reduction device can calculate the correlation degree between the two microphone signals according to the calculated number of sampling points with negative signals in the two microphone signals. For example, in a feasible implementation manner, define the correlation degree function based on x 2 :
[0095]
[0096] where o 12 and o 22 are the elements of the following matrix:
[0097]
[0098] where represents the number of sampling points with positive signals in the microphone signal 1 from time 0 to k, represents the number of points with negative signals in the microphone signal 2 from time 0 to k, k is the frame length of one frame of the microphone signal, and N = 2k.
[0099] Further, based on the above first and / or second embodiments, a third embodiment of the wind noise suppression method of the present invention is proposed. In this embodiment, the deep neural network includes an encoder, a recurrent neural network module, a decoder, and a fully connected layer. The step of performing noise cancellation processing on the low-frequency signal in the microphone signal by using a preset deep neural network in step S20 to obtain a first processed signal includes:
[0100] Step S201: Input the low-frequency signals of each frame of the microphone signal into the encoder for processing respectively to obtain first signal processing results corresponding to each frame of the microphone signal;
[0101] In this embodiment, the preset deep neural network may include an encoder, a recurrent neural network module, a decoder, and a fully connected layer. Among them, the encoder layer is used to extract data features and downsample the input microphone signal; the recurrent neural network module is used to process the result output by the encoder layer, and the result of processing the previous frame of the microphone signal will be used during the processing, so as to use the information of the historical frame to cancel the noise of the current frame and improve the wind noise suppression effect; the decoder is used to upsample the result output by the recurrent neural network; the fully connected layer is used to process the result output by the decoder and then output the signal after noise cancellation. This deep neural network can be pre-trained through a training data set, and the training method can adopt the conventional neural network training method, which will not be elaborated here.
[0102] In a feasible implementation manner, the encoder and the decoder can draw on the encoder-decoder structure in the U-net network, that is, the decoder is used to implement the cross-connection and upsampling of data features. As Figure 2 shown, a schematic diagram of the deep neural network structure in this implementation manner is drawn. In the figure, R_rnn represents the recurrent neural network. The encoder (enconde) may include multiple encoding layers ( Figure 2 three are drawn in the figure), and the decoder (decode) may include multiple decoding layers ( Figure 2 three are drawn in the figure), and each layer of the encoder and the decoder realizes cross-connection. Among them, the encoding layer can be implemented by one-dimensional convolution (1D-conv) + downsampling + activation function. The downsampling can adopt a 2*2 pooling layer, and the activation function can use LeakyRelu, which is defined as follows:
[0103]
[0104] In a specific implementation manner, the recurrent neural network can be implemented by models such as LSTM and GRU with gate structures to have a stronger ability to suppress the gradient vanishing problem and can more effectively learn the causal relationships in the data that are far apart in time.
[0105] The noise reduction device processes frame by frame according to the frame sequence. Taking one frame as an example for illustration below, the microphone signal of this frame is called the target microphone signal for distinction.
[0106] The noise reduction device inputs the low-frequency signal in the target microphone signal into the encoder for processing, and obtains the signal processing result corresponding to the target microphone signal (hereinafter referred to as the first signal processing result for distinction).
[0107] Step S202: Input the first signal processing results of each frame into the recurrent neural network module for processing respectively, to obtain the second signal processing results corresponding to the first signal processing results of each frame. Among them, when processing the target signal processing result through the recurrent neural network module, the result obtained by processing the previous frame of the first signal processing result of the target signal processing result is used, and the target signal processing result is any one of the first signal processing results of each frame;
[0108] The first signal processing result corresponding to the target microphone signal is called the target signal processing result for distinction. The noise reduction device inputs the target signal processing result into the recurrent neural network module for processing, and the obtained signal processing result is called the second signal processing result for distinction. It can be understood that when processing the target signal processing result through the recurrent neural network, the result obtained by processing the previous frame of the first signal processing result of the target signal processing result will be used, and the previous frame of the first signal processing result of the target signal processing result is the result obtained by processing the previous frame of the microphone signal of the target microphone signal through the encoder.
[0109] Step S203: Input the second signal processing results of each frame into the decoder for processing respectively, to obtain the third signal processing results corresponding to the second signal processing results of each frame;
[0110] The noise reduction device inputs the second signal processing result corresponding to the target signal processing result into the decoder for processing, and the obtained result is called the third signal processing result for distinction.
[0111] Step S204: Input the third signal processing results of each frame into the fully connected layer for processing respectively, to obtain the first processed signals corresponding to the microphone signals of each frame.
[0112] The noise reduction device inputs the third signal processing result corresponding to the target microphone signal into the fully connected layer for processing, to obtain the first processed signal corresponding to the target microphone signal.
[0113] Further, in a feasible implementation, the recurrent neural network module includes at least one recurrent neural network layer connected in series. The recurrent neural network layer includes a reset gate and a new memory gate. The step of inputting the target signal processing result into the recurrent neural network module in step S202 to obtain a second signal processing result corresponding to the target signal processing result includes:
[0114] Step S2021, input the target signal processing result into the recurrent neural network module, and after the series processing of each recurrent neural network layer, obtain a second signal processing result corresponding to the target signal processing result;
[0115] The recurrent neural network module includes at least one recurrent neural network layer connected in series. For example, it includes two recurrent neural network layers. Then, the noise reduction device inputs the target signal processing result into the first recurrent neural network layer for processing, and the output result is input into the second recurrent neural network layer for processing to obtain a second signal processing result corresponding to the target signal processing result.
[0116] It should be noted that since only the input data of the first recurrent neural network layer is the target signal processing result, and the input data of the subsequent recurrent neural network layers are all the output data of the previous recurrent neural network layer, in this implementation, the "input data corresponding to the target signal processing result in the recurrent neural network layer" is used to represent the input data of this recurrent neural network layer when the noise reduction device uses the recurrent neural network module to process the target signal processing result. For example, assuming that the recurrent neural network module includes two recurrent neural network layers, then the input data corresponding to the target signal processing result in the first recurrent neural network layer is the target signal processing result, and the input data corresponding to the target signal processing result in the second recurrent neural network layer is the result obtained by processing the target signal processing result through the first recurrent neural network layer. Similarly, in this implementation, the "output data corresponding to the target signal processing result in the recurrent neural network layer" is used to represent the output data of this recurrent neural network layer when the noise reduction device uses the recurrent neural network module to process the target signal processing result.
[0117] The following takes the processing process of a single recurrent neural network layer as an example for illustration, and this recurrent neural network layer is called the target recurrent neural network layer for distinction, and the input data corresponding to the target signal processing result in the target recurrent neural network layer is called the target input data for distinction.
[0118] In a feasible implementation, the recurrent neural network layer can be implemented using the structure as Figure 3 shown.
[0119] In the process of processing the target signal processing result through cascading of each layer of the recurrent neural network layer, the step of inputting the target signal processing result into the target recurrent neural network layer corresponding to the target input data in step S2021 to obtain the target output data corresponding to the target signal processing result in the target recurrent neural network layer includes:
[0120] Step S20211: Input the target input data and the output data corresponding to the previous frame of the first signal processing result of the target signal processing result in the target recurrent neural network layer into the reset gate of the target recurrent neural network layer to obtain the reset gate processing result corresponding to the target input data;
[0121] In a feasible implementation manner, when using a recurrent neural network layer as shown in Figure 3 the expression of the reset gate can be:
[0122] A1(t) = sigmoid(X(t) * W1 + Y(t - 1) * V1 + B1).
[0123] Where the symbol * represents matrix multiplication, A1(t) represents the reset gate processing result corresponding to the target input data, X(t) represents the target input data, Y(t - 1) represents the output data corresponding to the previous frame of the first signal processing result of the target signal processing result in the target recurrent neural network layer, and W1, V1, and B1 are parameters in the reset gate, which can be obtained in the model training stage.
[0124] Step S20212: Input the target input data, the reset gate processing result corresponding to the target input data, and the output data corresponding to the previous frame of the first signal processing result of the target signal processing result in the target recurrent neural network layer into the new memory gate of the target recurrent neural network layer to obtain the new memory gate processing result corresponding to the target input data;
[0125] In a feasible implementation manner, when using a recurrent neural network layer as shown in Figure 3 the expression of the new memory gate can be:
[0126]
[0127] Where the symbol c represents element-wise multiplication, and W2, V2, and B2 are parameters in the new memory gate, which can be obtained in the model training stage.
[0128] Step S20213, calculate the target output data according to the new memory gate processing result and the reset gate processing result corresponding to the target input data, and the output data corresponding to the previous frame of the first signal processing result of the target signal processing result in the target recurrent neural network layer.
[0129] In a feasible implementation, when adopting a recurrent neural network layer as Figure 3 shown, the target output result is expressed as Y(t), and can be calculated using the following expression:
[0130] Y(t) = (1 - A1(t)) * Y(t - 1) + A1(t) * A2(t).
[0131] In a feasible implementation, during the training process of a deep neural network adopting a recurrent neural network layer as Figure 3 shown, the gradients of each parameter can be calculated using backpropagation, and each parameter can be updated according to the gradients. The gradients of each parameter can be calculated in the following manner.
[0132] For the new memory gate:
[0133]
[0134]
[0135]
[0136] Here, calculate the gradient of w 2,k
[0137]
[0138] Similarly
[0139]
[0140]
[0141] Among them
[0142]
[0143] 2) For the reset gate
[0144]
[0145]
[0146] Here, calculate the gradient of w 1,k
[0147] Similarly
[0148]
[0149]
[0150] wherein
[0151]
[0152] In a feasible embodiment, the noise reduction device may perform wind noise suppression according to the process as Figure 4 shown.
[0153] 1. The input signals are respectively the time-domain microphone signal 1 (y1) and the time-domain microphone signal 2 (y2). The microphone signals can be one or multiple. Here, two signals are taken as an example;
[0154] 2. Perform time-frequency transformation on the input time-domain microphone signals. Here, the FFT (Fast Fourier Transform) is adopted to obtain the frequency-domain signals Y1(K) and Y2(K) respectively. According to the far-end signal (loudspeaker signal), perform echo cancellation processing on the two signals respectively;
[0155] 3. Perform beamforming on the two microphone signals to suppress the noise outside the directivity;
[0156] 4. Determine whether the currently processed signal frame contains wind noise or not through the two microphone signals;
[0157] 5. If it is determined that the current frame is a non-wind noise frame, perform traditional noise cancellation processing on the microphone signal;
[0158] 6. If it is determined that the current frame is a wind noise frame, perform DNN-based noise cancellation on the low-frequency signal and perform traditional noise processing on the high-frequency signal;
[0159] 7. Perform high-pass filtering on the time-domain microphone signal after traditional noise processing to obtain the output signal out1;
[0160] 8. Perform low-pass filtering on the signal after DNN-based noise processing to obtain the output signal out2;
[0161] 9. The fused signal out = k1 * out1 + k2 * out2, where k1 and k2 are weights preset according to requirements;
[0162] 10. Perform dynamic range control (DRC) on the signals under both wind noise and non-wind noise conditions;
[0163] 11. Output the final time-domain signal out.
[0164] In addition, an embodiment of the present invention also provides a wind noise suppression device, and the device includes:
[0165] A wind noise analysis module, configured to obtain a microphone signal, perform wind noise analysis on the microphone signal, and obtain a wind noise analysis result;
[0166] A noise cancellation module, configured to, when it is determined according to the wind noise analysis result that there is wind noise in the microphone signal, perform noise cancellation processing on the low-frequency signal in the microphone signal by using a preset deep neural network to obtain a first processed signal, and perform noise cancellation processing on the high-frequency signal in the microphone signal by using a preset non-neural network algorithm to obtain a second processed signal, where the high-frequency signal is a signal with a frequency greater than a first preset frequency, and the low-frequency signal is a signal with a frequency less than or equal to the first preset frequency;
[0167] A fusion module, configured to fuse the first processed signal and the second processed signal to obtain a wind noise suppression result.
[0168] Further, the wind noise analysis module is further configured to:
[0169] When there are two or more of the microphone signals, calculate the target correlation between each of the microphone signals;
[0170] Match to obtain a wind noise analysis result according to the target correlation and a preset correspondence between the correlation and the wind speed;
[0171] Or, calculate the target low-frequency energy of the signal with a frequency less than a second preset frequency in any one of the microphone signals;
[0172] Match to obtain a wind noise analysis result according to the target low-frequency energy and a preset correspondence between the low-frequency energy and the wind speed.
[0173] Further, the wind noise analysis module is further configured to:
[0174] Respectively calculate the number of sampling points with negative signals in two of the microphone signals;
[0175] Calculate the target correlation between the two microphone signals according to the number of sampling points.
[0176] Further, the deep neural network includes an encoder, a recurrent neural network module, a decoder, and a fully connected layer, and the noise cancellation module is further configured to:
[0177] Respectively input the low-frequency signals in each frame of the microphone signal into the encoder for processing to obtain first signal processing results corresponding to each frame of the microphone signal;
[0178] The first signal processing results of each frame are respectively input into the recurrent neural network module for processing, and second signal processing results corresponding to the first signal processing results of each frame are obtained. Among them, when the recurrent neural network module processes the target signal processing result, the result obtained by using the recurrent neural network module to process the first signal processing result of the previous frame of the target signal processing result is used. The target signal processing result is any one of the first signal processing results of each frame;
[0179] The second signal processing results of each frame are respectively input into the decoder for processing, and third signal processing results corresponding to the second signal processing results of each frame are obtained;
[0180] The third signal processing results of each frame are respectively input into the fully connected layer for processing, and the first processing signals corresponding to the microphone signals of each frame are obtained.
[0181] Further, the recurrent neural network module includes at least one recurrent neural network layer connected in series. The recurrent neural network layer includes a reset gate and a new memory gate. The noise cancellation module is further configured to:
[0182] Input the target signal processing result into the recurrent neural network module, and obtain the second signal processing result corresponding to the target signal processing result after the series processing of each recurrent neural network layer;
[0183] Among them, the target recurrent neural network layer is any one of the recurrent neural network layers. During the process of serially processing the target signal processing result through each recurrent neural network layer, the noise cancellation module is further configured to:
[0184] Input the target input data and the output data corresponding to the first signal processing result of the previous frame of the target signal processing result in the target recurrent neural network layer into the reset gate of the target recurrent neural network layer, and obtain the reset gate processing result corresponding to the target input data;
[0185] Input the target input data, the reset gate processing result corresponding to the target input data, and the output data corresponding to the first signal processing result of the previous frame of the target signal processing result in the target recurrent neural network layer into the new memory gate of the target recurrent neural network layer, and obtain the new memory gate processing result corresponding to the target input data;
[0186] The target output data is calculated based on the new memory gate processing result and the reset gate processing result corresponding to the target input data, and the output data corresponding to the first signal processing result of the previous frame of the target signal processing result in the target recurrent neural network layer.
[0187] Further, when there are two or more of the microphone signals, the noise cancellation module is further configured to:
[0188] Perform echo cancellation on each of the microphone signals using a far-end signal to obtain an echo-canceled signal;
[0189] Perform beamforming on each of the echo-canceled signals, and perform noise suppression in a preset direction on each of the echo-canceled signals based on the result of the beamforming to obtain a directional noise-suppressed signal;
[0190] Perform noise cancellation processing on the low-frequency signals in the directional noise-suppressed signal using a preset deep neural network to obtain a first processed signal, and perform noise cancellation processing on the high-frequency signals in the directional noise-suppressed signal using a preset non-neural network algorithm to obtain a second processed signal.
[0191] Further, when it is determined according to the wind noise analysis result that there is no wind noise in the microphone signal, the noise cancellation module is further configured to perform noise cancellation processing on the microphone signal using the non-neural network algorithm to obtain a noise suppression result.
[0192] In addition, an embodiment of the present invention further provides a wind noise suppression device, as Figure 5 shown, Figure 5 is a schematic structural diagram of a device of a hardware operating environment related to the solution of the embodiment of the present invention. It should be noted that the wind noise suppression device in the embodiment of the present invention may be a device such as an earphone, a smart phone, a personal computer, a server, etc., and no specific limitation is made here.
[0193] Such as Figure 5As shown in the figure, the wind noise suppression device may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0194] Those skilled in the art can understand that Figure 5 the device structure shown in the figure does not constitute a limitation on the wind noise suppression device, and it may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.
[0195] As Figure 5 shown, the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a wind noise suppression program. The operating system is a program for managing and controlling the hardware and software resources of the device, and supports the operation of the wind noise suppression program and other software or programs. In Figure 5 the device shown in the figure, the user interface 1003 is mainly used for data communication with the client; the network interface 1004 is mainly used for establishing a communication connection with the server; and the processor 1001 may be used to call the wind noise suppression program stored in the memory 1005 and perform the following operations:
[0196] Obtain a microphone signal, perform wind noise analysis on the microphone signal to obtain a wind noise analysis result;
[0197] When it is determined according to the wind noise analysis result that there is wind noise in the microphone signal, perform noise cancellation processing on the low-frequency signal in the microphone signal using a preset deep neural network to obtain a first processed signal, and perform noise cancellation processing on the high-frequency signal in the microphone signal using a preset non-neural network algorithm to obtain a second processed signal, where the high-frequency signal is a signal with a frequency greater than a first preset frequency, and the low-frequency signal is a signal with a frequency less than or equal to the first preset frequency;
[0198] Fuse the first processed signal and the second processed signal to obtain a wind noise suppression result.
[0199] Further, the operation of performing wind noise analysis on the microphone signal to obtain a wind noise analysis result includes:
[0200] When there are two or more of the microphone signals, calculating a target correlation degree between each of the microphone signals;
[0201] According to the target correlation degree and a preset corresponding relationship between the correlation degree and the wind speed, matching to obtain the wind noise analysis result;
[0202] Or, calculating a target low-frequency energy of a signal with a frequency less than a second preset frequency in any one of the microphone signals;
[0203] According to the target low-frequency energy and a preset corresponding relationship between the low-frequency energy and the wind speed, matching to obtain the wind noise analysis result.
[0204] Further, the operation of calculating the target correlation degree between two of the microphone signals includes:
[0205] Respectively calculating the number of sampling points with negative signals in two of the microphone signals;
[0206] Calculating the target correlation degree between two of the microphone signals according to the number of sampling points.
[0207] Further, the deep neural network includes an encoder, a recurrent neural network module, a decoder, and a fully connected layer. The operation of performing noise cancellation processing on the low-frequency signals in the microphone signals by using a preset deep neural network to obtain a first processed signal includes:
[0208] Respectively inputting the low-frequency signals in each frame of the microphone signals into the encoder for processing to obtain first signal processing results respectively corresponding to each frame of the microphone signals;
[0209] Respectively inputting each frame of the first signal processing results into the recurrent neural network module for processing to obtain second signal processing results respectively corresponding to each frame of the first signal processing results. Among them, when processing a target signal processing result through the recurrent neural network module, using the result obtained by processing the previous frame of the first signal processing result of the target signal processing result, and the target signal processing result is any one of the first signal processing results of each frame;
[0210] Respectively inputting each frame of the second signal processing results into the decoder for processing to obtain third signal processing results respectively corresponding to each frame of the second signal processing results;
[0211] Input the third signal processing results of each frame into the fully connected layer for processing to obtain the first processing signals corresponding to the microphone signals of each frame.
[0212] Further, the recurrent neural network module includes at least one recurrent neural network layer connected in series. The recurrent neural network layer includes a reset gate and a new memory gate. The operation of inputting the target signal processing result into the recurrent neural network module for processing to obtain the second signal processing result corresponding to the target signal processing result includes:
[0213] Input the target signal processing result into the recurrent neural network module, and after the series processing of each layer of the recurrent neural network layer, obtain the second signal processing result corresponding to the target signal processing result;
[0214] Among them, the target recurrent neural network layer is any one of each layer of the recurrent neural network layer. In the process of processing the target signal processing result through the series connection of each layer of the recurrent neural network layer, the operation of inputting the target input data corresponding to the target signal processing result in the target recurrent neural network layer into the target recurrent neural network layer for processing to obtain the target output data corresponding to the target signal processing result in the target recurrent neural network layer includes:
[0215] Input the target input data and the output data corresponding to the previous frame of the first signal processing result of the target signal processing result in the target recurrent neural network layer into the reset gate of the target recurrent neural network layer to obtain the reset gate processing result corresponding to the target input data;
[0216] Input the target input data, the reset gate processing result corresponding to the target input data, and the output data corresponding to the previous frame of the first signal processing result of the target signal processing result in the target recurrent neural network layer into the new memory gate of the target recurrent neural network layer to obtain the new memory gate processing result corresponding to the target input data;
[0217] Calculate the target output data according to the new memory gate processing result and the reset gate processing result corresponding to the target input data, and the output data corresponding to the previous frame of the first signal processing result of the target signal processing result in the target recurrent neural network layer.
[0218] Further, when there are two or more microphone signals, the operation of performing noise cancellation processing on the low-frequency signals in the microphone signals using a preset deep neural network to obtain the first processing signal, and performing noise cancellation processing on the high-frequency signals in the microphone signals using a preset non-neural network algorithm to obtain the second processing signal includes:
[0219] Perform echo cancellation on each of the microphone signals using the far - end signal to obtain echo - cancelled signals;
[0220] Perform beamforming on each of the echo - cancelled signals, and perform noise suppression in a preset direction on each of the echo - cancelled signals based on the result of beamforming to obtain a directional noise - suppressed signal;
[0221] Perform noise cancellation processing on the low - frequency signals in the directional noise - suppressed signal using a preset deep neural network to obtain a first processed signal, and perform noise cancellation processing on the high - frequency signals in the directional noise - suppressed signal using a preset non - neural network algorithm to obtain a second processed signal.
[0222] Further, after the operation of acquiring the microphone signal and performing wind noise analysis on the microphone signal to obtain a wind noise analysis result, the processor 1001 can also be used to call the wind noise suppression program stored in the memory 1005 and perform the following operations:
[0223] When it is determined according to the wind noise analysis result that there is no wind noise in the microphone signal, perform noise cancellation processing on the microphone signal using the non - neural network algorithm to obtain a noise suppression result.
[0224] In addition, an embodiment of the present invention also proposes a computer - readable storage medium, on which a wind noise suppression program is stored. When the wind noise suppression program is executed by a processor, the steps of the wind noise suppression method described below are implemented.
[0225] For each embodiment of the wind noise suppression device and the computer - readable storage medium of the present invention, reference can be made to each embodiment of the wind noise suppression method of the present invention, which will not be elaborated here.
[0226] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non - exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including that element.
[0227] The serial numbers of the above - mentioned embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0228] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0229] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall equally be included in the patent protection scope of the present invention.
Claims
1. A method for suppressing wind noise, characterized in that, The method includes the following steps: Obtain a microphone signal, perform wind noise analysis on the microphone signal to obtain a wind noise analysis result; When it is determined according to the wind noise analysis result that there is wind noise in the microphone signal, perform noise cancellation processing on the low-frequency signal in the microphone signal using a preset deep neural network to obtain a first processed signal, and perform noise cancellation processing on the high-frequency signal in the microphone signal using a preset non-neural network algorithm to obtain a second processed signal, where the high-frequency signal is a signal with a frequency greater than a first preset frequency, and the low-frequency signal is a signal with a frequency less than or equal to the first preset frequency; Fuse the first processed signal and the second processed signal to obtain a wind noise suppression result; The step of performing wind noise analysis on the microphone signal to obtain a wind noise analysis result includes: When there are two or more of the microphone signals, calculate the target correlation between each of the microphone signals; According to the target correlation and the preset corresponding relationship between the correlation and the wind speed, match to obtain a wind noise analysis result; Or, calculate the target low-frequency energy of the signal with a frequency less than a second preset frequency in any one of the microphone signals; According to the target low-frequency energy and the preset corresponding relationship between the low-frequency energy and the wind speed, match to obtain a wind noise analysis result.
2. The wind noise suppression method according to claim 1, wherein The step of calculating the target correlation between two of the microphone signals includes: Respectively calculate the number of sampling points where the signals in the two microphone signals are negative; Calculate the target correlation between the two microphone signals according to the number of sampling points.
3. The wind noise suppression method according to claim 1, characterized in that The deep neural network includes an encoder, a recurrent neural network module, a decoder, and a fully connected layer. The step of performing noise cancellation processing on the low-frequency signal in the microphone signal using a preset deep neural network to obtain a first processed signal includes: Respectively input the low-frequency signals in each frame of the microphone signal into the encoder for processing to obtain first signal processing results corresponding to each frame of the microphone signal; Respectively input the first signal processing results of each frame into the recurrent neural network module for processing to obtain second signal processing results corresponding to the first signal processing results of each frame, where when processing the target signal processing result through the recurrent neural network module, use the result obtained by processing the previous frame of the first signal processing result of the target signal processing result, and the target signal processing result is any one of the first signal processing results of each frame; Respectively input the second signal processing results of each frame into the decoder for processing to obtain third signal processing results corresponding to the second signal processing results of each frame; Respectively input the third signal processing results of each frame into the fully connected layer for processing to obtain the first processed signal corresponding to each frame of the microphone signal.
4. The wind noise suppression method according to claim 3, wherein The recurrent neural network module includes at least one recurrent neural network layer connected in series. The recurrent neural network layer includes a reset gate and a new memory gate. The step of inputting the target signal processing result into the recurrent neural network module for processing to obtain a second signal processing result corresponding to the target signal processing result includes: Inputting the target signal processing result into the recurrent neural network module, and obtaining a second signal processing result corresponding to the target signal processing result after series processing through each recurrent neural network layer; Wherein, the target recurrent neural network layer is any one of the recurrent neural network layers. In the process of series processing the target signal processing result through each recurrent neural network layer, the step of inputting the target input data corresponding to the target signal processing result in the target recurrent neural network layer into the target recurrent neural network layer for processing to obtain the target output data corresponding to the target signal processing result in the target recurrent neural network layer includes: Inputting the target input data and the output data corresponding to the previous frame of the first signal processing result of the target signal processing result in the target recurrent neural network layer into the reset gate of the target recurrent neural network layer to obtain a reset gate processing result corresponding to the target input data; Inputting the target input data, the reset gate processing result corresponding to the target input data, and the output data corresponding to the previous frame of the first signal processing result of the target signal processing result in the target recurrent neural network layer into the new memory gate of the target recurrent neural network layer to obtain a new memory gate processing result corresponding to the target input data; Calculating the target output data according to the new memory gate processing result and the reset gate processing result corresponding to the target input data, and the output data corresponding to the previous frame of the first signal processing result of the target signal processing result in the target recurrent neural network layer.
5. The wind noise suppression method according to claim 1, characterized in that, When there are two or more microphone signals, the steps of performing noise cancellation processing on the low-frequency signals in the microphone signals using a preset deep neural network to obtain a first processed signal, and performing noise cancellation processing on the high-frequency signals in the microphone signals using a preset non-neural network algorithm to obtain a second processed signal include: Performing echo cancellation on each microphone signal using a far-end signal to obtain an echo cancellation signal; Performing beamforming on each echo cancellation signal, and performing noise suppression in a preset direction on each echo cancellation signal based on the result of beamforming to obtain a directional noise suppression signal; Performing noise cancellation processing on the low-frequency signals in the directional noise suppression signal using a preset deep neural network to obtain a first processed signal, and performing noise cancellation processing on the high-frequency signals in the directional noise suppression signal using a preset non-neural network algorithm to obtain a second processed signal.
6. The wind noise suppression method according to any one of claims 1 to 5, characterized in that, After the step of obtaining the microphone signal and performing wind noise analysis on the microphone signal to obtain a wind noise analysis result, the following steps are further included: When it is determined that there is no wind noise in the microphone signal according to the wind noise analysis result, the non-neural network algorithm is used to perform noise cancellation processing on the microphone signal to obtain a noise suppression result.
7. A wind noise suppression device, characterized in that The device includes: A wind noise analysis module, configured to obtain a microphone signal, perform wind noise analysis on the microphone signal, and obtain a wind noise analysis result; A noise cancellation module, configured to, when it is determined that there is wind noise in the microphone signal according to the wind noise analysis result, perform noise cancellation processing on the low-frequency signal in the microphone signal by using a preset deep neural network to obtain a first processed signal, and perform noise cancellation processing on the high-frequency signal in the microphone signal by using a preset non-neural network algorithm to obtain a second processed signal, where the high-frequency signal is a signal with a frequency greater than a first preset frequency, and the low-frequency signal is a signal with a frequency less than or equal to the first preset frequency; A fusion module, configured to fuse the first processed signal and the second processed signal to obtain a wind noise suppression result; The wind noise analysis module is further configured to: When there are two or more of the microphone signals, calculate the target correlation between the microphone signals of each path; According to the target correlation and the corresponding relationship between the correlation and the wind speed preset, match to obtain the wind noise analysis result; Or, calculate the target low-frequency energy of the signal with a frequency less than a second preset frequency in any one of the microphone signals; According to the target low-frequency energy and the corresponding relationship between the low-frequency energy and the wind speed preset, match to obtain the wind noise analysis result.
8. A wind noise suppression device, characterized in that, The wind noise suppression device includes: a memory, a processor, and a wind noise suppression program stored on the memory and executable on the processor. When the wind noise suppression program is executed by the processor, the steps of the wind noise suppression method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, characterized in that, A wind noise suppression program is stored on the computer-readable storage medium. When the wind noise suppression program is executed by the processor, the steps of the wind noise suppression method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Call voice noise reduction method and device, equipment and storage medium
CN114302286A
Speech enhancement method and device, earphone equipment and computer readable storage medium
CN114822573A