A wind noise filtering method and device suitable for multiple microphones
Through the multi-microphone wind noise filtering method, multi-channel Wiener filtering and wind noise fitting filter combined with inverse short-time Fourier transform are used to solve the problem of wind noise filtering in complex wind noise environments, improve call quality and retain the harmonics of the voice signal.
Patent Information
- Application Number
- CN202410915390.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-09
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2044-07-09
AI Technical Summary
Existing technologies have difficulty effectively filtering out wind noise in complex wind noise environments, resulting in reduced call quality. This is especially true in low-computing resource devices such as headphones. Traditional methods are ineffective, and deep learning algorithms require large computing resources and may damage the human voice.
A multi-microphone wind noise filtering method is adopted. The estimated covariance matrix of the frequency domain signal is calculated by the sample moving average estimation method. A multi-channel Wiener filter and a wind noise fitting filter are constructed. Combined with the inverse short-time Fourier transform and harmonic suppression method, multiple wind noise filtering operations are performed.
Effectively filters out wind noise, improves call quality, adapts to different wind noise environments, retains the harmonics of the voice signal, and avoids damage to the human voice.
Smart Images

Figure CN118890583B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of signal processing, and in particular to a wind noise filtering method and device applicable to multiple microphones. Background Art
[0002] When using headphones in certain windy environments, both indoors and outdoors, the wind can create turbulence near the microphone, causing wind noise that can affect the quality of voice captured by the headphones. In the frequency spectrum, wind noise is primarily concentrated in the low frequencies, where it overlaps significantly with the voiced sounds of the human voice. Furthermore, wind noise is extremely uneven, mixing with the human voice and affecting the user's listening experience. Severe wind noise can even make it impossible to hear what's being said through the headphones, severely degrading the user experience.
[0003] In the existing technology, there are two main methods for filtering out wind noise: using physical methods to filter out wind noise and using noise reduction algorithms to filter out wind noise. Among them, physical methods such as using windshields and changing the physical structure of the equipment can effectively reduce wind noise, but are not suitable for portable products such as headphones; traditional noise reduction algorithms, such as spectral subtraction, are less effective in complex wind noise scenes or some specific scenes; noise reduction algorithms based on deep learning require more computing resources, and in devices with low computing resources such as headphones, it is difficult to meet the real-time requirements for voice signal processing, and the wind noise filtering effect is poor. In addition, there is a possibility that part of the human voice will be filtered out during the wind noise filtering process, thereby damaging the human voice and reducing the quality of voice calls. Therefore, how to effectively filter out wind noise in a complex wind noise environment and thus improve call quality is still a problem that needs to be solved in the existing technology. Summary of the Invention
[0004] The present invention provides a wind noise filtering method and device applicable to multiple microphones, so as to solve the technical problem that the prior art cannot effectively filter wind noise in a complex wind noise environment and thus improve call quality.
[0005] In a first aspect, the present application provides a wind noise filtering method applicable to multiple microphones, comprising:
[0006] Acquire a first frequency domain signal to be filtered of wind noise;
[0007] Calculating an estimated covariance matrix of the first frequency domain signal based on a sample moving average estimation method;
[0008] constructing a multi-channel Wiener filter based on the estimated covariance matrix, and substituting the first frequency domain signal into the multi-channel Wiener filter to obtain a second frequency domain signal;
[0009] constructing a wind noise fitting filter based on a preset wind noise fitting coefficient, and substituting the second frequency domain signal into the wind noise fitting filter to obtain a third frequency domain signal;
[0010] The third frequency domain signal is converted into a first time domain signal based on inverse short-time Fourier transform, and the first time domain signal is subjected to wind noise filtering according to a preset first harmonic suppression method to obtain and output a speech signal after wind noise filtering.
[0011] In this way, the first frequency domain signal is first obtained, and its estimated covariance matrix is calculated based on the sample moving average estimation method. Multi-channel Wiener filtering is performed based on the multi-channel Wiener filter constructed according to the estimated covariance matrix, and wind noise fitting filtering is performed based on the wind noise fitting filter constructed according to the preset wind noise fitting coefficient. The signal is then converted from the frequency domain to the time domain based on the inverse short-time Fourier transform, and then wind noise is filtered according to the first harmonic suppression method to obtain and output the voice signal after filtering out the wind noise. Performing wind noise filtering multiple times in this way can comprehensively utilize the different characteristics of wind noise to perform discriminative filtering on the input signal, better adapt to different wind noise environments, and thus more effectively filter wind noise, thereby improving call quality.
[0012] Furthermore, after obtaining the first frequency domain signal from which wind noise is to be filtered out, the method further includes:
[0013] Based on a preset speech model with wind noise and a preset steering vector, the first frequency domain signal is modeled and decomposed to obtain a first speech signal and a first wind noise signal.
[0014] By modeling and decomposing the input signal in this way, the input signal can be represented by a model as a superposition of a speech signal and a wind noise signal, thereby providing a corresponding basis for the subsequent construction of the filter.
[0015] Furthermore, constructing a multi-channel Wiener filter based on the estimated covariance matrix, and substituting the first frequency domain signal into the multi-channel Wiener filter to obtain a second frequency domain signal specifically includes:
[0016] Based on the minimum Frobenius norm, obtaining a first speech power spectral density and an estimated wind noise power spectral density according to the estimated covariance matrix and the model covariance matrix of the first frequency domain signal;
[0017] A multi-channel Wiener filter is constructed according to the first speech power spectral density and the estimated wind noise power spectral density, and the first frequency domain signal is substituted into the multi-channel Wiener filter to obtain a second frequency domain signal.
[0018] Furthermore, constructing a wind noise fitting filter based on a preset wind noise fitting coefficient, and substituting the second frequency domain signal into the wind noise fitting filter to obtain a third frequency domain signal specifically includes:
[0019] Calculating the low-frequency wind noise power spectral density based on the wind noise fitting coefficient;
[0020] A wind noise fitting filter is constructed based on the second speech power spectral density and the low-frequency wind noise power spectral density, and the second frequency domain signal is substituted into the wind noise fitting filter to obtain a third frequency domain signal; wherein, the second speech power spectral density is obtained based on the second frequency domain signal.
[0021] In this way, the low-frequency wind noise power spectrum density is first calculated according to the preset wind noise fitting coefficient, which can reasonably estimate the power spectrum density of the low-frequency wind noise and improve the fitting accuracy of the actual low-frequency wind noise. Then, a wind noise fitting filter is constructed according to the low-frequency wind noise power spectrum density and wind noise fitting filtering is performed, which can improve the accuracy of filtering out low-frequency wind noise.
[0022] Furthermore, the converting of the third frequency domain signal into a first time domain signal based on an inverse short-time Fourier transform, and performing wind noise filtering on the first time domain signal according to a preset first harmonic suppression method to obtain and output a speech signal after filtering out the wind noise specifically includes:
[0023] Converting the third frequency domain signal into a first time domain signal based on an inverse short-time Fourier transform;
[0024] Based on a preset fundamental frequency, construct a plurality of comb filters, and substitute the first time domain signal into the plurality of comb filters to obtain a second time domain signal; wherein the plurality of comb filters are connected end to end, the output of the previous comb filter serves as the input of the next comb filter, the input of the first comb filter is the first time domain signal, and the output of the last comb filter is the second time domain signal;
[0025] Based on a preset low-pass filter, filtering the second time domain signal to obtain a third time domain signal;
[0026] A difference between the first time domain signal and the third time domain signal is calculated, and the difference is output as a speech signal after filtering out wind noise.
[0027] In this way, the input signal is first converted from the frequency domain to the time domain, and several comb filters are constructed based on the preset fundamental frequency for comb filtering, which can suppress the harmonics in the voice signal, and then low-pass filtering is performed based on the preset low-pass filter to obtain low-frequency wind noise, and the difference between the input signal and the low-frequency wind noise is used as the output signal. It can filter out the low-frequency wind noise while retaining the harmonics of the voice signal, thereby not damaging the voice signal, effectively filtering out wind noise, and thus improving call quality.
[0028] In a second aspect, the present application provides a wind noise filtering device suitable for multiple microphones, comprising a signal acquisition module, a matrix calculation module, a multi-channel Wiener filtering module, a wind noise fitting filtering module, and a harmonic suppression module;
[0029] The signal acquisition module is used to acquire a first frequency domain signal to be filtered out of wind noise;
[0030] The matrix calculation module is used to calculate the estimated covariance matrix of the first frequency domain signal based on a sample moving average estimation method;
[0031] The multi-channel Wiener filtering module is used to construct a multi-channel Wiener filter based on the estimated covariance matrix, and substitute the first frequency domain signal into the multi-channel Wiener filter to obtain a second frequency domain signal;
[0032] The wind noise fitting filter module is configured to construct a wind noise fitting filter based on a preset wind noise fitting coefficient, and substitute the second frequency domain signal into the wind noise fitting filter to obtain a third frequency domain signal;
[0033] The harmonic suppression module is used to convert the third frequency domain signal into a first time domain signal based on an inverse short-time Fourier transform, and to filter out wind noise from the first time domain signal according to a preset first harmonic suppression method, to obtain and output a speech signal after filtering out wind noise.
[0034] Furthermore, the wind noise filtering device applicable to multiple microphones further includes a modeling and decomposition module;
[0035] The modeling and decomposition module is used to model and decompose the first frequency domain signal based on a preset speech model with wind noise and a preset steering vector to obtain a first speech signal and a first wind noise signal.
[0036] Furthermore, the multi-channel Wiener filtering module includes a power spectrum density calculation unit and a multi-channel Wiener filtering unit;
[0037] The power spectral density calculation unit is configured to obtain a first speech power spectral density and an estimated wind noise power spectral density based on the estimated covariance matrix and the model covariance matrix of the first frequency domain signal based on a minimum Frobenius norm;
[0038] The multi-channel Wiener filter unit is used to construct a multi-channel Wiener filter according to the first speech power spectral density and the estimated wind noise power spectral density, and substitute the first frequency domain signal into the multi-channel Wiener filter to obtain a second frequency domain signal.
[0039] Furthermore, the wind noise fitting filtering module includes a low-frequency wind noise power spectrum density calculation unit and a wind noise fitting filtering unit;
[0040] The low-frequency wind noise power spectrum density calculation unit is used to calculate the low-frequency wind noise power spectrum density based on the wind noise fitting coefficient;
[0041] The wind noise fitting filter unit is used to construct a wind noise fitting filter based on the second speech power spectrum density and the low-frequency wind noise power spectrum density, and substitute the second frequency domain signal into the wind noise fitting filter to obtain a third frequency domain signal; wherein, the second speech power spectrum density is obtained based on the second frequency domain signal.
[0042] Furthermore, the harmonic suppression module includes an inverse short-time Fourier transform unit, a comb filter unit, a low-pass filter unit and a signal output unit;
[0043] The inverse short-time Fourier transform unit is used to convert the third frequency domain signal into the first time domain signal based on the inverse short-time Fourier transform;
[0044] The comb filter unit is configured to construct a plurality of comb filters based on a preset fundamental frequency, and substitute the first time domain signal into the plurality of comb filters to obtain a second time domain signal; wherein the plurality of comb filters are connected end to end, the output of the preceding comb filter serves as the input of the succeeding comb filter, the input of the first comb filter is the first time domain signal, and the output of the last comb filter is the second time domain signal;
[0045] The low-pass filtering unit is configured to filter the second time domain signal based on a preset low-pass filter to obtain a third time domain signal;
[0046] The signal output unit is configured to calculate a difference between the first time domain signal and the third time domain signal, and output the difference as a speech signal after wind noise is filtered out.
[0047] In this way, the first frequency domain signal is first obtained, and its estimated covariance matrix is calculated based on the sample moving average estimation method. Multi-channel Wiener filtering is performed based on the multi-channel Wiener filter constructed according to the estimated covariance matrix, and wind noise fitting filtering is performed based on the wind noise fitting filter constructed according to the preset wind noise fitting coefficient. The signal is then converted from the frequency domain to the time domain based on the inverse short-time Fourier transform, and then wind noise is filtered according to the first harmonic suppression method to obtain and output the voice signal after filtering out the wind noise. Performing wind noise filtering multiple times in this way can comprehensively utilize the different characteristics of wind noise to perform discriminative filtering on the input signal, better adapt to different wind noise environments, and thus more effectively filter wind noise, thereby improving call quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 : A schematic flow chart of an embodiment of a wind noise filtering method applicable to multiple microphones provided by the present invention;
[0049] Figure 2 : A module structure diagram of an embodiment of a wind noise filtering device suitable for multiple microphones provided by the present invention. DETAILED DESCRIPTION
[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0051] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. In the description of this application, unless otherwise explicitly and specifically defined, "several" means one or more than one.
[0052] Wind noise, professionally known as aerodynamic noise, is generated by the interaction between moving objects in a flow field, or by the interaction between fluids caused by the turbulent motion of the fluid itself. The mechanisms of wind noise generation vary in different scenarios, but there are two main scenarios in daily life that are most affected by wind noise: communicating using a microphone in indoor and outdoor environments with high wind speeds, and communicating in the cabin of a high-speed car. Based on people's requirements for sound quality and auditory experience, wind noise needs to be reduced and filtered.
[0053] In the prior art, there are a variety of methods for filtering out wind noise when using a microphone. One is to use physical methods to filter out wind noise, such as using a wind shield, a windproof ball, or changing the physical structure of the device, or simply changing the microphone's location to a windless indoor environment or an environment with low wind speed; the second is to use a deep learning noise reduction algorithm to filter out wind noise, but this method requires large computing resources and cannot be used well in low-resource devices such as headphones. It is easy to lose the real-time performance of voice signal processing. At the same time, there is a possibility that part of the human voice is also filtered out during the wind noise filtering process, resulting in a decrease in human voice quality and call quality; the third is to use a traditional noise reduction algorithm to filter out wind noise, such as spectral subtraction, but the traditional algorithm provides poor noise reduction effects in special scenarios or complex wind noise scenarios, and cannot meet people's auditory requirements in these scenarios. In response to the above problems, the present invention proposes a wind noise filtering method and device suitable for multiple microphones, and provides some embodiments of the present invention as described below.
[0054] Example 1
[0055] Please refer to Figure 1, a wind noise filtering method applicable to multiple microphones provided in an embodiment of the present invention, including steps S1 to S5, each step is specifically as follows:
[0056] Step S1: Acquire a first frequency domain signal from which wind noise is to be filtered out.
[0057] In an optional embodiment, multiple time domain signals can be collected by M microphones (M≥2) Where t is the discrete time index and T is the matrix transpose. By presetting the frame length to L, the time domain signal x(t) can be converted into a frequency domain signal through short-time Fourier transform, thereby obtaining the first frequency domain signal x(f) = [x1(f), ..., x M (f)] T , where f∈{0, F-1}, F=L / 2+1, f is the frequency band index, and F is the number of frequency bands.
[0058] Furthermore, after obtaining the first frequency domain signal from which wind noise is to be filtered out, the method further includes:
[0059] Based on a preset speech model with wind noise and a preset steering vector, the first frequency domain signal is modeled and decomposed to obtain a first speech signal and a first wind noise signal.
[0060] In an optional embodiment, by presetting a speech model with wind noise and a speech steering vector, the first frequency domain signal can be modeled and decomposed into: x(f) = g(f)s(f) + v(f), where g(f) = [g1(f), ..., g M (f)] T is the preset steering vector, v(f)=[v1(f),…,v M (f)] T is the first wind noise signal, and s(f) is the first speech signal.
[0061] By modeling and decomposing the input signal in this way, the input signal can be represented by a model as a superposition of a speech signal and a wind noise signal, thereby providing a corresponding basis for the subsequent construction of the filter.
[0062] Step S2: Calculate the estimated covariance matrix of the first frequency domain signal based on a sample moving average estimation method.
[0063] In an optional embodiment, the covariance matrix of each frequency point of each frame is estimated by a sample moving average estimation method, and an estimated covariance matrix of the first frequency domain signal is obtained:
[0064]
[0065] in, represents the estimated covariance matrix of the f-th frequency point of the n-th frame of the first frequency domain signal, and α is a preset forgetting factor.
[0066] For the convenience of formula expression, the frequency band index f is omitted in the following formula.
[0067] Step S3: constructing a multi-channel Wiener filter based on the estimated covariance matrix, and substituting the first frequency domain signal into the multi-channel Wiener filter to obtain a second frequency domain signal.
[0068] Furthermore, constructing a multi-channel Wiener filter based on the estimated covariance matrix, and substituting the first frequency domain signal into the multi-channel Wiener filter to obtain a second frequency domain signal specifically includes:
[0069] Based on the minimum Frobenius norm, obtaining a first speech power spectral density and an estimated wind noise power spectral density according to the estimated covariance matrix and the model covariance matrix of the first frequency domain signal;
[0070] A multi-channel Wiener filter is constructed according to the first speech power spectral density and the estimated wind noise power spectral density, and the first frequency domain signal is substituted into the multi-channel Wiener filter to obtain a second frequency domain signal.
[0071] Furthermore, the obtaining of the first speech power spectral density and the estimated wind noise power spectral density based on the minimum Frobenius norm and the estimated covariance matrix and the model covariance matrix of the first frequency domain signal specifically includes:
[0072] Based on the minimum Frobenius norm, constructing a first norm solution expression according to the estimated covariance matrix and the model covariance matrix;
[0073] Based on preset parameters, converting the first norm solution expression to obtain a solution expression for the first speech power spectral density and a solution expression for the estimated wind noise power spectral density, and solving the solution expressions to obtain the first speech power spectral density and the estimated wind noise power spectral density;
[0074] The first norm solution expression is specifically:
[0075]
[0076] in, are the first speech power spectral density and the estimated wind noise power spectral density, are speech power spectral density and wind noise power spectral density respectively, are the estimated covariance matrix and model covariance matrix of the first frequency domain signal respectively, argmin is the function for finding the minimum value of the variable, is the square form of the Frobenius norm;
[0077] The solution expression for the first speech power spectrum density is specifically:
[0078]
[0079] The solution expression for estimating the wind noise power spectrum density is specifically:
[0080]
[0081] in, g, g H are the steering vector and its conjugate transpose, Γ, Γ H are the wind noise spatial coherence matrix and its conjugate transpose, tr(Γ H Γ), are Γ H Γ and traces.
[0082] Furthermore, constructing a multi-channel Wiener filter according to the first speech power spectral density and the estimated wind noise power spectral density, and substituting the first frequency domain signal into the multi-channel Wiener filter to obtain a second frequency domain signal specifically includes:
[0083] Constructing a wind noise covariance matrix according to a preset wind noise spatial coherence matrix and the estimated wind noise power spectral density;
[0084] Calculating an estimated noise power spectral density based on the steering vector and the wind noise covariance matrix;
[0085] Calculating a priori signal-to-noise ratio based on the estimated noise power spectral density and the first speech power spectral density;
[0086] A multi-channel Wiener filter is constructed according to the steering vector, the wind noise covariance matrix and the priori signal-to-noise ratio, and the first frequency domain signal is substituted into the multi-channel Wiener filter to obtain a second frequency domain signal.
[0087] In an alternative embodiment, when f<f L1 In the low frequency band, the preferred construction of the multi-channel Wiener filter is:
[0088]
[0089] Among them, f L1 is the preset first frequency, ω MWF is a multi-channel Wiener filter, R x is the model covariance matrix of the first frequency domain signal, is the speech power spectral density, is the wind noise covariance matrix, is the wind noise power spectral density, is the steering vector, Γ is the preset wind noise spatial coherence matrix, is the prior signal-to-noise ratio, To estimate the noise power spectral density.
[0090] In an optional embodiment, according to the characteristics of wind noise, the preferred solution of the wind noise spatial coherence matrix Γ is to set it as a unit matrix.
[0091] In this way, the first speech power spectral density and the estimated wind noise power spectral density are first calculated based on the minimum Frobenius norm, which can reasonably estimate the power spectral densities of speech and wind noise, improve the estimation accuracy of the power spectral density, and reduce the unreasonable filtering of speech in the subsequent filtering process. Then, based on the estimated wind noise power spectral density and the preset wind noise spatial coherence matrix, the wind noise covariance matrix is constructed and the estimated noise power spectral density is calculated, and then the prior signal-to-noise ratio is calculated and a multi-channel Wiener filter is constructed. Finally, multi-channel Wiener filtering is performed based on the multi-channel Wiener filter, which can effectively filter out wind noise and improve the effect of wind noise filtering.
[0092] Step S4: constructing a wind noise fitting filter based on a preset wind noise fitting coefficient, and substituting the second frequency domain signal into the wind noise fitting filter to obtain a third frequency domain signal.
[0093] Furthermore, constructing a wind noise fitting filter based on a preset wind noise fitting coefficient, and substituting the second frequency domain signal into the wind noise fitting filter to obtain a third frequency domain signal specifically includes:
[0094] Calculating the low-frequency wind noise power spectral density based on the wind noise fitting coefficient;
[0095] A wind noise fitting filter is constructed based on the second speech power spectral density and the low-frequency wind noise power spectral density, and the second frequency domain signal is substituted into the wind noise fitting filter to obtain a third frequency domain signal; wherein, the second speech power spectral density is obtained based on the second frequency domain signal.
[0096] In an optional embodiment, the low-frequency wind noise power spectrum density is preferably constructed as follows:
[0097]
[0098] in, is the low-frequency wind noise power spectral density, is the power spectrum of the preset reference frequency point, c(f) is the preset wind noise fitting coefficient, f L2 is the preset second frequency.
[0099] In an optional embodiment, the second speech power spectrum density is preferably constructed as the difference between the noisy speech power spectrum density and the low-frequency wind noise power spectrum density.
[0100] In an optional embodiment, the noisy speech power spectrum density can be estimated from the second frequency domain signal. Specifically, the noisy speech power spectrum density is preferably constructed as the square of the amplitude of the second frequency domain signal.
[0101] In an optional embodiment, the above c(f), f L1 , f L2 The value of can be pre-set by those skilled in the art according to actual conditions or based on the experience of those skilled in the art. It should be understood that other values set by those skilled in the art without creative work fall within the scope of protection of the present invention.
[0102] In an optional embodiment, the wind noise fitting filter is preferably constructed as follows:
[0103]
[0104] Among them, ω SWF is the wind noise fitting filter, is the low-frequency signal-to-noise ratio, is the second speech power spectrum density, is the power spectral density of noisy speech, is the power spectral density of low-frequency wind noise.
[0105] In this way, the low-frequency wind noise power spectrum density is first calculated according to the preset wind noise fitting coefficient, which can reasonably estimate the power spectrum density of the low-frequency wind noise and improve the fitting accuracy of the actual low-frequency wind noise. Then, a wind noise fitting filter is constructed according to the low-frequency wind noise power spectrum density and wind noise fitting filtering is performed, which can improve the filtering effect of the low-frequency wind noise.
[0106] Step S5: converting the third frequency domain signal into a first time domain signal based on inverse short-time Fourier transform, and filtering the first time domain signal for wind noise according to a preset first harmonic suppression method to obtain and output a speech signal after filtering out wind noise.
[0107] Furthermore, the converting of the third frequency domain signal into a first time domain signal based on an inverse short-time Fourier transform, and performing wind noise filtering on the first time domain signal according to a preset first harmonic suppression method to obtain and output a speech signal after filtering out the wind noise specifically includes:
[0108] Converting the third frequency domain signal into a first time domain signal based on an inverse short-time Fourier transform;
[0109] Based on a preset fundamental frequency, construct a plurality of comb filters, and substitute the first time domain signal into the plurality of comb filters to obtain a second time domain signal; wherein the plurality of comb filters are connected end to end, the output of the previous comb filter serves as the input of the next comb filter, the input of the first comb filter is the first time domain signal, and the output of the last comb filter is the second time domain signal;
[0110] Based on a preset low-pass filter, filtering the second time domain signal to obtain a third time domain signal;
[0111] A difference between the first time domain signal and the third time domain signal is calculated, and the difference is output as a speech signal after filtering out wind noise.
[0112] In an optional embodiment, according to actual conditions, those skilled in the art may set the number of the plurality of comb filters to one, two, or more than two.
[0113] In an optional embodiment, in order to avoid the stability problem of the filter, a preferred solution for constructing a plurality of comb filters is to replace the plurality of comb filters with a plurality of IIR (Infinite Impulse Response) notch filters.
[0114] In an optional embodiment, the preferred solution for constructing a plurality of comb filters is:
[0115] Based on the preset fundamental frequency f0, the center frequency of the first comb filter is set to f0, the center frequency of the second comb filter is set to 2f0, and so on, the center frequency of the last (preset Nth) comb filter is set to Nf0.
[0116] Comb filtering is considered to suppress the harmonics of the speech signal by constructing several comb filters based on the fundamental frequency. This is because the information in the speech signal is primarily concentrated in the harmonics. By suppressing the harmonics, wind noise between the speech harmonics can be removed. Based on the low-frequency, high-energy characteristics of wind noise, low-pass filtering is performed to obtain low-frequency, high-energy wind noise (the third time domain signal). The first and third time domain signals are then subtracted to obtain the speech signal after wind noise has been filtered out.
[0117] In this way, the input signal is first converted from the frequency domain to the time domain, and several comb filters are constructed based on the preset fundamental frequency for comb filtering, which can suppress the harmonics in the voice signal, and then low-pass filtering is performed based on the preset low-pass filter to obtain low-frequency wind noise, and the difference between the input signal and the low-frequency wind noise is used as the output signal. It can filter out the low-frequency wind noise while retaining the harmonics of the voice signal, thereby not damaging the voice signal, effectively filtering out wind noise, and thus improving call quality.
[0118] By first acquiring a first frequency domain signal and calculating its estimated covariance matrix based on the sample moving average estimation method, multi-channel Wiener filtering is performed using a multi-channel Wiener filter constructed based on the estimated covariance matrix, and wind noise fitting filtering is performed using a wind noise fitting filter constructed based on preset wind noise fitting coefficients. The signal is then converted from the frequency domain to the time domain based on an inverse short-time Fourier transform, and wind noise is then filtered using the first harmonic suppression method to obtain and output a voice signal after wind noise filtering. This multiple wind noise filtering process can comprehensively utilize the different characteristics of wind noise to perform discriminative filtering on the input signal, better adapt to different wind noise environments, and thus more effectively filter wind noise, thereby improving call quality.
[0119] Example 2
[0120] Please refer to Figure 2 , a wind noise filtering device for multiple microphones provided in an embodiment of the present invention, comprising a signal acquisition module 210, a matrix calculation module 220, a multi-channel Wiener filtering module 230, a wind noise fitting filtering module 240, and a harmonic suppression module 250;
[0121] The signal acquisition module 210 is used to acquire a first frequency domain signal to be filtered of wind noise;
[0122] The matrix calculation module 220 is used to calculate the estimated covariance matrix of the first frequency domain signal based on a sample moving average estimation method;
[0123] The multi-channel Wiener filtering module 230 is used to construct a multi-channel Wiener filter based on the estimated covariance matrix, and substitute the first frequency domain signal into the multi-channel Wiener filter to obtain a second frequency domain signal;
[0124] The wind noise fitting filter module 240 is configured to construct a wind noise fitting filter based on a preset wind noise fitting coefficient, and substitute the second frequency domain signal into the wind noise fitting filter to obtain a third frequency domain signal;
[0125] The harmonic suppression module 250 is used to convert the third frequency domain signal into a first time domain signal based on an inverse short-time Fourier transform, and to filter out wind noise from the first time domain signal according to a preset first harmonic suppression method, to obtain and output a speech signal after filtering out wind noise.
[0126] Furthermore, the wind noise filtering device applicable to multiple microphones further includes a modeling and decomposition module 260;
[0127] The modeling and decomposition module 260 is configured to model and decompose the first frequency domain signal based on a preset speech model with wind noise and a preset steering vector to obtain a first speech signal and a first wind noise signal.
[0128] Furthermore, the multi-channel Wiener filtering module 230 includes a power spectrum density calculation unit 231 and a multi-channel Wiener filtering unit 232;
[0129] The power spectrum density calculation unit 231 is configured to obtain a first speech power spectrum density and an estimated wind noise power spectrum density based on the estimated covariance matrix and the model covariance matrix of the first frequency domain signal based on a minimum Frobenius norm;
[0130] The multi-channel Wiener filter unit 232 is used to construct a multi-channel Wiener filter according to the first speech power spectrum density and the estimated wind noise power spectrum density, and substitute the first frequency domain signal into the multi-channel Wiener filter to obtain a second frequency domain signal.
[0131] Furthermore, the power spectrum density calculation unit 231 includes a norm solution construction subunit 2311 and a power spectrum density calculation subunit 2312;
[0132] The norm solution construction subunit 2311 is used to construct a first norm solution expression based on the minimum Frobenius norm according to the estimated covariance matrix and the model covariance matrix;
[0133] The power spectrum density calculation subunit 2312 is configured to convert the first norm solution expression into a solution expression for the first speech power spectrum density and a solution expression for the estimated wind noise power spectrum density based on preset parameters, and solve the solution expressions to obtain the first speech power spectrum density and the estimated wind noise power spectrum density;
[0134] The first norm solution expression is specifically:
[0135]
[0136] in, are the first speech power spectral density and the estimated wind noise power spectral density, are speech power spectral density and wind noise power spectral density respectively, are the estimated covariance matrix and model covariance matrix of the first frequency domain signal respectively, argmin is the function for finding the minimum value of the variable, is the square form of the Frobenius norm;
[0137] The solution expression for the first speech power spectrum density is specifically:
[0138]
[0139] The solution expression for estimating the wind noise power spectrum density is specifically:
[0140]
[0141] in, g, g H are the steering vector and its conjugate transpose, Γ, Γ H are the wind noise spatial coherence matrix and its conjugate transpose, tr(Γ H Γ) are Γ H Γ and traces.
[0142] Furthermore, the multi-channel Wiener filtering unit 232 includes a wind noise covariance matrix construction subunit 2321, an estimated noise power spectrum density calculation subunit 2322, a priori signal-to-noise ratio calculation subunit 2323 and a multi-channel Wiener filtering subunit 2324;
[0143] The wind noise covariance matrix construction subunit 2321 is configured to construct a wind noise covariance matrix according to a preset wind noise spatial coherence matrix and the estimated wind noise power spectrum density;
[0144] The estimated noise power spectrum density calculation subunit 2322 is used to calculate the estimated noise power spectrum density according to the steering vector and the wind noise covariance matrix;
[0145] The a priori signal-to-noise ratio calculation subunit 2323 is used to calculate the a priori signal-to-noise ratio according to the estimated noise power spectrum density and the first speech power spectrum density;
[0146] The multi-channel Wiener filter subunit 2324 is used to construct a multi-channel Wiener filter according to the steering vector, the wind noise covariance matrix and the priori signal-to-noise ratio, and substitute the first frequency domain signal into the multi-channel Wiener filter to obtain a second frequency domain signal.
[0147] Furthermore, the wind noise fitting filtering module 240 includes a low-frequency wind noise power spectrum density calculation unit 241 and a wind noise fitting filtering unit 242;
[0148] The low-frequency wind noise power spectrum density calculation unit 241 is used to calculate the low-frequency wind noise power spectrum density based on the wind noise fitting coefficient;
[0149] The wind noise fitting filter unit 242 is used to construct a wind noise fitting filter based on the second speech power spectrum density and the low-frequency wind noise power spectrum density, and substitute the second frequency domain signal into the wind noise fitting filter to obtain a third frequency domain signal; wherein, the second speech power spectrum density is obtained based on the second frequency domain signal.
[0150] Furthermore, the harmonic suppression module 250 includes an inverse short-time Fourier transform unit 251, a comb filter unit 252, a low-pass filter unit 253 and a signal output unit 254;
[0151] The inverse short-time Fourier transform unit 251 is used to convert the third frequency domain signal into the first time domain signal based on the inverse short-time Fourier transform;
[0152] The comb filter unit 252 is configured to construct a plurality of comb filters based on a preset fundamental frequency, and substitute the first time domain signal into the plurality of comb filters to obtain a second time domain signal; wherein the plurality of comb filters are connected end to end, and the output of the previous comb filter serves as the input of the next comb filter, the input of the first comb filter is the first time domain signal, and the output of the last comb filter is the second time domain signal;
[0153] The low-pass filtering unit 253 is configured to filter the second time domain signal based on a preset low-pass filter to obtain a third time domain signal;
[0154] The signal output unit 254 is configured to calculate a difference between the first time domain signal and the third time domain signal, and output the difference as a speech signal after wind noise is filtered out.
[0155] In this way, the first frequency domain signal is first obtained, and its estimated covariance matrix is calculated based on the sample moving average estimation method. Multi-channel Wiener filtering is performed based on the multi-channel Wiener filter constructed according to the estimated covariance matrix, and wind noise fitting filtering is performed based on the wind noise fitting filter constructed according to the preset wind noise fitting coefficient. The signal is then converted from the frequency domain to the time domain based on the inverse short-time Fourier transform, and then wind noise is filtered according to the first harmonic suppression method to obtain and output the voice signal after filtering out the wind noise. Performing wind noise filtering multiple times in this way can comprehensively utilize the different characteristics of wind noise to perform discriminative filtering on the input signal, better adapt to different wind noise environments, and thus more effectively filter wind noise, thereby improving call quality.
[0156] Correspondingly, the embodiment of the present invention also adaptively provides a terminal device and a computer-readable storage medium.
[0157] The terminal device includes: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor;
[0158] When the processor executes the computer program, the wind noise filtering method applicable to multiple microphones as described above is implemented.
[0159] The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the wind noise filtering method applicable to multiple microphones as described above.
[0160] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A wind noise filtering method applicable to multiple microphones, characterized in that: include: Acquire a first frequency domain signal to be filtered of wind noise; Calculating an estimated covariance matrix of the first frequency domain signal based on a sample moving average estimation method; constructing a multi-channel Wiener filter based on the estimated covariance matrix, and substituting the first frequency domain signal into the multi-channel Wiener filter to obtain a second frequency domain signal; constructing a wind noise fitting filter based on a preset wind noise fitting coefficient, and substituting the second frequency domain signal into the wind noise fitting filter to obtain a third frequency domain signal; The third frequency domain signal is converted into a first time domain signal based on inverse short-time Fourier transform, and the first time domain signal is subjected to wind noise filtering according to a preset first harmonic suppression method to obtain and output a speech signal after wind noise filtering.
2. The wind noise filtering method applicable to multiple microphones according to claim 1, characterized in that: After obtaining the first frequency domain signal to be filtered of wind noise, the method further includes: Based on a preset speech model with wind noise and a preset steering vector, the first frequency domain signal is modeled and decomposed to obtain a first speech signal and a first wind noise signal.
3. The wind noise filtering method applicable to multiple microphones according to claim 2, characterized in that: The step of constructing a multi-channel Wiener filter based on the estimated covariance matrix and substituting the first frequency domain signal into the multi-channel Wiener filter to obtain a second frequency domain signal specifically includes: Based on the minimum Frobenius norm, obtaining a first speech power spectral density and an estimated wind noise power spectral density according to the estimated covariance matrix and the model covariance matrix of the first frequency domain signal; A multi-channel Wiener filter is constructed according to the first speech power spectral density and the estimated wind noise power spectral density, and the first frequency domain signal is substituted into the multi-channel Wiener filter to obtain a second frequency domain signal.
4. The wind noise filtering method applicable to multiple microphones according to claim 3, characterized in that: The step of constructing a wind noise fitting filter based on a preset wind noise fitting coefficient and substituting the second frequency domain signal into the wind noise fitting filter to obtain a third frequency domain signal specifically includes: Calculating the low-frequency wind noise power spectral density based on the wind noise fitting coefficient; A wind noise fitting filter is constructed based on the second speech power spectral density and the low-frequency wind noise power spectral density, and the second frequency domain signal is substituted into the wind noise fitting filter to obtain a third frequency domain signal; wherein, the second speech power spectral density is obtained based on the second frequency domain signal.
5. The wind noise filtering method applicable to multiple microphones according to claim 1, characterized in that: The converting of the third frequency domain signal into a first time domain signal based on an inverse short-time Fourier transform, and filtering the first time domain signal for wind noise according to a preset first harmonic suppression method to obtain and output a speech signal after the wind noise is filtered out, specifically includes: Converting the third frequency domain signal into a first time domain signal based on an inverse short-time Fourier transform; Based on a preset fundamental frequency, construct a plurality of comb filters, and substitute the first time domain signal into the plurality of comb filters to obtain a second time domain signal; wherein the plurality of comb filters are connected end to end, the output of the previous comb filter serves as the input of the next comb filter, the input of the first comb filter is the first time domain signal, and the output of the last comb filter is the second time domain signal; Based on a preset low-pass filter, filtering the second time domain signal to obtain a third time domain signal; A difference between the first time domain signal and the third time domain signal is calculated, and the difference is output as a speech signal after filtering out wind noise.
6. A wind noise filtering device suitable for multiple microphones, characterized in that: It includes signal acquisition module, matrix calculation module, multi-channel Wiener filter module, wind noise fitting filter module and harmonic suppression module; The signal acquisition module is used to acquire a first frequency domain signal to be filtered out of wind noise; The matrix calculation module is used to calculate the estimated covariance matrix of the first frequency domain signal based on a sample moving average estimation method; The multi-channel Wiener filtering module is used to construct a multi-channel Wiener filter based on the estimated covariance matrix, and substitute the first frequency domain signal into the multi-channel Wiener filter to obtain a second frequency domain signal; The wind noise fitting filter module is configured to construct a wind noise fitting filter based on a preset wind noise fitting coefficient, and substitute the second frequency domain signal into the wind noise fitting filter to obtain a third frequency domain signal; The harmonic suppression module is used to convert the third frequency domain signal into a first time domain signal based on an inverse short-time Fourier transform, and to filter out wind noise from the first time domain signal according to a preset first harmonic suppression method, to obtain and output a speech signal after filtering out wind noise.
7. The wind noise filtering device applicable to multiple microphones according to claim 6, characterized in that: Also includes a modeling decomposition module; The modeling and decomposition module is used to model and decompose the first frequency domain signal based on a preset speech model with wind noise and a preset steering vector to obtain a first speech signal and a first wind noise signal.
8. The wind noise filtering device applicable to multiple microphones according to claim 7, characterized in that: The multi-channel Wiener filtering module includes a power spectrum density calculation unit and a multi-channel Wiener filtering unit; The power spectral density calculation unit is configured to obtain a first speech power spectral density and an estimated wind noise power spectral density based on the estimated covariance matrix and the model covariance matrix of the first frequency domain signal based on a minimum Frobenius norm; The multi-channel Wiener filter unit is used to construct a multi-channel Wiener filter according to the first speech power spectral density and the estimated wind noise power spectral density, and substitute the first frequency domain signal into the multi-channel Wiener filter to obtain a second frequency domain signal.
9. The wind noise filtering device applicable to multiple microphones according to claim 8, characterized in that: The wind noise fitting filtering module includes a low-frequency wind noise power spectrum density calculation unit and a wind noise fitting filtering unit; The low-frequency wind noise power spectrum density calculation unit is used to calculate the low-frequency wind noise power spectrum density based on the wind noise fitting coefficient; The wind noise fitting filter unit is used to construct a wind noise fitting filter based on the second speech power spectrum density and the low-frequency wind noise power spectrum density, and substitute the second frequency domain signal into the wind noise fitting filter to obtain a third frequency domain signal; wherein, the second speech power spectrum density is obtained based on the second frequency domain signal.
10. The wind noise filtering device applicable to multiple microphones according to claim 6, characterized in that: The harmonic suppression module includes an inverse short-time Fourier transform unit, a comb filter unit, a low-pass filter unit and a signal output unit; The inverse short-time Fourier transform unit is used to convert the third frequency domain signal into the first time domain signal based on the inverse short-time Fourier transform; The comb filter unit is configured to construct a plurality of comb filters based on a preset fundamental frequency, and substitute the first time domain signal into the plurality of comb filters to obtain a second time domain signal; wherein the plurality of comb filters are connected end to end, the output of the preceding comb filter serves as the input of the succeeding comb filter, the input of the first comb filter is the first time domain signal, and the output of the last comb filter is the second time domain signal; The low-pass filtering unit is configured to filter the second time domain signal based on a preset low-pass filter to obtain a third time domain signal; The signal output unit is configured to calculate a difference between the first time domain signal and the third time domain signal, and output the difference as a speech signal after wind noise is filtered out.
Citation Information
Patent Citations
Method for filtering sound noise
CN101442696A
Multi-channel speech enhancement method based on reference microphone optimization
CN113257270A