Method, apparatus, terminal device and storage medium for reducing wind noise
By performing multiple signal analysis on the audio signals collected by a single microphone, identifying wind noise categories and using target filters to process them, the problem of poor noise reduction effect of single microphone is solved, and more accurate wind noise recognition and noise reduction are achieved.
Patent Information
- Application Number
- CN202210327843.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-30
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-03-30
AI Technical Summary
In the prior art, it is difficult to effectively distinguish between human voice and wind noise in a single microphone, resulting in poor noise reduction effect and prone to missed or missed detection.
A variety of preset signal analysis algorithms are used to analyze the audio signal, and multiple wind noise identification identifiers are output. Based on these identifiers, the wind noise category of the audio signal is determined, and the corresponding target wind noise filter is used for noise reduction processing.
It improves the accuracy of audio signal classification recognition, effectively avoids missed or misdetection, and improves the noise reduction effect.
Smart Images

Figure CN114974287B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio signal processing, and in particular, to a method, device, terminal device and storage medium for reducing wind noise. Background Art
[0002] Wind noise exists in people's living and working scenarios, such as the sound of fans, air conditioners, and wind noise caused by walking. When a user makes a call, wind noise often affects the call quality, so it is necessary to reduce wind noise in the call audio.
[0003] Currently, the wind noise reduction scheme is mainly based on multiple microphones. By analyzing multiple audio signals collected by multiple microphones, the difference information between human voices and wind noise is analyzed, and the noise reduction gain of the multiple microphones is determined using this difference information. Finally, the input frequency-domain signal is noise-reduced according to the noise reduction gain. However, for a single microphone, the difference information between the human voice and wind noise that can be collected is less, and it is difficult to use this difference information to distinguish between the human voice and wind noise, which easily causes missed detection or false detection of wind noise. It can be seen that the current wind noise reduction scheme applied to a single microphone has the problem of poor noise reduction effect. Summary of the Invention
[0004] The present application provides a method, device, terminal device and storage medium for reducing wind noise to improve the wind noise recognition accuracy of a single microphone, thereby improving the noise reduction effect.
[0005] To solve the above technical problems, an embodiment of the present application provides a method for reducing wind noise, including:
[0006] Based on a variety of preset signal analysis algorithms, perform signal analysis on the audio signal collected by the microphone, and output multiple wind noise recognition identifiers. The wind noise recognition identifier is used to represent the wind noise recognition result corresponding to each preset signal analysis algorithm. The microphone includes a single microphone;
[0007] Based on multiple wind noise recognition identifiers, determine the wind noise category of the audio signal. The wind noise category includes pure wind noise and wind noise-containing human voice;
[0008] Based on the wind noise category of the audio signal, determine the target wind noise filter corresponding to the wind noise category;
[0009] Use the target wind noise filter to reduce the noise of the audio signal and output the target audio signal after noise reduction.
[0010] In this embodiment, by using a variety of preset signal analysis algorithms, the audio signal collected by the microphone is analyzed, and multiple wind noise recognition identifiers are output. The microphone includes a single microphone, and based on the multiple wind noise recognition identifiers, the wind noise category of the audio signal is determined, so as to classify and identify the wind noise category of the audio signal, which is convenient for subsequent targeted noise reduction processing of audio signals with different wind noise categories. At the same time, using a variety of preset signal analysis algorithms for signal analysis can effectively avoid missed detection or misdetection, improve the accuracy of audio signal classification and recognition, and then improve the noise reduction effect; and based on the wind noise category of the audio signal, a target wind noise filter corresponding to the wind noise category is determined, and the target wind noise filter is used to perform noise reduction on the audio signal, and the target audio signal after noise reduction is output, so as to be able to perform targeted noise reduction processing on audio signals with different wind noise categories and improve the noise reduction effect.
[0011] In one embodiment, the preset signal analysis algorithms include a low-frequency power ratio algorithm, a power spectrum difference algorithm, and an LPC analysis algorithm. Based on the multiple preset signal analysis algorithms, the audio signal collected by the microphone is analyzed, and multiple wind noise recognition identifiers are output, including:
[0012] Based on the low-frequency power ratio algorithm, perform low-frequency power ratio analysis on the audio signal, and output the first wind noise recognition identifier;
[0013] Based on the power spectrum difference algorithm, perform power spectrum analysis on the audio signal, and output the second wind noise recognition identifier;
[0014] Based on the LPC analysis algorithm, perform LPC analysis on the audio signal, and output the third wind noise recognition identifier.
[0015] In an alternative embodiment, based on the low-frequency power ratio algorithm, performing low-frequency power ratio analysis on the audio signal and outputting the first wind noise recognition identifier includes:
[0016] Based on the low-frequency power ratio algorithm, calculate the low-frequency energy and the full-band energy of the audio signal;
[0017] According to the low-frequency energy and the full-band energy, calculate the low-frequency energy power ratio of the audio signal;
[0018] According to the low-frequency energy power ratio, determine the first wind noise recognition identifier of the audio signal.
[0019] In an alternative embodiment, based on the power spectrum difference algorithm, performing power spectrum analysis on the audio signal and outputting the second wind noise recognition identifier includes:
[0020] Based on the power spectrum difference algorithm, determine the power spectrum of each frequency point of the audio signal within the preset frequency range;
[0021] Calculate the power spectrum difference of the audio signal within a preset frequency range according to the power spectra of each frequency point;
[0022] Determine the second wind noise recognition identifier of the audio signal according to the power spectrum difference.
[0023] In an alternative embodiment, perform LPC analysis on the audio signal based on the LPC analysis algorithm and output a third wind noise recognition identifier, including:
[0024] Determine the second-order LPC analysis resonance peak points of the audio signal based on the LPC analysis algorithm;
[0025] Input the second-order LPC analysis resonance peak points into a preset LPC analysis polynomial to obtain a polynomial value;
[0026] Determine the third wind noise recognition identifier of the audio signal according to the polynomial value.
[0027] In an embodiment, determine the wind noise category of the audio signal based on multiple wind noise recognition identifiers, including:
[0028] Identify the audio category of the audio signal based on multiple wind noise recognition identifiers, where the audio category includes pure human voice, pure wind noise, and human voice with wind noise;
[0029] If the audio category of the audio signal is not pure wind noise or human voice with wind noise, determine whether the wind noise recognition delay value is greater than a preset threshold, where the wind noise recognition delay value is used to characterize whether the wind noise category of the previous audio signal consecutive to the audio signal is pure wind noise or human voice with wind noise;
[0030] If the wind noise recognition delay value is greater than the preset threshold, determine that the wind noise category of the audio signal is pure wind noise or human voice with wind noise.
[0031] In an alternative embodiment, the wind noise recognition identifier includes a first wind noise recognition identifier analyzed based on the low-frequency power ratio algorithm, a second wind noise recognition identifier analyzed based on the power spectrum difference algorithm, and a third wind noise recognition identifier analyzed based on the LPC analysis algorithm. Identify the audio category of the audio signal based on multiple wind noise recognition identifiers, including:
[0032] If the first wind noise recognition identifier is not the first preset identifier, the second wind noise recognition identifier is not the second preset identifier, and the third wind noise recognition identifier is not the third preset identifier, then determine that the audio category of the audio signal is pure human voice;
[0033] If the first wind noise recognition identifier is the first preset identifier, the second wind noise recognition identifier is the second preset identifier, or the third wind noise recognition identifier is the third preset identifier, then determine that the audio category of the audio signal is human voice with wind noise;
[0034] If the first wind noise recognition identifier is the first preset identifier, the second wind noise recognition identifier is the second preset identifier, and the third wind noise recognition identifier is the third preset identifier, and the high-frequency energy ratio of the audio signal is less than the preset energy ratio, then it is determined that the audio category of the audio signal is pure wind noise.
[0035] In an alternative embodiment, after identifying the audio category of the audio signal based on multiple wind noise recognition identifiers, it further includes:
[0036] If the audio category of the audio signal is pure wind noise or wind noise with human voice, then add the wind noise recognition delay value to the preset value to obtain the latest wind noise recognition delay value.
[0037] In an alternative embodiment, if the audio category of the audio signal is not pure wind noise or wind noise with human voice, after determining whether the wind noise recognition delay value is greater than the preset threshold, it further includes:
[0038] If the wind noise recognition delay value is not greater than the preset threshold, then set the wind noise recognition delay value to the preset threshold to obtain the latest wind noise recognition delay value.
[0039] In an embodiment, based on the wind noise category of the audio signal, determining the target wind noise filter corresponding to the wind noise category includes:
[0040] If the wind noise category of the audio signal is pure wind noise, then determine that the target wind noise filter is a single wind noise filter, and the single wind noise filter is used to perform wind noise filtering on the low-frequency band of the audio signal;
[0041] If the wind noise category of the audio signal is wind noise with human voice, then determine that the target wind noise filter is a combined wind noise filter, and the combined wind noise filter is used to perform wind noise filtering on each frequency band of the audio signal with different attenuation degrees.
[0042] In an alternative embodiment, before determining that the target wind noise filter is a single wind noise filter if the wind noise category of the audio signal is pure wind noise, it further includes:
[0043] Perform Euler transformation on the Laplace domain of the preset LPC filter to obtain a single wind noise filter.
[0044] In a second aspect, an embodiment of the present application provides a wind noise reduction device, including:
[0045] An analysis module, configured to perform signal analysis on the audio signal collected by the microphone based on a variety of preset signal analysis algorithms, and output multiple wind noise recognition identifiers, where the wind noise recognition identifier is used to represent the wind noise recognition result corresponding to each preset signal analysis algorithm, and the microphone includes a single microphone;
[0046] A first determination module, configured to determine the wind noise category of the audio signal based on multiple wind noise recognition identifiers, where the wind noise category includes pure wind noise and wind noise with human voice;
[0047] A second determination module, configured to determine a target wind noise filter corresponding to the wind noise category based on the wind noise category of the audio signal;
[0048] A noise reduction module, configured to use the target wind noise filter to reduce the noise of the audio signal and output a target audio signal after noise reduction.
[0049] In a third aspect, an embodiment of the present application provides a terminal device, including a processor and a memory. The memory is used to store a computer program, and when the computer program is executed by the processor, the steps of the wind noise reduction method in the first aspect are implemented.
[0050] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the steps of the wind noise reduction method in the first aspect are implemented.
[0051] It should be noted that for the beneficial effects of the above second aspect to the fourth aspect, please refer to the relevant descriptions of the first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a schematic flowchart of a wind noise reduction method provided by an embodiment of the present application;
[0053] Figure 2 It is a schematic flowchart of a wind noise reduction method provided by another embodiment of the present application;
[0054] Figure 3 It is a schematic flowchart of a wind noise reduction method provided by still another embodiment of the present application;
[0055] Figure 4 It is a schematic flowchart of a wind noise reduction method provided by yet another embodiment of the present application;
[0056] Figure 5 It is a schematic structural diagram of a wind noise reduction device provided by an embodiment of the present application;
[0057] Figure 6 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0059] As described in the related art, for a single microphone, the difference information between the human voice and wind noise that can be collected is less, and it is difficult to use this difference information to distinguish between the human voice and wind noise, which easily causes missed detection or false detection of wind noise. It can be seen that the current wind noise reduction scheme applied to a single microphone has the problem of poor noise reduction effect.
[0060] Therefore, the embodiments of the present application provide a wind noise reduction method, device, terminal device and storage medium. By using a variety of preset signal analysis algorithms, signal analysis is performed on the audio signal collected by the microphone, and multiple wind noise recognition identifiers are output. The microphone includes a single microphone, and based on the multiple wind noise recognition identifiers, the wind noise category of the audio signal is determined, so as to classify and identify the wind noise category of the audio signal, which is convenient for subsequent targeted noise reduction processing of audio signals of different wind noise categories. At the same time, using a variety of preset signal analysis algorithms for signal analysis can effectively avoid missed detection or false detection, improve the accuracy of audio signal classification and recognition, and thus improve the noise reduction effect; and based on the wind noise category of the audio signal, a target wind noise filter corresponding to the wind noise category is determined, and the target wind noise filter is used to perform noise reduction on the audio signal, and the target audio signal after noise reduction is output, so as to be able to perform targeted noise reduction processing on audio signals of different wind noise categories and improve the noise reduction effect.
[0061] Please refer to Figure 1 , which is a schematic flowchart of a wind noise reduction method provided by the embodiments of the present application. The wind noise reduction method of the embodiments of the present application can be applied to a terminal device, and the terminal device includes but is not limited to devices such as smart phones, tablet computers, laptop computers and headphones with a built-in or external microphone. The microphone can be a single microphone or a multi-microphone. As Figure 1 shown, the wind noise reduction method of this embodiment includes steps S101 to S104, which are described in detail as follows:
[0062] Step S101, based on a variety of preset signal analysis algorithms, perform signal analysis on the audio signal collected by the microphone, and output multiple wind noise recognition identifiers. The wind noise recognition identifier is used to represent the wind noise recognition result corresponding to each preset signal analysis algorithm. The microphone includes a single microphone.
[0063] In this step, the audio signal is a continuous signal collected by the microphone. It can be understood that this embodiment can perform wind noise reduction for a single microphone and can also perform wind noise reduction for a multi-microphone.
[0064] The preset signal analysis algorithms include, but are not limited to, the low-frequency power ratio algorithm, the power spectrum difference algorithm, and the Linear Prediction Coefficients (LPC) analysis algorithm. The low-frequency power ratio algorithm is a method for identifying wind noise by using the proportion of low-frequency energy in the audio signal. The power spectrum difference algorithm is a method for identifying wind noise by using the different fluctuations of wind noise and human voices in a certain frequency range. The LPC analysis algorithm is a method for identifying wind noise by using the characteristic that wind noise has the same formants while human voices have different formants.
[0065] Optionally, based on the low-frequency power ratio algorithm, the power spectrum difference algorithm, and the LPC analysis algorithm, signal analysis is performed on the audio signal, and three corresponding wind noise identification identifiers are output. Exemplarily, based on the low-frequency power ratio algorithm, when signal analysis is performed on the audio signal, if wind noise is analyzed in the audio signal, the wind noise identification identifier wind_noise1 = 1 is output; if no wind noise is analyzed in the audio signal, the wind noise identification identifier wind_noise1 = 0 is output.
[0066] Step S102, based on the multiple wind noise identification identifiers, determine the wind noise category of the audio signal, where the wind noise category includes pure wind noise and wind noise-containing human voices.
[0067] In this step, the audio signal includes two cases: without wind noise and with wind noise. Among them, the wind noise categories when there is wind noise include pure wind noise and wind noise-containing human voices. Pure wind noise means the signal segment with only wind noise in the audio signal, and wind noise-containing human voices means the signal segment with both wind noise and human voices in the audio signal.
[0068] Optionally, through multiple wind noise identification identifiers, jointly determine whether there is wind noise in the audio signal. If there is wind noise, continue to determine the wind noise category of the audio signal according to the multiple wind noise identification identifiers.
[0069] Step S103, based on the wind noise category of the audio signal, determine the target wind noise filter corresponding to the wind noise category.
[0070] In this step, the target filter is a filter used to perform noise reduction filtering on the current audio signal. Since there is the user's voice in the wind noise-containing human voice, in order to ensure the quality of the user's voice after noise reduction, different filters are used to perform noise reduction on the audio signals of pure wind noise and wind noise-containing human voices respectively.
[0071] Optionally, the target wind noise filter corresponding to pure wind noise is a single wind noise filter, which mainly performs noise reduction filtering on a single frequency band with wind noise; the target wind noise filter corresponding to wind noise-containing human voices is a combined wind noise filter, which can perform noise reduction filtering with different attenuation degrees on multiple frequency bands.
[0072] Step S104: Use the target wind noise filter to reduce the noise of the audio signal and output the target audio signal after noise reduction.
[0073] In this step, the signal segment is denoised by the target wind noise filter corresponding to the wind noise category of a certain signal segment in the audio signal. After denoising all the wind noise signal segments of the audio signal with the corresponding target wind noise filters, operations such as inverse Fourier transform and shift addition are performed on the audio signal to output the target audio signal.
[0074] It can be understood that for signal segments without wind noise, that is, pure human voice signal segments, no noise reduction filtering is performed.
[0075] In one embodiment, Figure 2 The flowchart of the wind noise reduction method provided by another embodiment of the present application is shown. As Figure 2 shown, the above step S101 specifically includes steps S201 to S203. It can be understood that the Figure 1 same steps will not be elaborated here.
[0076] Step S201: Based on the low-frequency power ratio algorithm, perform low-frequency power ratio analysis on the audio signal and output the first wind noise identification flag.
[0077] In this step, since the wind noise energy is mainly concentrated in the low frequency (below 400 Hz), and the human voice energy is mainly evenly distributed in the mid-low frequency (below 800 Hz), the low-frequency energy ratio can be used to identify wind noise.
[0078] Optionally, based on the low-frequency power ratio algorithm, calculate the low-frequency energy and the full-band energy of the audio signal; according to the low-frequency energy and the full-band energy, calculate the low-frequency energy power ratio of the audio signal; according to the low-frequency energy power ratio, determine the first wind noise identification flag of the audio signal.
[0079] In this optional embodiment, calculate the low-frequency energy L_Energy (the sum of the squared spectral powers below 300 Hz) and the entire band energy T_Energy of the audio signal, then the low-frequency energy power ratio corresponding to the sub-frame can be obtained:
[0080] E r = L_Energy / T_Energy;
[0081] Compare E r with the preset low-frequency energy power ratio E the . If E r > E the , then the first wind noise identification flag wind_noise1 = 1. If E r < E the, then wind_noise1 = 0.
[0082] Step S202: Based on the power spectrum difference algorithm, perform power spectrum analysis on the audio signal and output a second wind noise recognition identifier.
[0083] In this step, since the wind noise spectrum has a small fluctuation within the frequency range of 2 kHz, and the human voice spectrum has a large fluctuation within the frequency range of 2 kHz, the sum of the spectrum differences within 2 kHz for each can be calculated to identify the wind noise.
[0084] Optionally, based on the power spectrum difference algorithm, determine the power spectrum of each frequency point of the audio signal within a preset frequency range; calculate the power spectrum difference of the audio signal within the preset frequency range according to the power spectra of each frequency point; determine the second wind noise recognition identifier of the audio signal according to the power spectrum difference.
[0085] In this optional embodiment, the power spectrum difference is the sum of the absolute differences between the frequency average power spectra in adjacent frequency units of the audio signal, specifically as follows:
[0086]
[0087] After simplification, we get:
[0088] Where
[0089] Φ n is the sum of the absolute differences of the average power spectra, n represents the number of frames of the currently processed audio signal, X n (l) represents the Fourier transform of the audio signal, k and l are the frequency points of the audio signal; Xbar n (k) represents the average value of the absolute value of the power spectrum at frequency point K, is the difference between the average absolute value of the power spectra of M frequency points between frequency point K and frequency point K - 1.
[0090] Since the power spectrum difference of the human voice is greater than that of the wind noise and is obvious within the preset frequency range (such as below 2 kHz), the above sum of the average absolute values of the power spectra can be accumulated within the first N frequency points to obtain Φ n . If Φ n is greater than the preset power spectrum difference θ Φ , it is determined that this frame is speech, that is, the second wind noise recognition identifier wind_noise2 = 0, otherwise it is wind noise, that is, wind_noise2 = 1.
[0091] Step S203: Based on the LPC analysis algorithm, perform LPC analysis on the audio signal and output a third wind noise recognition flag.
[0092] In this step, since the resonance point positions of the LPC analysis of different orders of wind noise are roughly the same, but the resonance point positions of the LPC analysis of different orders of human voices vary greatly. That is, wind noise under LPC analysis of different orders has the same resonance peak, while human voices have different resonance peaks. Therefore, this characteristic can be used to identify wind noise.
[0093] Optionally, based on the LPC analysis algorithm, determine the second-order LPC analysis resonance peak points of the audio signal; input the second-order LPC analysis resonance peak points into a preset LPC analysis polynomial to obtain a polynomial value; according to the polynomial value, determine the third wind noise recognition flag of the audio signal.
[0094] In this optional embodiment, the LPC analysis algorithm is a linear prediction technology for speech signals, which uses existing old signals to predict new signals through polynomial fitting. The second-order LPC analysis expression is A(z) = 1 + a1z -1 + a2z -2 , Z = e^(-jw), this polynomial is the LPC expression in the Laplace transform domain, w is the angular frequency, and j is the imaginary number; a is the LPC analysis coefficient, and the solution obtained when this polynomial is zero, that is, the resonance peak frequency point z0:
[0095]
[0096] Import the second-order resonance peak point z0 into LPC analysis polynomials of different orders. If it is a wind noise signal, this polynomial is close to zero, that is, the third wind noise recognition flag wind_noise3 = 1; if it is not close to zero, then wind_noise3 = 0.
[0097] In one embodiment, Figure 3 shows a schematic flow chart of a wind noise reduction method provided by another embodiment of the present application. As Figure 3 shown, the above step S102 specifically includes steps S301 to S303. It can be understood that the same steps Figure 1 will not be described herein again.
[0098] Step S301: Based on multiple wind noise recognition flags, identify the audio category of the audio signal, and the audio category includes pure human voice, pure wind noise, and wind noise-containing human voice.
[0099] In this step, the wind noise identification identifier includes a first wind noise identification identifier analyzed based on the low-frequency power ratio algorithm, a second wind noise identification identifier analyzed based on the power spectrum difference algorithm, and a third wind noise identification identifier analyzed based on the LPC analysis algorithm.
[0100] Optionally, if the first wind noise identification identifier is not the first preset identifier, the second wind noise identification identifier is not the second preset identifier, and the third wind noise identification identifier is not the third preset identifier, it is determined that the audio category of the audio signal is pure human voice; if the first wind noise identification identifier is the first preset identifier, the second wind noise identification identifier is the second preset identifier, or the third wind noise identification identifier is the third preset identifier, it is determined that the audio category of the audio signal is human voice with wind noise; if the first wind noise identification identifier is the first preset identifier, the second wind noise identification identifier is the second preset identifier, and the third wind noise identification identifier is the third preset identifier, and the high-frequency energy ratio of the audio signal is less than the preset energy ratio, it is determined that the audio category of the audio signal is pure wind noise.
[0101] In this alternative embodiment, if wind_noise1, wind_noise2, or wind_noise3 are all 0, it is determined that the audio category is pure human voice; if wind_noise1, wind_noise2, or wind_noise3 are all 1, it is determined that the audio category is human voice with wind noise; if wind_noise1, wind_noise2, and wind_noise3 are all 1, and the high-frequency energy ratio of the audio signal is less than the preset energy ratio, it is determined that the audio category is pure wind noise.
[0102] Step S302, if the audio category of the audio signal is not pure wind noise or human voice with wind noise, determine whether the wind noise identification delay value is greater than a preset threshold, where the wind noise identification delay value is used to characterize whether the wind noise category of the previous audio signal consecutive to the audio signal is pure wind noise or human voice with wind noise.
[0103] In this step, in order to avoid misprocessing pure human voice and retain the voice quality of human voice with wind noise, an identification delay logic is introduced to make up for the missed detection of wind noise. Since the wind noise that can affect the call quality must be continuous, the obvious wind starting point is detected, and then the identification delay logic is used to identify continuous wind noise. At the same time, the wind noise identification delay values (flag1 and flag2) change dynamically, which can timely avoid the misdetection of non-wind noise segments. Among them, flag1 is the wind noise identification delay value of human voice with wind noise, and it is considered to be human voice with wind noise during the stage when its value is not zero. flag2 is the wind noise identification delay value of pure wind noise, and it is considered to contain pure wind noise during the stage when its value is not zero.
[0104] Optionally, if the audio category of the audio signal is pure wind noise or wind-noise-containing human voice, add the wind noise recognition delay value to a preset value to obtain the latest wind noise recognition delay value, which is used for the wind noise delay determination of the next audio signal consecutive to the current audio signal. Among them, for pure wind noise, flag2 + flag2add = flag1, where flag2add is a second preset value; for wind-noise-containing human voice, flag1 + flag1add = flag1, where flag1add is a first preset value.
[0105] Step S303: If the wind noise recognition delay value is greater than a preset threshold, determine that the wind noise category of the audio signal is pure wind noise or wind-noise-containing human voice.
[0106] In this step, the preset threshold can be 0. For example, for the previous audio signal consecutive to the current audio signal with a wind noise category of pure wind noise, determine whether flag2 is greater than 0. If flag2 > 0, determine that the wind noise category of the current audio signal is pure wind noise, and at the same time, flag2 = flag2 - flag2add to update flag2. For the previous audio signal consecutive to the current audio signal with a wind noise category of wind-noise-containing human voice, determine whether flag1 is greater than 0. If flag1 > 0, determine that the wind noise category of the current audio signal is wind-noise-containing human voice, and at the same time, flag1 = flag1 - flag1add to update flag1.
[0107] Optionally, if the wind noise recognition delay value is not greater than the preset threshold, set the wind noise recognition delay value to the preset threshold to obtain the latest wind noise recognition delay value.
[0108] In this optional embodiment, for the previous audio signal consecutive to the current audio signal with a wind noise category of pure wind noise, determine whether flag2 is greater than 0. If flag2 ≤ 0, determine that there is no wind noise in the current audio signal, and at the same time, flag2 = 0 to update flag2. For the previous audio signal consecutive to the current audio signal with a wind noise category of wind-noise-containing human voice, determine whether flag1 is greater than 0. If flag1 ≤ 0, determine that there is no wind noise in the current audio signal, and at the same time, flag1 = 0 to update flag1.
[0109] In one embodiment, Figure 4 shows a schematic flowchart of a wind noise reduction method provided in another embodiment of the present application. As Figure 4 shown, the above step S103 specifically includes step S401 and step S402. It can be understood that the Figure 1 same steps will not be elaborated here.
[0110] Step S401: If the wind noise category of the audio signal is pure wind noise, determine that the target wind noise filter is a single wind noise filter, and the single wind noise filter is used to perform wind noise filtering on the low-frequency band of the audio signal.
[0111] In this step, since the filtering degree of the traditional LPC filter for wind noise is not easy to control, the wind noise cannot be completely filtered from the human voice. At the same time, because the filtering is inaccurate, it will affect the quality of the human voice. Therefore, the second-order LPC filter is simplified and improved to focus on processing low-frequency wind noise.
[0112] Optionally, perform an Euler transform on the Laplace domain of the preset LPC filter to obtain the single wind noise filter.
[0113] where the preset LPC filter is K represents the order of LPC analysis. In this embodiment, K = 2, Z is the Laplace domain, Z = e -jw , w is the angular frequency, and j is the imaginary number; perform an Euler transform on the Laplace domain, that is, e jw = cos(x) + j·sin(x). Through trigonometric simplification and operations such as only taking the real part first, the single wind noise filter is obtained as:
[0114]
[0115] L = 1 + a1×α×cos(w) + a2×α 2 ×(2cos(w)cos(w) - 1);
[0116] M = a1×α×sin(w) + a2×α 2 ×(2sin(w)cos(w));
[0117] N = 1 + a1×β×cos(w) + a2×β 2 ×(2cos(w)cos(w) - 1);
[0118] R = a1×β×sin(w) + a2×β 2 ×(2sin(w)cos(w));
[0119] α = 0.3 + w1×(0.6 - 0.3);
[0120] β = 0.6 + w2×(0.9 - 0.6);
[0121] where w1, w2, and μ are attenuation degree coefficients.
[0122] By optimizing the LPC filter, the computational complexity and memory occupancy of the single wind noise filter are reduced.
[0123] Step S402, if the wind noise category of the audio signal is wind noise-containing human voice, determine that the target wind noise filter is a combined wind noise filter, and the combined wind noise filter is used to perform wind noise filtering with different attenuation degrees on each frequency band of the audio signal.
[0124] In this step, in order to ensure the quality of high-frequency human voice and filter out low-frequency wind noise, a combined filter is designed using an improved LPC filter to perform noise reduction filtering with different attenuation degrees in different frequency bands, that is, perform frequency-band attenuation using a combined wind noise filter. Optionally, using the above single wind noise filter, modify the attenuation degree coefficient to reconstruct the corresponding combined wind noise filter, thereby reducing the computational amount and memory occupancy, and paying more attention to filtering out low-frequency wind noise with less attenuation of high-frequency energy, which can ensure the quality of human voice.
[0125] In order to execute the wind noise reduction method corresponding to the above method embodiment to achieve the corresponding functions and technical effects. Refer to Figure 5 , Figure 5 which shows a structural block diagram of a wind noise reduction device provided by an embodiment of the present application. For the convenience of description, only the parts related to this embodiment are shown. The wind noise reduction device provided by the embodiment of the present application includes:
[0126] An analysis module 501, configured to perform signal analysis on an audio signal collected by a microphone based on a plurality of preset signal analysis algorithms, and output a plurality of wind noise recognition identifiers, where the wind noise recognition identifiers are used to represent the wind noise recognition results corresponding to each of the preset signal analysis algorithms, and the microphone includes a single microphone;
[0127] A first determination module 502, configured to determine the wind noise category of the audio signal based on the plurality of wind noise recognition identifiers, where the wind noise category includes pure wind noise and wind noise-containing human voice;
[0128] A second determination module 503, configured to determine a target wind noise filter corresponding to the wind noise category based on the wind noise category of the audio signal;
[0129] A noise reduction module 504, configured to use the target wind noise filter to perform noise reduction on the audio signal and output a target audio signal after noise reduction.
[0130] In an embodiment, the analysis module 501 includes:
[0131] A first analysis unit, configured to perform low-frequency power ratio analysis on the audio signal based on the low-frequency power ratio algorithm and output a first wind noise recognition identifier;
[0132] A second analysis unit, configured to perform power spectrum analysis on the audio signal based on the power spectrum difference algorithm and output a second wind noise recognition identifier;
[0133] A third analysis unit, configured to perform LPC analysis on the audio signal based on the LPC analysis algorithm, and output a third wind noise recognition identifier.
[0134] In an optional embodiment, the first analysis unit includes:
[0135] A first calculation subunit, configured to calculate the low-frequency energy and the full-band energy of the audio signal based on the low-frequency power ratio algorithm;
[0136] A second calculation subunit, configured to calculate the low-frequency energy power ratio of the audio signal according to the low-frequency energy and the full-band energy;
[0137] A first determination subunit, configured to determine a first wind noise recognition identifier of the audio signal according to the low-frequency energy power ratio.
[0138] In an optional embodiment, the second analysis unit includes:
[0139] A second determination subunit, configured to determine the power spectrum of each frequency point of the audio signal within a preset frequency range based on the power spectrum difference algorithm;
[0140] A third calculation subunit, configured to calculate the power spectrum difference value of the audio signal within the preset frequency range according to the power spectrum of each frequency point;
[0141] A third determination subunit, configured to determine a second wind noise recognition identifier of the audio signal according to the power spectrum difference value.
[0142] In an optional embodiment, the third analysis unit includes:
[0143] A fourth determination subunit, configured to determine the second-order LPC analysis resonance peak points of the audio signal based on the LPC analysis algorithm;
[0144] An input subunit, configured to input the second-order LPC analysis resonance peak points into a preset LPC analysis polynomial to obtain a polynomial value;
[0145] A fifth determination subunit, configured to determine a third wind noise recognition identifier of the audio signal according to the polynomial value.
[0146] In an embodiment, the first determination module 502 includes:
[0147] A recognition unit, configured to recognize the audio category of the audio signal based on multiple wind noise recognition identifiers, where the audio category includes pure human voice, pure wind noise, and human voice with wind noise;
[0148] A first determination unit, configured to determine whether a wind noise recognition delay value is greater than a preset threshold if an audio category of the audio signal is not pure wind noise or wind noise-containing human voice, where the wind noise recognition delay value is used to represent whether a wind noise category of a previous audio signal consecutive to the audio signal is pure wind noise or wind noise-containing human voice;
[0149] A determination unit, configured to determine that the wind noise category of the audio signal is pure wind noise or wind noise-containing human voice if the wind noise recognition delay value is greater than the preset threshold.
[0150] In an alternative embodiment, the wind noise recognition identifier includes a first wind noise recognition identifier analyzed based on the low-frequency power ratio algorithm, a second wind noise recognition identifier analyzed based on the power spectrum difference algorithm, and a third wind noise recognition identifier analyzed based on the LPC analysis algorithm. The recognition unit includes:
[0151] A first determination subunit, configured to determine that the audio category of the audio signal is pure human voice if the first wind noise recognition identifier is not a first preset identifier, the second wind noise recognition identifier is not a second preset identifier, and the third wind noise recognition identifier is not a third preset identifier;
[0152] A second determination subunit, configured to determine that the audio category of the audio signal is wind noise-containing human voice if the first wind noise recognition identifier is the first preset identifier, the second wind noise recognition identifier is the second preset identifier, or the third wind noise recognition identifier is the third preset identifier;
[0153] A third determination subunit, configured to determine that the audio category of the audio signal is pure wind noise if the first wind noise recognition identifier is the first preset identifier, the second wind noise recognition identifier is the second preset identifier, and the third wind noise recognition identifier is the third preset identifier, and a high-frequency energy ratio of the audio signal is less than a preset energy ratio.
[0154] In an alternative embodiment, the first determination module 502 further includes:
[0155] An addition unit, configured to add the wind noise recognition delay value to a preset value to obtain an updated wind noise recognition delay value if the audio category of the audio signal is pure wind noise or wind noise-containing human voice.
[0156] In an alternative embodiment, the first determination module 502 further includes:
[0157] A setting unit, configured to set the wind noise recognition delay value to the preset threshold to obtain an updated wind noise recognition delay value if the wind noise recognition delay value is not greater than the preset threshold.
[0158] In an embodiment, a second determination module 503 includes:
[0159] A second determination unit, configured to determine that the target wind noise filter is a single wind noise filter if the wind noise category of the audio signal is pure wind noise, where the single wind noise filter is used to perform wind noise filtering on the low frequency band of the audio signal;
[0160] A third determination unit, configured to determine that the target wind noise filter is a combined wind noise filter if the wind noise category of the audio signal is wind noise-containing human voice, where the combined wind noise filter is used to perform wind noise filtering on each frequency band of the audio signal with different attenuation degrees.
[0161] In an optional embodiment, the second determination module 503 further includes:
[0162] A transformation unit, configured to perform an Euler transformation on the Laplace domain of a preset LPC filter to obtain the single wind noise filter.
[0163] The above wind noise reduction device can implement the wind noise reduction method in the above method embodiment. The optional items in the above method embodiment are also applicable to this embodiment, which will not be elaborated here. The remaining content of the embodiment of the present application can refer to the content of the above method embodiment and will not be repeated in this embodiment.
[0164] Figure 6 It is a schematic structural diagram of a terminal device provided in an embodiment of the present application. As Figure 6 shown, the terminal device 6 in this embodiment includes: at least one processor 60 ( Figure 6 only one is shown in the figure), a processor, a memory 61, and a computer program 62 stored in the memory 61 and executable on the at least one processor 60. When the processor 60 executes the computer program 62, the steps in any of the above method embodiments are implemented.
[0165] The terminal device 6 may be a device such as a smart phone, a tablet computer, a notebook computer, and a headset. The terminal device may include but is not limited to the processor 60 and the memory 61. Those skilled in the art can understand that Figure 6 merely examples of the terminal device 6, which do not constitute a limitation on the terminal device 6, and may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0166] The so-called processor 60 may be a Central Processing Unit (CPU), and the processor 60 may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0167] In some embodiments, the memory 61 may be an internal storage unit of the terminal device 6, such as the hard disk or memory of the terminal device 6. In other embodiments, the memory 61 may also be an external storage device of the terminal device 6, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the terminal device 6. Further, the memory 61 may also include both the internal storage unit and the external storage device of the terminal device 6. The memory 61 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of the computer program, etc. The memory 61 may also be used to temporarily store data that has been output or will be output.
[0168] In addition, an embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.
[0169] An embodiment of the present application provides a computer program product, and when the computer program product runs on a terminal device, the terminal device is caused to execute the steps in each of the above method embodiments when executed.
[0170] In several embodiments provided in this application, it can be understood that each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, the program segment, or the part of code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the block may occur in a different order from that marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved.
[0171] If the above functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a terminal device to execute all or part of the steps of the methods described in the various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs, Read-Only Memories), random access memories (RAMs, Random Access Memories), magnetic disks, or optical discs that can store program codes.
[0172] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of this application. It should be understood that the above description is only for the specific embodiments of this application and is not used to limit the protection scope of this application. In particular, it is pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application shall be included in the protection scope of this application.
Claims
1. A method for reducing wind noise, characterized in that, Including: Based on a variety of preset signal analysis algorithms, perform signal analysis on the audio signal collected by the microphone, and output multiple wind noise recognition identifiers. The wind noise recognition identifiers are used to represent the wind noise recognition results corresponding to each of the preset signal analysis algorithms. The microphone includes a single microphone. Based on the multiple wind noise recognition identifiers, determine the wind noise category of the audio signal. The wind noise category includes pure wind noise and wind noise-containing human voice. Among them, the determining the wind noise category of the audio signal based on the multiple wind noise recognition identifiers includes: based on the multiple wind noise recognition identifiers, identify the audio category of the audio signal. The audio category includes pure human voice, pure wind noise, and wind noise-containing human voice. If the audio category of the audio signal is not pure wind noise or wind noise-containing human voice, then determine whether the wind noise recognition delay value is greater than a preset threshold. The wind noise recognition delay value is used to characterize whether the wind noise category of the previous audio signal consecutive to the audio signal is pure wind noise or wind noise-containing human voice. If the wind noise recognition delay value is greater than the preset threshold, then determine that the wind noise category of the audio signal is the same wind noise category as the previous audio signal. Based on the wind noise category of the audio signal, determine a target wind noise filter corresponding to the wind noise category. Use the target wind noise filter to perform noise reduction on the audio signal and output a target audio signal after noise reduction.
2. The wind noise reduction method according to claim 1, wherein The preset signal analysis algorithms include a low-frequency power ratio algorithm, a power spectrum difference algorithm, and an LPC analysis algorithm. The performing signal analysis on the audio signal collected by the microphone based on a variety of preset signal analysis algorithms and outputting multiple wind noise recognition identifiers includes: Based on the low-frequency power ratio algorithm, perform low-frequency power ratio analysis on the audio signal and output a first wind noise recognition identifier. Based on the power spectrum difference algorithm, perform power spectrum analysis on the audio signal and output a second wind noise recognition identifier. Based on the LPC analysis algorithm, perform LPC analysis on the audio signal and output a third wind noise recognition identifier.
3. The wind noise reduction method according to claim 2, wherein The performing low-frequency power ratio analysis on the audio signal based on the low-frequency power ratio algorithm and outputting a first wind noise recognition identifier includes: Based on the low-frequency power ratio algorithm, calculate the low-frequency energy and the full-band energy of the audio signal. According to the low-frequency energy and the full-band energy, calculate the low-frequency energy power ratio of the audio signal. According to the low-frequency energy power ratio, determine the first wind noise recognition identifier of the audio signal.
4. The wind noise reduction method according to claim 2, wherein The performing power spectrum analysis on the audio signal based on the power spectrum difference algorithm and outputting a second wind noise recognition identifier includes: Based on the power spectrum difference algorithm, determine the power spectrum of each frequency point of the audio signal within a preset frequency range. According to the power spectra of each frequency point, calculate the power spectrum difference of the audio signal within the preset frequency range. According to the power spectrum difference, determine the second wind noise recognition identifier of the audio signal.
5. The wind noise reduction method according to claim 2, wherein, The performing LPC analysis on the audio signal based on the LPC analysis algorithm and outputting a third wind noise recognition identifier includes: Based on the LPC analysis algorithm, determine the second-order LPC analysis resonance peak points of the audio signal. Input the resonance peak points of the second-order LPC analysis into a preset LPC analysis polynomial to obtain a polynomial value; Determine a third wind noise recognition identifier of the audio signal according to the polynomial value.
6. The wind noise reduction method according to claim 1, characterized in that, The wind noise recognition identifier includes a first wind noise recognition identifier obtained by analyzing based on a low-frequency power ratio algorithm, a second wind noise recognition identifier obtained by analyzing based on a power spectrum difference algorithm, and a third wind noise recognition identifier obtained by analyzing based on an LPC analysis algorithm. Based on the multiple wind noise recognition identifiers, identifying the audio category of the audio signal includes: If the first wind noise recognition identifier is not a first preset identifier, the second wind noise recognition identifier is not a second preset identifier, and the third wind noise recognition identifier is not a third preset identifier, then determine that the audio category of the audio signal is pure human voice; If the first wind noise recognition identifier is the first preset identifier, the second wind noise recognition identifier is the second preset identifier, or the third wind noise recognition identifier is the third preset identifier, then determine that the audio category of the audio signal is human voice with wind noise; If the first wind noise recognition identifier is the first preset identifier, the second wind noise recognition identifier is the second preset identifier, and the third wind noise recognition identifier is the third preset identifier, and the high-frequency energy ratio of the audio signal is less than a preset energy ratio, then determine that the audio category of the audio signal is pure wind noise.
7. The method for reducing wind noise according to claim 1, wherein After identifying the audio category of the audio signal based on the multiple wind noise recognition identifiers, it further includes: If the audio category of the audio signal is pure wind noise or human voice with wind noise, then add the wind noise recognition delay value to a preset value to obtain an updated wind noise recognition delay value.
8. The wind noise reduction method according to claim 1, wherein, After determining whether the wind noise recognition delay value is greater than a preset threshold if the audio category of the audio signal is not pure wind noise or human voice with wind noise, it further includes: If the wind noise recognition delay value is not greater than the preset threshold, then set the wind noise recognition delay value to the preset threshold to obtain an updated wind noise recognition delay value.
9. The wind noise reduction method according to claim 1, wherein Determining a target wind noise filter corresponding to the wind noise category based on the wind noise category of the audio signal includes: If the wind noise category of the audio signal is pure wind noise, then determine that the target wind noise filter is a single wind noise filter, and the single wind noise filter is used to perform wind noise filtering on the low-frequency band of the audio signal; If the wind noise category of the audio signal is human voice with wind noise, then determine that the target wind noise filter is a combined wind noise filter, and the combined wind noise filter is used to perform wind noise filtering on each frequency band of the audio signal with different attenuation degrees.
10. The wind noise reduction method according to claim 9, wherein Before determining that the target wind noise filter is a single wind noise filter if the wind noise category of the audio signal is pure wind noise, it further includes: Perform an Euler transform on the Laplace domain of a preset LPC filter to obtain the single wind noise filter.
11. A wind noise reduction device, characterized in that, Including: An analysis module, configured to perform signal analysis on an audio signal collected by a microphone based on multiple preset signal analysis algorithms, and output multiple wind noise recognition identifiers. The wind noise recognition identifier is used to represent the wind noise recognition result corresponding to each preset signal analysis algorithm. The microphone includes a single microphone; A first determination module, configured to determine a wind noise category of the audio signal based on a plurality of the wind noise recognition identifiers, where the wind noise category includes pure wind noise and wind noise-containing human voice; wherein, determining the wind noise category of the audio signal based on a plurality of the wind noise recognition identifiers includes: recognizing an audio category of the audio signal based on a plurality of the wind noise recognition identifiers, where the audio category includes pure human voice, pure wind noise, and wind noise-containing human voice; if the audio category of the audio signal is not pure wind noise or wind noise-containing human voice, then determine whether a wind noise recognition delay value is greater than a preset threshold, where the wind noise recognition delay value is used to characterize whether the wind noise category of the previous audio signal consecutive to the audio signal is pure wind noise or wind noise-containing human voice; if the wind noise recognition delay value is greater than the preset threshold, then determine that the wind noise category of the audio signal is the same wind noise category as the previous audio signal. A second determination module, configured to determine a target wind noise filter corresponding to the wind noise category based on the wind noise category of the audio signal. A noise reduction module, configured to perform noise reduction on the audio signal by using the target wind noise filter and output a target audio signal after noise reduction.
12. A terminal device, characterized in that, It includes a processor and a memory, where the memory is used to store a computer program, and when the computer program is executed by the processor, the steps of the wind noise reduction method according to any one of claims 1 to 10 are implemented.
13. A computer-readable storage medium, characterized in that, It stores a computer program, and when the computer program is executed by a processor, the steps of the wind noise reduction method according to any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
Single-microphone wind noise suppression
US20100223054A1
Method and apparatus for masking wind noise
US20120191447A1