Wind noise processing method and device, storage medium, electronic equipment and chip

By selecting the channel with the least wind noise for wind noise suppression processing, combined with wind noise detection and adaptive echo cancellation, the stability problem of the wind noise suppression algorithm is solved, the voice signal-to-noise ratio and clarity are improved, and it is suitable for a variety of terminal devices.

CN120833796AActive Publication Date: 2025-10-24BEIJING X RING TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511318761.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-10-24
Estimated Expiration
2045-09-15

AI Technical Summary

Technical Problem

Existing technologies cannot effectively reduce the pressure on wind noise suppression algorithms, affecting the stability of the results after wind noise suppression. In particular, the impact of wind noise is significant in outdoor voice communications and intelligent voice recognition.

Method used

By acquiring channel sound signals collected by multiple microphones, the first channel whose wind noise pollution meets the preset conditions is selected, and wind noise suppression processing is performed based on this channel. Combined with wind noise detection and adaptive echo cancellation processing, the execution order of the wind noise suppression algorithm is optimized.

Benefits of technology

It effectively reduces the pressure on the wind noise suppression algorithm, ensures the stability of the wind noise suppression results, improves the signal-to-noise ratio and voice clarity, and is suitable for a variety of terminal devices such as mobile phones, tablets, wearable devices, and sports cameras.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833796A_ABST
    Figure CN120833796A_ABST
Patent Text Reader

Abstract

The invention provides a wind noise processing method and device, a storage medium, electronic equipment and a chip, and relates to the technical field of signal processing.The method comprises the steps that firstly, sound signals, collected by multiple microphones, of multiple channels are obtained; selecting a first channel with wind noise pollution meeting a preset condition from the plurality of channels, such as a channel with relatively low wind noise, under the condition that the existence of wind noise is determined according to the sound signals of the plurality of channels; and then performing wind noise suppression processing based on the sound signal of the first channel, for example, performing wind noise suppression processing based on the sound signal of the channel with smaller wind noise, compared with the related technology, the method can effectively reduce the pressure of a wind noise suppression algorithm, can ensure the stability of a result after wind noise suppression, and can improve the user experience when collecting user voice information through a plurality of microphones. By applying the wind noise processing scheme provided by the invention, the target voice with higher signal-to-noise ratio, definition and intelligibility can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of signal processing, and particularly relates to a wind noise processing method and device, a storage medium, an electronic device and a chip. BACKGROUND

[0002] With the evolution and development of technology, mobile communication and wearable devices gradually become popular. When making a voice communication, recording or interacting with a smart voice assistant outdoors, it is inevitable to encounter wind noise. Wind noise not only seriously affects the call quality and speech intelligibility, but also significantly reduces the recognition rate of the smart voice. SUMMARY The present disclosure provides a wind noise processing method and device, a storage medium, an electronic device and a chip, and mainly aims to solve the technical problem that related technologies cannot effectively reduce the pressure of wind noise suppression algorithm, affecting the stability of the result after wind noise suppression.

[0003] According to a first aspect of an embodiment of the present disclosure, a wind noise processing method is provided, comprising: obtaining a plurality of channel sound signals collected by a plurality of microphones; in a case where it is determined that wind noise exists according to the plurality of channel sound signals, selecting a first channel with wind noise pollution meeting a preset condition from the plurality of channels; performing wind noise suppression processing based on the sound signal of the first channel.

[0004] Optionally, the step of selecting the first channel with wind noise pollution meeting the preset condition from the plurality of channels comprises: selecting a channel with minimum wind noise from the plurality of channels as the first channel.

[0005] Optionally, the first channel is a channel with wind noise pollution meeting the preset condition selected based on a current frame; The step of performing wind noise suppression processing based on the sound signal of the first channel comprises: if a second channel with wind noise pollution meeting the preset condition selected based on a historical frame is the same as the first channel, performing wind noise suppression processing based on the sound signal of the first channel; if the second channel is different from the first channel, determining an amplitude difference between the first channel and the second channel based on a time domain signal of a current frame, and selecting a target channel for wind noise suppression processing from the first channel and the second channel according to the amplitude difference.

[0006] Optionally, the step of selecting the target channel for wind noise suppression processing from the first channel and the second channel according to the amplitude difference comprises: in a case where the amplitude difference is less than or equal to a preset amplitude difference threshold, performing wind noise suppression processing on the sound signal of the second channel; in a case where the amplitude difference is greater than the preset amplitude difference threshold, performing wind noise suppression processing on the sound signal of the first channel.

[0007] Optionally, the method further comprises: performing wind noise detection according to the sound signals of the multiple channels; performing adaptive echo cancellation processing on the sound signals after wind noise detection.

[0008] Optionally, performing wind noise detection according to the sound signals of the multiple channels comprises: obtaining a short-time wind noise existence probability of a current frame and a long-time wind noise existence probability of a previous frame according to the sound signals of the multiple channels; performing smoothing processing to obtain a long-time wind noise existence probability of the current frame according to the short-time wind noise existence probability of the current frame and the long-time wind noise existence probability of the previous frame; determining whether wind noise exists based on the long-time wind noise existence probability of the current frame.

[0009] Optionally, performing smoothing processing to obtain a long-time wind noise existence probability of the current frame according to the short-time wind noise existence probability of the current frame and the long-time wind noise existence probability of the previous frame comprises: performing smoothing processing to obtain a long-time wind noise existence probability of the current frame according to the short-time wind noise existence probability of the current frame and the long-time wind noise existence probability of the previous frame by using a smoothing coefficient; wherein the smoothing coefficient is non-fixed.

[0010] Optionally, the method further comprises: determining the smoothing coefficient according to the long-time wind noise existence probability of the previous frame and the short-time wind noise existence probability of the current frame.

[0011] Optionally, determining the smoothing coefficient according to the long-time wind noise existence probability of the previous frame and the short-time wind noise existence probability of the current frame comprises: if the long-time wind noise existence probability of the previous frame is greater than a first probability threshold and the short-time wind noise existence probability of the current frame is less than the first probability threshold, determining the smoothing coefficient as a first coefficient value; if the long-time wind noise existence probability of the previous frame is less than a second probability threshold and the short-time wind noise existence probability of the current frame is greater than the first probability threshold, determining the smoothing coefficient as a second coefficient value, the second coefficient value being less than the first coefficient value, and the second probability threshold being less than the first probability threshold; if none of the above conditions is met, determining the smoothing coefficient as a third coefficient value, the third coefficient value being an intermediate value between the first coefficient value and the second coefficient value.

[0012] Optionally, before determining whether wind noise exists based on the long-time wind noise existence probability of the current frame, the method further comprises: obtaining amplitudes and of the plurality of channels based on a time domain signal of the current frame; performing smoothing processing on the amplitudes and of the plurality of channels of the current frame according to the amplitudes and of the plurality of channels of the previous frame; obtaining a third channel with the smallest amplitude and after the smoothing processing; determining whether wind noise exists based on the long-time wind noise existence probability of the current frame, comprising: determining whether wind noise exists based on the long-time wind noise existence probability of the current frame and in combination with the amplitude and of the third channel after the smoothing processing.

[0013] Optionally, determining whether wind noise exists based on the long-time wind noise existence probability of the current frame and in combination with the amplitude and of the third channel after the smoothing processing, comprising: if the long-time wind noise existence probability of the current frame is greater than a third probability threshold value and the amplitude and of the second channel after the smoothing processing is greater than a preset threshold value, determining that wind noise exists.

[0014] Optionally, obtaining the short-time wind noise existence probability of the current frame according to the sound signals of the plurality of channels, comprising: analyzing a cross-correlation coefficient of amplitude squares in a frequency domain according to the sound signals of the plurality of channels; determining the short-time wind noise existence probability of the current frame by using the cross-correlation coefficient.

[0015] According to a second aspect of the embodiments of the present disclosure, a wind noise processing device is provided, comprising: an obtaining module configured to obtain sound signals of a plurality of channels collected by a plurality of microphones; a selecting module configured to, in a case where it is determined that wind noise exists according to the sound signals of the plurality of channels, select a first channel with wind noise pollution meeting a preset condition from the plurality of channels; a processing module configured to perform wind noise suppression processing based on the sound signals of the first channel.

[0016] Optionally, the selecting module is specifically configured to select a channel with the smallest wind noise from the plurality of channels as the first channel.

[0017] Optionally, the first channel is a channel with wind noise pollution meeting a preset condition selected based on a current frame. The processing module is specifically configured to, if a second channel selected based on a historical frame and meeting a preset condition is the same as the first channel, perform wind noise suppression processing based on a sound signal of the first channel; if the second channel is different from the first channel, determine an amplitude difference between the first channel and the second channel based on a time domain signal of a current frame, and select a target channel for wind noise suppression processing from the first channel and the second channel according to the amplitude difference.

[0018] Optionally, the apparatus further comprises a detection module. The detection module is configured to perform wind noise detection based on the sound signals of the multiple channels. The processing module is further configured to perform adaptive echo cancellation processing on the sound signal after wind noise detection.

[0019] Optionally, the detection module is configured to obtain a short-time wind noise existence probability of a current frame and a long-time wind noise existence probability of a previous frame based on the sound signals of the multiple channels, perform smoothing processing to obtain the long-time wind noise existence probability of the current frame based on the short-time wind noise existence probability of the current frame and the long-time wind noise existence probability of the previous frame, and determine whether wind noise exists based on the long-time wind noise existence probability of the current frame.

[0020] Optionally, the detection module is specifically configured to perform smoothing processing to obtain the long-time wind noise existence probability of the current frame based on the short-time wind noise existence probability of the current frame and the long-time wind noise existence probability of the previous frame by using a smoothing coefficient; the smoothing coefficient is non-fixed.

[0021] Optionally, the detection module is further configured to, if the long-time wind noise existence probability of the previous frame is greater than a first probability threshold and the short-time wind noise existence probability of the current frame is less than the first probability threshold, determine the smoothing coefficient as a first coefficient value. If the long-time wind noise existence probability of the previous frame is less than a second probability threshold and the short-time wind noise existence probability of the current frame is greater than the first probability threshold, determine the smoothing coefficient as a second coefficient value, the second coefficient value is less than the first coefficient value, and the second probability threshold is less than the first probability threshold. If none of the above conditions is met, determine the smoothing coefficient as a third coefficient value, the third coefficient value being an intermediate value of the first coefficient value and the second coefficient value.

[0022] Optionally, the detection module is specifically configured to obtain the amplitude sum corresponding to each of the plurality of channels based on the time domain signal of the current frame; perform smoothing processing on the amplitude sum corresponding to each of the plurality of channels of the current frame according to the amplitude sum corresponding to each of the plurality of channels of the previous frame; obtain a third channel with the smallest amplitude sum after the smoothing processing; and determine whether the wind noise exists based on the long-time wind noise existence probability of the current frame and in combination with the amplitude sum corresponding to the third channel after the smoothing processing.

[0023] Optionally, the determination module is specifically configured to determine that the wind noise exists if the long-time wind noise existence probability of the current frame is greater than a third probability threshold and the amplitude sum corresponding to the second channel after the smoothing processing is greater than a preset threshold.

[0024] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, including: a processor; a memory connected with the processor, and having a computer program stored thereon, the computer program being executed by the processor to implement the wind noise processing method in the first aspect.

[0025] According to a fourth aspect of the embodiments of the present disclosure, a computer readable storage medium is provided, having a computer program stored thereon, the computer program being executed by a processor to implement the wind noise processing method in the first aspect.

[0026] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, including a computer program, and the computer program is executed by a processor to implement the wind noise processing method in the first aspect.

[0027] According to a sixth aspect of the embodiments of the present disclosure, a chip is provided, including one or more interface circuits and one or more processors; the interface circuit is used to receive a signal from a memory of an electronic device and send the signal to the processor, the signal including computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device performs the wind noise processing method in the first aspect.

[0028] By means of the technical solutions, the wind noise processing method, device, storage medium, electronic device and chip are provided. Firstly, a plurality of channel sound signals collected by a plurality of microphones are acquired; in a case where it is determined that there is wind noise according to the plurality of channel sound signals, a first channel with wind noise pollution meeting a preset condition, such as a channel with less wind noise, is selected from the plurality of channels; and then wind noise suppression processing is performed based on the sound signal of the first channel, such as wind noise suppression processing based on the sound signal of the channel with less wind noise. Compared with related technologies, the wind noise suppression algorithm can be effectively relieved, and the stability of the result after wind noise suppression can be ensured. When collecting user voice information through a plurality of microphones, the wind noise processing scheme provided by the present disclosure can obtain target voice with higher signal-to-noise ratio, clarity and intelligibility.

[0029] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and are not limiting to the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0030] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure.

[0031] Figure 1 A flowchart of a wind noise processing method provided by an embodiment of the present disclosure is shown; Figure 2 A structural diagram of an example provided by an embodiment of the present disclosure is shown; Figure 3 A flowchart of another wind noise processing method provided by an embodiment of the present disclosure is shown; Figure 4 A structural diagram of another example provided by an embodiment of the present disclosure is shown; Figure 5 A flowchart of another wind noise processing method provided by an embodiment of the present disclosure is shown; Figure 6 A flowchart of an example provided by an embodiment of the present disclosure is shown; Figure 7 A flowchart of another wind noise processing method provided by an embodiment of the present disclosure is shown; Figure 8 A flowchart of another example provided by an embodiment of the present disclosure is shown; Figure 9 A structural diagram of a wind noise processing device provided by an embodiment of the present disclosure is shown; Figure 10 A structural diagram of an electronic device provided by an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0032] In order to more clearly understand the above-mentioned purposes, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0033] In the following description, a large number of specific details are set forth in order to facilitate a thorough understanding of the present disclosure. However, the present disclosure can be implemented in many different ways from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of the present disclosure, so the present disclosure is not limited to the specific implementation disclosed below.

[0034] The terms used in one or more embodiments of the present disclosure are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of the present disclosure. The singular forms "a", "said" and "the" used in one or more embodiments of the present disclosure and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present disclosure means and includes any or all possible combinations of one or more associated listed items.

[0035] It should be understood that although the terms first, second, etc. can be employed in one or more embodiments of the present disclosure to describe various information, these information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present disclosure, first can also be referred to as second, and similarly, second can also be referred to as first. Depending on the context, the word "if" as used herein can be interpreted as "when" or "upon determining" or "in response to determining".

[0036] In order to reduce the influence of wind noise on voice, work can be carried out in terms of microphone layout, microphone pipeline structure, wind noise detection algorithm, wind noise suppression algorithm, etc. For the sound signals collected by multiple microphones, wind noise needs to be detected before being suppressed, and corresponding noise reduction processing is performed when wind noise is detected. However, after wind noise detection in this way, channel selection is not involved, which cannot effectively reduce the pressure of wind noise suppression algorithm, and affects the stability of the result after wind noise suppression.

[0037] In order to solve the above technical problems, the present embodiment provides a wind noise processing method, as shown in the figure, the method can be applied to the terminal side of wind noise processing device or equipment, etc., the method comprises the following steps: Figure 1 As shown in the figure, the method can be applied to the terminal side of wind noise processing device or equipment, etc., the method comprises the following steps: Step 101, acquiring multiple channel sound signals collected by multiple microphones.

[0038] The multiple microphones can refer to two or more microphone devices deployed in an array form (such as a linear array, a ring array) or a scattered layout, for receiving sound from different spatial positions or angles. Each microphone independently acquires sound and outputs an electrical signal, and each electrical signal corresponds to a channel. For example, the signals acquired by three microphones form three-channel sound signals, and each channel signal carries sound information at the position of the corresponding microphone, which can include target sound, noise, environmental interference, and the like.

[0039] Step 102: In a case where it is determined that wind noise exists according to the multiple-channel sound signals, a first channel with wind noise pollution meeting a preset condition is selected from the multiple channels.

[0040] For the wind noise detection scheme, the correlation between the multiple-channel sound signals, the energy difference, and the high-to-low frequency energy ratio, and the like can be used to analyze the probability value of the existence of wind noise according to the physical mechanism and signal characteristics of wind noise. If the probability value of the existence of wind noise is greater than a certain threshold, it can be determined that wind noise exists.

[0041] The embodiments of the present disclosure can select a first channel with wind noise pollution meeting a preset condition from the multiple channels in a case where it is determined that wind noise exists. The wind noise pollution refers to the interference degree of wind noise on the sound signal of the channel. The preset condition can be a pre-set screening standard, for example, a smaller wind noise pollution degree, or the lowest wind noise pollution degree, or a wind noise pollution degree within a certain threshold range, and the like. The first channel is a channel meeting the preset condition after screening, such as a channel with a smaller wind noise pollution degree, or a channel with the lowest wind noise pollution degree, and the like.

[0042] In some embodiments, a channel with the smallest wind noise can be selected from the multiple channels as the first channel with wind noise pollution meeting the preset condition. The smallest wind noise means that the interference degree of the wind noise component in the sound signal of the channel is the lowest, for example, the energy, the frequency spectrum proportion, and the masking degree of the effective signal of the wind noise are at the lowest level among the channels. In this way, the wind noise suppression processing can be performed based on the sound signal of the channel with the smallest wind noise, the pressure of the wind noise suppression algorithm can be reduced to the greatest extent, and the stability of the result after wind noise suppression can be ensured. When collecting user voice information through a microphone, the target voice with higher signal-to-noise ratio, clarity, and intelligibility can be obtained by applying the wind noise processing scheme provided by the embodiments of the present disclosure.

[0043] Step 103: Perform wind noise suppression processing based on the sound signal of the first channel.

[0044] In some examples, a wind noise suppression algorithm (such as spectral subtraction, or Wiener filtering, or a deep learning-based noise reduction model, etc.) is used to process the sound signal of the first channel to weaken or eliminate the wind noise component therein. Subsequently, only the sound signal (after wind noise suppression processing) of the first channel can be output. No wind noise suppression processing is performed on the sound signals of other channels.

[0045] For example, when collecting user voice information (corresponding to sound signals of multiple channels) through multiple microphones, in a case where wind noise is determined to exist, a channel A with the least wind noise is selected from the multiple channels, wind noise suppression processing is performed based on the sound signal of the channel A, and no wind noise suppression processing is performed on the sound signals of other channels. Subsequently, only the sound signal of the channel A after wind noise suppression processing can be output, such as user voice playback or input into a voice recognition model for voice recognition.

[0046] The method provided by the embodiments of the present disclosure can effectively reduce the pressure of the wind noise suppression algorithm and ensure the stability of the result after wind noise suppression. When collecting user voice information through multiple microphones, the wind noise processing scheme provided by the embodiments of the present disclosure can be applied to obtain target voice with higher signal-to-noise ratio, clarity and intelligibility.

[0047] The method provided by the embodiments of the present disclosure can be applied to many scenes, for example, hands-free calling, recording or voice recognition based on two or more microphones. For example, as shown in Figure 2 In an outdoor recording scenario, the signal picked up by the dual microphones includes near-end voice, environmental noise and wind noise. In a calling scenario, the dual microphones also pick up echo signals transmitted from a far-end. The proportion of the echo signals is usually large in a hands-free scenario. The present scheme can be applied to many technical fields of real-time or non-real-time voice enhancement and is suitable for many terminal forms such as mobile phones, computers, wearable devices, sports cameras and recording pens.

[0048] In some embodiments, for a calling scenario, the influence of near-end echo needs to be considered in wind noise detection, that is, the near-end microphone picks up the far-end voice emitted by the near-end loudspeaker. In order to eliminate the echo from the microphone input signal as much as possible while maintaining linearity and not affecting the correlation calculation in wind noise detection, a related technical scheme usually performs adaptive echo cancellation (AEC) on the microphone signal using an echo reference signal before performing wind noise detection. Such an algorithm architecture design will cause two problems: on the one hand, the wind noise detection algorithm is very dependent on the performance of AEC. Once the output result of AEC appears whitening or spectral leakage due to errors, the correlation between the microphone signals will be reduced, which can cause the wind noise false detection rate to increase. On the other hand, the calculation complexity of the AEC algorithm itself is large. With the increase of the number of microphones and the bandwidth, the calculation complexity of the AEC part will increase exponentially.

[0049] In order to solve the above technical problems, combined Figure 1 The embodiment shown gives the following Figure 3 The specific implementation method shown includes the following steps: Step 201 : Before performing adaptive echo cancellation processing, perform wind noise detection based on sound signals of multiple channels.

[0050] Step 202: After wind noise detection, perform adaptive echo cancellation processing on the sound signal.

[0051] In some embodiments, for call scenarios, the disclosed embodiments may first perform wind noise detection based on the sound signals of multiple channels before performing adaptive echo cancellation. Then, adaptive echo cancellation is performed on the sound signals after wind noise detection. This approach makes wind noise detection independent of the performance of the AEC algorithm when there is an echo, reducing the risk of false wind noise detection and lowering computational complexity and power consumption in wind noise scenarios.

[0052] For example, the technical solution of the embodiment of the present disclosure is introduced by taking the dual-microphone hands-free call scenario as an example. Figure 4 As shown, the two microphone signals MIC0 and MIC1 first pass through a high-pass filter (HPF) to remove DC and suppress low-frequency noise. The two HPF output signals are then adaptively filtered using a time-aligned echo reference signal (REF) to eliminate the linear echo component in both microphone signals, effectively performing the AEC algorithm. The two AEC-processed signals are then simultaneously fed into the beamforming algorithm (BF) module, resulting in a single output. In general, other algorithm modules are connected after the BF in a call scenario, but this is not specifically limited in this embodiment.

[0053] Because the echoes received by MIC0 and MIC1 during hands-free calls are highly correlated, while the two channels are less correlated under the influence of wind noise, the two channels are well distinguishable. Therefore, the wind noise detection module (WND) in the disclosed embodiment uses the two signals before AEC processing and after HPF processing as algorithm inputs. It detects and outputs a flag (det_flag) in real time to indicate whether wind noise is present in the two channel signals in the current frame. If wind noise is detected, it outputs the index of the channel (channel_index) where wind noise contamination meets preset conditions. The wind noise flag and channel index serve as inputs to the BF module. If no wind noise is detected, the BF combines the two AEC output signals into a single beam for output. If wind noise is detected, the BF directly outputs the channel (corresponding to the channel index) where wind noise contamination meets preset conditions, allowing wind noise suppression to be performed only on that channel.

[0054] By applying the technical solutions of the embodiments of the present disclosure, in a call scenario, wind noise detection can be placed before AEC algorithm processing, so that wind noise detection when there is echo does not depend on the performance of the AEC algorithm, which can reduce the risk of wind noise false detection, and at the same time, in a wind noise scenario, it can reduce the computational complexity and power consumption.

[0055] Further, in order to illustrate the specific process of wind noise detection according to the sound signals of multiple channels in the embodiments of the present disclosure, a specific implementation method as shown in Figure 5 is given, which includes the following steps: Step 301, according to the sound signals of multiple channels, obtaining the short-time wind noise existence probability of the current frame and the long-time wind noise existence probability of the previous frame.

[0056] Among them, the sound signal will be segmented into continuous frames when processing, the current frame refers to the latest frame signal being processed in time. And the previous frame refers to the previous frame signal adjacent to the current frame.

[0057] The short-time wind noise existence probability refers to the possibility of the existence of wind noise for each frame signal, that is, it is calculated independently for a single frame, which belongs to the judgment in the short time scale. And the long-time wind noise existence probability refers to the wind noise existence probability fused with the long-time trend, that is, it is obtained by recursively accumulating multiple frames, which fuses the information of historical frames, and belongs to the judgment of whether the wind noise exists in a long time scale. In the embodiments of the present disclosure, each frame signal has its corresponding short-time wind noise existence probability and long-time wind noise existence probability.

[0058] In some embodiments, the amplitude square cross-correlation coefficient in the frequency domain can be analyzed according to the sound signals of multiple channels; and then the short-time wind noise existence probability of the current frame is determined by using the cross-correlation coefficient.

[0059] For example, based on Figure 4 is taken as an example to introduce the technical solutions of the embodiments of the present disclosure. First, define the amplitude square cross-correlation coefficient (MSC) in the frequency domain, as shown in the following formula:

[0060] Among them, is the frame index, is the frequency point index, is the self-power spectrum or cross-power spectrum after recursive smoothing, and the corresponding represents the cross-power spectrum after recursive smoothing of the two microphone signals x1 and x2, represents the self-power spectrum after recursive smoothing of the microphone signal x1, represents the self-power spectrum after recursive smoothing of the microphone signal x2, As shown in the following formula:

[0061] And denotes the short-time Fourier spectrum of the two-channel MIC signal filtered by HPF, is a smoothing coefficient.

[0062] The (short-time) wind noise presence probability WPP of the current frame is defined as shown in the following formula:

[0063] And are the lower and upper limits of the frequency range for intercepting the frequency points when calculating WPP, and since the main energy of wind noise is concentrated in low frequencies, it is recommended that be the corresponding frequency point below 1500 Hz. As can be seen from the above formula, the smaller the inter-channel cross-correlation coefficient MSC, the greater the WPP, that is, the greater the probability of the current wind noise presence.

[0064] Step 302, smoothing processing is performed according to the short-time wind noise presence probability of the current frame and the long-time wind noise presence probability of the previous frame to obtain the long-time wind noise presence probability of the current frame.

[0065] In some examples, the long-time wind noise presence probability of the current frame can be obtained by smoothing processing using a smoothing coefficient according to the short-time wind noise presence probability of the current frame and in combination with the long-time wind noise presence probability of the previous frame; wherein the smoothing coefficient is non-fixed.

[0066] For example, in order to obtain a more robust wind noise presence probability value, the long-time wind noise presence probability is introduced in the embodiments of the present disclosure.

[0067] wherein, denotes a smoothing coefficient, and the smoothing coefficient applied here can be non-fixed, and its specific value can be determined by some important logical judgments. For example, the smoothing coefficient may be determined according to the long-time wind noise presence probability of the previous frame and the short-time wind noise presence probability of the current frame.

[0068] In the process of performing wind noise detection on sound signals collected by multiple microphones, the related art uses a fixed smoothing coefficient for the smoothing of the cross-correlation coefficient, which has poor adaptability to wind noise changes and is difficult to balance between the robustness of the wind noise detection results and the fast tracking performance. In addition, when the signal picked up by the microphone is small, the cross-correlation coefficient between the channels will be significantly reduced, resulting in false detection of wind noise, such as in a quiet room or a blocked microphone scenario. In order to solve this technical problem, the smoothing coefficient used in the embodiment of the present disclosure can be non-fixed, and its specific value can be determined by some important logical judgments. The embodiment of the present disclosure uses a flexibly configured smoothing coefficient for the smoothing of the cross-correlation coefficient, which has better adaptability to wind noise changes and can achieve a better balance between the robustness of the wind noise detection results and the fast tracking performance. In addition, the problem of false detection of wind noise in scenarios such as quiet rooms or blocked microphones can be avoided through threshold design.

[0069] Exemplarily, in some embodiments, if the probability of long-term wind noise in the previous frame is greater than the first probability threshold, and the probability of short-term wind noise in the current frame is less than the first probability threshold, the smoothing coefficient is determined to be the first coefficient value; if the probability of long-term wind noise in the previous frame is less than the second probability threshold, and the probability of short-term wind noise in the current frame is greater than the first probability threshold, the smoothing coefficient is determined to be the second coefficient value, the second coefficient value is less than the first coefficient value, and the second probability threshold is less than the first probability threshold; if none of these conditions are met, the smoothing coefficient is determined to be the third coefficient value, and the third coefficient value is the middle value between the first coefficient value and the second coefficient value.

[0070] For example, Figure 6 As shown, (Condition 1) when the long-term wind noise of the previous frame exists, the probability When the probability of short-term wind noise in the current frame is greater than the threshold thr_max (i.e., the first probability threshold), and the probability of short-term wind noise in the current frame is less than the threshold thr_max, the probability of long-term wind noise is very high, while the probability of short-term wind noise is reduced. To ensure the robustness of subsequent wind noise judgment results and avoid large fluctuations in the judgment results, the probability of long-term wind noise in the previous frame is chosen to be more reliable. , give the smoothing coefficient a larger value beta_max (i.e. the first coefficient value), so The update speed slows down.

[0071] (Condition 2) If the above condition 1 is not met, then when there is a probability of long-term wind noise in the previous frame When the probability of short-term wind noise in the current frame is less than the threshold thr_min (i.e., the second probability threshold), and the probability of short-term wind noise in the current frame WPP is greater than the threshold thr_max, the probability of long-term wind noise is very low, while the probability of short-term wind noise is very high. It is considered that the possibility of wind changing from no wind to wind is very high, and it is necessary to quickly track this change. Therefore, a smaller smoothing coefficient value beta_min (i.e., the second coefficient value) is assigned to make Can be updated quickly.

[0072] If the above conditions 1 and 2 are not met, the smoothing coefficient is given an intermediate value beta_mid (i.e. the third coefficient value), so that Compared with a fixed-value smoothing coefficient, the wind noise existence probability calculated by the embodiment of the present disclosure is flexible and can obtain a more accurate and robust wind noise existence probability statistic.

[0073] Step 303: Determine whether wind noise exists based on the long-term wind noise existence probability of the current frame.

[0074] The embodiment of the present disclosure determines the probability of long-term wind noise existing in the current frame based on the above-mentioned method, and can accurately determine whether there is wind noise.

[0075] Under normal circumstances, when wind noise is present, the signal picked up by the microphone will be relatively small, at least higher than the normal noise floor. However, in environments such as listening rooms and anechoic chambers, or when the microphone is blocked, the noise floor can be significantly lower. In this case, the correlation between microphone signals is significantly reduced, which can lead to a higher probability of misjudging the presence of wind noise. To address this issue, it is necessary to avoid these low noise floor situations.

[0076] To this end, in some embodiments, the amplitude sums corresponding to the multiple channels can be obtained based on the time domain signal of the current frame; then, based on the amplitude sums corresponding to the multiple channels of the previous frame, the amplitude sums corresponding to the multiple channels of the current frame are smoothed; then, the third channel with the smallest amplitude sum after smoothing is obtained; and then, in the process of determining whether wind noise exists based on the probability of long-term wind noise existing in the current frame, it is determined whether wind noise exists based on the probability of long-term wind noise existing in the current frame and in combination with the amplitude sum corresponding to the smoothed third channel.

[0077] For example, based on Figure 4 As shown, the technical solution of the embodiment of the present disclosure is introduced by taking the dual-microphone hands-free call scenario as an example. First, calculate the time domain signal of the current frame The amplitude and is as follows:

[0078] Here, k = 1 and 2 are the channel identifiers of the two microphones. Then, to avoid large fluctuations, sum_amp is smoothed to obtain the following:

[0079] in, is the smoothing coefficient.

[0080] Will The minimum value of is defined as minval (i.e. the third channel), that is

[0081] The corresponding channel index is channel_index = 1 or 2.

[0082] The subsequent determination of whether wind noise exists can be based on the and accurate determination of whether wind noise exists.

[0083] In some examples, if the long-time wind noise existence probability of the current frame is greater than a third probability threshold, and the smoothed amplitude sum of the second channel is greater than a preset threshold, it is determined that wind noise exists.

[0084] For example, based on the technical solution of the embodiment of the disclosure is introduced in a dual-microphone hands-free call scenario as shown in the figure. When the long-time wind noise existence probability Figure 4 is greater than the threshold flag_thr (i.e., the third probability threshold), and minval is greater than level_thr (i.e., the preset threshold), it is determined that the current frame has wind noise, and the wind noise identifier det_flag = 1 is output; if the conditions are not met, it is determined that the current frame does not have wind noise, and the wind noise identifier det_flag = 0 is output, and the current frame channel index channel_index = 0 (not 1 or 2) is output. The design here can avoid the high probability of false wind noise in the low background noise scenarios such as quiet environment and blocked microphone, and select the channel with less wind noise, which reduces the pressure for subsequent wind noise suppression and helps to improve the signal-to-noise ratio, clarity and intelligibility of the target speech. Next, it is necessary to determine the current frame channel index channel_index when det_flag = 1. In actual application, there may be a problem of frequent jumping of channel_index between multiple channel indexes. Therefore, the embodiment of the disclosure also provides a specific implementation method as shown in the figure.

[0085] Figure 7 The method comprises the following steps: Step 401, obtaining a second channel with wind noise pollution meeting a preset condition based on historical frames, and a first channel with wind noise pollution meeting a preset condition based on a current frame.

[0086] The first channel is the channel with wind noise pollution meeting the preset condition detected based on the current frame, and the second channel is the channel with wind noise pollution meeting the preset condition corresponding to the last time wind noise was detected.

[0087] Step 402a, if the second channel is the same as the first channel, performing wind noise suppression processing based on the sound signal of the first channel.

[0088] ​In step 402b, which is parallel to step 402a, if the second channel is different from the first channel, the amplitude difference between the first channel and the second channel is determined based on the time domain signal of the current frame, and a target channel for wind noise suppression processing is selected from the first channel and the second channel according to the amplitude difference.

[0089] In some embodiments, when the amplitude difference between the first channel and the second channel is less than or equal to a preset amplitude difference threshold, wind noise suppression processing is performed based on the sound signal of the second channel.

[0090] In some embodiments, when the amplitude difference between the first channel and the second channel is greater than a preset amplitude difference threshold, wind noise suppression processing is performed based on the sound signal of the first channel.

[0091] For example, based on Figure 4 As shown, the technical solution of the embodiment of the present disclosure is introduced by taking the dual-microphone hands-free call scenario as an example. If there is wind noise in the historical frame, that is, channel_index_pre ≠ 0, then when the channel index channel_index of the current frame is inconsistent with the channel index channel_index_pre of the historical frame, it is necessary to determine whether the amplitude difference between the two channels of the current frame meets the conditions, and define and The decibel difference is as follows:

[0092] like Figure 8 As shown in the figure, if decrement is less than or equal to the threshold thr_dB (i.e., the preset amplitude difference threshold), it is determined that the wind noise energy difference between the two channels in the current frame is not large, and channel_index can be left unchanged and remain at the channel index channel_index_pre of the historical frame. If decrement is greater than the threshold thr_dB, it is determined that the wind noise energy difference between the two channels in the current frame is large, and channel_index needs to be switched. channel_index_pre can be switched to the channel index channel_index of the current frame. Finally, the channel index channel_index of the current frame can be assigned to the channel index channel_index_pre of the historical frame for use in the next frame.

[0093] By applying the technical solutions of the embodiments of the present disclosure, a method for channel selection and switching after wind noise detection is provided, which can effectively reduce the pressure on the subsequent wind noise suppression algorithm, and is conducive to obtaining target speech with higher signal-to-noise ratio, clarity and intelligibility, while avoiding frequent switching of selected channels and ensuring the stability of subsequent output results.

[0094] Figure 9is a wind noise processing device block diagram according to some embodiments of the present disclosure, which can be configured to perform Figures 1 to 8 The method shown is described with reference to Figure 9 The device comprises an acquisition module 51, a selection module 52 and a processing module 53.

[0095] The acquisition module 51 is configured to acquire a plurality of channel sound signals collected by a plurality of microphones; The selection module 52 is configured to select a first channel with wind noise pollution meeting a preset condition from the plurality of channels in a case where it is determined that there is wind noise according to the plurality of channel sound signals; The processing module 53 is configured to perform wind noise suppression processing based on the sound signal of the first channel.

[0096] In some embodiments of the present disclosure, the selection module 52 is specifically configured to select a channel with the least wind noise from the plurality of channels as the first channel.

[0097] In some embodiments of the present disclosure, the first channel is a channel with wind noise pollution meeting a preset condition selected based on a current frame; the processing module 53 is specifically configured to perform wind noise suppression processing based on the sound signal of the first channel if a second channel with wind noise pollution meeting a preset condition selected based on a historical frame is the same as the first channel; if the second channel is different from the first channel, determine an amplitude difference between the first channel and the second channel based on a time domain signal of the current frame, and select a target channel for wind noise suppression processing from the first channel and the second channel according to the amplitude difference.

[0098] In some embodiments of the present disclosure, the processing module 53 is specifically configured to perform wind noise suppression processing based on the sound signal of the second channel in a case where the amplitude difference is less than or equal to a preset amplitude difference threshold; and perform wind noise suppression processing based on the sound signal of the first channel in a case where the amplitude difference is greater than the preset amplitude difference threshold.

[0099] In some embodiments of the present disclosure, the device further comprises a detection module 54; The detection module 54 is configured to perform wind noise detection according to the plurality of channel sound signals; The processing module 53 is further configured to perform adaptive echo cancellation processing on the sound signal after wind noise detection.

[0100] In some embodiments of the present disclosure, the detection module 54 is specifically configured to obtain a short-time wind noise existence probability of a current frame and a long-time wind noise existence probability of a previous frame according to sound signals of the multiple channels; perform smoothing processing to obtain the long-time wind noise existence probability of the current frame according to the short-time wind noise existence probability of the current frame and the long-time wind noise existence probability of the previous frame; and determine whether wind noise exists based on the long-time wind noise existence probability of the current frame.

[0101] In some embodiments of the present disclosure, the detection module 54 is specifically configured to perform smoothing processing to obtain the long-time wind noise existence probability of the current frame by using a smoothing coefficient according to the short-time wind noise existence probability of the current frame and in combination with the long-time wind noise existence probability of the previous frame; and the smoothing coefficient is non-fixed.

[0102] In some embodiments of the present disclosure, the detection module 54 is specifically configured to determine the smoothing coefficient according to the long-time wind noise existence probability of the previous frame and the short-time wind noise existence probability of the current frame.

[0103] In some embodiments of the present disclosure, the detection module 54 is specifically configured to determine that the smoothing coefficient is a first coefficient value if the long-time wind noise existence probability of the previous frame is greater than a first probability threshold and the short-time wind noise existence probability of the current frame is less than the first probability threshold; determine that the smoothing coefficient is a second coefficient value if the long-time wind noise existence probability of the previous frame is less than a second probability threshold and the short-time wind noise existence probability of the current frame is greater than the first probability threshold, the second coefficient value being less than the first coefficient value and the second probability threshold being less than the first probability threshold. If the above conditions are not met, it is determined that the smoothing coefficient is a third coefficient value, the third coefficient value being an intermediate value between the first coefficient value and the second coefficient value.

[0104] In some embodiments of the present disclosure, the detection module 54 is specifically configured to obtain amplitude sums corresponding to the multiple channels respectively based on a time domain signal of the current frame; perform smoothing processing on the amplitude sums corresponding to the multiple channels of the current frame according to the amplitude sums corresponding to the multiple channels of the previous frame; obtain a third channel with the smallest amplitude sum after the smoothing processing; and determine whether wind noise exists based on the long-time wind noise existence probability of the current frame and in combination with the amplitude sum corresponding to the third channel after the smoothing processing.

[0105] In some embodiments of the present disclosure, the detection module 54 is specifically configured to determine that wind noise exists if the long-time wind noise existence probability of the current frame is greater than a third probability threshold and the amplitude sum corresponding to the second channel after the smoothing processing is greater than a preset threshold.

[0106] In some embodiments of the present disclosure, the detection module 54 is specifically configured to analyze the cross-correlation coefficient of the amplitude square in the frequency domain according to the sound signals of the plurality of channels; and determine the short-time wind noise existence probability of the current frame by using the cross-correlation coefficient.

[0107] As to the apparatus in the above-mentioned embodiments, the specific manners in which various modules perform operations have been described in detail in the embodiments of the method, and thus will not be described in detail here.

[0108] It should be noted that other corresponding descriptions of the functions of the units involved in the wind noise processing apparatus provided in the embodiments of the present disclosure can be referred to the corresponding descriptions in the Figures 1 to 8 , which will not be described here in detail.

[0109] Based on the method as shown in Figures 1 to 8 , correspondingly, the embodiments of the present disclosure also provide a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method as shown in Figures 1 to 8 .

[0110] Based on the method as shown in Figures 1 to 8 , correspondingly, the embodiments of the present disclosure also provide a computer program product, which includes a computer program, and the computer program is executed by a processor to implement the method as shown in Figures 1 to 8 .

[0111] Based on such understanding, the technical solutions of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.) and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods of various implementation scenarios of the present disclosure.

[0112] Based on the method as shown in Figures 1 to 8 , and the virtual apparatus embodiment as shown in Figure 9 , the embodiments of the present disclosure also provide a chip, which includes one or more interface circuits and one or more processors; the interface circuit is used to receive a signal from a memory of an electronic device and send the signal to the processor, and the signal includes computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device executes the method as shown in Figures 1 to 8 .

[0113] Figure 10A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.

[0114] As shown in Figure 10 The device 1000 includes a computing unit 1001 that can perform various appropriate actions and processes in accordance with a computer program stored in a ROM (Read-Only Memory) 1002 or a computer program loaded into a RAM (Random Access Memory) 1003 from the storage unit 1008. Various programs and data required for the operation of the device 1000 can also be stored in the RAM 1003. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An I / O (Input / Output) interface 1005 is also connected to the bus 1004.

[0115] Various components in the device 1000 are connected to the I / O interface 1005, including an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, an optical disk, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the device 1000 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0116] The computing unit 1001 can be various general purpose and / or special purpose processing components having processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, CPUs (Central Processing Units), GPUs (Graphic Processing Units), various specialized AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSPs (Digital Signal Processors), and any suitable processors, controllers, microcontrollers, and the like. The computing unit 1001 performs various methods and processes described above, such as the aforementioned methods. For example, in some embodiments, the aforementioned methods can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded onto the RAM 1003 and executed by the computing unit 1001, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, the computing unit 1001 can be configured to perform the aforementioned methods by any other suitable means, such as by means of firmware.

[0117] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a Field Programmable Gate Array (FPGA), an Application-Specific Integrated Circuit (ASIC), an Application Specific Standard Product (ASSP), a System on a Chip (SoC), a Complex Programmable Logic Device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0118] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0119] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable storage medium can include but are not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include one or more lines of electrical wire, portable computer diskette, hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, fiber optics, CD-ROM (Compact Disc Read-Only Memory), optical storage device, magnetic storage device, or any suitable combination of the foregoing.

[0120] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0121] The systems and techniques described herein can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described herein, or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a blockchain network.

[0122] The computer system can include clients and servers. The clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server is one of communication and distribution, with the server receiving requests from the client and transmitting responses via the communication network. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with a blockchain.

[0123] It should be noted that artificial intelligence is a discipline that studies enabling computers to simulate some thinking processes and intelligent behaviors of people (such as learning, reasoning, thinking, planning, etc.), both hardware and software technologies. Artificial intelligence hardware technology generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing, etc.; artificial intelligence software technology mainly includes computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, knowledge graph technology, etc. several major directions.

[0124] It should be understood that the various forms of flow shown above can be used to reorder, add or delete steps. For example, each step described in the present disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, which is not limited herein.

[0125] The above detailed description does not limit the scope of the disclosure. Various modifications, combinations, sub-combinations and alternatives can be made to the detailed description. Any modification, equivalent replacement and improvement etc. made within the spirit and principle of the disclosure shall be included in the scope of the disclosure.

Claims

1. A wind noise processing method, characterized by, The method comprises: acquiring a plurality of channel sound signals collected by a plurality of microphones; in a case where it is determined that wind noise exists according to the plurality of channel sound signals, selecting a first channel from the plurality of channels, wherein wind noise pollution of the first channel meets a preset condition; performing wind noise suppression processing based on the sound signal of the first channel; the first channel is a channel selected based on a current frame, wherein wind noise pollution of the channel meets the preset condition; the wind noise suppression processing based on the sound signal of the first channel comprises: if a second channel selected based on a historical frame, wherein wind noise pollution of the second channel meets the preset condition, is the same as the first channel, then performing wind noise suppression processing based on the sound signal of the first channel; if the second channel is different from the first channel, then determining an amplitude difference between the first channel and the second channel based on a time domain signal of the current frame, and selecting a target channel for performing wind noise suppression processing from the first channel and the second channel according to the amplitude difference.

2. The method of claim 1, wherein, the selection of the first channel from the plurality of channels, wherein wind noise pollution of the first channel meets the preset condition, comprises: selecting a channel with minimum wind noise from the plurality of channels as the first channel.

3. The method of claim 1, wherein, the selection of the target channel for performing wind noise suppression processing from the first channel and the second channel according to the amplitude difference comprises: in a case where the amplitude difference is less than or equal to a preset amplitude difference threshold, performing wind noise suppression processing based on the sound signal of the second channel; in a case where the amplitude difference is greater than the preset amplitude difference threshold, performing wind noise suppression processing based on the sound signal of the first channel.

4. The method according to any one of claims 1 to 3, characterized in that, The method further comprises: performing wind noise detection according to the plurality of channel sound signals; performing adaptive echo cancellation processing on the sound signal after wind noise detection.

5. The method of claim 4, wherein, the wind noise detection according to the plurality of channel sound signals comprises: acquiring a short-time wind noise existence probability of a current frame and a long-time wind noise existence probability of a previous frame according to the plurality of channel sound signals; performing smoothing processing to obtain a long-time wind noise existence probability of the current frame according to the short-time wind noise existence probability of the current frame and the long-time wind noise existence probability of the previous frame; determining whether wind noise exists based on the long-time wind noise existence probability of the current frame.

6. The method of claim 5, wherein, the smoothing processing to obtain the long-time wind noise existence probability of the current frame according to the short-time wind noise existence probability of the current frame and the long-time wind noise existence probability of the previous frame comprises: performing smoothing processing to obtain the long-time wind noise existence probability of the current frame according to the short-time wind noise existence probability of the current frame and the long-time wind noise existence probability of the previous frame by using a smoothing coefficient; wherein the smoothing coefficient is non-fixed.

7. The method of claim 6, wherein, The method further comprises: determining the smoothing coefficient according to the long-time wind noise existence probability of the previous frame and the short-time wind noise existence probability of the current frame.

8. The method of claim 7, wherein, the determination of the smoothing coefficient according to the long-time wind noise existence probability of the previous frame and the short-time wind noise existence probability of the current frame comprises: if the long-time wind noise existence probability of the previous frame is greater than a first probability threshold and the short-time wind noise existence probability of the current frame is less than the first probability threshold, then determining the smoothing coefficient as a first coefficient value. If the long-time wind noise existence probability of the previous frame is less than a second probability threshold, and the short-time wind noise existence probability of the current frame is greater than the first probability threshold, the smoothing coefficient is determined as a second coefficient value, the second coefficient value is less than the first coefficient value, and the second probability threshold is less than the first probability threshold. If the above conditions are not met, the smoothing coefficient is determined as a third coefficient value, the third coefficient value is an intermediate value between the first coefficient value and the second coefficient value.

9. The method of claim 5, wherein, Before determining whether wind noise exists based on the long-time wind noise existence probability of the current frame, the method further comprises: obtaining amplitudes of the multiple channels corresponding to the time domain signal of the current frame; performing smoothing processing on the amplitudes of the multiple channels corresponding to the current frame according to the amplitudes of the multiple channels corresponding to the previous frame; obtaining a third channel with the smallest amplitude after smoothing processing; determining whether wind noise exists based on the long-time wind noise existence probability of the current frame, comprising: determining whether wind noise exists based on the long-time wind noise existence probability of the current frame and the amplitude of the third channel after smoothing processing.

10. The method of claim 9, wherein, determining whether wind noise exists based on the long-time wind noise existence probability of the current frame and the amplitude of the third channel after smoothing processing, comprising: if the long-time wind noise existence probability of the current frame is greater than a third probability threshold, and the amplitude of the second channel after smoothing processing is greater than a preset threshold, it is determined that wind noise exists.

11. The method of claim 5, wherein, According to the sound signals of the multiple channels, the short-time wind noise existence probability of the current frame is obtained, comprising: analyzing the cross-correlation coefficient of the amplitude square in the frequency domain according to the sound signals of the multiple channels; determining the short-time wind noise existence probability of the current frame by using the cross-correlation coefficient.

12. A wind noise processing apparatus, characterized by, comprising: an obtaining module configured to obtain sound signals of multiple channels collected by multiple microphones; a selecting module configured to, in a case where it is determined that wind noise exists according to the sound signals of the multiple channels, select a first channel with wind noise pollution meeting a preset condition from the multiple channels; a processing module configured to perform wind noise suppression processing based on the sound signals of the first channel; the first channel is a channel with wind noise pollution meeting the preset condition selected based on the current frame; the processing module is configured to, if a second channel with wind noise pollution meeting the preset condition selected based on a historical frame is the same as the first channel, perform wind noise suppression processing based on the sound signals of the first channel; if the second channel is different from the first channel, determine an amplitude difference between the first channel and the second channel based on the time domain signal of the current frame, and select a target channel for wind noise suppression processing from the first channel and the second channel according to the amplitude difference.

13. The apparatus of claim 12, wherein the selecting module is specifically configured to select a channel with the minimum wind noise from the multiple channels as the first channel.

14. The apparatus of claim 12, wherein, The apparatus further comprises a detecting module. The detecting module is configured to perform wind noise detection according to the sound signals of the multiple channels. The processing module is further configured to perform adaptive echo cancellation processing on the sound signals after wind noise detection.

15. The apparatus of claim 14, wherein the detection module is configured to obtain a short-time wind noise existence probability of a current frame and a long-time wind noise existence probability of a previous frame according to sound signals of the multiple channels, perform smoothing processing on the long-time wind noise existence probability of the current frame according to the short-time wind noise existence probability of the current frame and the long-time wind noise existence probability of the previous frame, and determine whether wind noise exists based on the long-time wind noise existence probability of the current frame.

16. The apparatus of claim 15, wherein the detection module is specifically configured to perform smoothing processing on the long-time wind noise existence probability of the current frame according to the short-time wind noise existence probability of the current frame and the long-time wind noise existence probability of the previous frame by using a smoothing coefficient, and wherein the smoothing coefficient is non-fixed.

17. The apparatus of claim 16, wherein the detection module is further configured to determine the smoothing coefficient as a first coefficient value if the long-time wind noise existence probability of the previous frame is greater than a first probability threshold and the short-time wind noise existence probability of the current frame is less than the first probability threshold, determine the smoothing coefficient as a second coefficient value if the long-time wind noise existence probability of the previous frame is less than a second probability threshold and the short-time wind noise existence probability of the current frame is greater than the first probability threshold, and wherein the second coefficient value is less than the first coefficient value and the second probability threshold is less than the first probability threshold, and determine the smoothing coefficient as a third coefficient value if neither of the above conditions is met, and wherein the third coefficient value is an intermediate value between the first coefficient value and the second coefficient value.

18. The apparatus of claim 14, wherein the detection module is specifically configured to obtain amplitude sums corresponding to the multiple channels respectively based on a time domain signal of the current frame, perform smoothing processing on the amplitude sums corresponding to the multiple channels of the current frame according to amplitude sums corresponding to the multiple channels of the previous frame, obtain a third channel with a smallest amplitude sum after the smoothing processing, and determine whether wind noise exists based on the long-time wind noise existence probability of the current frame and the amplitude sum corresponding to the third channel after the smoothing processing.

19. The apparatus of claim 18, wherein the determination module is specifically configured to determine that wind noise exists if the long-time wind noise existence probability of the current frame is greater than a third probability threshold and the amplitude sum corresponding to the second channel after the smoothing processing is greater than a preset threshold. An apparatus includes: a processor; a memory connected with the processor, the memory having stored thereon a computer program, the computer program being executed by the processor to implement the method of any one of claims 1 to 11. The computer program, when executed by a processor, implements the method of any one of claims 1 to 11. The computer program, when executed by a processor, implements the method of any one of claims 1 to 11. ​ ​ 20. An electronic device, comprising: ​ ​ ​ 21. A computer readable storage medium having stored thereon a computer program, characterized in that, ​ 22. A computer program product comprising a computer program, characterized in that, ​ 23. A chip, characterized by An electronic device comprising one or more interface circuits and one or more processors; the interface circuit is configured to receive a signal from a memory of the electronic device, and send the signal to the processor, the signal comprising computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device is caused to perform the method of any one of claims 1 to 11.

Citation Information

Patent Citations

  • Wind noise pollution range estimation method, wind noise pollution range suppression method, wind noise pollution range estimation device, wind noise pollution range suppression device, medium and terminal

    CN115691532A

  • Wind noise pollution degree estimation method, wind noise suppression method, medium and terminal

    CN115691533A

  • Method and system for processing multi-channel voice signal containing wind noise

    CN116580722A

  • Wind noise filtering method and device suitable for multiple microphones

    CN118890583A

  • Voice detection method and related device thereof

    WO2024093460A1