Wind noise processing method and device, storage medium, electronic equipment and chip

By selecting the channel with the least wind noise for wind noise suppression processing, and combining wind noise detection and adaptive echo cancellation, the stability problem of wind noise suppression algorithm is solved, the voice signal-to-noise ratio and clarity are improved, and it is suitable for a variety of terminal devices.

CN120833796BActive Publication Date: 2026-01-16BEIJING X RING TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511318761.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2026-01-16
Estimated Expiration
2045-09-15

AI Technical Summary

Technical Problem

Existing technologies cannot effectively reduce the pressure on wind noise suppression algorithms, affecting the stability of the results after wind noise suppression. In particular, wind noise seriously affects call quality and speech intelligibility in outdoor voice communication and intelligent speech recognition.

Method used

By acquiring sound signals from multiple channels collected by multiple microphones, the first channel that meets the preset conditions for wind noise pollution is selected, and wind noise suppression processing is performed based on this channel. Combined with wind noise detection and adaptive echo cancellation processing, the execution order of the wind noise suppression algorithm is optimized to reduce computational complexity and false detection rate.

Benefits of technology

It effectively reduces the pressure on wind noise suppression algorithms, ensures the stability of the results after wind noise suppression, improves the signal-to-noise ratio and speech clarity, and is suitable for various terminal devices such as mobile phones, tablets, wearable devices and action cameras, thereby improving speech recognition performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833796B_ABST
    Figure CN120833796B_ABST
Patent Text Reader

Abstract

The present disclosure provides a wind noise processing method and device, a storage medium, an electronic device and a chip, and relates to the technical field of signal processing. The method comprises the following steps: first, acquiring sound signals of multiple channels collected by multiple microphones; in the case where it is determined that wind noise exists according to the sound signals of the multiple channels, selecting a first channel from the multiple channels, which meets a preset condition, such as a channel with less wind noise; and then performing wind noise suppression processing based on the sound signal of the first channel, such as performing wind noise suppression processing based on the sound signal of the channel with less wind noise. Compared with related technologies, the method can effectively reduce the pressure of the wind noise suppression algorithm and ensure the stability of the result after wind noise suppression. When collecting user voice information through multiple microphones, the wind noise processing scheme provided by the present disclosure can obtain target voice with higher signal-to-noise ratio, clarity and intelligibility.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of signal processing, and particularly relates to a wind noise processing method and device, a storage medium, an electronic device and a chip. BACKGROUND

[0002] With the evolution and development of technology, mobile communication and wearable devices gradually become popular. When making a voice communication, recording or interacting with a smart voice assistant outdoors, it is inevitable to encounter wind noise. Wind noise not only seriously affects the call quality and speech intelligibility, but also significantly reduces the recognition rate of the smart voice. SUMMARY

[0003] The present disclosure provides a wind noise processing method and device, a storage medium, an electronic device and a chip, and the main purpose is to solve the technical problem that the related technology cannot effectively reduce the pressure of the wind noise suppression algorithm, affecting the stability of the result after wind noise suppression.

[0004] According to a first aspect of an embodiment of the present disclosure, a wind noise processing method is provided, comprising:

[0005] obtaining a plurality of channel sound signals collected by a plurality of microphones;

[0006] in a case where it is determined that wind noise exists according to the plurality of channel sound signals, selecting a first channel with wind noise pollution meeting a preset condition from the plurality of channels;

[0007] performing wind noise suppression processing based on the sound signal of the first channel.

[0008] Optionally, the step of selecting a first channel with wind noise pollution meeting a preset condition from the plurality of channels comprises:

[0009] selecting a channel with the minimum wind noise from the plurality of channels as the first channel.

[0010] Optionally, the first channel is a channel with wind noise pollution meeting a preset condition selected based on a current frame;

[0011] The step of performing wind noise suppression processing based on the sound signal of the first channel comprises:

[0012] if a second channel with wind noise pollution meeting a preset condition selected based on a historical frame is the same as the first channel, performing wind noise suppression processing based on the sound signal of the first channel;

[0013] if the second channel is different from the first channel, determining an amplitude difference between the first channel and the second channel based on a time domain signal of a current frame, and selecting a target channel for wind noise suppression processing from the first channel and the second channel according to the amplitude difference.

[0014] Optionally, the method further comprises:

[0015] in a case that the amplitude difference is less than or equal to a preset amplitude difference threshold, performing wind noise suppression processing based on the sound signal of the second channel;

[0016] in a case that the amplitude difference is greater than the preset amplitude difference threshold, performing wind noise suppression processing based on the sound signal of the first channel.

[0017] Optionally, the method further comprises:

[0018] performing wind noise detection according to the sound signals of the multiple channels;

[0019] performing adaptive echo cancellation processing on the sound signals after the wind noise detection.

[0020] Optionally, the wind noise detection according to the sound signals of the multiple channels comprises:

[0021] obtaining a short-time wind noise existence probability of a current frame and a long-time wind noise existence probability of a previous frame according to the sound signals of the multiple channels;

[0022] performing smoothing processing to obtain a long-time wind noise existence probability of the current frame according to the short-time wind noise existence probability of the current frame and the long-time wind noise existence probability of the previous frame;

[0023] determining whether wind noise exists based on the long-time wind noise existence probability of the current frame.

[0024] Optionally, the performing smoothing processing to obtain a long-time wind noise existence probability of the current frame according to the short-time wind noise existence probability of the current frame and the long-time wind noise existence probability of the previous frame comprises:

[0025] performing smoothing processing to obtain a long-time wind noise existence probability of the current frame according to the short-time wind noise existence probability of the current frame and the long-time wind noise existence probability of the previous frame by using a smoothing coefficient; wherein the smoothing coefficient is non-fixed.

[0026] Optionally, the method further comprises:

[0027] determining the smoothing coefficient according to the long-time wind noise existence probability of the previous frame and the short-time wind noise existence probability of the current frame.

[0028] Optionally, the determining the smoothing coefficient according to the long-time wind noise existence probability of the previous frame and the short-time wind noise existence probability of the current frame comprises:

[0029] If the long-time wind noise existence probability of the previous frame is greater than the first probability threshold and the short-time wind noise existence probability of the current frame is less than the first probability threshold, the smoothing coefficient is determined as a first coefficient value;

[0030] If the long-time wind noise existence probability of the previous frame is less than a second probability threshold and the short-time wind noise existence probability of the current frame is greater than the first probability threshold, the smoothing coefficient is determined as a second coefficient value, the second coefficient value is less than the first coefficient value, and the second probability threshold is less than the first probability threshold;

[0031] If none of the above conditions is met, the smoothing coefficient is determined as a third coefficient value, the third coefficient value is an intermediate value between the first coefficient value and the second coefficient value.

[0032] Optionally, before determining whether wind noise exists based on the long-time wind noise existence probability of the current frame, the method further comprises:

[0033] obtaining amplitudes of the plurality of channels corresponding to the current frame based on time-domain signals of the current frame;

[0034] performing smoothing processing on the amplitudes of the plurality of channels corresponding to the current frame according to the amplitudes of the plurality of channels corresponding to the previous frame;

[0035] obtaining a third channel with the smallest amplitude after the smoothing processing;

[0036] determining whether wind noise exists based on the long-time wind noise existence probability of the current frame, comprising:

[0037] determining whether wind noise exists based on the long-time wind noise existence probability of the current frame and in combination with the amplitude of the third channel after the smoothing processing.

[0038] Optionally, determining whether wind noise exists based on the long-time wind noise existence probability of the current frame and in combination with the amplitude of the third channel after the smoothing processing, comprises:

[0039] if the long-time wind noise existence probability of the current frame is greater than a third probability threshold and the amplitude of the second channel after the smoothing processing is greater than a preset threshold, it is determined that wind noise exists.

[0040] Optionally, the short-time wind noise existence probability of the current frame is obtained according to the sound signals of the plurality of channels, comprising:

[0041] analyzing a cross-correlation coefficient of amplitude squares in a frequency domain according to the sound signals of the plurality of channels;

[0042] determining the short-time wind noise existence probability of the current frame by using the cross-correlation coefficient.

[0043] According to a second aspect of the embodiments of the present disclosure, a wind noise processing apparatus is provided, comprising:

[0044] an acquisition module configured to acquire a plurality of channel sound signals collected by a plurality of microphones;

[0045] a selection module configured to, in a case where it is determined that wind noise exists according to the plurality of channel sound signals, select a first channel from the plurality of channels, wherein wind noise pollution of the first channel meets a preset condition;

[0046] a processing module configured to perform wind noise suppression processing based on the sound signal of the first channel.

[0047] Optionally, the selection module is specifically configured to select a channel with minimum wind noise from the plurality of channels as the first channel.

[0048] Optionally, the first channel is a channel with wind noise pollution meeting the preset condition selected based on a current frame;

[0049] The processing module is specifically configured to, if a second channel with wind noise pollution meeting the preset condition selected based on a historical frame is the same as the first channel, perform wind noise suppression processing based on the sound signal of the first channel; if the second channel is different from the first channel, determine an amplitude difference between the first channel and the second channel based on a time domain signal of the current frame, and select a target channel for wind noise suppression processing from the first channel and the second channel according to the amplitude difference.

[0050] Optionally, the apparatus further comprises a detection module;

[0051] The detection module is configured to perform wind noise detection according to the plurality of channel sound signals.

[0052] The processing module is further configured to perform adaptive echo cancellation processing on the sound signal after wind noise detection.

[0053] Optionally, the detection module is configured to acquire a short-time wind noise existence probability of a current frame and a long-time wind noise existence probability of a previous frame according to the plurality of channel sound signals, perform smoothing processing to obtain a long-time wind noise existence probability of the current frame according to the short-time wind noise existence probability of the current frame and the long-time wind noise existence probability of the previous frame, and determine whether wind noise exists based on the long-time wind noise existence probability of the current frame.

[0054] Optionally, the detection module is specifically configured to perform smoothing processing to obtain a long-time wind noise existence probability of a current frame by using a smoothing coefficient according to a short-time wind noise existence probability of the current frame and in combination with a long-time wind noise existence probability of a previous frame; wherein the smoothing coefficient is non-fixed.

[0055] Optionally, the detection module is further configured to determine the smoothing coefficient as a first coefficient value if the long-time wind noise existence probability of the previous frame is greater than a first probability threshold and the short-time wind noise existence probability of the current frame is less than the first probability threshold.

[0056] If the long-time wind noise existence probability of the previous frame is less than a second probability threshold and the short-time wind noise existence probability of the current frame is greater than the first probability threshold, the smoothing coefficient is determined as a second coefficient value, the second coefficient value is less than the first coefficient value, and the second probability threshold is less than the first probability threshold.

[0057] If none of the above conditions is met, the smoothing coefficient is determined as a third coefficient value, the third coefficient value is an intermediate value between the first coefficient value and the second coefficient value.

[0058] Optionally, the detection module is specifically configured to obtain the amplitudes corresponding to the plurality of channels based on the time-domain signal of the current frame, perform smoothing processing on the amplitudes corresponding to the plurality of channels of the current frame based on the amplitudes corresponding to the plurality of channels of the previous frame, obtain a third channel with the smallest amplitude after the smoothing processing, and determine whether wind noise exists based on the long-time wind noise existence probability of the current frame and in combination with the amplitude corresponding to the third channel after the smoothing processing.

[0059] Optionally, the determination module is specifically configured to determine that wind noise exists if the long-time wind noise existence probability of the current frame is greater than a third probability threshold and the amplitude corresponding to the second channel after the smoothing processing is greater than a preset threshold.

[0060] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, including:

[0061] a processor;

[0062] a memory connected with the processor, the memory storing a computer program, and the computer program being executed by the processor to implement the wind noise processing method in the first aspect.

[0063] According to a fourth aspect of the embodiments of the present disclosure, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the wind noise processing method in the first aspect.

[0064] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, including a computer program, and the computer program is executed by a processor to implement the wind noise processing method in the first aspect.

[0065] According to a sixth aspect of the embodiments of the present disclosure, a chip is provided, comprising one or more interface circuits and one or more processors; the interface circuit is configured to receive a signal from a memory of an electronic device, and send the signal to the processor, the signal comprising computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device is caused to perform the wind noise processing method of the first aspect.

[0066] By means of the technical solutions described above, the wind noise processing method, device, storage medium, electronic device and chip provided by the present disclosure are provided. Firstly, a plurality of channel sound signals collected by a plurality of microphones are acquired; in the case where it is determined according to the plurality of channel sound signals that wind noise exists, a first channel with wind noise pollution meeting a preset condition, such as a channel with less wind noise, is selected from the plurality of channels; then, wind noise suppression processing is performed based on the sound signal of the first channel, such as wind noise suppression processing based on the sound signal of the channel with less wind noise. Compared with related technologies, the pressure of the wind noise suppression algorithm can be effectively reduced, and the stability of the result after wind noise suppression can be ensured. When collecting user voice information through a plurality of microphones, by applying the wind noise processing scheme provided by the present disclosure, the target voice with higher signal-to-noise ratio, clarity and intelligibility can be obtained.

[0067] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0068] The accompanying drawings, which are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure.

[0069] Figure 1 A flowchart of a wind noise processing method provided by an embodiment of the present disclosure is shown;

[0070] Figure 2 A structural schematic diagram of an example provided by an embodiment of the present disclosure is shown;

[0071] Figure 3 A flowchart of another wind noise processing method provided by an embodiment of the present disclosure is shown;

[0072] Figure 4 A structural schematic diagram of another example provided by an embodiment of the present disclosure is shown;

[0073] Figure 5 A flowchart of another wind noise processing method provided by an embodiment of the present disclosure is shown;

[0074] Figure 6 A flowchart of an example provided by an embodiment of the present disclosure is shown;

[0075] Figure 7 Fig. 6 shows a flow diagram of another wind noise processing method according to an embodiment of the present disclosure;

[0076] Figure 8 Fig. 7 shows a flow diagram of another example according to an embodiment of the present disclosure;

[0077] Figure 9 Fig. 8 shows a structure diagram of a wind noise processing device according to an embodiment of the present disclosure;

[0078] Figure 10 Fig. 9 shows a structure diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0079] In order to more clearly understand the above-mentioned purposes, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0080] In the following description, a lot of specific details are set forth in order to give a thorough understanding of the present disclosure. However, the present disclosure can be implemented in many different ways than those described herein, and those skilled in the art can make similar changes without departing from the spirit of the present disclosure, so the present disclosure is not limited to the specific implementations disclosed below.

[0081] The terms used in one or more embodiments of the present disclosure are merely for the purpose of describing specific embodiments and are not intended to limit one or more embodiments of the present disclosure. The singular forms "a", "an" and "the" used in one or more embodiments of the present disclosure and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present disclosure means and includes any or all possible combinations of one or more associated listed items.

[0082] It should be understood that although the terms first, second, etc. can be used in one or more embodiments of the present disclosure to describe various information, these information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present disclosure, first can also be referred to as second, and similarly, second can also be referred to as first. Depending on the context, the word "if" as used herein can be interpreted as "when" or "upon" or "in response to determining".

[0083] In order to reduce the influence of wind noise on voice, work can be carried out in terms of microphone layout, microphone pipeline structure, wind noise detection algorithm, wind noise suppression algorithm and the like. For the sound signals collected by multiple microphones, wind noise needs to be detected before being suppressed, and corresponding noise reduction processing is performed when it is detected that wind noise exists. However, after wind noise detection in this way, channel selection is not involved, which cannot effectively reduce the pressure of the wind noise suppression algorithm, and affects the stability of the result after wind noise suppression.

[0084] To solve the above technical problems, the embodiments of the present disclosure provide a wind noise processing method, as shown in the formula: Figure 1 The method can be applied to the terminal side of a wind noise processing device or equipment and the like, and the method comprises the following steps:

[0085] Step 101, obtaining multiple channel sound signals collected by multiple microphones.

[0086] Among them, multiple microphones can refer to two or more microphone devices deployed, which can be arranged in an array form (such as a linear array, a ring array) or a scattered layout, for receiving sound from different spatial positions or angles.

[0087] Each microphone independently collects sound and outputs an electric signal, and each electric signal corresponds to a channel. For example, the signals collected by 3 microphones will form 3 channel sound signals, and each channel signal carries the sound information at the position of the corresponding microphone, which can include target sound, noise, environmental interference and the like.

[0088] Step 102, in the case where it is determined that wind noise exists according to the multiple channel sound signals, selecting a first channel with wind noise pollution meeting a preset condition from the multiple channels.

[0089] For the wind noise detection scheme, according to the physical mechanism and signal characteristics of wind noise itself, the correlation between the multiple channel sound signals, the energy difference, and the high-low frequency energy ratio and the combination thereof can be used to analyze the probability value of the existence of wind noise. If the probability value of the existence of wind noise is greater than a certain threshold, it can be determined that wind noise exists.

[0090] The embodiments of the present disclosure can select a first channel with wind noise pollution meeting a preset condition from the multiple channels in the case where it is determined that wind noise exists. Among them, wind noise pollution refers to the interference degree of wind noise on the channel sound signal. The preset condition can be a pre-set screening standard, for example, a smaller wind noise pollution degree, or the lowest wind noise pollution degree, or the wind noise pollution degree being within a certain threshold range, etc. The first channel is a channel meeting the preset condition after screening, such as a channel with a smaller wind noise pollution degree, or a channel with the lowest wind noise pollution degree, etc.

[0091] In some embodiments, the channel with the least wind noise can be selected from the multiple channels as the first channel with wind noise pollution meeting the preset condition. The least wind noise means that the wind noise component in the sound signal of the channel has the lowest interference degree, for example, the energy, frequency spectrum proportion, and masking degree of the wind noise to the effective signal are at the lowest level among the channels. In this way, the wind noise suppression processing can be performed based on the sound signal of the channel with the least wind noise, the pressure on the wind noise suppression algorithm can be reduced to the greatest extent, and the stability of the result after wind noise suppression can be ensured. When collecting user voice information through a microphone, the target voice with higher signal-to-noise ratio, clarity, and intelligibility can be obtained by applying the wind noise processing scheme provided in the embodiments of the present disclosure.

[0092] In step 103, wind noise suppression processing is performed based on the sound signal of the first channel.

[0093] In some examples, the wind noise suppression algorithm (such as spectral subtraction, Wiener filtering, or a deep learning-based noise reduction model) is used to process the sound signal of the first channel to weaken or eliminate the wind noise component therein. Subsequently, only the sound signal (after wind noise suppression processing) of the first channel can be output. The sound signals of other channels are not subjected to wind noise suppression processing.

[0094] For example, when collecting user voice information (corresponding to the sound signals of multiple channels) through multiple microphones, in the case where wind noise exists, a channel A with the least wind noise is selected from the multiple channels, wind noise suppression processing is performed based on the sound signal of the channel A, and the sound signals of other channels are not subjected to wind noise suppression processing. Subsequently, only the sound signal of the channel A after wind noise suppression processing can be output, such as user voice playing or input into a voice recognition model for voice recognition.

[0095] The method provided in the embodiments of the present disclosure can effectively reduce the pressure on the wind noise suppression algorithm and ensure the stability of the result after wind noise suppression. When collecting user voice information through multiple microphones, the target voice with higher signal-to-noise ratio, clarity, and intelligibility can be obtained by applying the wind noise processing scheme provided in the embodiments of the present disclosure.

[0096] The method provided in the embodiments of the present disclosure can be applied to many scenarios, for example, hands-free calling, recording, voice recognition, and the like based on two or more microphones. For example, as shown in FIG. 1, in a double-microphone scenario, the sound signals of the two microphones include the voice of the near-end, environmental noise, and wind noise. In a calling scenario, the sound signals of the two microphones also include echo signals transmitted from the far-end, and the proportion of the echo signals is usually large in the hands-free scenario. The present scheme can be applied to many technical fields of real-time and non-real-time voice enhancement, and is suitable for many terminal forms such as mobile phones, computers, wearable devices, sports cameras, and recording pens. Figure 2 ​

[0097] In some embodiments, for a call type scene, the wind noise detection also needs to consider the influence of the near-end echo, i.e. the far-end voice emitted by the near-end speaker picked up by the near-end microphone. In order to eliminate the echo from the microphone input signal as much as possible while maintaining linearity and not affecting the correlation calculation in the wind noise detection, a related technical solution usually performs adaptive echo cancellation (AEC) on the microphone signal using an echo reference signal before performing wind noise detection. Such algorithm architecture design brings two problems: on the one hand, the wind noise detection algorithm is very dependent on the performance of the AEC, and once the output result of the AEC appears whitening or spectral leakage due to errors, the correlation between the microphone signals will be reduced, which can cause the wind noise false detection rate to increase; on the other hand, the calculation complexity of the AEC algorithm itself is relatively large, and with the increase of the number of microphones and the bandwidth, the calculation complexity of the AEC part will increase exponentially.

[0098] In order to solve the above technical problems, in combination with Figure 1 the embodiments shown, a specific implementation method as shown in Figure 3 is given, which includes the following steps:

[0099] Step 201, before performing adaptive echo cancellation processing, wind noise detection is performed according to the sound signals of multiple channels.

[0100] Step 202, after wind noise detection, adaptive echo cancellation processing is performed on the sound signals.

[0101] In some embodiments, for a call type scene, the embodiments of the present disclosure can first perform wind noise detection according to the sound signals of multiple channels before performing adaptive echo cancellation processing, and then perform adaptive echo cancellation processing on the sound signals after wind noise detection. In this way, the wind noise detection does not depend on the performance of the AEC algorithm when there is echo, reducing the risk of wind noise false detection, while reducing the calculation complexity and power consumption in the wind noise scene.

[0102] For example, the technical solutions of the embodiments of the present disclosure are introduced taking a dual-microphone hands-free call scene as an example. As shown in Figure 4 , two microphone signals MIC0 and MIC1 first pass through a high-pass filter (HPF) to remove direct current and suppress low-frequency noise, and then the echo reference signal (REF) after time alignment is used to perform adaptive filtering on the two HPF output signals, respectively, to eliminate the linear echo part in the two microphone signals, i.e. to perform AEC algorithm processing. Then the two signals after AEC algorithm processing are sent into a beamforming algorithm module (BF) at the same time, and finally one output is obtained. Generally, under a call scene, the BF will also be connected to other algorithm modules, which is not specifically limited by the embodiments of the present disclosure.

[0103] Because the echoes received by MIC0 and MIC1 have a high correlation under hands-free calling, while the correlation between the two channels under the influence of wind noise is low, the two have good distinguishability. Therefore, the wind noise detection module (WND) in this embodiment uses two signals, one before AEC algorithm processing and one after HPF processing, as the input of the algorithm. It detects and outputs a flag (det_flag) indicating whether wind noise exists in the two channels of the current frame. When wind noise is detected, it outputs the channel index (channel_index) indicating that the wind noise pollution meets the preset conditions. The wind noise flag and the channel index are used as the input of the BF module. If no wind noise is detected, the BF combines the output signals of the two AEC algorithm-processed channels into one beam for output. If wind noise is detected, the BF directly outputs the channel whose wind noise pollution meets the preset conditions (i.e., the channel corresponding to the channel index) from the two AEC algorithm-processed output signals, so that wind noise suppression processing is only performed on that channel.

[0104] By applying the technical solution of the present disclosure embodiments, in a call scenario, wind noise detection can be placed before AEC algorithm processing, so that wind noise detection in the presence of echo does not depend on the performance of AEC algorithm, thereby reducing the risk of false wind noise detection. At the same time, in wind noise scenarios, computational complexity and power consumption can be reduced.

[0105] Furthermore, to illustrate the specific process of wind noise detection based on multiple channel sound signals in the embodiments of this disclosure, the following is provided: Figure 5 The specific implementation method shown includes the following steps:

[0106] Step 301: Based on the sound signals from multiple channels, obtain the probability of short-term wind noise in the current frame and the probability of long-term wind noise in the previous frame.

[0107] In this process, the audio signal is divided into consecutive frames. The current frame refers to the latest frame of the signal being processed, while the previous frame refers to the frame of the signal that is adjacent to the current frame.

[0108] The short-term wind noise probability refers to the likelihood of wind noise presence in each frame of the signal; it is calculated independently for each frame and represents a judgment on a short time scale. The long-term wind noise probability, on the other hand, refers to the probability of wind noise presence by incorporating a long-term trend; it is obtained through recursive accumulation across multiple frames, incorporating information from historical frames, and represents a judgment on the presence of wind noise over a longer time scale. In this embodiment, each frame of the signal has its own corresponding short-term and long-term wind noise probability.

[0109] In some embodiments, the cross-correlation coefficient of the amplitude squares in the frequency domain can be analyzed based on the audio signals of multiple channels; then the cross-correlation coefficient can be used to determine the probability of the presence of short-term wind noise in the current frame.

[0110] For example, based on Figure 4 The technical solutions of the embodiments of the present disclosure are introduced by taking a dual-microphone hands-free call scenario as an example, as shown in the figure. First, the amplitude square cross-correlation coefficient (MSC) in the frequency domain is defined, as shown in the following formula:

[0111]

[0112] Wherein, is the frame index, is the frequency point index, is the recursive smoothed self-power spectrum or cross-power spectrum, and the corresponding represents the cross-power spectrum of the two microphone signals x1 and x2 after recursive smoothing, represents the self-power spectrum of the microphone signal x1 after recursive smoothing, represents the self-power spectrum of the microphone signal x2 after recursive smoothing, as shown in the following formula:

[0113]

[0114] and represents the short-time Fourier spectrum of the two MIC signals after HPF filtering, is the smoothing coefficient.

[0115] The (short-time) wind noise existence probability WPP of the current frame is defined, as shown in the following formula:

[0116]

[0117] and are the lower and upper limits of the frequency points in the frequency range intercepted when calculating WPP, and since the main energy of wind noise is concentrated in low frequency, it is recommended that be the corresponding frequency point below 1500Hz. As can be seen from the above formula, the smaller the cross-correlation coefficient MSC between channels, the greater the WPP, that is, the greater the probability of the existence of current wind noise.

[0118] Step 302, according to the short-time wind noise existence probability of the current frame and the long-time wind noise existence probability of the previous frame, smoothing processing is performed to obtain the long-time wind noise existence probability of the current frame.

[0119] In some examples, the long-time wind noise existence probability of the current frame can be obtained by using a smoothing coefficient to perform smoothing processing according to the short-time wind noise existence probability of the current frame and in combination with the long-time wind noise existence probability of the previous frame; wherein the smoothing coefficient is not fixed.

[0120] For example, in order to obtain a more robust wind noise presence probability value, embodiments of this disclosure introduce a long-term wind noise presence probability. :

[0121]

[0122] in, This represents the smoothing coefficient, the smoothing coefficient applied here. It can be non-fixed; its specific value can be determined by some important logical judgments. For example, the smoothing coefficient can be determined based on the probability of long-term wind noise in the previous frame and the probability of short-term wind noise in the current frame. .

[0123] In wind noise detection using multiple microphones, existing techniques employ fixed smoothing coefficients for cross-correlation coefficient smoothing, resulting in poor adaptability to wind noise variations and difficulty in balancing robustness and fast tracking performance. Furthermore, when the signal picked up by the microphones is low, the cross-correlation coefficient between channels decreases significantly, leading to false wind noise detections, such as in quiet rooms or microphone-blocking scenarios. To address this issue, the smoothing coefficient used in this embodiment is flexible, determined by key logical judgments. This embodiment employs a flexibly configurable smoothing coefficient for cross-correlation coefficient smoothing, offering better adaptability to wind noise variations and achieving a better balance between robustness and fast tracking performance. Additionally, threshold design can mitigate false wind noise detections in quiet rooms or microphone-blocking scenarios.

[0124] For example, in some embodiments, if the probability of long-term wind noise in the previous frame is greater than a first probability threshold and the probability of short-term wind noise in the current frame is less than the first probability threshold, then the smoothing coefficient is determined to be a first coefficient value; if the probability of long-term wind noise in the previous frame is less than a second probability threshold and the probability of short-term wind noise in the current frame is greater than the first probability threshold, then the smoothing coefficient is determined to be a second coefficient value, which is less than the first coefficient value, and the second probability threshold is less than the first probability threshold; if none of these conditions are met, then the smoothing coefficient is determined to be a third coefficient value, which is the midpoint between the first coefficient value and the second coefficient value.

[0125] For example, such as Figure 6 As shown, (Condition 1) when the long-term wind noise of the previous frame exists, the probability is... greater than the threshold value thr max (i.e. the first probability threshold value) and the short-time wind noise existence probability WPP of the current frame is less than the threshold value thr max, at this time, the long-time wind noise existence probability is very high, and the short-time wind noise existence probability decreases, in order to ensure the robustness of the subsequent wind noise judgment result and avoid large fluctuations in the judgment result, the long-time wind noise existence probability of the last frame is selected to be more trusted , a larger value beta max (i.e. the first coefficient value) is given to the smoothing coefficient, so that the update speed of the long-time wind noise existence probability is slower.

[0126] (Condition 2) If the above condition 1 is not met, when the long-time wind noise existence probability of the last frame is less than the threshold value thr min (i.e. the second probability threshold value) and the short-time wind noise existence probability WPP of the current frame is greater than the threshold value thr max, at this time, the long-time wind noise existence probability is very low, and the short-time wind noise existence probability is very high, it is considered that the possibility of changing from no wind to wind is very high, and it is necessary to quickly track such changes, so a smaller value beta min (i.e. the second coefficient value) is given to the smoothing coefficient, so that the long-time wind noise existence probability can be updated quickly.

[0127] If the above conditions 1 and 2 are not met, an intermediate value beta mid (i.e. the third coefficient value) is given to the smoothing coefficient, so that the update speed of the long-time wind noise existence probability is moderate. Compared with the smoothing coefficient with a fixed value, the wind noise existence probability calculated by the embodiment of the present disclosure is flexible, and a more accurate and robust wind noise existence probability statistical quantity can be obtained.

[0128] Step 303, determining whether wind noise exists based on the long-time wind noise existence probability of the current frame.

[0129] The long-time wind noise existence probability of the current frame determined based on the above-mentioned manner in the embodiment of the present disclosure can accurately determine whether wind noise exists.

[0130] Under normal circumstances, when wind noise exists, the signal picked up by the microphone will not be very small, and at least should be higher than the normal noise floor, but in environments similar to listening rooms, anechoic chambers and the like, and when the microphone is blocked, the noise floor will be significantly lower, at this time, the correlation between the microphone signals will be significantly reduced, which can lead to a higher probability of misjudging the existence of wind noise. In order to solve this problem, it is necessary to avoid the above-mentioned low noise floor situation.

[0131] Therefore, in some embodiments, the sum of amplitudes corresponding to the multiple channels can be obtained based on the time-domain signal of the current frame; then, based on the sum of amplitudes corresponding to the multiple channels of the previous frame, the sum of amplitudes corresponding to the multiple channels of the current frame can be smoothed; then, the third channel with the smallest sum of amplitudes after smoothing can be obtained; and then, in the process of determining whether wind noise exists based on the long-term wind noise existence probability of the current frame, the existence of wind noise can be determined based on the long-term wind noise existence probability of the current frame and in combination with the sum of amplitudes after smoothing corresponding to the third channel.

[0132] For example, based on Figure 4 As shown, the technical solution of this disclosure embodiment is introduced using a dual-microphone hands-free calling scenario as an example. First, the time-domain signal of the current frame is calculated. The sum of their magnitudes is shown in the following formula:

[0133]

[0134] Where k = 1, 2 are the channel identifiers of the two microphones, respectively. Then, to avoid large fluctuations, sum_amp is smoothed, resulting in the following:

[0135]

[0136] in, This is the smoothing coefficient.

[0137] Will The minimum value is defined as minval (i.e., the third channel), that is...

[0138]

[0139] The corresponding channel index is channel_index = 1 or 2.

[0140] Subsequent frames can be based on the current frame. and To accurately determine whether wind noise exists.

[0141] In some examples, if the probability of long-term wind noise in the current frame is greater than the third probability threshold, and the sum of the amplitudes after smoothing in the second channel is greater than a preset threshold, then wind noise is determined to exist.

[0142] For example, based on Figure 4 As shown, the technical solution of this disclosure embodiment is introduced using a dual-microphone hands-free calling scenario as an example. When long-term wind noise exists... When the flag_thr (i.e., the third probability threshold) is greater than the threshold value, and the minval is greater than the level_thr (i.e., the preset threshold), it is determined that the current frame has wind noise, and a wind noise identifier det_flag = 1 is output; if the condition is not met, it is determined that the current frame does not have wind noise, and a wind noise identifier det_flag = 0 is output, and a current frame channel index channel_index = 0 (not 1 or 2) is output. The design here can avoid the high probability of misjudging the existence of wind noise in the low noise environment, and select the channel with less wind noise, thereby reducing the pressure of subsequent wind noise suppression, and helping to improve the signal-to-noise ratio, clarity and intelligibility of the target speech.

[0143] Next, it is necessary to determine the current frame channel index channel_index when det_flag = 1. In actual application, there may be a problem that channel_index frequently jumps between multiple channel indexes. Therefore, the embodiment of the present disclosure also provides a specific implementation method as shown in Figure 7 The method includes the following steps:

[0144] Step 401: obtaining a second channel with wind noise pollution meeting a preset condition based on historical frames, and a first channel with wind noise pollution meeting a preset condition based on a current frame.

[0145] The first channel is a channel with wind noise pollution meeting a preset condition detected based on the current frame, and the second channel is a channel with wind noise pollution meeting a preset condition corresponding to the wind noise detected in the last time.

[0146] Step 402a: if the second channel is the same as the first channel, performing wind noise suppression processing based on the sound signal of the first channel.

[0147] Step 402b parallel to step 402a: if the second channel is different from the first channel, determining the amplitude difference between the first channel and the second channel based on the time domain signal of the current frame, and selecting the target channel for wind noise suppression processing from the first channel and the second channel according to the amplitude difference.

[0148] In some embodiments, in the case that the amplitude difference between the first channel and the second channel is less than or equal to a preset amplitude difference threshold, wind noise suppression processing is performed based on the sound signal of the second channel.

[0149] In some embodiments, in the case that the amplitude difference between the first channel and the second channel is greater than the preset amplitude difference threshold, wind noise suppression processing is performed based on the sound signal of the first channel.

[0150] For example, based on Figure 4As shown, the technical solutions of the embodiments of the present disclosure are introduced taking a dual-microphone hands-free call scenario as an example. If there is wind noise in the historical frame, i.e., channel_index_pre ≠ 0, when the channel index channel_index of the current frame is inconsistent with the channel index channel_index_pre of the historical frame, it is necessary to determine whether the amplitude difference of the two channels of the current frame satisfies the condition, and the decibel difference of the definition and is as follows:

[0151]

[0152] As shown in Figure 8 , if the decrement is less than or equal to the threshold thr_dB (i.e., the preset amplitude difference threshold), it is determined that the wind noise energy difference of the two channels of the current frame is not large, and the channel_index can not be switched, and continue to maintain the channel index channel_index_pre of the historical frame. If the decrement is greater than the threshold thr_dB, it is determined that the wind noise energy difference of the two channels of the current frame is large, and the channel_index needs to be switched, and the channel_index_pre can be switched to the channel index channel_index of the current frame. Finally, the channel index channel_index of the current frame can be assigned to the channel index channel_index_pre of the historical frame for use in the next frame.

[0153] By applying the technical solutions of the embodiments of the present disclosure, the channel selection and switching method after wind noise detection is given, which can effectively reduce the pressure of the subsequent wind noise suppression algorithm, is conducive to obtaining target speech with higher signal-to-noise ratio, clarity and intelligibility, while avoiding frequent switching of the selected channel, and ensuring the stability of the subsequent output results.

[0154] Figure 9 is a block diagram of a wind noise processing device according to some embodiments of the present disclosure, which can be configured to perform the method shown in Figures 1 to 8 . Referring to Figure 9 , the device includes an acquisition module 51, a selection module 52, and a processing module 53.

[0155] The acquisition module 51 is configured to acquire sound signals of multiple channels collected by multiple microphones;

[0156] The selection module 52 is configured to, in a case where it is determined that there is wind noise according to the sound signals of the multiple channels, select a first channel from the multiple channels, the wind noise pollution of which meets a preset condition;

[0157] The processing module 53 is configured to perform wind noise suppression processing based on the sound signal of the first channel.

[0158] In some embodiments of the present disclosure, the selection module 52 is specifically configured to select a channel with the least wind noise from the plurality of channels as the first channel.

[0159] In some embodiments of the present disclosure, the first channel is a channel with wind noise pollution meeting a preset condition selected based on a current frame; the processing module 53 is specifically configured to, if a second channel with wind noise pollution meeting the preset condition selected based on a historical frame is the same as the first channel, perform wind noise suppression processing based on the sound signal of the first channel; if the second channel is different from the first channel, determine an amplitude difference between the first channel and the second channel based on a time domain signal of the current frame, and select a target channel for wind noise suppression processing from the first channel and the second channel according to the amplitude difference.

[0160] In some embodiments of the present disclosure, the processing module 53 is specifically configured to, in a case where the amplitude difference is less than or equal to a preset amplitude difference threshold, perform wind noise suppression processing based on the sound signal of the second channel; and in a case where the amplitude difference is greater than the preset amplitude difference threshold, perform wind noise suppression processing based on the sound signal of the first channel.

[0161] In some embodiments of the present disclosure, the apparatus further comprises a detection module 54.

[0162] The detection module 54 is configured to perform wind noise detection based on the sound signals of the plurality of channels.

[0163] The processing module 53 is further configured to perform adaptive echo cancellation processing on the sound signal after wind noise detection.

[0164] In some embodiments of the present disclosure, the detection module 54 is specifically configured to obtain a short-time wind noise existence probability of a current frame and a long-time wind noise existence probability of a previous frame based on the sound signals of the plurality of channels; perform smoothing processing to obtain a long-time wind noise existence probability of the current frame according to the short-time wind noise existence probability of the current frame and the long-time wind noise existence probability of the previous frame; and determine whether wind noise exists based on the long-time wind noise existence probability of the current frame.

[0165] In some embodiments of the present disclosure, the detection module 54 is specifically configured to perform smoothing processing to obtain a long-time wind noise existence probability of a current frame by using a smoothing coefficient according to a short-time wind noise existence probability of the current frame and in combination with a long-time wind noise existence probability of a previous frame; wherein the smoothing coefficient is non-fixed.

[0166] In some embodiments of the present disclosure, the detection module 54 is specifically configured to determine the smoothing coefficient based on the long-time wind noise existence probability of the previous frame and the short-time wind noise existence probability of the current frame.

[0167] In some embodiments of the present disclosure, the detection module 54 is specifically configured to determine the smoothing coefficient as a first coefficient value if the long-time wind noise existence probability of the previous frame is greater than a first probability threshold and the short-time wind noise existence probability of the current frame is less than the first probability threshold; determine the smoothing coefficient as a second coefficient value if the long-time wind noise existence probability of the previous frame is less than a second probability threshold and the short-time wind noise existence probability of the current frame is greater than the first probability threshold, the second coefficient value being less than the first coefficient value, and the second probability threshold being less than the first probability threshold.

[0168] If none of the above conditions is met, the smoothing coefficient is determined as a third coefficient value, the third coefficient value being an intermediate value between the first coefficient value and the second coefficient value.

[0169] In some embodiments of the present disclosure, the detection module 54 is specifically configured to obtain the amplitude sum corresponding to each of the plurality of channels based on the time-domain signal of the current frame, perform smoothing processing on the amplitude sum corresponding to each of the plurality of channels of the current frame according to the amplitude sum corresponding to each of the plurality of channels of the previous frame, obtain a third channel with the smallest amplitude sum after smoothing processing, and determine whether wind noise exists based on the long-time wind noise existence probability of the current frame and in combination with the amplitude sum corresponding to the third channel after smoothing processing.

[0170] In some embodiments of the present disclosure, the detection module 54 is specifically configured to determine that wind noise exists if the long-time wind noise existence probability of the current frame is greater than a third probability threshold and the amplitude sum corresponding to the second channel after smoothing processing is greater than a preset threshold.

[0171] In some embodiments of the present disclosure, the detection module 54 is specifically configured to analyze the cross-correlation coefficient of the amplitude square in the frequency domain according to the sound signal of the plurality of channels, and determine the short-time wind noise existence probability of the current frame by using the cross-correlation coefficient.

[0172] As to the apparatus in the above-mentioned embodiments, the specific manners in which various modules perform operations have been described in detail in the embodiments of the method, and will not be described in detail here.

[0173] It should be noted that other corresponding descriptions of the functions of the units involved in the wind noise processing apparatus provided in the embodiments of the present disclosure can be referred to the corresponding descriptions in the Figures 1 to 8 , which will not be described here in detail.

[0174] Based on the method as shown in Figures 1 to 8 , correspondingly, the embodiments of the present disclosure also provide a computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the method as shown in Figures 1 to 8 .

[0175] Based on the above method as Figures 1 to 8 shown, the embodiments of the present disclosure also provide a computer program product, comprising a computer program which, when executed by a processor, implements the above method as Figures 1 to 8 shown.

[0176] Based on such understanding, the technical solutions of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB, a mobile hard disk, etc.) and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods of various implementation scenarios of the present disclosure.

[0177] Based on the above method as Figures 1 to 8 shown, and Figure 9 the virtual device embodiment as shown, the embodiments of the present disclosure also provide a chip, comprising one or more interface circuits and one or more processors; the interface circuit is used to receive a signal from the memory of an electronic device and send the signal to the processor, the signal includes computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device executes the above method as Figures 1 to 8 shown.

[0178] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present disclosure described and / or claimed in this document.

[0179] As Figure 10As shown, the device 1000 includes a computing unit 1001 that can perform various appropriate actions and processes in accordance with a computer program stored in a ROM (Read-Only Memory) 1002 or a computer program loaded into a RAM (Random Access Memory) 1003 from the storage unit 1008. Various programs and data required for the operation of the device 1000 can also be stored in the RAM 1003. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An I / O (Input / Output) interface 1005 is also connected to the bus 1004.

[0180] A plurality of components in the device 1000 are connected to the I / O interface 1005, including an input unit 1006 such as a keyboard, a mouse, and the like, an output unit 1007 such as various types of displays, speakers, and the like, a storage unit 1008 such as a magnetic disk, a magneto-optical disk, and the like, and a communication unit 1009 such as a network card, a modem, a wireless communication transceiver, and the like. The communication unit 1009 allows the device 1000 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0181] The computing unit 1001 can be various general and / or special purpose processing components having processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, and the like. The computing unit 1001 performs various methods and processes described above, such as the aforementioned methods. For example, in some embodiments, the aforementioned methods can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, the computing unit 1001 can be configured to perform the aforementioned methods by any other appropriate means, such as by means of firmware.

[0182] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a Field Programmable Gate Array (FPGA), an Application-Specific Integrated Circuit (ASIC), an Application Specific Standard Product (ASSP), a System on a Chip (SOC), a Complex Programmable Logic Device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0183] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as part of a separate software package, and partially on a remote machine or server.

[0184] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory), or flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0185] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0186] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a blockchain network.

[0187] The computer system can include clients and servers. This relationship can be between a client and a server that are typically remote from each other and typically interact through a communication network. The relationship between client and server exists by virtue of computer programs running on the respective computer systems and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with a blockchain.

[0188] It should be noted that artificial intelligence is a discipline that studies enabling computers to simulate some thinking processes and intelligent behaviors of people (such as learning, reasoning, thinking, planning, etc.), both hardware and software technologies. Artificial intelligence hardware technology generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing, etc.; artificial intelligence software technology mainly includes computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, knowledge graph technology, etc. several major directions.

[0189] It should be understood that the various forms of flow shown above can be used to reorder, add or delete steps. For example, each step described in the present disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, which is not limited herein.

[0190] The above detailed description does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A wind noise processing method, characterized by, The method comprises: acquiring a plurality of channel sound signals collected by a plurality of microphones; in a case where it is determined that wind noise exists according to the plurality of channel sound signals, selecting a first channel from the plurality of channels, wherein wind noise pollution of the first channel meets a preset condition; performing wind noise suppression processing based on the sound signal of the first channel; the first channel is a channel selected based on a current frame, wherein wind noise pollution of the channel meets the preset condition; the wind noise suppression processing based on the sound signal of the first channel comprises: if a second channel selected based on a historical frame, wherein wind noise pollution of the second channel meets the preset condition, is the same as the first channel, then performing wind noise suppression processing based on the sound signal of the first channel; if the second channel is different from the first channel, then determining an amplitude difference between the first channel and the second channel based on a time domain signal of the current frame, and selecting a target channel for performing wind noise suppression processing from the first channel and the second channel according to the amplitude difference.

2. The method of claim 1, wherein, the selection of the first channel from the plurality of channels, wherein wind noise pollution of the first channel meets the preset condition, comprises: selecting a channel with minimum wind noise from the plurality of channels as the first channel.

3. The method of claim 1, wherein, the selection of the target channel for performing wind noise suppression processing from the first channel and the second channel according to the amplitude difference comprises: in a case where the amplitude difference is less than or equal to a preset amplitude difference threshold, performing wind noise suppression processing based on the sound signal of the second channel; in a case where the amplitude difference is greater than the preset amplitude difference threshold, performing wind noise suppression processing based on the sound signal of the first channel.

4. The method according to any one of claims 1 to 3, characterized in that, The method further comprises: performing wind noise detection according to the plurality of channel sound signals; performing adaptive echo cancellation processing on the sound signal after wind noise detection.

5. The method of claim 4, wherein, the wind noise detection according to the plurality of channel sound signals comprises: acquiring a short-time wind noise existence probability of a current frame and a long-time wind noise existence probability of a previous frame according to the plurality of channel sound signals; performing smoothing processing to obtain a long-time wind noise existence probability of the current frame according to the short-time wind noise existence probability of the current frame and the long-time wind noise existence probability of the previous frame; determining whether wind noise exists based on the long-time wind noise existence probability of the current frame.

6. The method of claim 5, wherein, the smoothing processing to obtain the long-time wind noise existence probability of the current frame according to the short-time wind noise existence probability of the current frame and the long-time wind noise existence probability of the previous frame comprises: performing smoothing processing to obtain the long-time wind noise existence probability of the current frame according to the short-time wind noise existence probability of the current frame and the long-time wind noise existence probability of the previous frame by using a smoothing coefficient; wherein the smoothing coefficient is non-fixed.

7. The method of claim 6, wherein, The method further comprises: determining the smoothing coefficient according to the long-time wind noise existence probability of the previous frame and the short-time wind noise existence probability of the current frame.

8. The method of claim 7, wherein, the determination of the smoothing coefficient according to the long-time wind noise existence probability of the previous frame and the short-time wind noise existence probability of the current frame comprises: if the long-time wind noise existence probability of the previous frame is greater than a first probability threshold and the short-time wind noise existence probability of the current frame is less than the first probability threshold, then determining the smoothing coefficient as a first coefficient value. If the long-time wind noise existence probability of the previous frame is less than a second probability threshold, and the short-time wind noise existence probability of the current frame is greater than the first probability threshold, the smoothing coefficient is determined as a second coefficient value, the second coefficient value is less than the first coefficient value, and the second probability threshold is less than the first probability threshold. If the above conditions are not met, the smoothing coefficient is determined as a third coefficient value, the third coefficient value is an intermediate value between the first coefficient value and the second coefficient value.

9. The method of claim 5, wherein, Before determining whether wind noise exists based on the long-time wind noise existence probability of the current frame, the method further comprises: obtaining amplitudes of the multiple channels corresponding to the time domain signal of the current frame; performing smoothing processing on the amplitudes of the multiple channels corresponding to the current frame according to the amplitudes of the multiple channels corresponding to the previous frame; obtaining a third channel with the smallest amplitude after smoothing processing; determining whether wind noise exists based on the long-time wind noise existence probability of the current frame, comprising: determining whether wind noise exists based on the long-time wind noise existence probability of the current frame and the amplitude of the third channel after smoothing processing.

10. The method of claim 9, wherein, determining whether wind noise exists based on the long-time wind noise existence probability of the current frame and the amplitude of the third channel after smoothing processing, comprising: if the long-time wind noise existence probability of the current frame is greater than a third probability threshold, and the amplitude of the second channel after smoothing processing is greater than a preset threshold, it is determined that wind noise exists.

11. The method of claim 5, wherein, According to the sound signals of the multiple channels, the short-time wind noise existence probability of the current frame is obtained, comprising: analyzing the cross-correlation coefficient of the amplitude square in the frequency domain according to the sound signals of the multiple channels; determining the short-time wind noise existence probability of the current frame by using the cross-correlation coefficient.

12. A wind noise processing apparatus, characterized by, comprising: an obtaining module configured to obtain sound signals of multiple channels collected by multiple microphones; a selecting module configured to, in a case where it is determined that wind noise exists according to the sound signals of the multiple channels, select a first channel with wind noise pollution meeting a preset condition from the multiple channels; a processing module configured to perform wind noise suppression processing based on the sound signals of the first channel; the first channel is a channel with wind noise pollution meeting the preset condition selected based on the current frame; the processing module is configured to, if a second channel with wind noise pollution meeting the preset condition selected based on a historical frame is the same as the first channel, perform wind noise suppression processing based on the sound signals of the first channel; if the second channel is different from the first channel, determine an amplitude difference between the first channel and the second channel based on the time domain signal of the current frame, and select a target channel for wind noise suppression processing from the first channel and the second channel according to the amplitude difference.

13. The apparatus of claim 12, wherein the selecting module is specifically configured to select a channel with the minimum wind noise from the multiple channels as the first channel.

14. The apparatus of claim 12, wherein, The apparatus further comprises a detecting module. The detecting module is configured to perform wind noise detection according to the sound signals of the multiple channels. The processing module is further configured to perform adaptive echo cancellation processing on the sound signals after wind noise detection.

15. The apparatus of claim 14, wherein the detection module is configured to obtain a short-time wind noise existence probability of a current frame and a long-time wind noise existence probability of a previous frame according to sound signals of the multiple channels, perform smoothing processing on the long-time wind noise existence probability of the current frame according to the short-time wind noise existence probability of the current frame and the long-time wind noise existence probability of the previous frame, and determine whether wind noise exists based on the long-time wind noise existence probability of the current frame.

16. The apparatus of claim 15, wherein the detection module is specifically configured to perform smoothing processing on the long-time wind noise existence probability of the current frame according to the short-time wind noise existence probability of the current frame and the long-time wind noise existence probability of the previous frame by using a smoothing coefficient, and wherein the smoothing coefficient is non-fixed.

17. The apparatus of claim 16, wherein the detection module is further configured to determine the smoothing coefficient as a first coefficient value if the long-time wind noise existence probability of the previous frame is greater than a first probability threshold and the short-time wind noise existence probability of the current frame is less than the first probability threshold, determine the smoothing coefficient as a second coefficient value if the long-time wind noise existence probability of the previous frame is less than a second probability threshold and the short-time wind noise existence probability of the current frame is greater than the first probability threshold, and wherein the second coefficient value is less than the first coefficient value and the second probability threshold is less than the first probability threshold, and determine the smoothing coefficient as a third coefficient value if neither of the above conditions is met, and wherein the third coefficient value is an intermediate value between the first coefficient value and the second coefficient value.

18. The apparatus of claim 14, wherein the detection module is specifically configured to obtain amplitude sums corresponding to the multiple channels respectively based on a time domain signal of the current frame, perform smoothing processing on the amplitude sums corresponding to the multiple channels of the current frame according to amplitude sums corresponding to the multiple channels of the previous frame, obtain a third channel with a smallest amplitude sum after the smoothing processing, and determine whether wind noise exists based on the long-time wind noise existence probability of the current frame and the amplitude sum corresponding to the third channel after the smoothing processing.

19. The apparatus of claim 18, wherein the detection module is specifically configured to determine that wind noise exists if the long-time wind noise existence probability of the current frame is greater than a third probability threshold and the amplitude sum corresponding to the second channel after the smoothing processing is greater than a preset threshold. An apparatus includes: a processor; a memory connected with the processor, the memory having stored thereon a computer program, the computer program being executed by the processor to implement the method of any one of claims 1 to 11. The computer program, when executed by a processor, implements the method of any one of claims 1 to 11. The computer program, when executed by a processor, implements the method of any one of claims 1 to 11. ​ ​ 20. An electronic device, comprising: ​ ​ ​ 21. A computer readable storage medium having stored thereon a computer program, characterized in that, ​ 22. A computer program product comprising a computer program, characterized in that, ​ 23. A chip, characterized by An electronic device comprising one or more interface circuits and one or more processors; the interface circuit is configured to receive a signal from a memory of the electronic device, and send the signal to the processor, the signal comprising computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device is caused to perform the method of any one of claims 1 to 11.

Citation Information

Patent Citations

  • Wind noise pollution degree estimation method, wind noise suppression method, medium and terminal

    CN115691533A

  • Wind noise filtering method and device suitable for multiple microphones

    CN118890583A