Enhancing audio signals

By applying filter coefficients based on noise covariance and aggressiveness parameters to multiple audio channels, the method addresses the challenge of wind noise reduction in audio recordings, preserving spatial and spectral characteristics and enhancing the audio experience.

WO2025160096A1PCT designated stage expired Publication Date: 2025-07-31DOLBY LABORATORIES LICENSING CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/012473
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-26
Filing Date
2025-01-21
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing audio recording technologies struggle to effectively reduce wind noise while preserving the spatial and spectral characteristics of the audio recording, with conventional methods like the minimum variance distortionless response (MVDR) algorithm producing mono outputs that lose important audio characteristics.

Method used

A method involving multiple microphones to capture audio signals, applying filter coefficients based on noise covariance and aggressiveness parameters to reduce wind noise on a per-channel basis, allowing for dynamic filtering and mixing to preserve spatial characteristics, and optionally upmixing the signals to enhance the audio experience.

Benefits of technology

The method effectively reduces wind noise in multiple audio channels while maintaining spatial and spectral characteristics, enabling accurate reconstruction of the original audio scene and reducing artifacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025012473_31072025_PF_FP_ABST
    Figure US2025012473_31072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein are techniques for enhancing audio signals. In some embodiments, the techniques may involve obtaining multiple audio signals associated with multiple microphones of one or more devices. The techniques may further involve responsive to detecting wind in at least one audio signal of the multiple audio signals, filtering the multiple audio signals to generate multiple filtered audio signals by filtering each audio signal based on filter coefficients determined for each audio signal that weights each audio signal based on a noise covariance associated with that audio signal. The techniques may further involve mixing the multiple filtered audio signals to generate a multi-channel wind-reduced audio signal, wherein each multi-channel wind-reduced audio signal comprises a summation of the multiple filtered audio signals.
Need to check novelty before this filing date? Find Prior Art

Description

ENHANCING AUDIO SIGNALSCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority from U.S. Provisional Application No. 63 / 623484, filed on 22 January 2024, and European Patent Application No. 24172718.9, filed 26 April 2024, each of which is incorporated by reference herein in its entirety.TECHNICAL FIELD

[0002] This disclosure pertains to systems, methods, and media for enhancing audio signals.BACKGROUND

[0003] Users, e.g., content creators, are interested in reducing wind noise in recorded audio content. For example, audio content recorded outdoors may include various degrees of wind noise which may interfere with the quality of the audio recording. However, it can be difficult to reduce wind noise in a manner that preserves various characteristics of the audio recording and remains faithful to the recorder’s intent.NOTATION AND NOMENCLATURE

[0004] Throughout this disclosure, including in the claims, the terms “speaker,” “loudspeaker” and “audio reproduction transducer” are used synonymously to denote any sound-emitting transducer (or set of transducers). A typical set of headphones includes two speakers. A speaker may be implemented to include multiple transducers (e.g., a woofer and a tweeter), which may be driven by a single, common speaker feed or multiple speaker feeds. In some examples, the speaker feed(s) may undergo different processing in different circuitry branches coupled to the different transducers.

[0005] Throughout this disclosure, including in the claims, the expression performing an operation “on” a signal or data (e.g., filtering, scaling, transforming, or applying gain to, the signal or data) is used in a broad sense to denote performing the operation directly on the signal or data, or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or pre-processing prior to performance of the operation thereon).

[0006] Throughout this disclosure including in the claims, the expression “system” is used in a broad sense to denote a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system including such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, in which the subsystem generates M of the inputs and the other X - M inputs are received from an external source) may also be referred to as a decoder system.

[0007] Throughout this disclosure including in the claims, the term “processor” is used in a broad sense to denote a system or device programmable or otherwise configurable (e.g., with software or firmware) to perform operations on data (e.g., audio, or video or other image data). Examples of processors include a field-programmable gate array (or other configurable integrated circuit or chip set), a digital signal processor programmed and / or otherwise configured to perform pipelined processing on audio or other sound data, a programmable general purpose processor or computer, and a programmable microprocessor chip or chip set.SUMMARY

[0008] Disclosed herein are systems, methods, and media for enhancing audio signals. In some embodiments, a method may involve obtaining multiple audio signals associated with multiple microphones of one or more devices. The method may further involve responsive to detecting wind in at least one audio signal of the multiple audio signals, filtering the multiple audio signals to generate multiple filtered audio signals by filtering each audio signal based on filter coefficients determined for each audio signal that weights each audio signal based on a noise covariance associated with that audio signal. The method may further involve mixing the multiple filtered audio signals to generate a multi-channel wind-reduced audio signal, wherein each multi-channel wind-reduced audio signal comprises a summation of the multiple filtered audio signals.

[0009] In some examples, weighting is determined on a frame-by-frame basis based on frequency information associated with each audio signal.

[0010] In some examples, the filter coefficients represent an instantaneous complex gain to be applied to each audio signal on a per- frequency band basis.

[0011] In some examples, the filter coefficients cause each audio signal to be filtered in a manner that is optimized for each audio signal of the multiple audio signals.

[0012] In some examples, the method may further involve upmixing the multi-channel wind- reduced audio signal.

[0013] In some examples, weighting each audio signal comprises determining a bias term for each audio signal that is determined on a per- frequency band basis. In some examples, the bias term is based at least in part on an aggressiveness parameter that represents a degree to which wind noise is to be reduced. In some examples, the bias term for a given audio signal is based at least in part on a metric that indicates a diffuseness associated with the given audio signal. In some examples, the metric that indicates a diffuseness associated with the given audio signal is normalized based on level. In some examples, a value of the bias term for a given audio signal increases with a value of the center of the frequency band.

[0014] In some examples, the method further involves determining a frequency threshold above which the multiple filtered audio signals are not mixed, wherein the frequency threshold is determined based on level differences at multiple frequency bands between the multiple audio signals.

[0015] In some examples, the method further involves upmixing the multi-channel wind- reduced audio signal; causing the upmixed signal to be rendered to generate a rendered signal; and causing the rendered signal to be played back.

[0016] Some or all of the operations, functions and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. Accordingly, some innovative aspects of the subject matter described in this disclosure can be implemented via one or more non-transitory media having software stored thereon.

[0017] At least some aspects of the present disclosure may be implemented via an apparatus. For example, one or more devices may be capable of performing, at least in part, the methods disclosed herein. In some implementations, an apparatus is, or includes, an audio processing system having an interface system and a control system. The control system may include one or more general purpose single- or multi-chip processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or combinations thereof.

[0018] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, the drawings, and the claims. Note that the relative dimensions of the following figures may not be drawn to scale.BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figures 1A, IB, and 1C are a diagram illustrating an example audio capture device in accordance with some embodiments.

[0020] Figure 2 is a schematic diagram of a system for reducing wind noise in accordance with some embodiments.

[0021] Figure 3 is a flowchart of an example process for reducing wind noise in accordance with some embodiments.

[0022] Figure 4 is a graph that illustrates the effect of a tuning parameter for reducing wind noise in accordance with some embodiments.

[0023] Figure 5 is a diagram that illustrates a device shadowing effect in accordance with some embodiments.

[0024] Figure 6 is a graph that illustrates the effect of an early mixing stop in accordance with some embodiments.

[0025] Figure 7 shows a block diagram that illustrates examples of components of an apparatus capable of implementing various aspects of this disclosure.

[0026] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION OF EMBODIMENTS

[0027] Recorded audio content, such as that recorded using audio devices such as mobile phones, tablet computers, cameras, etc. may include wind noise. In general, as used herein, “wind noise” may refer to “wind-induced microphone noise,” e.g., the noise captured by a microphone due to the presence of wind or other air movement. Wind noise can cause determinantal effects on the recorded audio content. For example, for a mobile device recording audio content without a physical wind screen protection cover or component, even a gentle breeze may overpower speech and other sounds that are intended to be captured.

[0028] The minimum variance distortionless response (MVDR) algorithm is a beamforming technique that has been conventionally used to reduce wind noise. However, the MVDR algorithm produces a mono output (e.g., a single channel output). A mono output is not suitable for rendering spatial characteristics of recorded audio sound, e.g., as a rendered spatial scene. Moreover, a mono output is not suitable for use with subsequent upmixing algorithms that may generate upmixed audio signals used to render such an audio spatial scene. In other words, the conventional MVDR algorithm, while reducing wind noise to at least some degree, causes important audio characteristics that yield a spatial perception in the rendered and played back audio content to be lost.

[0029] Disclosed herein are techniques for performing wind noise reduction on multiple audio signals in a manner such that wind noise is reduced for each audio signal (e.g., each audio signalcorresponding to a different microphone) to preserve multiple audio signals while reducing wind noise. Using the techniques disclosed herein, an n audio signals may be captured from n microphones, and wind noise may be reduced on the n audio signals to generate n enhanced audio signals rather than a single mono signal or channel. These n enhanced audio signals may contain spatial characteristics, and optionally may be upmixed to form n+m audio signals.

[0030] The multiple audio signals may be filtered to generate a set of wind-reduced audio signals by filtering each audio signal using filter coefficients that weight each audio signal based on a noise covariance associated with that audio signal. The weights, sometimes referred to herein as “bias terms,” may be frequency-dependent, allowing the audio signals to not only be filtered without aggregating the audio signals to a mono output, but may additionally be filtered in a frequency dependent manner, allowing the frequency-dependent nature of wind noise to be considered in the wind noise reduction on a per-channel basis. The weights, or bias terms, may additionally consider additional factors, such as an aggressiveness parameter that indicates a degree to which wind noise is to be reduced, a purity metric that indicates a diffuseness of the signal, or the like. The weights, or bias terms, may be determined on a frame-by-frame basis, thereby allowing the multiple audio signals to be filtered in a dynamic manner that optimizes for wind noise reduction based on the audio content of each signal and frequency information.

[0031] Note that, after filtering each input audio signal to generate multiple filtered audio signals, the multiple filtered audio signals may be mixed to generate the multiple wind-reduced audio signals. In particular, each wind-reduced audio signal may be a summed version of the filtered audio signals from each of the microphones. By way of example, in an instance in which there is no wind present on any input signals from any microphone, filtering will result in passing through each input audio signal such that each of the wind-reduced audio signals substantially correspond to the corresponding untouched audio signal. In contrast, in an instance in which there is wind present for one microphone and the corresponding input audio signal, the audio signal with wind present may be filtered such that the input audio signal for that microphone is replaced with a mix of the other remaining (i.e., non-windy input audio signals). In other words, in some embodiments, each wind-reduced audio signal may be a filtered and summed result (e.g., a mix) of all of the input audio signals, where the filtration is dynamically performed, e.g., based on the weights or bias terms as described above.

[0032] Figure 1A illustrates an example audio capture device in accordance with some embodiments. In particular, the audio capture device is a mobile phone 100. Front side 102 ofmobile phone 100 is associated with two microphones, top microphone 106 and bottom microphone 108. Back side 104 of mobile phone 100 includes a third microphone, back microphone 110. Note that each microphone may be configured to capture audio content from a direction associated with the location of the microphone with respect to mobile phone 100. Additionally, as described below, the relative locations of each microphone may be used to determine noise covariances and filter a multi-microphone audio signal based on aggregated noise covariances. Figure IB illustrates a side view of mobile phone 100 that includes top microphone 106, and Figure 1C illustrates a side view of mobile phone 100 that includes bottom microphone 108.

[0033] In some embodiments, wind noise may be reduced on a multi-channel set of audio signals such that each of the audio signals are filtered and wind noise reduction results in a filtered set of the multiple audio signals, rather than a single (e.g., mono) audio signal. In some embodiments, filtering may be performed using a modified minimum variance distortionless response (MVDR) algorithm where filter coefficients are determined separately for each audio signal in the multichannel set of audio signals by determining weights for each audio signal. The weights may be determined on a per-frequency band basis. The weights for a given audio signal may be determined based on a noise covariance associated with that audio signal. In some embodiments, the weights may be determined based on a bias term for each audio signal. The bias term may be dependent on an aggressiveness parameter that represents a degree to which wind noise is to be reduced, based at least in part on a purity metric that indicates a degree of diffuseness associated with the audio signal, or any combination thereof. Using the techniques disclosed herein, wind reduction may be performed dynamically and optimized based on characteristics of each audio signal of a set of multi-channel audio signals. By performing wind noise reduction on a multi-channel audio signal in a manner in which each audio signal is preserved rather than being combined to a single mono signal, the multi-channel audio signal may be rendered and played back in a manner that preserves spatial characteristics, spectral characteristics, or the like.

[0034] Figure 2 is a schematic diagram of an example system for performing wind noise reduction in accordance with some embodiments. As illustrated, raw microphone signals are provided to a wind noise reduction block 202. Note that although three microphone signals are shown in Figure 2, in some embodiments, the number of microphone signals may be two, four, five, etc. Additionally, it should be understood that the microphone signals may be obtained from multiple microphones of a single device (e.g. a single mobile phone, as shown in and describedabove in connection with Figures 1A-1C), or from multiple microphones of multiple devices (e.g., a mobile phone and a headset, a mobile phone and a tablet computer, etc.).

[0035] Wind noise reduction block 202 may receive the raw microphone signals as well as a wind detection flag, which may be obtained from wind noise detection block 204. As illustrated, wind noise detection block 204 may receive the raw microphone signals and may be configured to generate, as an output, a flag indicative of the presence of wind in the raw microphone signals. For example, in some embodiments, the flag may be a number (e.g., between 0 and 1, between 0 and 10, between 0 and 100, between 1 and 100, etc.) indicating a likelihood that wind is present. Note that the flag may vary for different frequency bands (e.g., the flag may be set on a per- frequency band basis). In some embodiments, wind noise reduction block 202 may implement wind noise reduction techniques responsive to the flag passed from wind noise detection block 204 indicating that wind noise is present in the raw microphone signals. Responsive to the flag indicating that wind noise is not present in the raw microphone signals, in some embodiments, wind noise reduction block 202 may pass the raw microphone signals untouched to a next stage of the system (e.g., upmix block 206).

[0036] As described below in more detail in connection with Figure 3, wind noise reduction block 202 may be configured to determine filter coefficients to be used to filter the microphone signals to reduce wind noise. Note that the output of wind noise reduction block 202 is a multichannel audio signal, e.g., the same number of audio signals as those input to wind noise reduction block 202. Unlike conventional wind noise reduction techniques which take, as input, multiple audio signals and generate a single mono output, wind noise reduction block 202 filters each audio signal to generate a set of multiple, enhanced audio signals with wind noise reduced individually for each audio signal in the set. In some embodiments, wind noise reduction block 202 may determine the filter coefficients for each audio signal based on a noise covariance associated with each audio signal. In some embodiments, the filter coefficients may be determined based on a weighting that is dependent on an aggressiveness parameter that represents a degree to which wind noise is to be reduced, a purity metric that indicates diffuseness of the signal, or any combination therein. Note that the output of block 202 is sometimes referred to herein as “a multi-channel wind-reduced audio signal,” and may represent filtered (e.g., based on the determined filter coefficients) and mixed multiple audio signals such that each audio signal included in the multichannel wind-reduced audio signal is a filtered and summed version of the audio signals.

[0037] The enhanced audio signals generated by filtering the multiple microphone signals based on the filter coefficients determined by wind noise reduction block 202 may be provided to optional upmix block 206. Note that, because upmixing recreates an audible scene using some combination of the microphone signals, maintaining the multiple audio signals prior to upmixing (rather than generating a single mono signal as a result of wind noise reduction) enables spatial scene reconstruction that is more faithful to the original spatial scene of the microphone signals compared to conventional techniques. Note that, in some embodiments, the multiple microphone signals may correspond to multiple audio signals of a spatially encoded set of audio signals, such as left and right audio signals for a stereo cardioid microphone, Ambisonic signals associated with an Ambisonic microphone array, or the like. In such embodiments, upmix block 202 may be omitted because no upmixing is needed prior to rendering the audio signals. However, in embodiments, in which the multiple audio signals are to be upmixed, the enhanced audio signals may be provided to upmix block 206 which may perform upmixing to generate any suitable number of a set of upmixed, spatially encoded set of audio signals. These upmixed signals may then be rendered and played back.

[0038] Figure 3 is a flowchart of an example process 300 for performing wind noise reduction in accordance with some embodiments. In some embodiments, blocks of process 300 may be performed by one or more processors or one or more control systems of one or more devices, such as a mobile phone, a tablet computer, a desktop computer, a laptop computer, a camera, or the like. An example of a control system of a device is shown in and described below in connection with control system 710 of Figure 7. In some embodiments, blocks of process 300 may be performed in an order other than what is shown in Figure 3. In some implementations, two or more blocks of process 300 may be performed substantially in parallel. In some embodiments, one or more blocks of process 300 may be omitted.

[0039] Process 300 can begin at 302 by obtaining multiple audio signals from multiple microphones of one or more devices. An example of a device and the multiple microphones is shown in and described above in connection with Figures 1 A- 1C. Note that, in some embodiments, the multiple audio signals may be obtained from multiple microphones of multiple devices, such as a mobile phone and a headset, a mobile phone and a camera, two mobile phones, etc. Note that, in some embodiments, process 300 may transform the multiple audio signals from a time domain representation to a frequency domain representation such that one or more other blocks of process 300 are performed in the frequency domain.

[0040] At 304, process 300 can determine whether wind noise is detected in any of the multiple audio signals. In some embodiments, wind noise detection may be performed on a per-frequency band basis such that a wind detection algorithm is applied at different frequency bands to each audio signal. An example of a wind detection block is shown in and described above in connection with Figure 2. In some embodiments, wind noise may be detected based on inter-microphone information that indicates, e.g., differences between the audio signals obtained via the different microphones (e.g., level differences at particular frequency bands, correlations between the signals at different frequency bands, or the like).

[0041] If, at 304, process 300 determines that wind noise is not detected in the multiple audio signals (“no” at 304), process 300 can pass the multiple audio signals without performing any wind noise reduction. For example, as shown in Figure 3, process 300 can proceed to optional block 308 and can upmix the multiple audio signals to generate upmixed signals. Alternatively, in an instance in which upmixing is not performed, process 300 can proceed to render the multiple audio signals, store the multiple audio signals, etc.

[0042] Conversely, if, at 304, process 300 determines that wind is detected in at least one audio signal of the multiple audio signals (“yes” at 304), process 300 can proceed to block 306 and can filter the multiple audio signals to generate a set of wind-reduced audio signals by filtering each audio signal based on filter coefficients determined for each audio signal. The filter coefficients may be determined by weighting each audio signal based on a noise covariance associated with that audio signal. Note that the filter coefficients described herein may be determined based on a variation of the MVDR algorithm using the optimization techniques described below. In general, the variation of the MVDR algorithm described herein produces filter coefficients, which, when applied, generate the equivalent of filtering and mixing multiple (e.g., two, three, five, etc.) microphone signals together.

[0043] By way of example, filter coefficients for a two-microphone signal may be determined by:

[0044] In the equation given above, Gi represents filter coefficients to be applied to the first audio signal, < >NI represents the noise covariance of the first audio signal, and < >N2 represents the noise covariance of the second audio signal. In other words, the filter coefficients may bedetermined, for a given audio signal, by weighting the noise covariances of the multiple audio signals based on a weighting factor, represented herein as w. Note that the equation given above may be extended to a system with more than two microphone signals, e.g., by mixing the multiple microphone signals weighted inversely by noise power.

[0045] The weighting factor, w, is sometimes referred to herein as a “bias term.” The weighting factor may be any suitable value greater than or equal to 1, such as 1, 1.5, 10, 50, 1000, etc. In some embodiments, the weighting factor may be determined on a per-frequency band basis, e.g., such that the weighting factor for a given audio signal varies as a function of frequency. The weighting factor may be dependent on: the center frequency of a given frequency band, a tuning parameter that controls the aggressiveness of the wind noise reduction, a purity metric that indicates on a per-frequency band basis the degree of diffuseness of the audio signal, and / or the wind level associated with the audio signal. In some embodiments, the aggressiveness tuning parameter may be a value from 0 to 1 (e.g., 0, 0.2, 0.5, 0.8, 1, etc.), and may be determined based on user input, may be application-dependent, and / or may be a pre-configured setting. In some embodiments, the purity metric may be a normalized metric. More detailed example techniques for determining a purity metric are described below. In some embodiments, wind level may be determined across all frequency bands (e.g., as a full-band measurement). The wind level may be a value between 0 and 1, where 0 indicates no wind.

[0046] By way of example, in some embodiments, the weighting factor for a given audio signal and at a given frequency band may be determined by: w = 10.0

[0047] In the equation given above, “fband” represents the frequency of the center of a given frequency band, “aggressiveness” represents the aggressiveness parameter, “purity” represents the purity metric indicative of a degree of diffuseness of the signal, and “wind level” represents the level of wind determined to be in the audio signal.

[0048] Turning to Figure 4, a graph of the weighting factor, or bias term, as a function of frequency band (e.g., the center frequency of the fband variable in the equation given above) is shown in accordance with some embodiments. In the graph of Figure 4, the y-axis represents the value of the weighting factor (represented in the equation above as “w”), or bias term, and the x- axis indicates frequency. Each curve depicted in the graph of Figure 4 represents a different value of the aggressiveness parameter. For example, for relatively low values of the aggressivenessparameter, as depicted by curve 402, the weighting factor will be relatively high above a certain frequency (e.g., above 400 Hz, or about 1000 Hz), and will thereby cause the signal to be generally passed through above this frequency without wind reduction. Conversely, for relatively higher values of the aggressiveness parameter, as depicted by curve 404, the weighting factor will be relatively low. For example, note that for every frequency band, the weighting factor is lower for curve 404 than for curve 402, which will in turn cause a higher degree of wind noise reduction at each frequency band compared to the weighting factors depicted in curve 402.

[0049] Referring back to Figure 3, after determining the weighting factor, the filter coefficients may be determined, and the audio signals may be filtered using the filter coefficients to generate the wind-reduced audio signals.

[0050] Note that, after filtering each input audio signal to generate multiple filtered audio signals, the multiple filtered audio signals may be mixed to generate the multiple wind-reduced audio signals. In particular, each wind-reduced audio signal may be a summed version of the filtered audio signals from each of the microphones. By way of example, in an instance in which there is no wind present on any input signals from any microphone, filtering will result in passing through each input audio signal such that each of the wind-reduced audio signals substantially correspond to the corresponding untouched audio signal. In contrast, in an instance in which there is wind present for one microphone and the corresponding input audio signal, the audio signal with wind present may be filtered such that the input audio signal for that microphone is replaced with a mix of the other remaining (i.e., non- windy input audio signals). Note that, the filtered and mixed audio signals are sometimes referred to herein as “a multi-channel wind-reduced audio signal.”

[0051] At 308, process 300 can optionally upmix the wind-reduced audio signals (e.g., the multichannel wind-reduced audio signal) to generate upmixed audio signals. Any suitable upmixing technique(s) may be used. Note that, because wind reduction is performed on all audio signals without aggregating the multiple audio signals to a single mono signal, upmixing may allow for reconstruction of a spatial scene with increased accuracy and reduced artifacts relative to conventional techniques.

[0052] After upmixing the signal, in some embodiments, the upmixed audio signals may be rendered. The rendered signals may be played back and / or stored.

[0053] Note that, in instances in which upmixing is not performed (e.g., due to the audio signals being left / right stereo signals of a set of cardioid microphones and / or Ambisonic signals associatedwith an Ambisonic microphone array), block 308 may be omitted, and the wind-reduced signals may optionally be directly rendered without upmixing.

[0054] As described above, the purity metric may indicate or represent a degree of diffuseness of the audio signal. The purity metric may be dependent on the distribution of the eigenvalues of the covariance matrix. For example, the purity metric, generally represented herein as y, may be conventionally determined by:

[0055] In the equation above, represents the ntheigenvalue of the covariance matrix having dimensions n x n.

[0056] However, determining the purity metric based directly on the eigenvalues of the covariance matrix may, in some cases, cause issues for some covariance matrices, particularly in cases with wind noise in the audio signals. In particular, as described above, the purity metric generally indicates a degree of diffuseness of the audio signal, however, in instances of high levels of wind noise and / or an occluded microphone, there may be a signal that is highly directional but that has a low purity metric, or vice versa. By way of example, consider a covariance matrix such as: r0.99 0 \V 0 0.01 /

[0057] Given the above covariance matrix, the purity value may be relatively high considering the sum of the diagonal values of 0.99 and 0.01. The high purity value would generally indicate a highly directional audio signal. However, the correlation between the signals is relatively low, indicating that the audio signal is not highly directional. Accordingly, for such cases, the purity metric as conventionally determined based on the covariance matrix may not be suitable for audio signals containing wind noise. It should be noted that, in general, purity may be determined by the Frobenius norm of the matrix (e.g., the square root of the sum of squares of all elements). In the example given above, the purity is 0.9802, despite a low correlation between channels.

[0058] In some embodiments, the purity metric as used herein (e.g., as described above in connection with Figure 3) may utilize a normalized covariance matrix where the normalized covariance matrix has been scaled such that the trace is 1. In other words, the covariance matrixmay be normalized with respect to amplitude such that the sum of the diagonal elements is 1. The normalized covariance matrix, for a covariance matrix represented by C may be determined by:C = C . / sqrt diagonal C) * transpose diagonal C))

[0059] In the equation given above, the output of diagonal(C) is a column vector whose elements are the diagonal of C, “transpose ” is a matrix transpose, andrepresents an element wise division operation. Division by the diagonal of the covariance matrix may allow for the low correlation of the audio signals to be accounted for in the calculation of the purity metric. In other words, the normalized covariance matrix may no longer be impacted by overall power differences between channels.

[0060] In some instances, the techniques disclosed herein may reduce or mitigate a device shadowing effect. For example, with a device shadowing effect, microphone response may become more directional at higher frequencies. Figure 5 illustrates the device shadowing effect. Diagram 502 illustrates microphone signal levels at 500 Hz, a relatively low frequency. Note that there is little level difference across the different microphone signals, depicting an omni, or diffuse signal rather than a highly directional signal. Conversely, diagram 504 illustrates microphone signal levels at 4000 Hz. Due to the relatively higher frequency, there are level differences of up to 10 dB present across the different microphones. Conventional techniques which minimize output power in each sub-band may cause the enhanced signal (e.g., with wind noise reduced) to sound low-pass filtered due to this device shadowing effect. However, using the techniques disclosed herein (e.g., as described above in connection with Figures 2 and 3), the device shadowing effect and the low-pass filtered sound quality that results due to the device shadowing effect, may be mitigated by dynamically determining filter coefficients for each frequency band individually and for each audio signal of a multi-microphone set of audio signals individually. In particular, the device shadowing effect may be mitigated through control of the frequencydependent bias term (e.g., as shown in and described in connection with Figure 4). Moreover, due to normalization of the powers of the microphone signals in instances in which a normalized covariance matrix is used to calculate the purity metric, the device shadowing effect may be less prominent.

[0061] As described above, the filter coefficients determined, when applied, may be equivalent to filtering and mixing multiple microphone signals together. In some embodiments, the techniques disclosed herein may allow for a dynamic determination of a frequency at which mixingthe multiple audio signals is to be terminated. In particular, conventional techniques typically utilize a fixed frequency threshold at which mixing is not performed. However, the frequency at which wind noise is prevalent in the signal varies widely from case to case. In some embodiments, a frequency threshold above which mixing is not performed may be dynamically determined (e.g., on a frame by frame basis for each frame of the audio signal). In some embodiments, the frequency threshold may be determined based on a level difference between the multiple audio signals of the set of microphone signals on a per-frequency band basis. By way of example, in a case of relatively low wind noise, there may be relatively little level difference between the audio signals at frequencies as low as 1000 Hz, whereas in a case of relatively high wind noise, there may be substantial level differences between the audio signals at much higher frequencies, e.g., up to 2 KHz, up to 4 KHz, up to 5 KHz, etc. Continuing with this example, by utilizing sub-band level differences to determine the frequency threshold at which mixing is to stop, mixing may be performed on a dynamic basis, which may in turn reduce artifacts which may be perceptible by listeners. Because the bias term described herein is frequency dependent, the frequency threshold at which mixing stops may be dynamically determined, and the bias term may effectively prevent mixing above the dynamically-determined frequency threshold by forcing mixing weight parameters to be low, thereby preventing mixing of microphone signals and allowing a single microphone signal to be passed through without mixing. Unlike conventional techniques which may yield a sudden step-like change in response based on the pre-selected frequency threshold regardless of wind noise, the techniques disclosed herein may allow for a smooth rolloff in mixing even at higher mixing frequencies.

[0062] Figure 6 is a graph of sound level for each of three microphones of a multi-microphone audio signal as a function of frequency. For example, each of curves 602, 604, and 606 illustrate the level as a function of frequency for a different microphone. Note that below about 3000 or 4000 Hz, there are large level differences across the different microphones, and that the level differences reduce considerably above 4000 Hz. As described above, a frequency threshold at which mixing is to stop may be dynamically determined (e.g., on a frame-by-frame basis) based on the level differences as a function of frequency. For example, referring to the example shown in Figure 6, mixing may stop at, e.g., 3500 Hz, 4000 Hz, 5000 Hz, or other frequency in the range of 3000 Hz - 5500 Hz due to the reduction in level differences across the microphone signals within that frequency range.

[0063] Figure 7 is a block diagram that shows examples of components of an apparatus capable of implementing various aspects of this disclosure. As with other figures provided herein, thetypes and numbers of elements shown in Figure 7 are merely provided by way of example. Other implementations may include more, fewer and / or different types and numbers of elements. According to some examples, the apparatus 700 may be configured for performing at least some of the methods disclosed herein. In some implementations, the apparatus 700 may be, or may include, a television, one or more components of an audio system, a mobile device (such as a cellular telephone), a laptop computer, a tablet device, a smart speaker, or another type of device.

[0064] According to some alternative implementations the apparatus 700 may be, or may include, a server. In some such examples, the apparatus 700 may be, or may include, an encoder. Accordingly, in some instances the apparatus 700 may be a device that is configured for use within an audio environment, such as a home audio environment, whereas in other instances the apparatus 700 may be a device that is configured for use in “the cloud,” e.g., a server.

[0065] In this example, the apparatus 700 includes an interface system 705 and a control system 710. The interface system 705 may, in some implementations, be configured for communication with one or more other devices of an audio environment. The audio environment may, in some examples, be a home audio environment. In other examples, the audio environment may be another type of environment, such as an office environment, an automobile environment, a train environment, a street or sidewalk environment, a park environment, etc. The interface system 705 may, in some implementations, be configured for exchanging control information and associated data with audio devices of the audio environment. The control information and associated data may, in some examples, pertain to one or more software applications that the apparatus 700 is executing.

[0066] The interface system 705 may, in some implementations, be configured for receiving, or for providing, a content stream. The content stream may include audio data. The audio data may include, but may not be limited to, audio signals. In some instances, the audio data may include spatial data, such as channel data and / or spatial metadata. In some examples, the content stream may include video data and audio data corresponding to the video data.

[0067] The interface system 705 may include one or more network interfaces and / or one or more external device interfaces (such as one or more universal serial bus (USB) interfaces). According to some implementations, the interface system 705 may include one or more wireless interfaces. The interface system 705 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system and / or a gesture sensor system. In some examples, the interface system 705 may include one or moreinterfaces between the control system 710 and a memory system, such as the optional memory system 715 shown in Figure 7. However, the control system 710 may include a memory system in some instances. The interface system 705 may, in some implementations, be configured for receiving input from one or more microphones in an environment.

[0068] The control system 710 may, for example, include a general purpose single- or multichip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.

[0069] In some implementations, the control system 710 may reside in more than one device. For example, in some implementations a portion of the control system 710 may reside in a device within one of the environments depicted herein and another portion of the control system 710 may reside in a device that is outside the environment, such as a server, a mobile device (e.g., a smartphone or a tablet computer), etc. In other examples, a portion of the control system 710 may reside in a device within one environment and another portion of the control system 710 may reside in one or more other devices of the environment. For example, a portion of the control system 710 may reside in a device that is implementing a cloud-based service, such as a server, and another portion of the control system 710 may reside in another device that is implementing the cloudbased service, such as another server, a memory device, etc. The interface system 705 also may, in some examples, reside in more than one device. In some implementations, a portion of a control system may reside in or on an earbud.

[0070] In some implementations, the control system 710 may be configured for performing, at least in part, the methods disclosed herein. According to some examples, the control system 710 may be configured for implementing methods of detecting wind in audio signals, filtering audio signals based on determined filter coefficients based on noise covariances, upmixing audio signals, or the like.

[0071] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non- transitory media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. The one or more non-transitory media may, for example, reside in the optional memory system 715 shown in Figure 7 and / or in the control system 710. Accordingly, various innovative aspects of the subject matter described in this disclosure can be implemented in one or more non-transitorymedia having software stored thereon. The software may, for example, filter audio signals to reduce noise, determine filter coefficients to perform such filtering, etc. The software may, for example, be executable by one or more components of a control system such as the control system 710 of Figure 7.

[0072] In some examples, the apparatus 700 may include the optional microphone system 720 shown in Figure 7. The optional microphone system 720 may include one or more microphones. In some implementations, one or more of the microphones may be part of, or associated with, another device, such as a speaker of the speaker system, a smart audio device, etc. In some examples, the apparatus 700 may not include a microphone system 720. However, in some such implementations the apparatus 700 may nonetheless be configured to receive microphone data for one or more microphones in an audio environment via the interface system 710. In some such implementations, a cloud-based implementation of the apparatus 700 may be configured to receive microphone data, or a noise metric corresponding at least in part to the microphone data, from one or more microphones in an audio environment via the interface system 710.

[0073] According to some implementations, the apparatus 700 may include the optional loudspeaker system 725 shown in Figure 7. The optional loudspeaker system 725 may include one or more loudspeakers, which also may be referred to herein as “speakers” or, more generally, as “audio reproduction transducers.” In some examples (e.g., cloud-based implementations), the apparatus 700 may not include a loudspeaker system 725. In some implementations, the apparatus 700 may include headphones. Headphones may be connected or coupled to the apparatus 700 via a headphone jack or via a wireless connection (e.g., BLUETOOTH).

[0074] Some aspects of present disclosure include a system or device configured (e.g., programmed) to perform one or more examples of the disclosed methods, and a tangible computer readable medium (e.g., a disc) which stores code for implementing one or more examples of the disclosed methods or steps thereof. For example, some disclosed systems can be or include a programmable general purpose processor, digital signal processor, or microprocessor, programmed with software or firmware and / or otherwise configured to perform any of a variety of operations on data, including an embodiment of disclosed methods or steps thereof. Such a general purpose processor may be or include a computer system including an input device, a memory, and a processing subsystem that is programmed (and / or otherwise configured) to perform one or more examples of the disclosed methods (or steps thereof) in response to data asserted thereto.

[0075] Some embodiments may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) that is configured (e.g., programmed and otherwise configured) to perform required processing on audio signal(s), including performance of one or more examples of the disclosed methods. Alternatively, embodiments of the disclosed systems (or elements thereof) may be implemented as a general purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include an input device and a memory) which is programmed with software or firmware and / or otherwise configured to perform any of a variety of operations including one or more examples of the disclosed methods. Alternatively, elements of some embodiments of the inventive system are implemented as a general purpose processor or DSP configured (e.g., programmed) to perform one or more examples of the disclosed methods, and the system also includes other elements (e.g., one or more loudspeakers and / or one or more microphones). A general purpose processor configured to perform one or more examples of the disclosed methods may be coupled to an input device (e.g., a mouse and / or a keyboard), a memory, and a display device.

[0076] Another aspect of present disclosure is a computer readable medium (for example, a disc or other tangible storage medium) which stores code for performing (e.g., coder executable to perform) one or more examples of the disclosed methods or steps thereof.

[0077] While specific embodiments of the present disclosure and applications of the disclosure have been described herein, it will be apparent to those of ordinary skill in the art that many variations on the embodiments and applications described herein are possible without departing from the scope of the disclosure described and claimed herein. It should be understood that while certain forms of the disclosure have been shown and described, the disclosure is not to be limited to the specific embodiments described and shown or the specific methods described.

[0078] Various aspects of the present disclosure may be appreciated from the following Enumerated Example Embodiments (EEEs):

[0079] EEE 1. A method of enhancing audio signals, the method comprising: obtaining multiple audio signals associated with multiple microphones of one or more devices; responsive to detecting wind in at least one audio signal of the multiple audio signals, filtering the multiple audio signals to generate multiple filtered audio signals by filtering each audio signal based on filter coefficients determined for each audio signal that weights each audio signal based on a noise covariance associated with that audio signal; and mixing the multiple filteredaudio signals to generate a multi-channel wind-reduced audio signal, wherein each multi-channel wind-reduced audio signal comprises a summation of the multiple filtered audio signals.

[0080] EEE 2. The method of EEE 1, wherein weighting is determined on a frame-by-frame basis based on frequency information associated with each audio signal.

[0081] EEE 3. The method of any one of EEEs 1 or 2, wherein the filter coefficients represent an instantaneous complex gain to be applied to each audio signal on a per-frequency band basis.

[0082] EEE 4. The method of any one of EEEs 1-3, wherein the filter coefficients cause each audio signal to be filtered in a manner that is optimized for each audio signal of the multiple audio signals.

[0083] EEE 5. The method of any one of EEEs 1-4, further comprising upmixing the multichannel wind-reduced audio signal.

[0084] EEE 6. The method of any one of EEEs 1-5, wherein weighting each audio signal comprises determining a bias term for each audio signal that is determined on a per-frequency band basis.

[0085] EEE 7. The method of EEE 6, wherein the bias term is based at least in part on an aggressiveness parameter that represents a degree to which wind noise is to be reduced.

[0086] EEE 8. The method of any one of EEEs 6 or 7, wherein the bias term for a given audio signal is based at least in part on a metric that indicates a diffuseness associated with the given audio signal.

[0087] EEE 9. The method of EEE 8, wherein the metric that indicates a diffuseness associated with the given audio signal is normalized based on level.

[0088] EEE 10. The method of any one of EEEs 6-9, wherein a value of the bias term for a given audio signal increases with a value of the center of the frequency band.

[0089] EEE 11. The method of any one of EEEs 1-10, further comprising determining a frequency threshold above which the multiple filtered audio signals are not mixed, wherein the frequency threshold is determined based on level differences at multiple frequency bands between the multiple audio signals.

[0090] EEE 12. The method of any one of EEEs 1-11, further comprising: upmixing the multichannel wind-reduced audio signal; causing the upmixed signal to be rendered to generate a rendered signal; and causing the rendered signal to be played back.

[0091] EEE 13. A system comprising: one or more processors; and a non-transitory computer- readable medium storing instructions that, upon execution by the one or more processors, cause the one or more processors to perform operations of EEEs 1-12.

[0092] EEE 14. A non-transitory computer-readable medium storing instructions that, upon execution by one or more processors, cause the one or more processors to perform operations of EEEs 1-13.

Claims

CLAIMS1. A method of enhancing audio signals, the method comprising: obtaining multiple audio signals associated with multiple microphones of one or more devices; responsive to detecting wind in at least one audio signal of the multiple audio signals, filtering the multiple audio signals to generate multiple filtered audio signals by filtering each audio signal based on filter coefficients determined for each audio signal that weights each audio signal based on a noise covariance associated with that audio signal; and mixing the multiple filtered audio signals to generate a multi-channel wind-reduced audio signal, wherein each multi-channel wind-reduced audio signal comprises a summation of the multiple filtered audio signals.

2. The method of claim 1, wherein weighting is determined on a frame-by-frame basis based on frequency information associated with each audio signal.

3. The method of any one of claims 1 or 2, wherein the filter coefficients represent an instantaneous complex gain to be applied to each audio signal on a per- frequency band basis.

4. The method of any one of claims 1-3, wherein the filter coefficients cause each audio signal to be filtered in a manner that is optimized for each audio signal of the multiple audio signals.

5. The method of any one of claims 1-4, further comprising upmixing the multi-channel wind-reduced audio signal.

6. The method of any one of claims 1-5, wherein weighting each audio signal comprises determining a bias term for each audio signal that is determined on a per-frequency band basis.

7. The method of claim 6, wherein the bias term is based at least in part on an aggressiveness parameter that represents a degree to which wind noise is to be reduced.

8. The method of any one of claims 6 or 7, wherein the bias term for a given audio signal is based at least in part on a metric that indicates a diffuseness associated with the given audio signal.

9. The method of claim 8, wherein the metric that indicates a diffuseness associated with the given audio signal is normalized based on level.

10. The method of any one of claims 6-9, wherein a value of the bias term for a given audio signal increases with a value of the center of the frequency band.

11. The method of any one of claims 1-10, further comprising determining a frequency threshold above which the multiple filtered audio signals are not mixed, wherein the frequency threshold is determined based on level differences at multiple frequency bands between the multiple audio signals.

12. The method of any one of claims 1-11, further comprising: upmixing the multi-channel wind-reduced audio signal; causing the upmixed signal to be rendered to generate a rendered signal; and causing the rendered signal to be played back.

13. A system comprising: one or more processors; and a non-transitory computer-readable medium storing instructions that, upon execution by the one or more processors, cause the one or more processors to perform operations of claims 1-12.

14. A non-transitory computer-readable medium storing instructions that, upon execution by one or more processors, cause the one or more processors to perform operations of claims 1-13.

Citation Information

Patent Citations

  • Detection and removal of wind noise

    US11217264B1

  • System and method for wind detection and suppression

    US20130308784A1

  • Method and system for acoustic source enhancement using acoustic sensor array

    US20180115855A1

  • Spatial audio wind noise detection

    US20220199100A1