Enhancing audio signals
By integrating beamforming and wind noise reduction using a shared MVDR algorithm, computational demands are reduced, enabling efficient audio enhancement on edge devices.
Patent Information
- Application Number
- PCT/US2025/012335
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-05
- Filing Date
- 2025-01-21
- Publication Date
- 2025-07-31
AI Technical Summary
Existing audio enhancement techniques, such as beamforming and wind noise reduction, are computationally intensive and require excessive memory, making them difficult to implement on edge devices like mobile phones.
Integrating beamforming and wind noise reduction using a shared instance of the Minimum Variance Distortionless Response (MVDR) algorithm, combining beamforming and wind noise covariance matrices to reduce computational complexity and resource usage.
This approach reduces computational complexity and resource usage while maintaining effective audio enhancement, allowing for flexible control over interference and wind noise suppression.
Smart Images

Figure US2025012335_31072025_PF_FP_ABST
Abstract
Description
ENHANCING AUDIO SIGNALS TECHNICAL FIELD
[0001] This disclosure pertains to systems, methods, and media for enhancing audio signals. CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of priority from International Patent Application No. PCT / CN2024 / 073489, filed 22 January 2024, U.S. Provisional Application No.63 / 554,743, filed 16 February 2024, and European Patent Application No. 2416138.6, filed 5 March 2024, each of which is hereby incorporated by reference herein in its entirety. BACKGROUND
[0003] Audio signals may be enhanced, e.g., to reduce noise, to perform beamforming to enhance a directionality of the signal, or the like. In some cases, multiple enhancement techniques may be performed on a given audio signal, which may be computationally intensive, and, e.g., require an excessive amount of computational memory. Such enhancement techniques may accordingly be difficult to perform on edge devices, such as mobile phones, due to the computational requirements. NOTATION AND NOMENCLATURE
[0004] Throughout this disclosure, including in the claims, the terms “speaker,” “loudspeaker” and “audio reproduction transducer” are used synonymously to denote any sound-emitting transducer (or set of transducers). A typical set of headphones includes two speakers. A speaker may be implemented to include multiple transducers (e.g., a woofer and a tweeter), which may be driven by a single, common speaker feed or multiple speaker feeds. In some examples, the speaker feed(s) may undergo different processing in different circuitry branches coupled to the different transducers.
[0005] Throughout this disclosure, including in the claims, the expression performing an operation “on” a signal or data (e.g., filtering, scaling, transforming, or applying gain to, the signal or data) is used in a broad sense to denote performing the operation directly on the signal or data, or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or pre-processing prior to performance of the operation thereon).
[0006] Throughout this disclosure including in the claims, the expression “system” is used in a broad sense to denote a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system including such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, in which the subsystem generates M of the inputs and the other X − M inputs are received from an external source) may also be referred to as a decoder system.
[0007] Throughout this disclosure including in the claims, the term “processor” is used in a broad sense to denote a system or device programmable or otherwise configurable (e.g., with software or firmware) to perform operations on data (e.g., audio, or video or other image data). Examples of processors include a field-programmable gate array (or other configurable integrated circuit or chip set), a digital signal processor programmed and / or otherwise configured to perform pipelined processing on audio or other sound data, a programmable general purpose processor or computer, and a programmable microprocessor chip or chip set. SUMMARY
[0008] Methods, systems, and media for enhancing audio signals are provided. In some embodiments, a method for enhancing audio signals involves obtaining an input signal obtained using multiple microphones of one or more audio capture devices. The method may further involve determining a beamforming noise covariance and a wind reduction noise covariance associated with the input signal. The method may further involve determining a weighted average of the beamforming noise covariance and the wind reduction noise covariance to determine an aggregate noise covariance. The method may further involve generating an output signal at least in part by filtering a representation of the input signal using filter weights determined based on the aggregate noise covariance, such that the output signal is enhanced relative to the input signal with respect to directionality and wind noise reduction.
[0009] In some examples, the method may further involve generating the filter weights by providing the aggregate noise covariance as an input parameter to a minimum variance distortionless response (MVDR) algorithm.
[0010] In some examples, the beamforming noise covariance is determined based at least in part on a direction of arrival associated with the input signal.
[0011] In some examples, the beamforming noise covariance is determined based at least in part on a metric that indicates a diffuseness associated with the input signal.
[0012] In some examples, the wind reduction noise covariance is determined based at least in part on a wind level metric associated with the input signal.
[0013] In some examples, determining the weighted average of the beamforming noise covariance and the wind reduction noise covariance is based at least in part on a weighting factor, wherein the weighting factor is determined based on at least one of: a metric that indicates a diffuseness associated with the input signal, a pre-determined value of the weighting factor, user input, or a ratio of a power associated with the beamforming noise covariance to a power associated with a wind reduction noise covariance. In some examples, the user input is received via a user interface. In some examples, the input signal is pre-recorded audio signal, and wherein the user interface is used to present the pre-recorded audio signal.
[0014] In some examples, the method may further involve transforming the input signal to a frequency domain representation of the input signal, and wherein the filter weights are used to filter the frequency domain representation of the input signal. In some examples, the beamforming noise covariance, the wind reduction noise covariance, and the aggregate noise covariance are determined on a per-frequency band basis for the frequency domain representation of the input signal. In some examples, the beamforming noise covariance is determined based at least in part on a metric that indicates a diffuseness associated with the input signal, and wherein a degree to which the beamforming noise covariance is weighted relative to the wind reduction noise covariance depends on the frequency band and the metric that indicates the diffuseness associated with the input signal.
[0015] In some examples, the output signal comprises stereo signals. In some examples, the beamforming noise covariance is determined based on beamforming associated with two different looking directions associated with each signal of the stereo signals.
[0016] In some examples, the one or more audio capture devices is one of, or a combination of: a mobile phone, a portable microphone, earbuds, a camera device, or an augmented reality (AR) or virtual reality (VR) headset.
[0017] In some examples, the method may further involve storing the output signal, the beamforming noise covariance, and / or the wind reduction noise covariance.
[0018] In some examples, the method may further involve: causing the output signal to be rendered to generate a rendered signal; and causing the rendered signal to be played back.
[0019] Some or all of the operations, functions and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. Accordingly, some innovative aspects of the subject matter described in this disclosure can be implemented via one or more non-transitory media having software stored thereon.
[0020] At least some aspects of the present disclosure may be implemented via an apparatus. For example, one or more devices may be capable of performing, at least in part, the methods disclosed herein. In some implementations, an apparatus is, or includes, an audio processing system having an interface system and a control system. The control system may include one or more general purpose single- or multi-chip processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or combinations thereof.
[0021] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, the drawings, and the claims. Note that the relative dimensions of the following figures may not be drawn to scale. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figures 1A, 1B, and 1C are a diagram illustrating an example audio capture device in accordance with some embodiments.
[0023] Figure 2 is a diagram illustrating a technique for wind noise reduction and beamforming in accordance with some embodiments.
[0024] Figure 3 is a diagram illustrating a system for performing wind noise reduction and beamforming using aggregated noise covariances in accordance with some embodiments.
[0025] Figure 4 is a graph that illustrates determination of a weighting factor based on a purity, or diffuseness, metric, in accordance with some embodiments.
[0026] Figure 5 is a flowchart of an example process for performing wind noise reduction and beamforming using aggregated noise covariances in accordance with some embodiments.
[0027] Figure 6 is a flowchart of an example process for generating enhanced audio signals using aggregated noise covariances in accordance with some embodiments.
[0028] Figure 7 shows a block diagram that illustrates examples of components of an apparatus capable of implementing various aspects of this disclosure.
[0029] Like reference numbers and designations in the various drawings indicate like elements. DETAILED DESCRIPTION OF EMBODIMENTS
[0030] Beamforming and wind noise reduction are two techniques that can be used to enhance audio signals.
[0031] Beamforming is a spatial filtering technique used in microphone arrays to enhance sound coming from a desired direction while suppressing unwanted noise and interference. Beamforming exploits the phase differences between microphones which result from the varying time delays of sound waves arriving at different positions in a microphone array.
[0032] Wind noise reduction refers to the process of minimizing or suppressing unwanted noise caused by wind turbulence affecting a microphone or microphone array. Wind noise is a common issue in outdoor audio recording, teleconferencing, and hearing aids, as it can significantly degrade audio quality by masking desired sounds such as speech or music.
[0033] The Delay-and-Sum Beamformer (DSB) is a fundamental method that delays and sums signals from multiple microphones to focus on a target sound source. While straightforward and effective in low-noise environments, DSB lacks adaptability to complex noise scenarios and provides limited interference suppression.
[0034] The minimum variance distortionless response (MVDR) beamformer addresses these limitations by adaptively adjusting the filter weights to minimize output noise power while preserving the desired signal. It achieves this by estimating a spatial noise covariance matrix and applying optimal weights to suppress interfering sources. Compared to delay-and-sum beamforming, MVDR provides improved interference rejection and higher directivity, making it suitable for noisy and reverberant environments. However, MVDR’s performance depends on accurate estimation of the noise covariance matrix.
[0035] Conventionally, both beamforming and wind noise reduction algorithms utilize the minimum variance distortionless response (MVDR) algorithm. Beamforming and wind noise reduction algorithms are typically performed separately such that beamforming and wind noisereduction algorithms each make use of the MVDR algorithm. For example, beamforming and wind noise reduction algorithms may be performed sequentially, where each of the beamforming and wind noise reduction algorithms utilize the MVDR algorithm at least once. The MVDR algorithm may be computationally intensive to perform, e.g., to generate filter taps. For example, the MVDR algorithm may require substantial memory in order to generate the filter taps, and this computational complexity and computational resource usage impact is magnified when both beamforming and wind noise reduction are performed in a way that the MVDR algorithm is utilized separately for each.
[0036] Since both beamformer and wind noise reduction aim to minimize output noise power while preserving the desired signal, the two algorithms can be integrated into a single processing framework. Disclosed herein are techniques for performing beamforming and wind noise reduction using a common instance of MVDR. Implementation of the techniques disclosed herein improves computational performance by reducing computational complexity and computational resource usage (e.g., memory usage), because use of the MVDR algorithm is shared for both beamforming and wind noise reduction.
[0037] The sharing of a single processing framework such as MVDR also provides flexibility to control and balance the suppression of interference and wind noise. For example, wind noise could be reduced when it is determined that wind noise is present, and interference could be reduced when it is determined that no wind is present. In mild wind conditions, when wind noise does not dominate the whole spectrum, wind noise reduction can be applied in the low frequency region while interference is reduced in the high frequency region.
[0038] In some embodiments, the techniques disclosed herein may determine a beamforming noise covariance and a wind noise reduction noise covariance. For brevity and clarity, the term ‘wind noise reduction noise covariance’ is shortened herein to ‘wind reduction noise covariance’, but it should be understood to mean a noise covariance associated with the reduction of wind noise. Analogously, ‘beamforming noise covariance’ refers to a noise covariance associated with beamforming. The beamforming noise covariance and the wind reduction noise covariance may be combined by generating a weighted average of the beamforming noise covariance and the wind reduction noise covariance. The weighted average may be determined using a weighting factor that controls the degree to which the beamforming noise covariance is considered in the aggregate noise covariance relative to the wind reduction noise covariance, and vice versa. In other words, the weighting factor may allow for control of prioritization of beamforming over wind noise reduction, and vice versa. The weighting factor may be determined based on a purity metricassociated with the audio signal that indicates a diffuseness of the input signal at different frequency bands, based on a power ratio of predicted beamforming output to wind noise reduction output. The power ratio may be determined using the filter weights of the last frame and the covariance matrices of the current frame, and / or based on a user setting.
[0039] Figure 1A illustrates an example audio capture device in accordance with some embodiments. In particular, the audio capture device is a mobile phone 100. Front side 102 of mobile phone 100 is associated with two microphones, top microphone 106 and bottom microphone 108. Back side 104 of mobile phone 100 includes a third microphone, back microphone 110. Note that each microphone may be configured to capture audio content from a direction associated with the location of the microphone with respect to mobile phone 100. Additionally, as described below, the relative locations of each microphone may be used to determine noise covariances and filter a multi-microphone audio signal based on aggregated noise covariances. Figure 1B illustrates a side view of mobile phone 100 that includes top microphone 106, and Figure 1C illustrates a side view of mobile phone 100 that includes bottom microphone 108.
[0040] Conventional techniques typically involve multi-channel wind noise reduction (MWNR) and beamforming operating sequentially. For example, MWNR may be performed by using raw microphone signals, detecting wind noise using inter-channel information, and reducing wind noise by mixing input channels and applying a dynamic high pass filter. A beamforming algorithm may take the MWNR outputs as an input and generate a mono output. The beamforming algorithm may be configured to extract signals from a region of interest and minimize signals from other directions. Note that both the MWNR algorithm and the beamforming algorithm may conventionally utilize the MVDR algorithm. Accordingly, conventional techniques that implement MWNR and beamforming separately implement multiple instances of the MVDR algorithm.
[0041] Figure 2 is a diagram that illustrates conventional techniques for performing MWNR and beamforming sequentially. As illustrated, a multi-channel microphone signal is processed to transform a time-domain signal to a frequency-domain signal using short time Fourier transform (STFT) block 202. The frequency domain signal is then provided to beamforming block 204, which is configured to extract signals from a region of interest and minimize signals from other directions. The frequency domain signal may additionally or alternatively be provided to MWNR block 206, which is configured to detect and reduce wind noise. The output of MWNR block 206 is a single channel mono output, which may then be converted to a time domain signal to form anenhanced signal that has been enhanced with respect to wind noise reduction and beamforming. Because each of beamforming block 204 and MWDR block 206 take as input a multi-channel signal and generate a mono output, selection switch 208 may be configured to select between these two parallel paths. Note that in some embodiments, as an alternative to a selection switch as shown in FIG. 2, a cross-fade may be used. For example, responsive to wind noise being detected, selection switch 208 may be configured to select the output of MWDR block 206 as the output signal, and conversely, responsive to wind noise not being detected, selection switch 208 may be configured to select the output of beamforming block 204. Note that each of beamforming block 204 and MWNR block 206 may implement one or more instances of the MVDR algorithm, which is computationally intensive. In particular, each instance of the MVDR algorithm utilizes multiple filter taps, which requires computing power and memory. Additionally, it should be noted that although Figure 2 depicts a three-channel microphone input, the same techniques may be applied for any number of microphone channels, e.g., two, four, five, ten, etc.
[0042] The techniques disclosed herein may combine wind noise reduction and beamforming in a manner that utilizes the MVDR once, thereby reducing computational resources (e.g., memory) and computational complexity. For example, the number of filter taps utilized to implement MVDR when shared by the wind noise reduction and beamforming algorithms is reduced relative to the number of filter taps utilized to implement MVDR for the wind noise reduction and beamforming algorithms separately (e.g., as shown in and described above in connection with Figure 2). Moreover, the techniques disclosed herein may allow for balancing wind noise reduction with beamforming such that in some instances, wind noise reduction may be prioritized over beamforming, and vice versa, with user control of the balancing and tradeoffs. A difference between beamforming algorithms and wind noise reduction algorithms is how the spatial noise covariance matrix is calculated. In beamforming, the noise covariance matrix (a Hermitian matrix) is estimated when the signal is estimated as not coming from the region of interest. In multi-channel wind noise reduction (MWNR), the noise covariance matrix (a diagonal matrix) is manipulated so that microphone signals can be optimally combined to reduce the wind noise while retaining other sound. While the covariance matrices are estimated differently, they are both NxN matrices, where N is the number of microphones, meaning that the matrices can be aggregated to determine a single covariance matrix which is used to obtain an MVDR weight vector. When wind does not exist, the weight vector serves as a spatial filter to suppress interference and diffuse noise while retaining the signal from the desired direction. When windexists, the weight vector works as mixing gain to minimize the output signal power while preserving the quality of other sounds such as speech, music, etc.
[0043] In some embodiments, the techniques disclosed herein may obtain an input signal obtained using multiple microphones of one or more audio capture devices (e.g., a mobile phone, a camera, a tablet computer, etc.). The relative locations of the multiple microphones may be known. Note that the techniques disclosed herein may be implemented using a single audio capture device, or multiple audio capture device, provided the relative locations of the multiple microphones is known. A beamforming noise covariance and a wind reduction noise covariance may be determined for the input signal. A weighted average of the beamforming noise covariance and the wind reduction noise covariance may be determined to generate an aggregate noise covariance. The weighting factor used to determine the weighted average may be determined based on various factors, such as a metric that indicates diffuseness of the input signal (generally referred to herein as “purity”), a user-selected value of the weighting factor, a ratio of power associated with the predicted beamforming output to a power associated with a wind noise output in the output signal, or any combination thereof. An enhanced output signal may be generated by filtering a representation of the input signal (e.g., a frequency domain representation of the input signal) using filter weights determined based on the aggregate noise covariance. Note that the output signal may be enhanced relative to the input signal with respect to both directionality and wind noise reduction, and that a tradeoff between enhancing directionality (e.g., via beamforming) and reducing wind noise may be controlled at least partially by the weighting factor used to determine the aggregate noise covariance.
[0044] Figure 3 illustrates an example system for implementing techniques described herein for enhancing audio signals in accordance with some embodiments. As illustrated, a multi- microphone input signal may be provided as an input to STFT block 302, which may convert the input signal to a frequency domain signal. Based on the frequency domain representation of the input signal, spatial covariance information associated with the input signal may be generated by spatial covariance block 304. The spatial covariance information may be used by block 305 to determine a direction of arrival of the input signal (e.g., using direction of arrival block 306), a purity of the input signal (e.g., using purity block 308), and / or a wind level associated with the input signal (e.g., using wind level block 310). Note that the purity associated with the input signal may generally indicate a diffuseness of sound. An audio signal with a high purity value indicates a highly directional signal, and, conversely, an audio signal with a low purity value indicates a highly diffuse signal with low directionality. In general, outputs of block 305 (e.g., the directionof arrival, the purity, and / or the wind level) may be used to determine which portions of the microphone signal should be used to aggregate the covariances of undesired signals to be minimized and / or rejected.
[0045] The spatial covariance information as well as the direction of arrival and the purity may be provided as input to beamforming noise covariance block 312, which may be configured to generate a beamforming noise covariance, generally referred to herein as Rb. The beamforming noise covariance may indicate the degree to which the current frame of the audio signal contains diffuse sound and / or the direction of arrival is outside a region of interest. In other words, using the direction of arrival and the purity metric, the direct portions of the audio signal within the region of interest are omitted from the beamforming noise covariance. The region of interest (ROI) refers to a pre-defined target region, the centre direction of which is the beamforming looking direction.
[0046] The beamforming noise covariance ^^^ is a matrix of size ^^ ൈ ^^ for each frequency binat each frame, where ^^ is the number of microphones. To estimate this matrix, a noise time- frequency (T-F) mask ^^^^^^^^^^^is estimated. ^^^^^^^^^^^ranges from 0 to 1, with a value of 0 meaning that the time-frequency bin is dominated by the direct sound source in the region of interest, and a value of 1 indicating that diffuse noise and / or interferences outside the region of interest dominate the T-F bin.
[0047] ^^^^^^^^^^^can be estimated by first estimating a ‘content mask’ using a trained neural network- or Digital Signal Processor-based methods. Examples include LensNet, and the well- known Improved Minima Controlled Recursive Averaging (IMCRA) noise spectrum estimation. Other techniques which are known in the art may be used. The resulting mask indicates whether a time-frequency bin is dominated by the target contents, such as speech, direct sound, etc. A direction-of-arrival estimation can then be applied to the contents detected by this content mask. Any suitable direction-of-arrival estimation technique can be used: include neural network-based methods, the GCC-PHAT algorithm, and the MUSIC (Multiple Signal Classification) algorithm. Then, given the region of interest, ^^^^^^^^^^^can be derived from the content mask and the estimated direction of arrival. If a givenby the target contents located in the region of interest, ^^^^^^^^^^^for that bin is close to 1, otherwise it is close to zero.
[0048] The beamforming noise covariance matrix can then be estimated as follows: ^^^^^^,^^^ ൌ ^^^^^,^^^^^^^^^^,^^ െ 1^ ^ ൫1 െ ^^^^^,^^^൯^^^^^, ^^^^^^^^^^,^^^,where ^^ is a bin index, ^^ is a frame index, ^^ is a noisy signal, ^^^is a constant smoothing factor,and ^^^^^, ^^^ is a weighting factor calculated as follows:^^^^^,^^^ ൌ 1 ^ ^^^^ െ 1^^^^^^^^^^^^.
[0049] In one example, ^^^takes a value of 0.9.
[0050] Similarly, the spatial covariance information, the purity, and the wind level may be provided as input to wind noise covariance block 314. Wind noise covariance block 314 may be configured to generate a wind noise covariance, generally referred to herein as Rw.
[0051] The wind reduction noise covariance ^^௪refers to an estimated spatial wind reduction noise covariance matrix which is used according to the methods described herein to reduce wind noise. Wind reduction noise covariance is computed as a diagonal covariance matrix indicating the power spectral density (PSD) of each microphone signal when wind exists. Given that wind signal between microphones has little correlation, the off-diagonal elements of this matrix can be set to zero. Therefore, the wind noise covariance matrix can be estimated as a real diagonal matrix ofsize ^^ ൈ ^^ for each frequency bin or band ^^ at each frame ^^.
[0052] To estimate ^^௪, the spatial covariance matrix R of the input noisy signal is computed as follows: ^^^^^,^^^ ൌ ^^^^^^^^,^^ െ 1^ ^ ^1 െ ^^^^^^^^^,^^^^^^^^^^, ^^^,with the wind noise covariance ^^௪given by: ^^^^^^,^^^ ൌ ^^^^^,^^^diag൫^^^^^,^^^൯,where diag(*) is an operator that when applied to a matrix extracts the diagonal entries of the given matrix, and ^^^^^,^^^ is a weighting factor for the k-th bin or band at the n-th frame.
[0053] The weighting factor ^^^^^,^^^ can be derived from purity and wind level (both explained above) according to the following formula: ^^^^^, ^^^ ൌ ൫1 െ ^^^^^^^^^^^^^^^, ^^^൯^^^^^^^^_^^^^^^^^^^^^^^.
[0054] The beamforming noise covariance and the wind noise covariance may be provided as inputs to weighting block 316. Weighting block 316 may be configured to generate an aggregate noise covariance, generally referred to herein as R. In one example, the aggregate noise covariance may be determined by: ^^ ൌ ^^ ∗ ^^௪ ^ ^1 െ ^^^ ∗ ^^^,where ^^ is a weighting factor which indicates the weighting of the wind noise covariance relative to the beamforming noise covariance.
[0055] The aggregate noise covariance, R, may be provided to MVDR block 318, which is configured to generate filter weights using the MVDR algorithm. The filter weights may be determined based on the steering vector, generally represented herein as γ. The steering vector may be determined by steering vector block 320. The steering vector may be determined based on a front direction for the audio capture device, based on a speech classifier that estimates the steering vector based on the direction associated with detected speech, and / or by user input. For example, in some embodiments, a steering vector may be determined based on a user clicking on a person in a video who is speaking to associate the speech-containing audio signal with the directionality of the selected person. As a more particular example, the direction of interest may be determined based on the user clicking the screen, and the steering vector may be determined based on the direction of interest and a dataset that specifies steering vector parameters for the particular user device. For example, the dataset may be obtained by measuring parameters associated with the user device, and / or by calculating steering vector parameters based on a simulation model using, e.g., finite element modeling.
[0056] In some embodiments, filter weights may be determined based on the aggregate noise covariance and the steering vector. In one example, filter weights are determined by: ^^ି^^^ ∗ ^^
[0057] The determined filter may used to filter the frequency-domain representation of the input signal by filter block 322. The filtered signal may then be transformed back to the time domain (e.g., using an inverse STFT) to generate the enhanced output signal.
[0058] As described above, in some embodiments, the weighting factor may be determined based on a metric that indicates diffuseness of the signal, generally referred to herein as “purity.” Purity and the weighting factor may be determined on a per-frequency band basis. In someembodiments, the weighting factor may be determined such that the reciprocal of the weighting factor (e.g., 1-α) may increase monotonically with the purity for a given frequency band. In other words, because the reciprocal of the weighting factor indicates the weighting of the beamforming noise covariance, increased purity at a given frequency band may cause the beamforming noise covariance to be weighted more for the frequency band relative to covariance in determining the filter weights, because the increased purity value indicates less wind noise present in the signal. In some embodiments, the weighting factor (and therefore, the reciprocal of the weighting factor) may vary for a given purity value for different frequency bands. For example, the reciprocal of the weighting factor (1-α) may be relatively high for high frequency bands even at low purity levels because wind noise is less likely to be present in high frequency bands, and therefore prioritizing wind noise reduction is less important and beneficial in high frequency bands. Conversely, the reciprocal of the weighting factor (1-α) may be relatively low for low frequency bands even at relatively higher purity levels, because wind noise is more likely to be present in low frequency bands, and therefore, prioritizing wind noise reduction may be more important and beneficial in low frequency bands relative to high frequency bands.
[0059] Figure 4 is a graph of example weighting factor reciprocals (e.g., 1-α) as a function of purity value for different frequency bands. In the legend of the graph of Figure 4, note that the frequency values indicate the center of each frequency band. Note that, for all frequency bands, the reciprocal of the weighting factor increases monotonically with purity value. This indicates that when input signals are more directional (e.g., less diffuse), directionality is to be prioritized via performing beamforming relative to wind noise reduction. As illustrated, the reciprocal of the weighting factor differs for a given purity value for different frequency bands. For example, curve 402 illustrates the reciprocal of the weighting factor as a function of purity value for the 24 kHz frequency band, and curve 404 illustrates the reciprocal of the weighting factor as a function of purity value for the 50 Hz band. Note that, for all but the smallest purity values (e.g., below 0.1), the reciprocal of the weighting factor for the 24 kHz frequency band is relatively high, indicating that beamforming is to be highly prioritized relative to wind noise reduction, regardless of the diffuseness of the signal, because wind noise is not likely to be present in high frequency bands regardless of purity. Conversely, note that the reciprocal of the weighting factor increases more gradually for the 50 Hz band as a function of purity value, because wind noise is more likely to be prevalent in the 50 Hz band regardless of diffuseness of the signal.
[0060] Figure 5 is a flowchart of an example process 500 for enhancing audio signals with respect to beamforming and wind noise reduction in accordance with some embodiments. In someimplementations, the techniques disclosed herein may be performed by an audio capture device, such as a mobile phone, a tablet computer, a desktop computer, a camera, etc. For example, the techniques may be performed by one or more processors or control systems of such an audio capture device. In some embodiments, in instances in which multiple audio capture devices are utilized to capture input audio signals, the techniques may be performed by one or more processors or control systems of one of the audio capture devices, or by processors and / or control systems distributed across multiple audio capture devices. An example of a processor or control system that may be utilized is shown in and described below in connection with Figure 7. It should be understood that, in some embodiments, process 500 may loop through multiple times for each frame of the input signal, and what is described below is generally with respect to one frame or a block of frames of the input signal.
[0061] Process 500 can begin at 502 by obtaining an input signal associated with multiple microphones of one or more audio capture devices. The relative locations e.g., distances between the microphones of the multiple microphones may be known. Examples of audio capture devices include mobile phones, tablet computers, augmented reality / virtual reality headsets, cameras, laptop computers, desktop computers, etc. Note that in instances in which multiple audio capture devices are utilized, the audio capture devices may be of the same type, or of different types. An example of an audio capture device is shown in and described above in connection with Figure 1A.
[0062] At 504, process 500 can transform the input signal to a frequency domain input signal. Process 500 may transform the input signal to the frequency domain using a STFT or any other suitable transform.
[0063] At 506, process 500 can determine a spatial covariance of the frequency domain input signal for each frequency channel or frequency band. The spatial covariance may indicate the degree of the spatial relationship at different frequency bands between the signals associated with each microphone. Spatial covariance may be represented as a square matrix, e.g., as an N x N matrix.
[0064] At 508, process 500 can determine, based on the spatial covariance, a direction of arrival (DOA), a purity metric, and a wind level associated with the input signal. Note that the purity metric and the wind level may be determined on a per-frequency band basis. The purity metric and the wind level are described above in more detail above in connection with Figure 4. For example, as described above, the purity metric may indicate a degree of diffuseness of the signalwithin a given frequency band, where a high purity value indicates low diffuseness (e.g., high directionality), and a low purity value indicates high diffuseness (e.g., low directionality). In some embodiments, the DOA may indicate whether the current frame of the input signal is within or outside of a region of interest. In general, the DOA, the purity metric, and the wind level may be used to determine which portions of the input signal should be used to aggregate the covariances of the undesired signals that are to be attenuated, as will be described below in connection with blocks 510-512.
[0065] At 510, process 500 can determine a beamforming noise covariance and a wind reduction noise covariance based on the DOA, purity, the wind level, and / or the spatial covariance of the frequency domain input signal. Note that the beamforming noise covariance and the wind reduction noise covariance may be determined on a per-frequency band basis. As shown in and described above in connection with Figure 3, the beamforming noise covariance may be determined based on the spatial covariance, the DOA, and / or the purity. The wind reduction noise covariance may be determined based on the spatial covariance, the purity, and / or the wind level. The beamforming noise covariance may indicate covariance associated with signals that are to be attenuated on the basis of directionality. The wind reduction noise covariance may indicate covariance associated with signals that are to be attenuated on the basis of wind presence in the signal. Note that, in some implementations, the beamforming noise covariance and / or the wind reduction noise covariance may be stored for future use. For example, the beamforming noise covariance and / or the wind reduction noise covariance may be stored such that, at a subsequent time, a different tradeoff between wind noise reduction and beamforming may be implemented using, e.g., a different weighting factor between the two covariances without re-calculating the beamforming noise covariance and / or the wind reduction noise covariance. This may be advantageous for allowing more flexible control for an end user while reducing computational complexity or computational resources used.
[0066] At 512, process 500 can determine a weighted average of the beamforming noise covariance and the wind reduction noise covariance to determine an aggregate noise covariance. For example, given a beamforming noise covariance of Rband a wind reduction noise covariance of Rw, the aggregate noise covariance, R, may be determined by: ^^ ൌ ^^^^^௪^ ^ ^1 െ ^^^ ∗ ^^^
[0067] In the equation given above, α represents the weighting factor that indicates how the beamforming noise covariance is to be weighted relative to the wind reduction noise covariance.In other words, the weighting factor allows for control of the tradeoff between performing beamforming and reducing wind noise. The weighting factor may be determined separately for each frequency band.
[0068] The weighting factor α may be determined in any suitable manner. For example, in some embodiments, the weighting factor may be determined based on the purity metric. As a more particular example, as shown in and described above in connection with Figure 4, the weighting factor may decrease monotonically with the purity value, or conversely, as shown in Figure 4, the reciprocal of the weighting factor (1-α) may increase monotonically with the purity value. Note that the impact of purity on the weighting factor may differ across different frequency bands, as shown in and described above in connection with Figure 4.
[0069] In some embodiments, the weighting factor may be determined based on user input or a pre-configured setting. For example, the weighting factor (or a weighting factor for each frequency band) may be specified by a user using a user interface. Alternatively, the weighting factor(s) may be a pre-configured setting associated with a particular application executing on the audio capture device.
[0070] As another example, in some embodiments, the weighting factor may be determined based on a ratio of the output power of the beamforming signal to the output power of the wind noise reduction signal. Note that, the use of the ratio of the powers may involve iteratively adjusting the weighting factor to achieve a target criteria. The output power of the beamforming signal may be determined based on the beamforming noise covariance and filter weights associated with the previous frame of the audio signal, and the output power of the wind noise reduction signal may be determined based on the wind reduction noise covariance and the filter weights of the previous frame. If the ratio of the powers exceeds a predetermined threshold, the weighting factor may be adjusted (e.g., increased) for the current frame.
[0071] In one example, power associated with the beamforming noise covariance may be determined by: ^^^^^^^^^^ ு^^^^ ൌ ^^ ^^^^^
[0072] The power associated with the wind reduction noise covariance may be determined by: ^^^^^^^^^^௪^^ௗ ൌ ^^ு^^௪^^
[0073] In the equations given above, H represents the conjugate transpose operator.
[0074] By way of example, in an instance in which the previous frame had no wind noise, the filter weights associated with the previous frame may be those that apply beamforming with no wind noise reduction. Continuing with this example, by analyzing the ratio of the output power of the beamforming signal to the output power of the wind noise reduction signal using the previous filter weights, in an instance in which the current frame includes wind noise, the deficiencies of the filter weights of the previous frame when applied to the current frame are used to adjust the filter weights that are to be used with the current frame in a desirable manner. Note that, by utilizing the ratio of the beamforming power to the wind noise reduction power, the filter weights may be determined without performing matrix inversion, which may reduce computational complexity.
[0075] It should be noted that in some embodiments, the beamforming noise covariance, the wind reduction noise covariance, and / or the aggregate noise covariance may be stored for future use, e.g., to calculate filter weights for a subsequent frame with reduced computational resources and / or more quickly.
[0076] At 514, process 500 can determine filter weights based on the aggregate noise covariance, R. The filter weights may additionally depend on a steering vector. The steering vector may be a pre-configured setting (e.g., that indicates a front-facing direction for the audio capture device). As another example, in some embodiments, the steering vector may be estimated based on a speech detection algorithm that detects speech in the input signal and assigns a directionality to the detected speech. As yet another example, in some embodiments, the steering vector may be user-selected. For example, in an instance in which the input signal is associated with video contact, the steering vector may be determined based on user input that involves the user selecting a person speaking who is to be associated with the estimated steering vector.
[0077] In some embodiments, the filter weights may be generated using the MVDR algorithm, the aggregate noise covariance, and the steering vector. By way of example, the filter weights may be determined by: ^^^^ି^ ∗ ^^
[0078] In the equation given transpose operator.
[0079] At 516, process 500 can filter the frequency domain input signal using the filter weights. For example, process 500 can apply the filter weights for each frequency band of the frequency domain signal.
[0080] At 518, process 500 can transform the filtered signal to a time domain output signal for each frequency channel. For example, the filtered signal may be transformed using an inverse STFT. The output signal may be enhanced with respect to both wind noise reduction and beamforming while allowing for a balancing and tradeoff of beamforming with respect to wind noise reduction.
[0081] It should be noted that, in some embodiments, the output signal may be a mono signal that includes one channel. Alternatively, in some embodiments, the output signal may be a stereo signal. In instances in which the output signal is a stereo signal, the beamforming covariance may be substituted with a fixed beam covariance such that Rb is the noise covariance of the fixed beam. The fixed steering vector ^^ may be determined by determining the weight of each measured or simulated microphone impulse response for a particular pattern or shape, such as a cardioid. The noise covariance Rb may then be calculated for each direction. The weighted noise covariance may be determined based on the impulse response in each direction, and an additional regularization term can be added to the noise covariance to improve robustness.
[0082] In some embodiments, process 500 may render the output signal. In some embodiments, process 500 may cause the rendered signal to be stored and / or played back, e.g., via headphones, one or more speakers, or the like. Note that, in some embodiments, the covariances may be stored, e.g., for use in determining filter weights for future frames, or for any other suitable purpose.
[0083] Figure 6 is a flowchart of an example process 600 for enhancing audio signals with respect to beamforming and wind noise reduction in accordance with some embodiments. In some implementations, the techniques disclosed herein may be performed by an audio capture device, such as a mobile phone, a tablet computer, a desktop computer, a camera, etc. For example, the techniques may be performed by one or more processors or control systems of such an audio capture device. In some embodiments, in instances in which multiple audio capture devices are utilized to capture input audio signals, the techniques may be performed by one or more processors or control systems of one of the audio capture devices, or by processors and / or control systems distributed across multiple audio capture devices. An example of a processor or control system that may be utilized is shown in and described below in connection with Figure 7. It should be understood that, in some embodiments, process 600 may loop through multiple times for eachframe of the input signal, and what is described below is generally with respect to one frame or a block of frames of the input signal.
[0084] Process 600 can begin at 602 by obtaining an input signal associated with multiple microphones of one or more audio capture devices. The relative locations e.g., distances between the microphones of the multiple microphones may be known. Examples of audio capture devices include mobile phones, tablet computers, augmented reality / virtual reality headsets, cameras, laptop computers, desktop computers, etc. Note that in instances in which multiple audio capture devices are utilized, the audio capture devices may be of the same type, or of different types. An example of an audio capture device is shown in and described above in connection with Figure 1A.
[0085] At 604, process 600 can determine a beamforming noise covariance and a wind reduction noise covariance associated with the input signal. In some embodiments, the beamforming noise covariance and the wind reduction noise covariance may be determined based on a spatial covariance of a representation of the input signal. For example, the representation of the input signal may be a frequency domain representation of the input signal, and the spatial covariance may be determined based on the frequency domain representation. In some embodiments, the beamforming noise covariance may be determined based on a DOA and / or a purity metric associated with the input signal. In some embodiments, the wind reduction noise covariance may be determined based on a purity metric associated with the input signal and / or a wind level associated with the input signal.
[0086] At 606, process 600 can determine a weighted average of the beamforming noise covariance and the wind reduction noise covariance to determine an aggregate noise covariance. For example, in some embodiments, the weighted average may utilize a weighting factor to control the degree to which the beamforming noise covariance contributes to the aggregate noise covariance relative to the wind reduction noise covariance, and vice versa. As described above, in some embodiments, the weighting factor may be determined on a per-frequency band basis such that the tradeoff between beamforming noise covariance and wind reduction noise covariance is on a per-frequency band basis. In some embodiments, the weighting factor may be determined based on a purity metric associated with the input signal that indicates, for each frequency band, the diffuseness of the signal at that band. In some embodiments, the weighting factor may be determined based on a ratio of beamforming noise power to wind noise reduction power. In some embodiments, the weighting factor may be specified by a user or may be a pre-configured setting, e.g., for a given media content application.
[0087] At 608, process 600 can generate an output signal at least in part by filtering a representation of the input signal using filter weights determined based on the aggregate noise covariance such that the output signal is enhanced relative to the input signal with respect to directionality and wind noise reduction. For example, in some embodiments, process 600 can generate filter weights using the aggregate noise covariance applied to the MVDR algorithm. Continuing with this example, process 600 can filter the representation of the input signal (e.g., the frequency domain representation of the input signal) using the filter weight.
[0088] In some embodiments, process 600 can then render the output signal (or a time domain representation of the output signal). Process 600 may then playback and / or store the rendered output signal.
[0089] Figure 7 is a block diagram that shows examples of components of an apparatus capable of implementing various aspects of this disclosure. As with other figures provided herein, the types and numbers of elements shown in Figure 7 are merely provided by way of example. Other implementations may include more, fewer and / or different types and numbers of elements. According to some examples, the apparatus 700 may be configured for performing at least some of the methods disclosed herein. In some implementations, the apparatus 700 may be, or may include, a television, one or more components of an audio system, a mobile device (such as a cellular telephone), a laptop computer, a tablet device, a smart speaker, or another type of device.
[0090] According to some alternative implementations the apparatus 700 may be, or may include, a server. In some such examples, the apparatus 700 may be, or may include, an encoder. Accordingly, in some instances the apparatus 700 may be a device that is configured for use within an audio environment, such as a home audio environment, whereas in other instances the apparatus 700 may be a device that is configured for use in “the cloud,” e.g., a server.
[0091] In this example, the apparatus 700 includes an interface system 705 and a control system 710. The interface system 705 may, in some implementations, be configured for communication with one or more other devices of an audio environment. The audio environment may, in some examples, be a home audio environment. In other examples, the audio environment may be another type of environment, such as an office environment, an automobile environment, a train environment, a street or sidewalk environment, a park environment, etc. The interface system 705 may, in some implementations, be configured for exchanging control information and associated data with audio devices of the audio environment. The control information and associated datamay, in some examples, pertain to one or more software applications that the apparatus 700 is executing.
[0092] The interface system 705 may, in some implementations, be configured for receiving, or for providing, a content stream. The content stream may include audio data. The audio data may include, but may not be limited to, audio signals. In some instances, the audio data may include spatial data, such as channel data and / or spatial metadata. In some examples, the content stream may include video data and audio data corresponding to the video data.
[0093] The interface system 705 may include one or more network interfaces and / or one or more external device interfaces (such as one or more universal serial bus (USB) interfaces). According to some implementations, the interface system 705 may include one or more wireless interfaces. The interface system 705 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system and / or a gesture sensor system. In some examples, the interface system 705 may include one or more interfaces between the control system 710 and a memory system, such as the optional memory system 715 shown in Figure 7. However, the control system 710 may include a memory system in some instances. The interface system 705 may, in some implementations, be configured for receiving input from one or more microphones in an environment.
[0094] The control system 710 may, for example, include a general purpose single- or multi- chip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.
[0095] In some implementations, the control system 710 may reside in more than one device. For example, in some implementations a portion of the control system 710 may reside in a device within one of the environments depicted herein and another portion of the control system 710 may reside in a device that is outside the environment, such as a server, a mobile device (e.g., a smartphone or a tablet computer), etc. In other examples, a portion of the control system 710 may reside in a device within one environment and another portion of the control system 710 may reside in one or more other devices of the environment. For example, a portion of the control system 710 may reside in a device that is implementing a cloud-based service, such as a server, and another portion of the control system 710 may reside in another device that is implementing the cloud- based service, such as another server, a memory device, etc. The interface system 705 also may,in some examples, reside in more than one device. In some implementations, a portion of a control system may reside in or on an earbud.
[0096] In some implementations, the control system 710 may be configured for performing, at least in part, the methods disclosed herein. According to some examples, the control system 710 may be configured for implementing methods of determining noise covariances, determining filter weights based on one or more noise covariances, generating filtered output signals based on determined filter weights, or the like.
[0097] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non- transitory media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. The one or more non-transitory media may, for example, reside in the optional memory system 715 shown in Figure 7 and / or in the control system 710. Accordingly, various innovative aspects of the subject matter described in this disclosure can be implemented in one or more non-transitory media having software stored thereon. The software may, for example, be executable by one or more components of a control system such as the control system 710 of Figure 7.
[0098] In some examples, the apparatus 700 may include the optional microphone system 720 shown in Figure 7. The optional microphone system 720 may include one or more microphones. In some implementations, one or more of the microphones may be part of, or associated with, another device, such as a speaker of the speaker system, a smart audio device, etc. In some examples, the apparatus 700 may not include a microphone system 720. However, in some such implementations the apparatus 700 may nonetheless be configured to receive microphone data for one or more microphones in an audio environment via the interface system 710. In some such implementations, a cloud-based implementation of the apparatus 700 may be configured to receive microphone data, or a noise metric corresponding at least in part to the microphone data, from one or more microphones in an audio environment via the interface system 710.
[0099] According to some implementations, the apparatus 700 may include the optional loudspeaker system 725 shown in Figure 7. The optional loudspeaker system 725 may include one or more loudspeakers, which also may be referred to herein as “speakers” or, more generally, as “audio reproduction transducers.” In some examples (e.g., cloud-based implementations), the apparatus 700 may not include a loudspeaker system 725. In some implementations, the apparatus700 may include headphones. Headphones may be connected or coupled to the apparatus 700 via a headphone jack or via a wireless connection (e.g., BLUETOOTH).
[0100] Some aspects of present disclosure include a system or device configured (e.g., programmed) to perform one or more examples of the disclosed methods, and a tangible computer readable medium (e.g., a disc) which stores code for implementing one or more examples of the disclosed methods or steps thereof. For example, some disclosed systems can be or include a programmable general purpose processor, digital signal processor, or microprocessor, programmed with software or firmware and / or otherwise configured to perform any of a variety of operations on data, including an embodiment of disclosed methods or steps thereof. Such a general purpose processor may be or include a computer system including an input device, a memory, and a processing subsystem that is programmed (and / or otherwise configured) to perform one or more examples of the disclosed methods (or steps thereof) in response to data asserted thereto.
[0101] Some embodiments may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) that is configured (e.g., programmed and otherwise configured) to perform required processing on audio signal(s), including performance of one or more examples of the disclosed methods. Alternatively, embodiments of the disclosed systems (or elements thereof) may be implemented as a general purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include an input device and a memory) which is programmed with software or firmware and / or otherwise configured to perform any of a variety of operations including one or more examples of the disclosed methods. Alternatively, elements of some embodiments of the inventive system are implemented as a general purpose processor or DSP configured (e.g., programmed) to perform one or more examples of the disclosed methods, and the system also includes other elements (e.g., one or more loudspeakers and / or one or more microphones). A general purpose processor configured to perform one or more examples of the disclosed methods may be coupled to an input device (e.g., a mouse and / or a keyboard), a memory, and a display device.
[0102] Another aspect of present disclosure is a computer readable medium (for example, a disc or other tangible storage medium) which stores code for performing (e.g., coder executable to perform) one or more examples of the disclosed methods or steps thereof.
[0103] Various aspects of the present disclosure may be appreciated from the following Enumerated Example Embodiments (EEEs):
[0104] EEE 1. A method of enhancing audio signals, the method comprising: obtaining an input signal obtained using multiple microphones of one or more audio capture devices; determining a beamforming noise covariance and a wind reduction noise covariance associated with the input signal; determining a weighted average of the beamforming noise covariance and the wind reduction noise covariance to determine an aggregate noise covariance; and generating an output signal at least in part by filtering a representation of the input signal using filter weights determined based on the aggregate noise covariance, such that the output signal is enhanced relative to the input signal with respect to directionality and wind noise reduction.
[0105] EEE 2. The method of EEE 1, further comprising generating the filter weights by providing the aggregate noise covariance as an input parameter to a minimum variance distortionless response (MVDR) algorithm.
[0106] EEE 3. The method of any one of EEEs 1 or 2, wherein the beamforming noise covariance is determined based at least in part on a direction of arrival associated with the input signal.
[0107] EEE 4. The method of any one of EEEs 1-3, wherein the beamforming noise covariance is determined based at least in part on a metric that indicates a diffuseness associated with the input signal.
[0108] EEE 5. The method of any one of EEEs 1-4, wherein the wind reduction noise covariance is determined based at least in part on a wind level metric associated with the input signal.
[0109] EEE 6. The method of any one of EEEs 1-5, wherein determining the weighted average of the beamforming noise covariance and the wind reduction noise covariance is based at least in part on a weighting factor, wherein the weighting factor is determined based on at least one of: a metric that indicates a diffuseness associated with the input signal, a pre-determined value of the weighting factor, user input, or a ratio of a power associated with the beamforming noise covariance to a power associated with a wind reduction noise covariance.
[0110] EEE 7. The method of EEE 6, wherein the user input is received via a user interface.
[0111] EEE 8. The method of EEE 7, wherein the input signal is pre-recorded audio signal, and wherein the user interface is used to present the pre-recorded audio signal.
[0112] EEE 9. The method of any one of EEEs 1-8, further comprising transforming the input signal to a frequency domain representation of the input signal, and wherein the filter weights are used to filter the frequency domain representation of the input signal.
[0113] EEE 10. The method of EEE 9, wherein the beamforming noise covariance, the wind reduction noise covariance, and the aggregate noise covariance are determined on a per-frequency band basis for the frequency domain representation of the input signal.
[0114] EEE 11. The method of EEE 10, wherein the beamforming noise covariance is determined based at least in part on a metric that indicates a diffuseness associated with the input signal, and wherein a degree to which the beamforming noise covariance is weighted relative to the wind reduction noise covariance depends on the frequency band and the metric that indicates the diffuseness associated with the input signal.
[0115] EEE 12. The method of any one of EEEs 1-11, wherein the output signal comprises stereo signals.
[0116] EEE 13. The method of EEE 12, wherein the beamforming noise covariance is determined based on beamforming associated with two different looking directions associated with each signal of the stereo signals.
[0117] EEE 14. The method of any one of EEEs 1-13, wherein the one or more audio capture devices is one of, or a combination of: a mobile phone, a portable microphone, earbuds, a camera device, or an augmented reality (AR) or virtual reality (VR) headset.
[0118] EEE 15. The method of any one of EEEs 1-14, further comprising storing the output signal, the beamforming noise covariance, and / or the wind reduction noise covariance.
[0119] EEE 16. The method of any one of EEEs 1-15, further comprising: causing the output signal to be rendered to generate a rendered signal; and causing the rendered signal to be played back.
[0120] EEE 17. A system comprising: one or more processors; and a non-transitory computer- readable medium storing instructions that, upon execution by the one or more processors, cause the one or more processors to perform operations of EEEs 1-16.
[0121] EEE 18. A non-transitory computer-readable medium storing instructions that, upon execution by one or more processors, cause the one or more processors to perform operations of EEEs 1-16.
[0122] While specific embodiments of the present disclosure and applications of the disclosure have been described herein, it will be apparent to those of ordinary skill in the art that many variations on the embodiments and applications described herein are possible without departing from the scope of the disclosure described and claimed herein. It should be understood that while certain forms of the disclosure have been shown and described, the disclosure is not to be limited to the specific embodiments described and shown or the specific methods described.
Claims
CLAIMS 1. A method of enhancing audio signals, the method comprising: obtaining an input signal obtained using multiple microphones of one or more audio capture devices; determining a beamforming noise covariance and a wind reduction noise covariance associated with the input signal; determining a weighted average of the beamforming noise covariance and the wind reduction noise covariance to determine an aggregate noise covariance; and generating an output signal at least in part by filtering a representation of the input signal using filter weights determined based on the aggregate noise covariance, such that the output signal is enhanced relative to the input signal with respect to directionality and wind noise reduction.
2. The method of claim 1, further comprising generating the filter weights by providing the aggregate noise covariance as an input parameter to a minimum variance distortionless response (MVDR) algorithm.
3. The method of any one of claims 1 or 2, wherein the beamforming noise covariance is determined based at least in part on a direction of arrival associated with the input signal.
4. The method of any one of claims 1-3, wherein the beamforming noise covariance is determined based at least in part on a metric that indicates a diffuseness associated with the input signal.
5. The method of any one of claims 1-4, wherein the wind reduction noise covariance is determined based at least in part on a wind level metric associated with the input signal.
6. The method of any one of claims 1-5, wherein determining the weighted average of the beamforming noise covariance and the wind reduction noise covariance is based at least in part on a weighting factor, wherein the weighting factor is determined based on at least one of: a metric that indicates a diffuseness associated with the input signal, a pre- determined value of the weighting factor, user input, or a ratio of a power associated withthe beamforming noise covariance to a power associated with a wind reduction noise covariance.
7. The method of claim 6, wherein the user input is received via a user interface.
8. The method of claim 7, wherein the input signal is pre-recorded audio signal, and wherein the user interface is used to present the pre-recorded audio signal.
9. The method of any one of claims 1-8, further comprising transforming the input signal to a frequency domain representation of the input signal, and wherein the filter weights are used to filter the frequency domain representation of the input signal.
10. The method of claim 9, wherein the beamforming noise covariance, the wind reduction noise covariance, and the aggregate noise covariance are determined on a per- frequency band basis for the frequency domain representation of the input signal.
11. The method of claim 10, wherein the beamforming noise covariance is determined based at least in part on a metric that indicates a diffuseness associated with the input signal, and wherein a degree to which the beamforming noise covariance is weighted relative to the wind reduction noise covariance depends on the frequency band and the metric that indicates the diffuseness associated with the input signal.
12. The method of claim 11, wherein the beamforming noise covariance is determined based on beamforming associated with two different looking directions associated with each signal of the stereo signals.
13. The method of any one of claims 1-12, wherein the one or more audio capture devices is one of, or a combination of: a mobile phone, a portable microphone, earbuds, a camera device, or an augmented reality (AR) or virtual reality (VR) headset.
14. The method of any one of claims 1-13, further comprising storing the output signal, the beamforming noise covariance, and / or the wind reduction noise covariance.
15. The method of any one of claims 1-14, further comprising: causing the output signal to be rendered to generate a rendered signal; andcausing the rendered signal to be played back.
16. The method of claim 1, wherein the beamforming noise covariance indicates a degree to which the current frame of the audio signal contains diffuse sound and / or the direction of arrival is outside a region of interest, and / or wherein the wind reduction noise covariance indicates a power spectral density of each microphone signal of the input signal when wind exists.
17. The method of any preceding claim, wherein the output signal comprises a stereo signals.
18. A system comprising: one or more processors; and a non-transitory computer- readable medium storing instructions that, upon execution by the one or more processors, cause the one or more processors to perform operations of any preceding claim.
19. A non-transitory computer-readable medium storing instructions that, upon execution by one or more processors, cause the one or more processors to perform operations of any one of claims 1-18.
Citation Information
Patent Citations
Globally optimized least-squares post-filtering for speech enhancement
US20170221502A1
Cited By
A sound source positioning method and system
CN122525492A