Beamforming-related adaptation for frequency equalization

By using spatial beamforming and decorrelation processing of audio devices, combined with a frequency equalizer, the performance limitations of existing audio beamforming methods in multi-source noise environments are solved, achieving efficient and low-complexity audio source separation and quality improvement.

CN121241393APending Publication Date: 2025-12-30KONINKLIJKE PHILIPS NV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202480035606.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-05-30
Filing Date
2024-05-16
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Existing audio beamforming methods suffer from poor performance in many scenarios and applications, especially when there are multiple audio and noise sources. They are difficult to effectively separate the desired audio source, and they also suffer from high complexity, resource conflicts, and a lack of flexibility and adaptability.

Method used

An audio device is used for spatial beamforming and decorrelation processing. Combined with a frequency equalizer, the beamforming and decorrelation parameters are adaptively adjusted to match the frequency equalization to improve the separation effect of the audio signal, reduce noise sensitivity, and compensate for frequency distortion through the equalizer.

Benefits of technology

It improves the capture quality and separation accuracy of audio signals, reduces frequency distortion, achieves efficient audio source extraction with low complexity, adapts to different environments and scenarios, and provides more natural sound capture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121241393A_ABST
    Figure CN121241393A_ABST
Patent Text Reader

Abstract

An audio device comprises a receiver (101) arranged to receive an audio signal, an audio beamformer (103) generating a beamformed output audio signal from the audio signal. The beamformer (103) comprises a signal processor (201) that applies signal processing to the audio signal. Signal processing includes spatial beamforming and spatial decorrelation, which in many cases is adaptive spatial decorrelation. The decorrelation adapter (203) may adapt the decorrelation parameters according to the audio signal, and the beamforming adapter (205) adapts the beamforming parameters according to the beamformed output audio signal. An equalizer (105) applies frequency equalization to the beamformed output audio signal to generate an output signal, and an equalization adapter (107) adapts the frequency equalization according to beamforming parameters and generally also according to decorrelation parameters. The method may provide improved fidelity of the desired audio source and may attenuate interfering audio sources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to apparatus and methods for generating audio output signals, and specifically, but not exclusively, to extracting audio from a desired audio source, such as a desired speaker, using beamforming. Background Technology

[0002] Audio capture, especially speech, has become increasingly important over the past few decades. For example, capturing speech or other audio is becoming increasingly crucial for a wide range of applications, including telecommunications, teleconferencing, gaming, and audio user interfaces. However, a problem in many scenarios and applications is that the desired audio source is often not the only audio source in the environment. Instead, in a typical audio environment, there are many other audio / noise sources captured by the microphone. Audio processing is commonly used to improve audio capture, and in particular, post-processing of the captured audio time intervals improves the resulting audio signal.

[0003] In many embodiments, audio can be represented by multiple different audio signals reflecting the same audio scene or environment. In particular, in many practical applications, audio is captured by multiple microphones at different locations. For example, a linear array of multiple microphones is often used to capture audio from an environment, such as a room. The use of multiple microphones allows for the capture of spatial information about the audio. Many different applications can utilize such spatial information, thereby enabling improved and / or new services.

[0004] A frequently used approach is to attempt to separate audio sources by applying beamforming to form a beam pointing in the direction of arrival of audio from a particular audio source. However, while this can provide favorable performance in many scenarios, it is not optimal in all cases. For example, in some situations it may not provide optimal source separation, and indeed, in some applications, such spatial beamforming may not provide the ideal audio properties for further processing to achieve a given effect.

[0005] Therefore, while spatial audio source separation, especially audio beamforming-based separation, is highly advantageous in many scenarios and applications, improvements in the performance and operation of this approach are desired. However, low complexity and / or resource usage (e.g., computational resources) are also generally desired, and these preferences often conflict with each other.

[0006] Therefore, improved approaches will be advantageous, and specifically, approaches that allow for reduced complexity, increased flexibility, facilitated implementation, reduced costs, improved audio capture, improved audio source differentiation, improved audio source separation, improved audio / voice application support, reduced reliance on known or static acoustic properties, improved flexibility and customization for different audio environments and scenarios, improved audio beamforming, and improved trade-offs between performance and complexity / resource usage, and / or improved performance, will be advantageous. Summary of the Invention

[0007] Therefore, the present invention seeks to mitigate, alleviate or eliminate one or more of the aforementioned disadvantages, preferably alone or in any combination.

[0008] According to one aspect of the present invention, an audio device is provided, comprising: a receiver arranged to receive a first set of audio signals, the first set of audio signals including audio signals capturing audio of a scene from different locations; an audio beamformer arranged to generate a beamformed output audio signal based on the first set of audio signals, the audio beamformer comprising: a signal processor arranged to generate the beamformed output audio signal based on signal processing of the first set of audio signals, the signal processing including spatial beamforming and spatial decorrelation for the first set of audio signals, the spatial decorrelation depending on a set of decorrelation parameters and the spatial beamforming depending on a set of beamforming parameters, the beamforming parameters being parameters describing at least one of filtering and weighting of an input signal for a combination of beamforming; a beamforming adapter arranged to adapt the beamforming parameter set to the beamforming output audio signal; an equalizer arranged to apply frequency equalization to the beamformed output audio signal to generate an output signal; and an equalization adapter arranged to adapt the frequency equalization to the beamforming parameter set.

[0009] The method can provide improved operation and / or performance in many embodiments. Specifically, it can allow for improved beamforming to focus on specific audio sources in a scene. It can allow for improved extraction / separation of audio from specific sources even when other audio and noise sources are present in the scene.

[0010] This method allows the operation of audio devices to be effectively adapted to current conditions, particularly the acoustic and spatial properties of the audio source and the scene. It can provide reduced sensitivity to noise and unwanted audio sources in the scene and captured audio.

[0011] This method allows for efficient operation and low complexity in many embodiments. Different adapters can work collaboratively to provide improved separation of the desired audio source from other captured audio in the scene. Furthermore, using multiple adapters can provide improved operation while allowing the use of lower-complexity adapter algorithms and standards.

[0012] This method can provide a higher quality output audio signal. In many embodiments, the method can reduce or mitigate frequency distortion. This method can result in a more natural sound capture of the desired audio source. In many embodiments, higher capture audio fidelity can be achieved.

[0013] An equalizer can be configured to apply frequency equalization to a beamformed output audio signal by applying its frequency response. An equalizer can also be configured to apply frequency equalization to a beamformed output audio signal by filtering it to the beamformed output audio signal, which has a non-constant amplitude frequency response. An equalizer adapter can be configured to adapt to the frequency response.

[0014] According to an optional feature of the invention, the equalization adapter is arranged to adapt frequency equalization based on a set of decorrelation parameters.

[0015] This can provide improved performance and / or operation in many embodiments. In particular, in many embodiments, it can generate an output audio signal that more accurately captures the desired audio source in an audio scene. In many scenarios, it can allow for improved adaptation to frequency equalization, thereby reducing frequency distortion.

[0016] According to an optional feature of the invention, the signal processor includes: a spatial decorrelation unit arranged to receive a first set of audio signals and generate a decorrelated first set of audio signals by performing spatial decorrelation; and a beamforming circuit arranged to perform spatial beamforming by combining the decorrelated first set of audio signals, the combination depending on a set of beamforming parameters.

[0017] This can provide improved performance and / or operation in many embodiments. In many embodiments, it can provide particularly accurate sound capture while facilitating implementation and maintaining low complexity.

[0018] According to an optional feature of the invention, the equalization adapter is arranged to adapt the frequency equalization to have a minimum phase frequency response.

[0019] This can provide improved performance and / or operation in many embodiments.

[0020] According to an optional feature of the invention, the equalizer adapter is arranged to adapt the frequency equalization to have a low-pass frequency response.

[0021] This can provide improved performance and / or operation in many embodiments. In particular, in many embodiments, it can generate an output audio signal that more accurately captures the desired audio source in an audio scene. In many scenarios, it can allow for improved frequency equalization adaptation, thereby reducing frequency distortion.

[0022] According to an optional feature of the invention, the equalizer is arranged to switch from a first frequency equalizer to a second frequency equalizer by: determining a first intermediate sample of the output signal for the first frequency equalizer and a second intermediate sample of the output signal for the second frequency equalizer, and generating a sample of the output signal as a weighted combination of the first intermediate sample and the second intermediate sample, wherein the equalizer is arranged to gradually change the relative weights of the first intermediate sample and the second intermediate sample during the switching time interval.

[0023] This can provide improved performance and / or operation in many embodiments. It can reduce audio artifacts in many scenarios.

[0024] According to an optional feature of the invention, spatial decorrelation is adaptive spatial decorrelation, and the audio device further includes a decorrelation adapter arranged to adapt the decorrelation parameter set according to the first audio signal set.

[0025] This can provide improved performance and / or operation in many embodiments.

[0026] According to an optional feature of the invention, the audio device further includes an audio detector arranged to determine a set of active time intervals during which the audio source is active and a set of inactive time intervals during which the audio source is inactive; and wherein at least one of the adaptation to the set of beamforming parameters and the adaptation to the set of decorrelation coefficients is different for the set of active time intervals and the set of inactive time intervals.

[0027] This can provide improved performance and / or operation in many embodiments. It can typically provide improved adaptation, which often results in improved extraction of the desired audio source.

[0028] The active time interval can be the expected / desired / voice audio source active time interval, and the inactive time interval can be the expected / desired / voice audio source inactive time interval.

[0029] In some embodiments, the beamforming adapter is arranged to adapt the first filter set and the second filter set at a higher adaptation rate during the active time interval set compared to the inactive time interval set.

[0030] This can provide improved performance and / or operation in many embodiments. It can typically provide improved adaptation, which often results in improved extraction of the desired audio source.

[0031] In some embodiments, the beamforming adapter is arranged to adapt the first filter set and the second filter set during only one of the active time interval set and the inactive time interval set.

[0032] This can provide improved performance and / or operation in many embodiments. It can typically provide improved adaptation, which often results in improved extraction of the desired audio source.

[0033] In some embodiments, the beamforming adapter is arranged to adapt to the first filter set and the second filter set during an active time interval set rather than during an inactive time interval set.

[0034] In some embodiments, the decorrelation adapter is arranged to adapt decorrelation coefficients at a higher adaptation rate during an inactive time interval set compared to an active time interval set.

[0035] This can provide improved performance and / or operation in many embodiments. It can typically provide improved adaptation, which often results in improved extraction of the desired audio source.

[0036] In some embodiments, the decorrelation adapter is configured to adapt decorrelation coefficients only during a set of time intervals in the set of inactive time intervals and the set of active time intervals.

[0037] This can provide improved performance and / or operation in many embodiments. It can typically provide improved adaptation, which often results in improved extraction of the desired audio source.

[0038] In some embodiments, the adaptive coefficient processor is configured to adapt the decorrelation coefficients during the inactive time interval set rather than during the active time interval set.

[0039] According to an optional feature of the invention, the signal processor includes: a first filter set arranged to filter a first audio signal set; a combiner arranged to combine the outputs of the first filter set to generate the beamformed output audio signal; a feedback circuit including a second filter set arranged to generate a second audio signal set based on filtering the beamformed output audio signal, each filter in the second filter set having a frequency response that is the complex conjugate of the filters in the first filter set; a first spatial filter set arranged to apply a first spatial filter to at least one of the first audio signal set and the second audio signal set, the first spatial filter set having coefficients determined according to decorrelation coefficients included in the decorrelation parameters; and wherein the beamformer adapter is arranged to adapt the first filter set and the second filter set in response to a comparison of the first audio signal set and the second audio signal set; and the decorrelation adapter is arranged to determine decorrelation coefficients for the spatial decorrelation filter set to generate a decorrelated output signal based on the first audio signal set, the decorrelation adapter being arranged to adapt the decorrelation coefficients in response to an update value determined based on the first audio signal set.

[0040] This can provide improved performance and / or operation in many embodiments. It can provide particularly advantageous beamforming for specific methods and can interoperate efficiently with spatial decorrelation and equalization.

[0041] Each filter in the first filter set is linked to a filter in the second filter set, wherein the filter has a complex conjugate frequency response that is the frequency response of the filter in the first filter set.

[0042] The first and second filter sets can have an equal number of linked paired filters with complex conjugate frequency responses. For each audio signal in the first audio signal set, there can be one filter in the first filter set, one linked / paired filter (with complex conjugate frequency response) in the second filter set, and one audio signal in the second audio signal set. Comparisons can be between linked / paired audio signals in the first and second audio signal sets. Adaptation to a given filter in the first filter set and a given linked filter in the second filter set can be responsive to / dependent on (possibly only) a comparison of the signal in the first audio signal set filtered by the given filter in the first filter set with the signal in the second audio signal set generated by filtering the beamformed output audio signal using the given linked filter in the second filter set. Adaptation can reduce the differences.

[0043] Each audio signal in the second set of audio signals can be an estimate of the contribution of the audio captured from the beamforming output audio signal to the linked audio signals in the first set of audio signals, and therefore can typically be an estimate of the contribution from the desired / expected source.

[0044] A spatial decorrelation filter set may include a spatial filter for each signal in a first set of signals. Each filter in the spatial decorrelation filter set may generate a filtered / modified version of an audio signal in the first set of audio signals. The spatial decorrelation filter set may together generate a modified first set of audio signals with a high degree of decorrelation. The spatial decorrelation filter may perform filtering on the first set of audio signals. The output of the decorrelation filter (the corresponding modified audio signal in the first set of audio signals) may depend on multiple audio signals in the (unmodified) first set of audio signals. The spatial decorrelation filter may specifically be a frequency domain filter, and the decorrelation coefficients may be frequency domain coefficients. For a given spatial decorrelation filter, the output value for a given frequency bin at a given time may be a weighted combination of multiple values ​​for a given frequency bin in the (unmodified) first set of audio signals at a given time. The decorrelated output signal may have reduced normalized cross-channel signal correlation relative to the first set of input signals of the spatial decorrelation filter set.

[0045] A spatial filter set may include one spatial filter for each signal in a first signal set / second signal set. Each filter in the spatial filter set may generate a filtered / modified version of an audio signal from the first or second audio signal set. The spatial filter set may together generate a modified first or second audio signal set. The spatial filter may perform filtering on the first audio signal set. The output of the spatial filter (the corresponding modified audio signal in the first or second audio signal set) may depend on multiple audio signals in the (unmodified) first or second audio signal set. The spatial filter may specifically be a frequency domain filter, and the coefficients may be frequency domain coefficients. For a given spatial filter, the output value for a given frequency bin at a given time may be a weighted combination of multiple values ​​for the given frequency bin from the (unmodified) first or second audio signal set at a given time.

[0046] According to an optional feature of the invention, the first set of spatial filters is arranged to filter the first set of audio signals.

[0047] This can provide improved performance and / or operation in many embodiments. In many embodiments and scenarios, this can provide particularly attractive performance and / or implementation methods.

[0048] In some embodiments, the first spatial filter set is arranged to have coefficients that are set to decorrelation coefficients determined for the spatial decorrelation filter set.

[0049] This can provide improved performance and / or operation in many embodiments. In many embodiments and scenarios, this can provide particularly attractive performance and / or implementation methods.

[0050] According to an optional feature of the invention, the first filter set is arranged to filter the first audio signal set after filtering by the first spatial filter set, and the beamforming adapter is arranged to perform a comparison using the first audio signal set before filtering by the first spatial filter set.

[0051] This can provide improved performance and / or operation in many embodiments. In many embodiments and scenarios, this can provide particularly attractive performance and / or implementation methods.

[0052] In some embodiments, the first spatial filter set is arranged to have coefficients that match the coefficients of the cascaded spatial filters that are two decorrelation filters in the decorrelation filter set.

[0053] This can provide improved performance and / or operation in many embodiments. In many embodiments and scenarios, this can provide particularly attractive performance and / or implementation methods.

[0054] According to an optional feature of the invention, the first set of spatial filters is arranged to filter the second set of audio signals.

[0055] This can provide improved performance and / or operation in many embodiments. In many embodiments and scenarios, this can provide particularly attractive performance and / or implementation methods.

[0056] In some embodiments, the first spatial filter set is arranged to have coefficients determined in response to the inverse spatial decorrelation filter set, which is the inverse filter in the spatial decorrelation filter set.

[0057] This can provide improved performance and / or operation in many embodiments. In many embodiments and scenarios, this can provide particularly attractive performance and / or implementation methods.

[0058] In some embodiments, a first set of spatial filters is arranged to have coefficients that match the coefficients of a spatial filter, which is a cascade of two sets of spatial inverse filters, each set of spatial inverse filters including an inverse filter from a set of spatial decorrelation filters.

[0059] This can provide improved performance and / or operation in many embodiments. In many embodiments and scenarios, this can provide particularly attractive performance and / or implementation methods.

[0060] According to an optional feature of the invention, the decorrelation adapter is arranged to determine an adaptive spatial decorrelation as a set of spatial decorrelation filters by performing the following steps: each output audio signal of a filter in the set of spatial decorrelation filters is linked to an input audio signal in a first set of audio signals; the first set of audio signals is segmented into time segments, and for at least some time segments, the following steps are performed: generating a frequency bin representation of the first set of audio signals, each frequency bin in the frequency bin representation of the first set of audio signals including a frequency bin value for each audio signal in the first set of audio signals; generating a frequency bin representation of an output set of signals, each frequency bin in the frequency bin representation of the output set including The frequency bin value of each output signal in the output signal, the frequency bin value of a given output signal for a given frequency bin in the set of output signals is generated as a weighted combination of the frequency bin values ​​of the first audio signal set for a given frequency bin, the weighted combination having a decorrelation coefficient as a weight; and in response to a correlation measure between a first previous frequency bin value of the first output signal linked to the first input audio signal for a first frequency bin and a second previous frequency bin value of the second output signal for a first frequency bin, the first weight of the contribution of the second frequency bin value of the first frequency bin for the second input audio signal linked to the second output signal to the first frequency bin value of the first frequency bin for the first output signal.

[0061] This can provide improved performance and / or operation in many embodiments. In many embodiments and scenarios, this can provide particularly attractive performance and / or implementation methods.

[0062] This can provide favorable generation of the output audio signal, which typically exhibits increased decorrelation compared to the input signal. In many embodiments, the method can provide efficient adaptation of the operation, thereby improving decorrelation. Adaptation can typically be implemented with low complexity and / or resource usage. The method can specifically apply local adaptations to individual weights while still achieving efficient global adaptation.

[0063] The generation of the output signal set can be adapted to provide increased decorrelation relative to the input signal, which can provide improved audio processing, particularly beamforming, in many embodiments and for many applications.

[0064] The first and second output audio signals can typically be different output audio signals.

[0065] According to an optional feature of the invention, the decorrelation adapter is arranged to update a first weight in response to the product of a first value and a second value, the first value being one of a first previous frequency bin value and a second previous frequency bin value, and the second value being the complex conjugate of the other of the first previous frequency bin value and the second previous frequency bin value.

[0066] This can provide improved performance and / or operation in many embodiments. It can typically provide improved adaptation, resulting in increased decorrelation of the output audio signal in many scenarios.

[0067] In some embodiments, the audio device may be arranged to update a second weight of the contribution of a third frequency bin value, which is a first frequency bin value for a first frequency bin value, to the first frequency bin value in response to the magnitude of a first previous frequency bin value.

[0068] This can provide improved performance and / or operation in many embodiments. It can generally provide improved adaptation, resulting in increased decorrelation of the output audio signal in many scenarios. It can particularly provide improved adaptation to the generated output signal. In many embodiments, the update of the weights reflecting the contribution of the input signal from the link to the output signal can depend on the signal amplitude / vibration of the input signal of the link. For example, the update may seek to compensate for the weights used for the level of the input signal to generate a normalized output signal.

[0069] This method allows for normalization / signal compensation / level compensation to provide, for example, the desired output level.

[0070] In some embodiments, the audio device may be arranged to set a predetermined value for the weight of the contribution of a third frequency bin value, which is a frequency bin value of a first input audio signal for a first frequency bin, to the first frequency bin value.

[0071] This can provide improved performance and / or operation in many embodiments. It can typically provide improved adaptation, resulting in increased decorrelation of the output audio signal in many scenarios. It can provide improved adaptation in many embodiments while ensuring that the adaptation converges towards a non-zero signal level. It can interact very effectively with adaptations of weights on signals not used for linking.

[0072] In many embodiments, the adapter can be arranged to keep the weights constant and not adjust or update the weights.

[0073] In some embodiments, the audio device may be arranged to constrain the weights of the contribution of a third frequency bin value to a first frequency bin value to a real value, the third frequency bin value being the frequency bin value of the first frequency bin for a first input audio signal.

[0074] This can provide improved performance and / or operation in many embodiments. It can typically provide improved adaptation, resulting in increased decorrelation of the output audio signal in many scenarios.

[0075] The weights between the input / output signals of a link can be advantageously determined as / constrained as real-valued weights. This can lead to improved performance and fit, thereby ensuring convergence on non-zero level solutions.

[0076] In some embodiments, the audio device may be arranged to set a second weight as the complex conjugate of a first weight, the second weight being a weight for the contribution of a fourth frequency bin value from a first input audio signal to a first frequency bin for a second output audio signal.

[0077] This can provide improved performance and / or operation in many embodiments. It can typically provide improved adaptation, resulting in increased decorrelation of the output audio signal in many scenarios.

[0078] In many embodiments, the two weights used for two pairs of input / output signals can be complex conjugates of each other.

[0079] In some embodiments, the weights used for the weighted combination of input audio signals other than the first input audio signal are complex weights.

[0080] This can provide improved performance and / or operation in many embodiments. The use of complex values ​​for the weights of the non-linked input signals provides improved frequency domain operation.

[0081] In some embodiments, the audio device may be arranged to determine the output bin value for a given frequency bin ω from the following formula: Where y(ω) is a vector containing the frequency bin values ​​for the output audio signal at a given frequency bin ω; x(ω) is a vector containing the frequency bin values ​​for the input audio signal at a given frequency bin ω; and It is a matrix with rows containing weights that include a weighted combination of the output audio signals.

[0082] This can provide improved performance and / or operation in many embodiments. It can typically provide improved adaptation, resulting in increased decorrelation of the output audio signal in many scenarios.

[0083] matrix Hermitian can be advantageously used. In many embodiments, the matrix... The diagonal can be constrained to real values, can be set to predetermined values, and / or can be left unupdated / adapted, but can be kept as fixed values. Weights / coefficients outside the diagonal can typically be complex values.

[0084] In some embodiments, the audio device may be arranged to adapt the matrix according to the following formula. weight w ij : Where i is a matrix The row index, j is the matrix. The column index, k is the time segment index, ω represents the frequency bin, and... These are scaling parameters used to adapt to different speeds.

[0085] This can provide improved performance and / or operation in many embodiments. It can typically provide improved adaptation, resulting in increased decorrelation of the output audio signal in many scenarios.

[0086] In some embodiments, the audio device may be arranged to compensate for the correlation value of the signal level for the first frequency compartment.

[0087] This can provide improved performance and / or operation in many embodiments. It can typically provide improved adaptation, resulting in increased decorrelation of the output audio signal in many scenarios. It can allow for compensation for the update rate in response to signal variations.

[0088] In some embodiments, the audio device may be arranged to initialize the weights used for weighted combination to include at least one zero-value weight and one non-zero-value weight.

[0089] This can provide improved performance and / or operation in many embodiments. It can typically provide improved adaptation, resulting in increased decorrelation of the output audio signal in many scenarios. It can allow for more efficient and / or faster adaptation and convergence towards favorable decorrelation. In many embodiments, the matrix... Initialization can be performed using zero values ​​for the weights or coefficients of the non-linked signals and fixed non-zero real values ​​for the linked signals. Typically, for the weights on the diagonal, the weights can be set to, for example, 1, and all other weights can initially be set to zero.

[0090] In some embodiments, the weighted combination includes applying a time-domain window to the frequency representation of the weights formed by weights for the first and second input audio signals for different frequency bins.

[0091] This can provide improved performance and / or operation in many embodiments. It can typically provide improved adaptation, resulting in increased decorrelation of the output audio signal in many scenarios.

[0092] Applying time-domain windowing to the frequency representation of weights can include: converting the frequency representation of weights to a time-domain representation of weights; applying a window to the time-domain representation to generate a modified time-domain representation; and converting the modified time-domain representation to the frequency domain.

[0093] According to one aspect of the present invention, an operating method for an audio device is provided, the method comprising: receiving a first set of audio signals, the first set of audio signals including audio signals capturing audio of a scene from different locations; generating a beamformed output audio signal based on the first set of audio signals, wherein generating the beamformed output audio signal comprises: generating the beamformed output audio signal based on signal processing of the first set of audio signals, the signal processing including spatial beamforming and adaptive spatial decorrelation for the first set of audio signals, the adaptive spatial decorrelation depending on a decorrelation parameter set, and the spatial beamforming depending on a beamforming parameter set, the beamforming parameters being parameters describing at least one of filtering and weighting of an input signal for a combination of beamforming; adapting the decorrelation parameter set according to the first set of audio signals; adapting the beamforming parameter set according to the beamformed output audio signal; applying frequency equalization to the beamformed output audio signal to generate an output signal; and adapting the frequency equalization according to the beamforming parameter set.

[0094] These and other aspects, features, and advantages of the invention will become apparent from the embodiments described below, and will be illustrated with reference to these embodiments. Attached Figure Description

[0095] Embodiments of the invention will be described by way of example only with reference to the accompanying drawings, wherein...

[0096] Figure 1 Examples of elements of an audio device according to some embodiments of the present invention are shown;

[0097] Figure 2 Examples of elements for an audio beamformer for an audio device according to some embodiments of the present invention are shown;

[0098] Figure 3 Examples of elements for an audio beamformer for an audio device according to some embodiments of the present invention are shown;

[0099] Figure 4 Examples of elements for an audio beamformer for an audio device according to some embodiments of the present invention are shown;

[0100] Figure 5 Examples of elements for an audio beamformer for an audio device according to some embodiments of the present invention are shown;

[0101] Figure 6 Examples of elements for an audio beamformer for an audio device according to some embodiments of the present invention are shown;

[0102] Figure 7 Examples of elements for an audio beamformer for an audio device according to some embodiments of the present invention are shown;

[0103] Figure 8 Examples of elements for an audio beamformer for an audio device according to some embodiments of the present invention are shown; and

[0104] Figure 9 Some elements of a processor for implementing an audio device are shown in some embodiments of the present invention. Detailed Implementation

[0105] The following description focuses on embodiments of the invention applicable to audio capture (e.g., voice capture for teleconferencing devices). However, it should be understood that the method is applicable to many other audio signals, audio processing systems, and scenarios for capturing and / or processing audio.

[0106] Figure 1 An example of an audio device arranged to generate an output audio signal based on a first set of audio signals representing audio captured from different locations is shown.

[0107] Figure 1 The audio device specifically includes a receiver 101 that receives a first set of audio signals, and in a specific example, receives a set of microphone signals from a set of microphones capturing an audio scene from different locations. The microphones may, for example, be arranged in a linear array and close to each other. For example, in many embodiments, the maximum distance between the points used to capture the audio signals may not exceed 1 meter, 50 cm, 25 cm, or even in some cases, not exceed 10 cm.

[0108] The input audio signal can be received from various sources, including internal or external sources. Hereinafter, an embodiment in which receiver 101 is coupled to multiple microphones (e.g., a linear array of microphones) that provide a set of input audio signals in the form of microphone signals will be described.

[0109] Receiver 101 is coupled to audio beamformer 103, which is arranged to generate a beamformed output audio signal based on a first set of audio signals. As will be described in more detail later, audio beamformer 103 is arranged to generate the beamformed output audio signal by performing signal processing, including spatial beamforming and spatial decorrelation, on the first set of audio signals. Thus, both beamforming and decorrelation are implemented by audio beamformer 103 to generate the beamformed output audio signal, which can be specifically guided to extract or isolate desired audio sources in the audio scene. As will be described in more detail later, audio beamformer 103 is an adaptive beamformer, wherein both the decorrelation function and the beamforming function are adapted, and audio beamformer 103 is specifically arranged to adapt its processing in response to the first set of input audio signals and the generated beamformed output audio signal.

[0110] The audio beamformer 103 accordingly provides processing to achieve the effects of both spatial decorrelation and adaptive beamforming. Spatial decorrelation is controlled by a set of decorrelation parameters, and spatial beamforming is based on a set of beamforming parameters. As will be described in more detail later, the decorrelation parameters can specifically be decorrelation coefficients used for a set of spatial filters arranged to generate N output signals from N input signals, where each output signal is generated as a weighted combination of samples from the N input signals. The decorrelation coefficients can specifically be weights used for such weighted combinations.

[0111] Beamforming parameters can specifically be parameters that describe or define the combination of signals to generate a beamformed output audio signal. Beamforming may include filtering multiple signals and combining the resulting filtered output signals. Beamforming parameters can specifically define filters (and specifically filter coefficients) and / or combinations (e.g., weights for each signal). It should be understood that applying weights to signals can be equivalently considered as part of filtering and / or as part of a combination operation. For example, if the beamformed output audio signal is generated by weighting the input signals (e.g., by complex weights) and then summing them, then the weights can be considered as filtering and / or the combination can be considered as including the weighting (in which case beamforming may not include any (other) filtering of the input signals). Specifically, the weights used for the input signals can be considered as filter coefficients.

[0112] The audio beamformer 103 can adapt spatial decorrelation by adapting decorrelation parameters and beamforming operation by adapting beamforming parameters.

[0113] Audio beamformer 103 is coupled to equalizer 105, which is arranged to apply frequency equalization to the beamformed output audio signal to generate an output signal for the audio device. Specifically, the equalizer may utilize a filter with a given frequency response to apply filtering to the beamformed output audio signal. The given frequency response is not fixed, constant, or predetermined, but can be dynamically adapted during operation.

[0114] The audio device includes an equalizer adapter 107, which is arranged to adapt frequency equalization according to a set of beamforming parameters used by the audio beamformer 103. The equalizer adapter 107 is coupled to the audio beamformer 103 and receives the currently applied beamforming parameters to determine the frequency response of the equalizer. Therefore, the equalizer adapter 107 is arranged to adapt the frequency response of the equalizer 105 according to the beamforming parameters (currently) applied by the audio beamformer 103.

[0115] In many embodiments, the adaptation of the frequency response of equalizer 105 may also depend on decorrelation parameters determined by audio beamformer 103 for adaptive decorrelation. Therefore, in many embodiments, equalizer adapter 107 also receives decorrelation parameters and continues to evaluate these decorrelation parameters to determine / adapt the current frequency response.

[0116] The inventors have recognized that the combination of spatial decorrelation and audio beamforming can provide improved performance, and in particular, allows for improved extraction of the desired audio source in the presence of strong and potentially dominant noise sources. Beamforming and decorrelation operations can work together to suppress unwanted sources for beamforming, while still allowing coherence to be maintained for effective beamforming toward the desired source.

[0117] However, the inventors also recognize that while this approach can provide effective performance and advantageous operation in many scenarios, the combined effects can also affect the overall frequency response and introduce some frequency distortion. Specifically, different frequency intervals can have different correlations, and the combined effect of decorrelation seeking to remove correlations and beamforming operations seeking to utilize correlations by combining desired signal components can lead to a non-flat frequency response. The inventors also recognize that the frequency effects are not constant or fixed, but rather depend significantly on the operations performed by the audio beamformer 103.

[0118] Figure 1The audio device reflects these understandings, and specifically, equalizer 105 can provide frequency compensation by applying a frequency response that compensates for the frequency distortion of audio beamformer 103. The frequency response of equalizer 105 can be determined such that the combined frequency response of audio beamformer 103 and equalizer 105 is flatter than the frequency response of audio beamformer 103. Equalizer adapter 107 adapts equalizer 105 such that the resulting compensated frequency response is dynamically changed to reflect the current operation of audio beamformer 103. Specifically, the compensated frequency response (the frequency response of equalizer 105) is adapted based on beamforming parameters and therefore on parameters of the signal combination as part of the combination. Advantageously, in many embodiments, the compensated frequency response can also be adapted based on decorrelation parameters. Thus, in many embodiments, the compensated frequency response can be adapted to reflect the current operation of both the decorrelation function and the beamforming function.

[0119] Therefore, the compensated frequency response is not determined based on the evaluation of signal components or the frequency content / distribution of the signal, or the frequency analysis or comparison of the signal components, but is controlled by parameters that determine the operation of the audio beamformer 103. These parameters are typically much slower than changes in signal properties and allow for much lower complexity operation with significantly reduced computational resources. However, it has been found that this method provides efficient and high-performance operation from which improved output signals can be generated. Specifically, the output signal can provide speech or other audio that sounds more natural and closer to the actual sound in the audio scene. Typically, frequency distortion can be reduced or even substantially removed from the output signal.

[0120] In many embodiments, the equalizer adapter 107 may be arranged to adapt the frequency equalization to have a low-pass frequency response. In many embodiments, the equalizer adapter 107 may be arranged to generate a compensated frequency response having an amplitude that monotonically decreases with increasing frequency in frequency intervals from 50 Hz or 100 Hz to 500 Hz, 1 kHz, 2 kHz, or 4 kHz. The equalizer adapter 107 may generate a frequency response to provide a low-frequency boost to the beamforming output audio signal.

[0121] The inventors have recognized that decorrelation can have a frequency effect on beamforming output because it can lead to decorrelation related to the frequency of the desired signal from which beamforming is adapted to extract it. For example, decorrelation of a noise source may also tend to affect lower frequencies of the desired audio source because lower frequencies tend to have higher correlations compared to higher frequencies due to their longer wavelengths. Therefore, for different audio sources and signals, lower frequencies may tend to have similar correlations across different microphones at different locations. Thus, decorrelation of one audio source (e.g., noise or the dominant source) will also tend to decorrelate other sources. This is more pronounced at lower frequencies compared to higher frequencies.

[0122] This decorrelation can tend to reduce beamforming efficiency because it attempts to utilize the correlation of the desired signal components in different captured signals, and thus may result in a reduction in amplitude. Since the desired signal decorrelation may be more pronounced at lower frequencies, this effect may cause a frequency-dependent attenuation of the audio captured by the beamformed output audio signal. The equalizer adapter 107 can accordingly control the equalizer 105 to provide bass enhancement (low-pass characteristics) to compensate for this attenuation.

[0123] In many embodiments, equalizer 105 may be arranged to determine the frequency equalization / compensation frequency response to have / be the minimum phase frequency response. In many embodiments, this can provide improved performance and perceived quality. It can reduce and specifically minimize the latency of equalizer 105, and therefore reduce and specifically minimize the overall latency of capture operation and audio devices.

[0124] In many embodiments, the equalization adapter 107 can be arranged to first determine the amplitude-compensated frequency response based on beamforming parameters (and optionally decorrelation parameters), and then proceed to determine a minimum-phase filter for that amplitude response. This approach will be described in more detail later.

[0125] It should be understood that the specific operations performed by the audio beamformer 103, as well as the specific decorrelation and beamforming combinations, will differ in different embodiments. The impact of beamforming operations on the overall frequency response will therefore depend on the specific details of the implementation, and accordingly, the specific adaptation for compensating the frequency response and its dependence on beamforming parameters will also vary and be implementation-specific.

[0126] It has been found that the combination of adaptive beamformer and spatial decorrelation function provides very good results, but at the cost of some (linear) frequency distortion. Due to decorrelation, the frequency response of the desired audio / speech in the beamformer output changes relative to the case without the decorrelation function. Typically, and depending on the location of both the desired audio source / speaker and the point noise source, as well as the size of the microphone array, decorrelation can lead to attenuation at low frequencies. At low frequencies, the wavelength of sound waves is large (e.g., greater than 3 meters at 100 Hz), and the input covariance matrices of the point noise source and the desired audio / speech source will not differ significantly at low frequencies. The decorrelation function decorrelates the noise signal, and therefore the noise signal will be attenuated, but as a side effect, the desired audio / speech signal will also be (partially) attenuated at low frequencies. However, the signal-to-noise ratio is generally significantly increased, and acceptable results are still obtained for some applications, such as speech recognition and wake word detection. However, for other applications, frequency distortion may be highly undesirable. For example, in voice communication, an improved signal-to-noise ratio results in good speech intelligibility, but the beamformer output will tend to sound unnatural due to the lack of low frequencies. An equalizer can be configured to compensate for this by boosting low frequencies. Since the desired frequency characteristics of the equalizer cannot be known a priori, as it depends on the beamforming and decorrelation parameters that vary with the fit, an adaptive equalizer is employed.

[0127] Figure 2 An example of the elements of an audio beamformer 103 is shown. The audio beamformer 103 includes a signal processor 201, which is arranged to generate a beamformed output audio signal by applying a process that includes the effects of both spatial beamforming and spatial decorrelation to a first set of audio signals.

[0128] In some embodiments, spatial decorrelation may be non-adaptive spatial decorrelation (e.g., having parameters that are manually set at initialization or set once (e.g., based on microphone distance measurements or other measurements)). However, in many embodiments, spatial decorrelation is advantageously adaptive spatial decorrelation, and the following description will focus on embodiments employing adaptive spatial decorrelation.

[0129] To accommodate adaptive spatial decorrelation, audio beamformer 103 includes a decorrelation adapter 203 arranged to adapt adaptive spatial decorrelation parameters based on the input signal to audio beamformer 103 (i.e., based on a first set of audio signals). Audio beamformer 103 also includes a beamforming adapter 205 arranged to adapt beamforming parameters, and specifically, to determine beamforming / combining parameters of a beamforming filter. The adaptation is based on the generated beamformed output audio signal, and in many embodiments, the beamforming parameters are adapted to maximize the signal level of the output signal. Various techniques for performing adaptive spatial decorrelation and beamforming are known and can be used without departing from the invention.

[0130] In many embodiments, the signal processor 201 is arranged to first perform decorrelation on a first set of audio signals (i.e., the input signals to the signal processor 201) to generate a decorrelated first set of audio signals, and then apply a beamforming function to the decorrelated first set of audio signals to generate a beamformed output audio signal. Figure 3 This method is illustrated, wherein a signal processor 201 is shown to include a spatial decorrelector, in a specific example, an adaptive spatial decorrelector 301. The adaptive spatial decorrelector 301 receives a first set of audio signals and is arranged to generate a decorrelated first set of audio signals by performing adaptive spatial decorrelation, such as specifically by filtering the first set of audio signals with a set of spatial decorrelation filters to produce the decorrelated first set of audio signals. The adaptive spatial decorrelector 301 may perform a decorrelation function based on decorrelation parameters as determined by a decorrelation adapter 203, and specifically, the decorrelation adapter 203 may determine filter parameters, such as filter coefficients, for the spatial decorrelation filters.

[0131] Therefore, the first set of decorrelated audio signals is the set of audio signals with reduced correlation / coherence (at least one audio source) relative to the first set of audio signals.

[0132] Decorrelation adapter 203 may specifically include a set of spatial filters and may seek to adapt the spatial filters to spatial decorrelation filters, which seek to generate an output signal (a first set of decorrelated audio signals) corresponding to the input signal, but with increased decorrelation of the signals, particularly for a (primary) audio source. The output audio signal is generated with increased spatial decorrelation, wherein the cross-correlation between the audio signals is lower for the output audio signal compared to the input audio signal. Specifically, the output signal may be generated with the same combined energy / power as the input signal (or with a given scaling), but with increased decorrelation (reduced correlation) between the signals. The output audio signal may be generated to include all audio / signal components of the input signal in the output audio signal, but with redistribution to different signals to achieve increased decorrelation.

[0133] A decorrelation filter can be specifically configured to generate an output signal with lower coherence / normalized correlation compared to the input signal. Therefore, the output signal of a decorrelation filter can have lower coherence / normalized correlation compared to the coherence in the input signal.

[0134] In some embodiments, the decorrelation adapter 203 may be specifically arranged to seek decorrelation / decoherence between a noise / interference / undesired audio source and the captured signal. It may be specifically arranged to adapt decorrelation when the noise / interference / undesired audio source is active. In many embodiments, it may be arranged not to adapt during the activity of a desired audio source. The beamforming adapter may be arranged to adapt when the desired audio source is active, but not when the noise / interference / undesired audio source is active. Therefore, in many embodiments, the decorrelation adapter 203 may be arranged to adapt decorrelation for noise / interference / undesired audio sources, and the beamforming adapter 205 may be arranged to adapt beamforming for desired audio sources.

[0135] It will be understood that in different embodiments, different algorithms and functions may be used to decorrelate the input signal, and in particular, different methods may be used to determine the appropriate filter coefficients for the decorrelation filter.

[0136] In some embodiments, the decorrelation adapter 203 may include, for example, a neural network arranged to generate updated values ​​for modifying the filter coefficients of 301 based on input samples. The trained network may have been trained, for example, on training data including many different cases and associated updated values, the associated updated values ​​of which have been manually determined to modify the weights toward increasing decorrelation.

[0137] As another example, the decorrelated adapter 203 may include a location processor that is configured to generate updated values ​​for modifying the filter coefficients based on visual cues associated with the (changed) location.

[0138] As another example, adapter 203 can determine the cross-correlation matrix of the input signal and compute its eigenvalue decomposition. Eigenvectors and eigenvalues ​​can be used to construct decorrelation.

[0139] As another example, the decorrelation adapter 203 may include a microphone distance processor arranged to generate updated values ​​for modifying the filter coefficients for the case of diffuse noise only. For diffuse noise, the cross-correlation matrix of the input signal depends only on the distance between the microphones and can then be calculated. In the specific case of diffuse noise only, and with a fixed distance between the microphones (static array), it only needs to be calculated during initialization. Therefore, in this case, a static or non-adaptive spatial decorrelation adapter 301 can be used.

[0140] exist Figure 3 In the example, the second set of output signals is fed to beamforming circuit 303, which is arranged to perform spatial beamforming by combining the first set of decorrelated audio signals, wherein the combination depends on a set of beamforming parameters. The combination may specifically include filtering each audio signal in the first set of decorrelated audio signals and then combining the filter outputs by, for example, (weighted) summation.

[0141] It will be understood that many different beamforming algorithms are known, and the audio beamformer 103 can use any suitable method or algorithm.

[0142] For example, beamformer adapter 205 can determine the cross-correlation matrix of its input signal and calculate its eigenvalue decomposition, and determine the eigenvectors belonging to the largest eigenvalues ​​(principal component analysis). The eigenvectors can be used to construct the beamforming parameters of beamforming circuit 303.

[0143] Beamforming circuit 303 is arranged to receive a set of multiple audio signals and generate a single signal based on that set. Beamforming circuit 303 attempts to combine the input signals such that contributions from a given audio source are complementarily combined. Specifically, beamforming circuit 303 combines signals to perform beamforming on a set of spatial audio signals (such as a set of signals from a microphone array) of audio in the capture environment. Beamforming circuit 303 is arranged to generate a beamformed output audio signal based on a first set of audio signals.

[0144] Figure 4An example of a possible beamforming circuit 303 is shown, which may be particularly suitable and advantageous for the described audio device.

[0145] Figure 4 The beamforming circuit 303 specifically includes a first filter set. , The first filter set filters the first set of audio signals provided to the beamforming circuit 303 (which, as described later, may specifically be a decorrelated set of first audio signals or may be directly captured microphone signals). Each filter is arranged to filter one audio signal from the audio signals. Furthermore, the filters are adaptive filters that are dynamically adapted to achieve adaptive beamforming. The first filter set will hereafter also be referred to as beamformer filters.

[0146] The audio arrangement also includes a feedback circuit 401, which includes a second filter set that generates a second audio signal set based on filtering the beamforming output audio signal. This second filter set will also be referred to as a feedback filter. The input signal to the beamforming circuit 303 can also be referred to as the beamformer input signal, and the signal generated by the feedback circuit can also be referred to as the feedback signal. The feedback circuit 401 serves as part of an adaptation to beamforming operation and can be considered part of the beamforming adapter 205. However, for clarity, the feedback circuit 401 is... Figure 4 (and the following figures) are shown as separate from the beamforming adapter 205.

[0147] The first and second filter sets are specifically matched / linked such that for each filter in the first set (i.e., for each beamformer filter), there exists a filter in the second set whose frequency response is the complex conjugate of the corresponding beamformer filter (i.e., a feedback filter). Therefore, the filters of beamforming circuit 303 and feedback circuit 401 are adapted together such that the filter coefficients provide a complex conjugate frequency response (equivalent to / corresponding to the time-reverse filter impulse response).

[0148] Therefore, based on the beamforming output audio signal, the feedback circuit 401 generates a set of feedback signals, wherein each feedback signal is generated from the beamforming output audio signal by time-reverse / complex conjugate filtering performed by a matched filter.

[0149] Beamforming adapter 205 is arranged to adapt to a first set of filters and a second set of filters, namely beamformer filters and feedback filters. Beamforming adapter 205 is arranged to perform adaptation based on a comparison of the first set of audio signals and the second set of audio signals, and specifically based on a comparison of the beamformer input signal and the feedback signal. Specifically, for each matched / linked beamformer input signal and feedback signal (beamformer filters and feedback filters for signals with their complex conjugate frequency responses), a difference metric can be determined, and the matched filter can be adapted to reduce / minimize the difference metric.

[0150] The arrangement of beamforming circuitry 303, feedback circuitry 401, and beamforming adapter 205 can interact to provide efficient adaptive beamforming. For a set of spatial audio signals, adaptive beamforming can adapt filters to form a beam toward the audio source, where the beamforming output audio signal provides the audio captured in the formed beam. Adapting filters to minimize the difference metric allows the beamformer arrangement to detect and track a given audio source. Specifically, ideally, each beamformer filter would concentrate all the energy from a given audio source captured in one of the beamformer input signals into a single signal value. Ideally, this is achieved by beamforming filters with an impulse response, which is the reciprocal of the time of the acoustic impulse response from the audio source to the microphone capturing the signal. Ideally, this is achieved for all signals / filters, resulting in all the energy captured from a particular audio source being combined into a single value by summing the outputs of the beamformer filters, thereby generating a sample value of the beamformed output audio signal that maximizes the energy captured from the audio source.

[0151] Furthermore, since each feedback filter is the complex conjugate frequency response of the corresponding beamformer filter, ideally, each feedback filter is identical to the acoustic impulse response from the audio source to the microphone (signal). Therefore, by filtering the beamformed output audio signal with feedback filters, a feedback signal is generated that reflects the signal captured at the microphone from the audio source, as represented by the beamformed output audio signal. In the ideal case where there is only one audio source and the filter is ideally fitted, the difference between the generated feedback signal and the captured microphone signal is zero. In the presence of other audio sources and / or noise, there may be differences reflected by a difference metric. However, for uncorrelated noise / audio, the difference, and therefore the difference metric, is typically zero on average and will therefore tend not to hinder effective fitting. Therefore, by fitting the filter for a given beamformer input signal to minimize the difference metric for the corresponding feedback signal, an optimal beamformer filter can usually be achieved.

[0152] This beamforming method is described in US 7,146,012. A more detailed description and analysis of this arrangement will be provided below. This description will be based on... Figure 4 An example where the adaptive beamformer arrangement consists of filter sections, wherein the microphone signal is filtered by two filters. and The filter is then applied, and the outputs are summed to obtain the output signal. In the update section, the output signal... The signal is fed into two adaptive filters that use the corresponding microphone signal as a reference. and .

[0153] Using this configuration, we have the ability to constrain... The adaptive beamformer maximizes output power. This constraint is implied by the configuration, particularly because of the filters in the update section. and The conjugate of is copied to the filter section. The adaptive filter in the update section is a "normal" unconstrained adaptive filter.

[0154] If we look at the optimal solution, i.e., when both residual signals are zero, we can see that the configuration implies constraints. Suppose we have a speech signal. and the transfer function from the speech source to the microphone and The microphone signal is then transmitted by... and Provided. Then the output signal for: . In the update section, after convergence, we obtain two equations, one of which is omitted for convenience. : ( (1) ( (2) From (1) and (2), follow (3): (3) Inserting (3) into (2) yields: Similarly, we obtain: Satisfying constraints . The solution is given by the following equation: , in It is a universal term, that is, | |=1, phase undetermined. Similarly, for : Since (3) must still be maintained, The so-called public universal item. More generally, and using matrix notation, where in It corresponds to The microphone's beamforming filter, and in From the source arrive From the microphone's transfer function, we obtain the output. (4) In the update section, after convergence, we have the following equation: (5) This leads to (because) and (is a scalar) (6) in It is a complex scalar. Substituting (4) into (5), we can use (6) to give: (7) in It is a public universal item. Substituting (7) into (6) finally yields the converged solution (8): (8) in This indicates the complex conjugate transpose. Regarding the constraints, we then obtain: To understand that the solution given by Equation 8 is also the optimal solution that maximizes the output power, we examine the expected value of the update after convergence, which includes the residual signal ( Figure 4 In and The correlation between the input signal and the conjugate of the adaptive filter: in in yes The microphone signal. This is equivalent to: use And therefore We obtain: in It is the input covariance matrix. From this, we know... and These are the eigenvectors and eigenvalues ​​of the input covariance matrix, respectively. It can be shown that when... When the eigenvalue is the largest eigenvalue, a stable solution is obtained. According to matrix theory, we know that for a Hermitian matrix, the maximum value of the Rayleigh coefficient is obtained for the eigenvector that belongs to the largest eigenvalue (omitted). ) This proves that beamformers are effective in constraining... Maximize its output by setting it to 1.

[0155] When the microphone signal contains only the desired signal (such as a specific voice), where ,but Maximizing corresponds to maximizing the desired signal / speech in the output.

[0156] If we consider the following situation , Where we assume It is uncorrelated noise with equal variance in all microphones, i.e. It can be written as: Maximizing corresponds to (using Maximize: This means that maximization corresponds to maximizing speech and maximizing the signal-to-noise ratio. Now assume Including relevant noise, such that: in It is the covariance matrix of noise with non-zero off-diagonal elements. Then Maximizing leads to the maximization of the following equation:

[0157] This means Maximizing this does not necessarily lead to a better signal-to-noise ratio, because for The choice of [aspect] determines not only the amount of speech in the output, but also the amount of noise in the output.

[0158] As described, improved performance can be achieved by including an adaptive spatial decorrelation function, for example, specifically by including an adaptive spatial decorrelation unit 301 at the input to decorrelate the first set of audio signals. Figure 5 An example is shown where the reference signal is further enhanced by introducing adaptive spatial filtering of the input signal. Figure 4 The described beamforming method.

[0159] Spatial filtering is performed across the signal and is based on decorrelation coefficients / weights, which are determined based on the input audio signal and adapted to the current audio properties. As will be described in more detail below, spatial filtering can be placed at different locations within the beamforming loop structure, including as inverse spatial decorrelation filtering in the feedback loop or as spatial decorrelation filtering directly on the input signal used for beamforming. It has been found that this method can provide significantly improved operation in many practical scenarios, such as, for example, significantly improved source separation. Regarding the above analysis, this method can particularly address noise contribution. Can be independent of The choice.

[0160] Figure 5 An audio device that can provide improved performance is shown. In this example, the audio device includes, as referenced... Figure 4 The beamforming structure described. However, this method is further enhanced by operating on signals that have already been spatially decorrelated by the adaptive spatial decorrelator 205 using this structure.

[0161] Figure 2The audio beamformer 103 receives a first set of audio signals, and in a specific example, receives a set of microphone signals from a set of microphones capturing an audio scene from different locations. The microphones may, for example, be arranged in a linear array and close to each other. For example, in many embodiments, the maximum distance between the capture points for the audio signals may not exceed 1 meter, 50 cm, 25 cm, or even in some cases, not exceed 10 cm.

[0162] The input audio signal can be received from various sources, including internal or external sources. An embodiment in which the audio beamformer 103 is coupled to multiple microphones (such as a linear array of microphones) will be described below to provide a set of input audio signals in the form of microphone signals.

[0163] However, Figure 5 The audio device is configured to apply adaptive spatial decorrelation to the audio signal prior to beamforming and adaptive operation, rather than performing it directly on the microphone signal. Figure 4 Spatial decorrelation is an adaptive decorrelation method that adapts to the input audio signal, allowing decorrelation to be continuously adapted to provide increased decorrelation.

[0164] Specifically, Figure 5 The audio device includes an adaptive spatial decorrelation unit 301 in the form of a set of spatial (decorrelated) filters, which applies spatial filtering to a first set of audio signals. Therefore, after the spatial filters have applied the filtering, the first set of audio signals is modified to have increased decorrelation / decoherence (for at least one audio source) compared to before the spatial decorrelation filtering.

[0165] The spatial filter is based on coefficients adapted by decorrelation adapter 203, which in this example dynamically adapts and updates the filter coefficients used by the filter set of the adaptive spatial decorrelation adapter 301. The spatial filter is thus configured with coefficients determined by decorrelation / coefficient adapter 203.

[0166] Beamforming, decorrelation, and adaptation can typically be performed in the frequency domain. Receiver 101 may include a segmenter arranged to segment the input audio signal set into time segments. In many embodiments, the segmentation may typically be fixed segments to time segments of fixed and equal duration, such as, for example, to a division of time segments / intervals having a fixed duration between 10 and 20 milliseconds. In some embodiments, the segmentation may be adaptive to have varying durations. For example, the input audio signal may have a varying sampling rate, and the segmentation may be determined to include a fixed number of samples.

[0167] Segmentation can typically be a segmentation of a given fixed number of time-domain samples of the input signal. For example, in many embodiments, the segmenter can be arranged to divide the input signal into consecutive segments of, for example, 256 or 512 samples.

[0168] Receiver 101 may be arranged to generate a frequency bin representation of the input audio signal, and the first set of input signals for further processing is typically represented in the frequency domain by the frequency bin representation. The audio device may be arranged to perform frequency domain processing on the frequency domain representation of the input audio signal. Signal representation and processing are based on frequency bins, and therefore the signal is represented by values ​​of the frequency bins, and these values ​​are processed to generate frequency bin values ​​for the output signal. In many embodiments, the frequency bins have the same size and therefore cover frequency intervals of the same size. However, in other embodiments, the frequency bins may have different bandwidths, and, for example, perceptually weighted bin frequency intervals may be used.

[0169] In some embodiments, the input audio signal may already be provided in frequency representation and no further processing or manipulation is required. However, in some cases, it may be desirable to rearrange it into a suitable segmented representation, including, for example, using interpolation between frequency values ​​to align the frequency representation with the time segments.

[0170] In other embodiments, a filter bank, such as a quadrature mirror filter (QMF), can be applied to the time-domain input signal to generate a frequency bin representation. However, in many embodiments, a discrete Fourier transform (DFT), and in particular a fast Fourier transform (FFT), can be applied to generate the frequency representation.

[0171] exist Figure 5 In the audio device, the spatial filter of the adaptive spatial decorrelation unit 301 specifically processes the audio signal in the frequency domain. In the following description, the first set of audio signals may also be referred to as the input audio signal before filtering (to the spatial filter), and the resulting signal may also be referred to as the output audio signal (from the spatial filter).

[0172] For each frequency bin, the output frequency bin value is generated based on one or more input frequency bin values ​​of one or more input signals, as will be described in more detail below. At least for an audio source, which may be, for example, the primary audio source, the output signal is generated to (typically / on average) reduce the correlation between the signals relative to the correlation of the input signals.

[0173] A set of spatial filters is arranged to filter the input audio signal. The filtering is spatial because, for a given output signal, the output value is determined based on multiple, and typically all, of the input audio signals (for the same time / segmentation and for the same frequency bin). Specifically, spatial filtering is performed on a frequency bin basis, such that the frequency bin value for a given frequency bin of the output signal is generated based on the frequency bin value of the input signal for that frequency bin. The filtering / weighting combination is across the signal, rather than typical time / frequency filtering.

[0174] Specifically, the frequency bin value for a given frequency bin is determined as a weighted combination of the frequency bin values ​​of the input signal for that frequency bin. The combination can specifically be a summation, and the frequency bin value can be determined as a weighted sum of the frequency bin values ​​of the input signal for that frequency bin. The determination of the bin value for a given frequency bin can be determined as a vector multiplication of a vector of weights / coefficients for the weighted summation and a vector including the bin values ​​of the input signal: Where y is the output bin value, w1-w3 are the weights of the weighted combination, and x1-x3 are the input signal bin values.

[0175] Represent the bin value for a given frequency bin ω as a vector. The output signal can be determined as follows: Where the matrix This represents the weights / coefficients for different output signals in the weighted summation, and It is a vector that includes the input signal values. For example, for a case with only three input signals and an output signal, the output bin value for the frequency bin ω can be given by the following formula: Where y n Indicates the output bin value, w ij This represents the weight of the weighted combination, and x m This indicates the input signal value.

[0176] The decorrelation adapter 203 may adapt a spatial filter to a spatial decorrelation filter, which seeks to generate an output signal corresponding to the input signal but with increased signal decorrelation. The output audio signal is generated with increased spatial decorrelation, where the cross-correlation between audio signals is lower for the output audio signal compared to that for the input audio signal. Specifically, the output signal may be generated with the same combined energy / power (or with a given scaling) as the input signal, but with increased decorrelation (reduced correlation line) between the signals. The output audio signal may be generated to include all audio / signal components of the input signal in the output audio signal, but with a redistribution of signals for the different signals to achieve increased decorrelation.

[0177] A decorrelation filter can be specifically arranged to generate an output signal with lower coherence / normalized decorrelation compared to the input signal. The output signal of the decorrelation filter can therefore have lower coherence / normalized decorrelation compared to the coherence in the input signal of the decorrelation filter.

[0178] The decorrelation adapter 203 is arranged to determine updated values ​​for the weights of the weighted combination forming the decorrelation filter set. Therefore, specifically, it can be configured for the matrix... Determine the update value. Go to the relevant adapter 203, and then you can update the weights of the weighted combination based on the update value.

[0179] The decorrelation adapter 203 is configured to apply an adaptive method for determining update values ​​that allow the output signal of the decorrelation filter set 301 to represent the audio of the input signal of the decorrelation filter set 301, but the output signal is generally more decorrelational than the input signal.

[0180] The decorrelated adapter 203 is configured to adapt weights using a specific method based on the generated output signals. This operation is based on each output audio signal being linked to an input audio signal. The exact link between the output and input signals is not required, and many different (in principle, random) links / pairs can be used for each output signal and input signal. However, the processing of weights that reflect the contribution from the output signal linked to the weighted output signal / input signal paired with that output signal differs from the processing of weights that reflect the contribution from the output signal not linked to the weighted output signal / input signal paired with that output signal. For example, in some embodiments, the weights for linked signals (i.e., for input signals linked to the output signal generated by a weighted combination including weights) can be set to fixed values ​​and not updated, and / or the weights for linked signals can be restricted to real-valued weights, while other weights are typically complex-valued.

[0181] The relevant adapter 203 uses an adaptation / update method, where an update value is determined for a given weight, representing the contribution of the given unlinked input signal to the given output signal's value, based on a correlation measure between the output bin value for a given output signal and the output bin value for an output signal linked to the same given (unlinked) input signal. The update value can then be applied to modify the given weight in subsequent segments, or to determine the update value for the weight in a given segment in response to (and typically directly) two output bin values ​​from a previous segment, where the two output values ​​represent the input and output signals to which the weight is involved.

[0182] The described method is typically applied to multiple, and usually all, weights used when determining the output bin value based on a non-linked input signal. For weights associated with the input signal linked to the weighted output signal, alternative considerations can be used, such as setting the weights to fixed values, as will be described in more detail later.

[0183] Specifically, the updated value can be determined based on the product of the weighted output bin value and the complex conjugate of the output bin value linked to the weighted input signal, or equivalently based on the product of the complex conjugate of the weighted output bin value and the output bin value linked to the weighted input signal.

[0184] As a concrete example, the update value for segment k+1 for the frequency bin ω can be determined based on the correlation metric given by the following formula: Or equivalently determined by the following formula: in It refers to the output bin value of the output signal i determined based on weights; and It is the output bin value of the output signal j linked to the input signal that determines the contribution from it (i.e., the input signal bin value multiplied by a weight to determine the contribution of the output bin value for signal i).

[0185] measure The (or conjugate) value indicates the correlation of the time-domain signal within a given segment. In a specific example, this value can then be used to update and adapt the weights w. i,j (k+1,ω).

[0186] As mentioned earlier, the decorrelation filter set can be arranged to determine the frequency bin for a given frequency according to the following formula. Output block value for the output signal: in It includes a warehouse for a given frequency. A vector of frequency values ​​for the output signal; It includes a warehouse for a given frequency. A vector of frequency binaries for the input audio signal; and It is a matrix with rows that include weights for a weighted combination of the output audio signal.

[0187] In this example, the de-correlation adapter 203 can be specifically arranged to adapt the matrix according to the following formula. At least some weights in : Where i is a matrix The row index, j is the matrix. The column index, k is a time-segmented index. Indicates frequency compartment, and These are scaling parameters used to adapt the speed of the adapter. Typically, the decorrelated adapter 203 can be configured to adapt all weights that do not associate the input signal with the linked output signal (i.e., "cross signal" weights).

[0188] In some embodiments, the decorrelation adapter 203 may be arranged to adapt to an appropriate update rate / speed for the weights. For example, in some embodiments, the adapter may be arranged to compensate for the correlation metric for a given weight based on the signal level of an output bin value that determines its contribution to the weight.

[0189] As a specific example, compensation value You can output the warehouse value The compensation can be made using the signal level. For example, compensation could be included to normalize the update step size value to a signal level that is less dependent on the generated decorrelation signal.

[0190] In many embodiments, this compensation or normalization can be specifically performed on a frequency bin basis; that is, the compensation can differ across different frequency bins. This can improve operation in many scenarios and often leads to an adaptation that generates improved weights for the decorrelated signal.

[0191] Compensation can, for example, be constructed into the scaling parameters of the previously updated equations. Therefore, in many embodiments, the decorrelated adapter 203 can be arranged to adapt / change scaling parameters differently in different frequency modules. .

[0192] In many embodiments, the input signal vector and output signal vector The arrangement ensures that the linked signals are at the same position in the corresponding vectors, that is, specifically... and Link, and , and Links, etc. In this case, the weights of the signals for the links are in the weight matrix. On the diagonal. In many embodiments, the diagonal value can be set to a fixed real value, such as, for example, specifically, a constant value of 1.

[0193] In many embodiments, the weight / spatial filter / weighting combination can be such that the weight for the contribution of the first input signal (not linked to the first output signal) to the first output signal is the complex conjugate of the contribution of the second input signal linked to the first input signal to the second output signal linked to the first input signal. Therefore, the two weights for the two pairs of linked input / output signals are complex conjugates.

[0194] The weights for the input and output signals of the link are arranged in a weight matrix. In the example on the diagonal, this results in a Hermitian matrix. In fact, in many embodiments, the weight matrix... It is a Hermitian matrix. Specifically, the weight matrix. The coefficients / weights can meet the following criteria: .

[0195] As mentioned earlier, the weights for the contributions of the input signals from the linking signals to the output signal bin value (corresponding to the weight matrix in a specific example) are... The values ​​of the diagonal (the weights for the non-linked input signals) are treated differently from the weights for the non-linked input signals. For simplicity, the weights for the linked input signals will be referred to as linked weights hereinafter, and for simplicity, the weights for the non-linked input signals will be referred to as non-linked weights hereinafter. Therefore, in a specific example, the weight matrix... It will be a Hermitian matrix, which includes link weights on the diagonal and non-link weights outside the diagonal.

[0196] In many methods, fitting non-link weights aims to reduce the correlation metric. Specifically, each updated value can be determined to decrease the correlation metric. In general, the fit will seek to reduce the cross-correlation between the output signals. However, link weights are determined differently to ensure the output signals maintain appropriate audio energy / power / level. In practice, if link weights are fitted to seek to reduce the autocorrelation of the output signals with respect to the weights, there is a high risk that the fit will converge to a solution where all weights and therefore the output signal is essentially zero (since this would practically result in the lowest possible correlation). Furthermore, audio devices are arranged to attempt to generate signals with less cross-correlation but not to reduce autocorrelation.

[0197] Therefore, in many embodiments, link weights can be set to ensure that the output signal is generated with the desired (combined) energy / power / level.

[0198] In some cases, the de-relevant adapter 203 can be configured to adapt to link weights, and in other cases, the adapter can be configured not to adapt to link weights.

[0199] For example, in some embodiments, link weights can simply be set to fixed, constant values ​​that are not adapted. For example, in many embodiments, link weights can be set to constant scalar values, specifically, a value of 1 (i.e., a single gain is applied to the input signal of the link). For example, a weight matrix... The weights on the diagonal can be set to 1.

[0200] Therefore, in many embodiments, the weight for the contribution of the input signal frequency bin value from the link to a given output signal frequency bin value can be set to a predetermined value. In many embodiments, this value can remain constant without any adaptation.

[0201] This method has been found to provide very efficient performance and may lead to a global adaptation, which has been found to provide the output signal to be generated, which provides a highly accurate representation of the original audio of the input signal in a set of output signals with increased decorrelation.

[0202] In some embodiments, link weights may also be adapted, but in a different way than non-link weights. In particular, in many embodiments, link weights may be adapted based on the output signal.

[0203] Specifically, in many embodiments, the link weights for the first input signal and the linked output signal can be adapted based on the generated output bin value of the linked audio signal and specifically based on the amplitude of the output bin value.

[0204] This method can, for example, allow for the normalization and / or setting of the desired energy level of the signal.

[0205] In many embodiments, link weights are constrained to real-valued weights, while non-link weights are typically complex-valued. Specifically, in many embodiments, the weight matrix... It can be a Hermitian matrix with real values ​​on the diagonal and complex values ​​outside the diagonal.

[0206] This approach can provide specific advantageous operations and adaptations in many scenarios and implementations. It has been found that this approach provides efficient space decorrelation while maintaining relatively low complexity and computational resources.

[0207] The fit can gradually adapt the weights to increase decorrelation between signals. In many embodiments, the fit can be arranged toward the appropriate weight matrix. Convergence occurs regardless of the initial values, and in some cases, the fit can actually be initialized with random values ​​for the weights.

[0208] However, in many embodiments, the fitting can begin with favorable initial values, which may, for example, lead to faster fitting and / or make the fitting more likely to converge toward better weights for decorrelation of the signal.

[0209] In particular, in many embodiments, the weight matrix Weights can be arranged as multiple zero weights, but at least some weights are non-zero. In many embodiments, the number of substantially zero weights can be no less than 2, 3, 5, or 10 times the number of weights set to non-zero values. This has been found to tend to provide improved fit in many scenarios.

[0210] Specifically, in many embodiments, the decorrelation adapter 203 can be arranged to initialize the weights, wherein the link weights are set to non-zero values, such as typically predetermined non-zero real values, while the non-link weights are set to substantially zero. Thus, in the example described above where the linked signals are positioned at the same locations in the vector, this will result in an initial weight matrix... It has a non-zero value on the diagonal and (essentially) a zero value outside the diagonal.

[0211] This initialization can provide particularly advantageous performance in many embodiments and scenarios. It reflects the tendency for the input signal to be slightly decorrelated, since audio signals typically represent audio at different locations. Therefore, assuming a fully correlated starting point for the input signal is generally advantageous and will result in faster and generally improved adaptation.

[0212] It should be understood that weights, especially non-link weights, may not necessarily be exactly zero, but in some embodiments may be set to low values ​​close to zero. However, the initial non-zero value may be at least 5, 10, 20, or 100 times higher than the initial substantially zero value.

[0213] The described method provides an efficient adaptive spatial decorrelational amplifier that generates an output signal representing the same audio as the input signal but with increased decorrelation. In practice, this method has been found to provide efficient adaptation in a wide variety of scenarios, many different acoustic environments, and for many different audio sources. For example, it has been found to provide efficient decorrelation of speaker signals in environments with multiple speakers.

[0214] Furthermore, the adaptation method is computationally efficient and, in particular, allows for local and individual adaptation of each weight based solely on two signals closely related to the weights (and specifically, only on two frequency bin values). However, this process leads to efficient spatial filtering and often essentially global optimization, especially of the weight matrix. In many embodiments, it has been found that local adaptation leads to highly advantageous global adaptation.

[0215] A key advantage of this method is that it can be used to decorrelate convolutionally mixed signals, and not just instantaneously mixed signals. For convolutionally mixed signals, the full impulse response determines how signals from different audio sources are combined at the microphone (i.e., delay / timing characteristics are significant), while for instantaneously mixed signals, the scalar representation is sufficient to determine how the audio sources are combined at the microphone (i.e., delay / timing characteristics are not significant). By transforming the convolutionally mixed signal into the frequency domain, the mixed signal can be considered as a complex-valued instantaneously mixed signal per frequency compartment.

[0216] The decorrelation adapter 203 thus determines the coefficients of the set of spatial decorrelation filters in the adaptive spatial decorrelation adapter 301, such that these coefficients modify the first set of audio signals to represent the same audio, but with increased decorrelation / decoherence for at least one of the audio sources. This decorrelation of the audio signal to which the beamforming method previously described is then applied can significantly improve overall performance in many scenarios. In fact, in many scenarios, it can lead to improved separation and selection of specific audio sources, such as specific speakers. Thus, contrary to intuition, decorrelation of a signal can provide improved beamforming performance, even though beamforming is inherently based on utilizing and adapting the correlation between audio signals from different locations to extract / separate audio sources by spatially forming a beam toward the desired source. In practice, decorrelation inherently disrupts the link between the audio signal and a specific location in the audio scene that is typically utilized by beamforming operations. However, the inventors have recognized that, despite this, decorrelation can provide highly advantageous effects and improved performance in many scenarios. For example, in the presence of strong noise sources, this method can facilitate and / or improve the extraction / isolation of specific desired audio sources, such as, in particular, the speaker.

[0217] The specific fit described above can provide a highly advantageous method in many embodiments. It can generally provide a low-complexity yet highly accurate fit to produce a spatial decorrelation filter that generates a highly decorrelational signal. In particular, it can allow local fits to individual weights / filter coefficients to result in efficient global decorrelation of a first set of audio signals.

[0218] However, it should be understood that in other embodiments, other methods for adapting spatial filters / decorrelation filters may be used.

[0219] In the following text, the export will be... Figure 5 The output z of the audio beamformer 103 ( In this example, which will also be referred to below as Configuration A, the spatial filter set is placed directly behind the microphone, entirely outside the beamformer. This has the advantage of not requiring changes to the beamformer algorithm: it will still have constraints. Only the input is different.

[0220] A suitable decorrelation can be used as described above. This decorrelation will reduce the noise covariance matrix. Transformed into: (9) in It is a diagonal matrix, and It is the adaptive spatial decorrelation 301. nmics X nmics Remove the correlation matrix. Replace maximization. The beamformer now maximizes It can be written as

[0221] if It is a scaled version of the identity matrix and can be written as , Transform into And the noise contribution is independent of the The choice, and Maximizing this leads to maximizing the signal-to-noise ratio in the output.

[0222] In the adaptive spatial decorrelation 301 The choice of [option name] does not affect the SNR, but rather the level of the speech in the output. Let's assume we choose [option name]. And if the noise level at the input is increased by a factor of 2, then All coefficients are scaled by a factor of 0.5, thus reducing the speech level at the output of the adaptive spatial decorrelector 301. The adaptive spatial decorrelector 301 is updated based on the input level (per frequency cell). As a result, the beamformer is also updated. It can be shown that... That is, it is linearly related to the power of the noise source. It depends solely on the acoustic transfer function between the noise source and the microphone. Since the transfer function typically changes much more slowly than the noise characteristics, this is also advantageous for beamformers, as they see a "stable" path. The solution specifically is where... . Elements that are linearly dependent on the input The drawback is that it is not... All elements are equal, and therefore, the noise in the output depends on Although this is not ideal, in practice it has been found that what is often more important is... It does not contain off-diagonal elements, while all elements on the diagonal are equal. Utilizing elements with... The combined solution of the decorrelation and beamformer has proven to be very effective.

[0223] Equation (8) can be used instead of the output of the spatial decorrelation. To find the optimal coefficients for the beamformer: .

[0224] Use with Hermione and therefore By using a decorrelation, we obtain the optimal solution: The combined output is given by the following formula: (Equation 10) (Equation 11) (Equation 12) Substituting (12) and (11) into (10) gives the output: .

[0225] Note that when there is equal variance uncorrelated noise at the input, using [the following]... And therefore With the decorrelation, we obtain the same solution as without the decorrelation and with only the beamformer.

[0226] Beamforming output audio signal The determination of this can be used to determine the optimal compensated frequency response of the equalizer 105 for a given desired frequency response. Specifically, for the beamforming output audio signal... The result is given as: It can be concluded that when only equally variable uncorrelated noise exists in the first audio signal set / microphone signal, the output can be written as: (Because the removal of correlated pairs has no effect on uncorrelated noise). In the presence of correlated noise, we expect speech to have the same frequency characteristics as in the absence of noise or with equally variable uncorrelated noise. This means we want to have the following compensated frequency response amplitude response. : It is unknown, but as previously shown, the optimal filter frequency response... Only depends on and And therefore, according to and Write Specifically, we obtain the optimal magnitude response: therefore, (Where the index opt indicates the optimal adapted value) represents the adapted frequency response of the beamforming filter, and is therefore given by the beamforming parameters. Furthermore, And therefore This represents the decorrelation coefficient, and is therefore given by the decorrelation parameter. Thus, based on the current decorrelation parameter and beamforming parameter, the equalizer adapter 107 can continue to determine the amplitude compensation frequency response given above, and control the equalizer 105 to apply this compensation frequency response to the beamformed output audio signal. Specifically, replacement To be used From the equation, we get: .

[0227] This demonstrates that the method results in the desired compensated frequency response as given above.

[0228] Phase characteristics can be determined to provide suitable performance, and in many embodiments, they can be determined as minimum phase characteristics to minimize delay. Different algorithms and methods for creating minimum phase characteristics from a given amplitude spectrum are known and will not be described further here for the sake of brevity. (References, e.g., AVOppenheim and RW Shafer, “Digital Signal Processing,” Englewood Cliffs (New York, USA): Prentice-Hall, 1975 or AVOppenheim and RW Shafer, “Discrete-Time Signal Processing,” Englewood Cliffs (New York, USA): Prentice-Hall, 1989).

[0229] In the previous example, the spatial decorrelation function and beamforming function were implemented by an adaptive spatial decorrelationer 301 and a beamforming circuit 303. The adaptive spatial decorrelationer 301 applied decorrelation to a first set of audio signals, and the beamforming circuit 303 then operated on the resulting output signal of the adaptive spatial decorrelationer 301. However, the same overall functionality can be achieved in other ways, including performing decorrelation or inverse decorrelation at different parts of the beamforming feedback or adapter circuitry.

[0230] exist Figure 5 In the example, the decorrelation and spatial decorrelation filters are applied directly to the first audio signal set before being fed to the beamforming circuit 303 and beamforming adapter 205. However, other methods can be used to adapt the operation by applying spatial / cross-signal filtering based on coefficients derived from the determined decorrelation coefficients.

[0231] Specifically, Figure 6 An example, also referred to as Configuration B, is shown where the audio beamformer 103 is arranged to perform spatial filtering on the first audio signal set before it is filtered by the first set of filters; that is, beamforming is based on the first audio signal set after it has been filtered by a set of spatial filters 601. However, in this example, the beamforming adapter 205 receives the first audio signal set before any filtering by the set of filters 601; that is, the second set of filters is applied only to the beamformer path and not to the adapter path.

[0232] The coefficients used for the spatial filter set 601 are determined based on the decorrelation coefficients already determined by the decorrelation adapter 203. In practice, in some embodiments, reference values ​​can be directly applied. Figure 5 The configuration described in the method, and the resulting second filter set can be applied (only) to the signal of the beamforming path.

[0233] However, in many embodiments, Figure 5 The spatial filter set 301 in the configuration example will be different. Figure 6 The configuration example shows a spatial filter set 601. Specifically, in many embodiments, the spatial filter set will be modified to correspond to a cascade of two decorrelation filter sets as determined by decorrelation adapter 203. Therefore, the filter coefficients of spatial filter set 601 have coefficients that match those of the cascaded spatial filters, which are two of the decorrelation filter sets. This can generally be considered equivalent to double / repetitive filtering of the first audio signal set by the decorrelation filter set as determined by decorrelation adapter 203.

[0234] Specifically, the decorrelation adapter 203 can determine the weight matrix used for the decorrelation filter. The weights / coefficients of the decorrelation filter can be determined based on... The first set of audio signals is decorrelated as described above.

[0235] In this case, the spatial filter set 601 can be generated as a cascaded application corresponding to two such filters; that is, it can be arranged as corresponding to .

[0236] Therefore, in this example, the relevant adapter 203 can continue to perform as for... Figure 5 The example described above describes an adaptation to determine suitable coefficients for a set of spatial decorrelation filters used to decorrelate a first set of audio signals. In some cases, such filters can then be derived from... Figure 6 The spatial filter set 601 is applied. However, in many embodiments, instead of using the determined decorrelation filter directly, the coefficients for the spatial filter set 601 are determined based on the coefficients determined for the decorrelation filter. In this case, as part of the determination / adaptation of the coefficients, the decorrelation adapter 203 can implement the decorrelation filter and apply it to the first signal set (e.g., so as to be based on...). (Determine the updated value). However, in certain examples, such a filtered signal may only be used in the adaptation process and may not be further used in beamforming / processing. Instead, the applied set of spatial filters 601 can be based on the decorrelation coefficient / weight matrix. Generate, specifically .

[0237] This method can provide improved overall performance, including an improved beamforming experience. In fact, it can be shown that, for which... Figure 6 The spatial filter set 601 is set to be with For example of the corresponding coefficients, this can lead to... Figure 5 The examples demonstrate the same performance, results, and output signals. It can be seen that these methods can produce the same optimal solution.

[0238] Figure 7 The diagram illustrates another possible example of applying spatial filtering by means of a set of spatial filters determined by the decorrelation coefficients determined by the decorrelation adapter 203, which will also be referred to as configuration C. In this example, the first audio signal set is not filtered by the set of spatial filters; specifically, the second audio signal set generated by the feedback circuit 401 is filtered by the set of spatial filters 701. Therefore, in this example, the feedback signal, rather than the input of the audio beamformer 103, is filtered by the set of spatial filters.

[0239] Furthermore, the set of spatial filters can be determined as a set of inverse spatial decorrelation filters, wherein each inverse spatial decorrelation filter may include the inverse filtering of the determined decorrelation filter.

[0240] For example, if the spatial decorrelation filter is composed of a weight matrix If we express this as an expression, then the inverse space decorrelation filter can be derived from the weight matrix. Determined. Therefore, in many embodiments, the set of spatial filters applied to the second set of audio signals may include, with The corresponding filters, where The spatial decorrelation is determined by the decorrelation adapter 203.

[0241] In many embodiments, the set of spatial filters applied to the second set of audio signals can be determined to have filter coefficients corresponding to the coefficients of the spatial filters, which are cascaded sets of two sets of spatial inverse filters, each set of spatial inverse filters being an inverse of a spatial decorrelation filter determined by decorrelation adapter 203.

[0242] Specifically, the decorrelation adapter 203 can therefore determine the weight matrix for the decorrelation filter that can decorrelate the first audio signal set according to the following formula. Weights / coefficients:

[0243] In this case, the spatial filter set 301 can be generated as a cascaded application corresponding to two such filters, that is, it can be arranged as corresponding to: in This represents the set of second audio signals generated by the feedback circuit 401.

[0244] Therefore, in this example, the relevant adapter 203 can continue to perform as for... Figure 5 The example described above describes an adaptation to determine suitable coefficients for a spatial decorrelation filter used to decorrelate a first set of audio signals. Instead of directly using the determined decorrelation filter coefficients, the coefficients for the spatial filter set 701 are determined based on the determined coefficients. In this case, as part of the determination / adaptation of the coefficients, the decorrelation adapter 203 can implement the decorrelation filter and apply it to the first signal set (e.g., so as to...). (Determine the updated value). However, in certain examples, such a filtered signal may only be used in the adaptation process and may not be used further in beamforming / processing.

[0245] This method can provide improved overall performance, including an improved beamforming experience. In fact, it can be seen that for one of them... Figure 7 The spatial filter set 701 is set to be with For example of the corresponding coefficients, this can lead to... Figure 5 The examples have the same performance and operation, and specifically, these can produce the same optimal solution.

[0246] An audio device achieves a highly adaptive method, which may be particularly well-suited for adaptation to extract specific audio sources, such as a desired speaker. This method may be especially advantageous when, for example, strong noise sources are present in the captured audio environment. The adaptive audio device includes different adaptations to a set of spatial filters (via adaptation to decorrelation filters / decorrelation coefficients) and adaptation to beamforming filters. Decorrelation adapter 203 and beamforming adapter 205 work together to provide a highly favorable and high-performance adaptation to the current audio properties of the scene.

[0247] For configuration B (such as...) Figure 6 As shown), first note that decorrelation can reduce the noise covariance matrix. Transformed into: in and They are positive definite, their inverses exist, and we can write them as: if It is a scaled version of the identity matrix and can be written as Then we get The reverse: Use with The decorrelation can ultimately be written as: It can be viewed as the inverse of the normalized covariance matrix. Note that in the decorrelation adapter 203, the calculation... In the filtering part, a square matrix is ​​used. .

[0248] We will now prove that this combination maximizes the speech-to-noise ratio in the output. First, we will derive the converged filter coefficients, as we did previously for the beamformer. At the output of the filtering section, we have: (Equation 13) In the update section, we have: (Equation 14) because and It is a scalar, which we can write as: Substituting equation 13 into equation 14, we get: in It is a common universal term. Finally, we obtain: Will Substitution get: Therefore, the constraints are satisfied. Next, we examine the expected value of the update after convergence: (Equation 15) in Where we use and . Then, for the update equation (Equation 15), we get: (Equation 16) in, and .

[0249] Equation 16 shows that, after convergence, It is an eigenvector, and It is a matrix The corresponding eigenvalues. As mentioned earlier, it can be seen that if If it is the largest eigenvalue, then the solution is only stable. Multiply the left and right sides by ,get: Note that the left side includes the output power of the beamformer. Now consider: The adaptive decorrelation has already adapted to the noise source. Then, for the output power... We have: .

[0250] Noise contribution and The choice of FSB is irrelevant, and therefore maximizing the output of FSB corresponds to maximizing the speech-to-noise ratio in the output, rather than the speech output itself.

[0251] After convergence, the optimal solution for solutions A and B is obtained by Related. This means that the outputs of the two solutions will be equal.

[0252] for Figure 7 In solution C, the matrix block can be placed after the filter in feedback circuit 401. Here, the constraint equals =1, where It is the coherence matrix, which is equal to the normalized covariance matrix. We want to use the decorrelation again and we use... . therefore, It equals the normalized covariance matrix, and we can use... Instead And the constraint becomes . Similar to before, we find the filter coefficients by considering the signal flow in the filtering and updating sections. The optimal solution. In the filtering section, we have: (Equation 17) And in the update section: (Equation 18) (Equation 19) Where we use and It is a complex scaler. Substituting (19) into (18) using (17) yields: Using (Equation 19), we obtain: in It is a universal term. Will Substitution get , The constraints are satisfied. Finally, we examine the expected value updated after convergence. in : (Equation 20) in and . Multiply both sides get: .

[0253] After convergence, It is an eigenvector and It is a matrix The corresponding eigenvalues. As mentioned earlier, it can be seen that if If it is the largest eigenvalue, then the solution is only stable.

[0254] Multiply the left and right sides of equation 20 by the equation get: Note that the left-hand side includes the output power of the FSB. Now consider The adaptive decorrelation has already adapted to the noise source. Then, for the output power... We have = .

[0255] Maximizing the output of the FSB corresponds to maximizing the speech-to-noise ratio in the output, not necessarily the speech output itself.

[0256] After convergence, the optimal solutions for solutions C and B are related by the following equation: .

[0257] This means that the outputs of the two solutions will be equal.

[0258] Therefore, the different configurations / solutions can be summarized as follows:

[0259] It can be seen that by appropriately selecting the coefficients, the exact same (optimal) beamforming output audio signal can be generated.

[0260] However, it can be seen that the coefficients chosen for the set of spatial filters to achieve this output may differ.

[0261] The choice of method to be used will depend on the preferences and requirements of each embodiment.

[0262] The advantage of Solution A is that it reduces the interaction between the decorreductor and the beamformer, in a sense that synergistic effects can be achieved by reducing modifications to the beamformer, feedback circuitry, and beamforming adapter.

[0263] The advantage of Solution B is that it is easier to do this for this configuration if you want to use beamformer coefficients to calculate the direction of arrival (DOA) of the audio source.

[0264] For example, US 6774934 describes a method for calculating DOA for a “normal” beamformer with added potentially irrelevant equal-variable noise. For this use case, the optimal coefficients are given by: .

[0265] Part of this process involves working with at least one pair of filter coefficients (e.g.) and Calculate the cross-power spectrum. This gives: Where we use Because a is a universal term.

[0266] Phase difference, rather than amplitude, is a more important factor. This means that for solution B, we can directly use the beamformer coefficients, but for methods A and C, we will respectively compare the coefficients with... and Multiply beforehand.

[0267] Different configurations can correspondingly result in generating the exact same beamforming output audio signal. Therefore, equalizer 105 can also advantageously apply the same compensation frequency response. The same considerations and methods as in configuration A can be used to determine the appropriate method for determining the suitable compensation frequency response based on the beamforming and decorrelation parameters.

[0268] For configuration B, the magnitude of the compensated frequency response can be determined by equalizer 107 as follows: In this case, we have Substitution From the previous equation, we get: This is the desired amplitude-compensated frequency response as shown above.

[0269] It should be noted that in this case, the compensated frequency response depends only on the beamforming parameters, and not on the decorrelation parameters.

[0270] For configuration C, the magnitude of the compensated frequency response can be determined by equalizer 107 as follows: In this case, we have Substitution From the previous equation, we get: This is the desired amplitude-compensated frequency response as shown above.

[0271] In many embodiments, the audio device can be arranged to adapt to different adapters at different times. Therefore, the decorrelation adapter 203 and the beamforming adapter 205 can be controlled to perform their respective adaptations at different time periods, and in many embodiments, these time periods may not overlap. Thus, at a given time, either the decorrelation adapter 203 or the beamforming adapter 105 may adapt, but not both (although for some time periods, neither may adapt).

[0272] For example, such as Figure 8 As shown, Figure 5 (or accordingly,) Figure 6 or Figure 7 The audio device may include an audio detector 801, which is arranged to detect when an audio source is active and when it is inactive. The audio detector 801 may specifically be arranged to divide time into a set of active time intervals during which the audio source is active and a set of inactive time intervals during which the audio source is inactive. The audio source may specifically be a desired audio source, such as a desired speaker.

[0273] In some embodiments, low-complexity audio detection can be used to determine whether a source is active. For example, an audio device can be used to extract the dominant audio source in the environment, such as the loudest speaker. For instance, typically in teleconference applications, only one person speaks at a time, and the audio device can be used to extract the current speaker's voice from background or ambient noise.

[0274] Audio detector 801 can, for example, detect audio source activity simply based on a level / energy exceeding a given threshold in such applications. If the signal level is above the given threshold (which could be a dynamic threshold set, for example, based on a longer-term average signal level), audio detector 801 can consider the audio source / speaker to be active; otherwise, it can consider the audio source / speaker to be inactive. Therefore, audio detector 801 can divide time into active time intervals when the captured audio is above the threshold (because the audio source is active) and inactive time intervals when the signal level is below the threshold (because the audio source is inactive).

[0275] In many embodiments, more sophisticated detection methods can be used, and techniques for separating different audio sources can indeed be employed. For example, in many embodiments, detection may include considering whether the audio signal has a desired / expected property that matches a particular audio source. For instance, audio detector 801 may distinguish speech from other types of audio based on assessing whether the captured audio has properties that match speech.

[0276] It should be understood that many different techniques and algorithms are known for speech / voice activity detection, and more generally for detecting that an audio source is active, and any suitable method can be used without departing from the present invention.

[0277] In some embodiments, a trained artificial neural network can be used, and in practice it has been found to perform effective speech / voice activity detection (and / or noise detection). The artificial neural network can be trained using speech and all types of non-speech, and can provide an indication of whether it is noise or speech for each frame.

[0278] In many embodiments, detection can be relatively fast, and the audio detector 801 can be arranged to designate relatively short time intervals as active or inactive time intervals. For example, in some embodiments, the audio detector 801 can be arranged to detect silence pauses / intervals during normal speech and designate such intervals as inactive time intervals. For example, in some embodiments, time intervals less than, for example, 5 milliseconds, 10 milliseconds, or 20 milliseconds can be identified / designated as active or inactive time intervals.

[0279] In many embodiments, the audio detector 801 can directly detect activity of an audio source based on a received microphone signal / a first set of audio signals / beamformer input signal / beamformed output audio signal. For example, the total captured audio energy can be determined and compared to a threshold. However, in other embodiments, other signals, such as a second set of audio signals, can be considered. In many embodiments, the audio detector 801 can make activity detection based on the generated beamformed output audio signal. For example, level detection or speech detection can be directly applied to the beamformed output audio signal. In many embodiments, this can provide improved performance because the output signal can be specifically generated to focus on a desired signal, such as the desired speaker.

[0280] The audio detector 801 can be arranged to control the adaptation based on the detection of audio source activity, and specifically, the adaptation of the decorrelation coefficients and / or beamforming filter coefficients can differ during the active time interval from that during the inactive time interval.

[0281] In practice, in some embodiments, the adaptation of decorrelation coefficients may not be performed during the active time interval, but only during the inactive time interval. The audio detector 801 may specifically seek to detect whether a desired audio source / speaker is active. It may further control the decorrelation adapter 203 to adapt the decorrelation coefficients only when the audio source is inactive. Therefore, the decorrelation adapter 203 may be arranged to adapt to provide a decorrelation filter that seeks to decorrelate unwanted audio. Thus, instead of updating the decorrelation filter for the entire captured audio, the adaptation is designed to reduce the correlation of unwanted audio. This approach can provide very efficient and improved operation and performance in many cases. The decorrelation method for unwanted audio can allow for improved beamforming, where the suppression of unwanted audio is improved by beamforming, while allowing beamforming to provide efficient and high-performance extraction of the desired signal.

[0282] In some embodiments, the decorrelation adapter 203 may be arranged to adapt decorrelation coefficients in both active and inactive time intervals, but at a higher rate during the inactive time interval than during the active time interval. Therefore, instead of adapting decorrelation coefficients only during the inactive time interval, some adaptation may also occur during the active time interval, but this adaptation is slower, for example, coefficients not less than 2, 5, 10, or 100. This may be advantageous in some scenarios, such as when the active time interval is much longer than the inactive time interval. In some embodiments, the update rate during the active and / or inactive time intervals may be dynamically adapted based on, for example, the time interval or a property of the input signal. For example, if the duration of the active time interval exceeds a given duration, the decorrelation adapter 203 may switch from updating only during the inactive time interval to updating also during the active time interval (at a lower rate).

[0283] In many embodiments, the adaptation of the beamforming filter may not be performed during inactive time intervals, but only during active time intervals. The audio detector 801 may specifically seek to detect whether a desired or desired audio source / speaker is active and control the beamforming adapter 205 to adapt the beamforming filter only when that audio source is active. Therefore, the beamforming adapter 205 is arranged to specifically adapt beamforming to specifically (virtually) form a beam toward the desired audio source.

[0284] In some embodiments, the beamforming adapter 205 may be arranged to adapt beamformer coefficients in both active and inactive time intervals, but at a higher rate during the active time interval than during the inactive time interval. Therefore, instead of adapting the beamforming filter only during the active time interval, some adaptation may also occur during the inactive time interval, but this adaptation is slower, for example, with coefficients not less than 2, 5, 10, or 100. This may be advantageous in some scenarios, such as, for example, where the audio detector 801 is configured to provide detection with a very low risk of false detection of an active audio source, making the risk of not detecting an active audio source relatively high. In this case, it may be desirable to still adapt during the inactive time interval, but with a (typically) much lower update rate. Thus, a low update rate can be used during time intervals where the desired source may or may not be active, and a high update rate can be used during time intervals where the desired source is almost certainly present.

[0285] In many embodiments, the audio device can be arranged to adapt the decorrelation coefficient (only) during inactive time intervals and the beamforming filter (only) during active time intervals. This can provide efficient performance and, in particular, can provide significantly improved extraction / separation of the desired audio source / speaker, especially in the presence of noise / undesired audio (such as, in particular, primary and related noise sources).

[0286] In many practical applications, such as many speech capture and processing applications, adaptation control of both beamforming and decorrelation is important for the efficient operation of audio devices. It may be desirable for the beamformer to adapt to speech that is not always active and to decorrelate only the noisy portions. In use cases where noise is continuously available, the beamformer can only adapt if noise is also present. Where the noise is uncorrelated or decorrelated by decorrelation, if the noise is isotropic, the noise does not affect the adaptation. If the noise is not isotropic, the beamformer will only diverge during periods of noise, but if the near end becomes active, the beamformer can typically quickly (almost imperceptibly) tune to the desired speaker. Depending on the application, either no further adaptation control is used, allowing the beamformer to quickly find new sources, or a speech activity detector is used, which can vary from simple energy-based detectors to more complex detectors that also use, for example, pitch-based speech characteristics.

[0287] The above description focuses on a scenario where the active time interval is the desired / voice audio source active time interval and the inactive time interval is the unwanted source / noise / interference active time interval. In this case, the adaptation rate for the beamformer filter is higher during the active time interval than during the inactive time interval, and specifically, the adaptation rate for the beamformer filter can be performed only during the active time interval and not during the inactive time interval. Equivalently, the adaptation rate for the decorrelation coefficient is higher during the inactive time interval than during the active time interval, and specifically, the adaptation rate for the filter can be performed only during the inactive time interval and not during the active time interval.

[0288] Such time intervals can be determined directly, for example, by using a speech detector, which detects and determines speech activity time intervals, which can be considered active time intervals. The remaining time intervals, i.e., non-speech activity time intervals, can be considered inactive time intervals.

[0289] In some embodiments, detection of unwanted audio properties, such as the detection of activity of specific interference or noise sources, can be used. For example, detection of activity of undesired audio sources, such as music detectors, silence detectors, or specific noise detectors, can be used. In this case, the detected time interval can be considered an inactive time interval. The remaining time interval can be considered an active time interval.

[0290] It should be understood that if, conversely, the active time interval is considered to correspond to the time interval in which the unwanted audio source is active, and the inactive time interval is considered to correspond to the time interval in which the unwanted audio source is inactive, then the previously described adaptation will be reversed (the adaptation rate of the filter will be higher during the inactive time interval, the adaptation rate of the decorrelation coefficient will be higher during the active time interval, and specifically, in many embodiments, the adaptation of the beamforming filter will only be during the inactive time interval, and the adaptation of the decorrelation coefficient will only be during the active time interval).

[0291] In this method, the decorrelation function and beamforming function are dynamically updated, and accordingly, the equalizer 105 is dynamically adapted to have a variable compensated frequency response that follows the adaptation to decorrelation and beamforming. In many embodiments, the operation is performed in segments or blocks, and for each block, new and updated parameters are determined.

[0292] However, abrupt changes in the compensation frequency response have been found to produce audio artifacts in many scenarios and applications. In many embodiments, the equalizer adapter 107 is arranged to gradually change the compensation frequency response, and the compensation frequency response does not undergo sudden and abrupt substantial changes. Specifically, the compensation frequency response changes not only between different segments, but also gradually during the segments.

[0293] In many embodiments, the equalizer adapter 107 can be arranged to interpolate between the current compensated frequency response and the new compensated frequency response, wherein the interpolation gradually changes from the former to the latter. Specifically, the beamforming output audio signal can undergo both the current frequency equalization and the new frequency equalization, and specifically, samples of the beamforming output audio signal can be filtered by both the current compensated frequency response and the new compensated frequency response. Therefore, samples are generated for the two signals corresponding to the current frequency equalization and the new frequency equalization, respectively. The equalizer 105 can then generate an output signal as an interpolation between these compensated frequency responses, and specifically, samples of the output signal can be generated as a weighted combination of samples of the two signals (at the same sample time). The weighting can then gradually change from the higher weight of the sample used for the current frequency equalization to the higher weight of the sample used for the new frequency equalization. The rate of change of the combination / interpolation can be selected to provide suitable performance for a particular embodiment.

[0294] The equalizer adapter 107 can therefore control the equalizer 105 to switch from one frequency equalizer / compensated frequency response to another frequency equalizer / compensated frequency response by determining intermediate samples of the output signals for two frequency equalizers and then generating samples of the output signals as a weighted combination of intermediate samples for different frequency equalizers, wherein the equalizer adapter 107 is then arranged to gradually change the relative weights of the samples of different frequency equalizers during the switching time interval.

[0295] Therefore, when When updated, a smooth transition is likely preferred. For a transition frame (typically a block of 256 samples), this can be achieved by calculating the old... and new The outputs of both are then calculated, and the output with a smooth transition is then computed to obtain good results.

[0296] Will utilize the old The processed frame samples are represented as "inold" and will utilize new... If the sample of the processed frame is represented as "innew", then it can be processed using the following method (in pseudocode) to... and for example Use this as the starting value to determine the output:

[0297] Although the foregoing description focuses on adaptive spatial decorrelation performed by adaptive spatial decorrelation 301, it should be understood that in other embodiments, non-adaptive spatial decorrelation may be performed.

[0298] The audio device can be specifically implemented in one or more appropriately programmed processors. Different functional blocks can be implemented in separate processors and / or, for example, in the same processor. Examples of suitable processors are provided below.

[0299] Figure 9This is a block diagram illustrating an example processor 900 according to an embodiment of the present disclosure. Processor 900 can be used to implement one or more processors, which implement the apparatus or elements thereof as described above (particularly including one or more artificial neural networks). Processor 900 can be any suitable processor type, including but not limited to microprocessors, microcontrollers, digital signal processors (DSPs), field-programmable gate arrays (FPGAs) (wherein the FPGA has been programmed to form a processor), graphics processing units (GPUs), application-specific integrated circuits (ASICs) (wherein the ASIC has been designed to form a processor), or combinations thereof.

[0300] Processor 900 may include one or more cores 902. Core 902 may include one or more arithmetic logic units (ALUs) 904. In some embodiments, in addition to or instead of ALU 904, core 902 may also include a floating-point logic unit (FPLU) 906 and / or a digital signal processing unit (DSPU) 908.

[0301] Processor 900 may include one or more registers 912 communicatively coupled to core 902. Registers 912 may be implemented using dedicated logic gates (e.g., flip-flops) and / or any memory technology. In some embodiments, registers 912 may be implemented using static memory. Registers may provide data, instructions, and addresses to core 902.

[0302] In some embodiments, processor 900 may include one or more levels of cache memory 910 communicatively coupled to core 902. Cache memory 910 may provide computer-readable instructions to core 902 for execution. Cache memory 910 may provide data for processing by core 902. In some embodiments, the computer-readable instructions may have already been provided to cache memory 910 by local memory (e.g., local memory attached to external bus 916). Cache memory 910 may be implemented using any suitable cache memory type, such as metal-oxide-semiconductor (MOS) memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), and / or any other suitable memory technology.

[0303] Processor 900 may include controller 914, which controls inputs to and / or outputs from other processors and / or components included in the system to other processors and / or components included in the system. Controller 914 may control data paths in ALU 904, FPLU 906, and / or DSPU 908. Controller 914 may be implemented as one or more state machines, data paths, and / or dedicated control logic. The gates of controller 914 may be implemented as stand-alone gates, FPGAs, ASICs, or any other suitable technology.

[0304] Register 912 and cache 910 can communicate with controller 914 and core 902 via internal connections 920A, 920B, 920C, and 920D. These internal connections can be implemented as buses, multiplexers, cross switches, and / or any other suitable connection technology.

[0305] Inputs and outputs for processor 900 may be provided via bus 916, which may include one or more conductive lines. Bus 916 may be communicatively coupled to one or more components of processor 900, such as controller 914, cache 910, and / or register 912. Bus 916 may be coupled to one or more components of the system.

[0306] Bus 916 may be coupled to one or more external memories. The external memory may include read-only memory (ROM) 932. ROM 932 may be a masked ROM, electrically programmable read-only memory (EPROM), or any other suitable technology. The external memory may include random access memory (RAM) 933. RAM 933 may be static RAM, battery-backed static RAM, dynamic RAM (DRAM), or any other suitable technology. The external memory may include electrically erasable programmable read-only memory (EEPROM) 935. The external memory may include flash memory 934. The external memory may include a magnetic storage device such as a disk 936. In some embodiments, external memory may be included in the system.

[0307] It should be understood that, for clarity, the above description has referenced various functional circuits, units, and processors in describing embodiments of the invention. However, it will be apparent that any suitable functional distribution among different functional circuits, units, or processors can be used without departing from the invention. For example, a function shown to be performed by a separate processor or controller may be performed by the same processor or controller. Therefore, references to specific functional units or circuits are to be considered merely as references to suitable units used to provide the described functions, and not as indications of a strict logical or physical structure or organization.

[0308] This invention can be implemented in any suitable form, including hardware, software, firmware, or any combination of these. The invention can optionally be implemented, at least in part, as computer software running on one or more data processors and / or digital signal processors. The elements and components of embodiments of the invention can be implemented physically, functionally, and logically in any suitable manner. In practice, functionality can be implemented in a single unit, in multiple units, or as part of other functional units. Therefore, the invention can be implemented in a single unit or can be physically and functionally distributed among different units, circuits, and processors.

[0309] Although the invention has been described in conjunction with some embodiments, it is not intended to be limited to the specific forms set forth herein. Rather, the scope of the invention is defined only by the appended claims. Furthermore, while features may appear to have been described in conjunction with specific embodiments, those skilled in the art will recognize that various features of the described embodiments can be combined according to the invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.

[0310] Furthermore, although listed separately, multiple units, elements, circuits, or method steps can be implemented by, for example, a single circuit, unit, or processor. Additionally, although individual features may be included in different claims, these features can be advantageously combined, and inclusion in different claims does not imply that such a combination of features is not feasible and / or advantageous. Furthermore, including a feature in one class of claims does not imply a limitation on that class, but rather indicates that the feature is equally applicable to other claim classes, as appropriate. Moreover, the order of features in a claim does not imply any particular order in which the features must operate, and in particular, the order of steps in a method claim does not imply that the steps must be performed in that order. Rather, these steps can be performed in any suitable order. Furthermore, singular references do not exclude plural. Therefore, references to “a,” “an,” “first,” “second,” etc., do not exclude plural. Reference numerals in the claims are provided only as clarifying examples and should not be construed as limiting the scope of the claims in any way.

[0311] Typically, examples of audio devices, methods of operating audio devices, and computer programs that implement such methods are indicated by the following embodiments. Example: Example 1: An audio device, comprising: A receiver (101) is arranged to receive a first set of audio signals, the first set of audio signals including audio signals of the scene captured from different locations; An audio beamformer (103) is arranged to generate a beamformed output audio signal based on the first set of audio signals, the audio beamformer (103) comprising: A signal processor (201) is arranged to generate the beamforming output audio signal based on signal processing of the first audio signal set, the signal processing including spatial beamforming and spatial decorrelation for the first audio signal set, the spatial decorrelation depending on a decorrelation parameter set and the spatial beamforming depending on a beamforming parameter set. A beamforming adapter (205) is arranged to adapt the beamforming parameter set according to the beamforming output audio signal. An equalizer (105) is arranged to apply frequency equalization to the beamforming output audio signal to generate an output signal; and An equalization adapter (107) is arranged to adapt the frequency equalization according to the set of beamforming parameters. Example 2: The audio device according to Example 1, wherein the equalizer adapter (107) is arranged to adapt frequency equalization according to the set of decorrelation parameters. Example 3: The audio device according to Example 1, wherein the signal processor (201) includes: Spatial decorrelation unit (301), configured to receive the first set of audio signals and generate a decorrelated first set of audio signals by performing the spatial decorrelation; and A beamforming circuit (303) is arranged to perform spatial beamforming by combining the first set of decorrelated audio signals, the combination depending on the set of beamforming parameters. Example 4: An audio device according to any of the foregoing embodiments, wherein the equalizer adapter (107) is arranged to adapt the frequency equalizer to have a minimum phase frequency response. Example 5: An audio device according to any of the foregoing embodiments, wherein the equalizer adapter (107) is arranged to adapt the frequency equalizer to have a low-pass frequency response. Example 6. An audio device according to any of the foregoing embodiments, wherein an equalizer adapter (107) is arranged to switch from a first frequency equalizer to a second frequency equalizer by: determining a first intermediate sample of the output signal for the first frequency equalizer and a second intermediate sample of the output signal for the second frequency equalizer, and generating a sample of the output signal as a weighted combination of the first intermediate sample and the second intermediate sample, wherein the equalizer adapter (107) is arranged to gradually change the relative weights of the first intermediate sample and the second intermediate sample during a switching time interval. Example 7: The audio device according to Example 1, wherein the spatial decorrelation is adaptive spatial decorrelation, and the audio device further includes: A decorrelation adapter (203) is configured to adapt the decorrelation parameter set according to the first set of audio signals. Example 8: The audio device according to Example 7 further includes an audio detector (801) arranged to determine an active time interval set during which the audio source is active and an inactive time interval set during which the audio source is inactive; and wherein at least one of the adaptation to the first filter set (303) and the second filter set (401) and the adaptation to the decorrelation parameter set (203) is different for the active time interval set and the inactive time interval set. Example 9: The audio device according to Example 7, wherein the signal processor (201) includes: A first filter set (303) is arranged to filter the first audio signal set, and A combiner (303) is arranged to combine the outputs of the first filter set (303) to generate the beamforming output audio signal; Feedback circuit (401) includes a second filter set arranged to generate a second audio signal set by filtering the beamforming output audio signal, each filter in the second filter set having a frequency response that is the complex conjugate of a filter in the first filter set (303). A first set of spatial filters (301, 601, 701) is arranged to apply a first spatial filter to at least one of the first set of audio signals and the second set of audio signals, the first set of spatial filters (301, 601, 701) having coefficients determined according to the decorrelation coefficients included in the decorrelation parameters; And among them The beamforming adapter (205) is arranged to adapt the first filter set (303) and the second filter set in response to a comparison of the first audio signal set and the second audio signal set; and The decorrelation adapter (203) is arranged to determine decorrelation coefficients for a set of spatial decorrelation filters to generate a decorrelated output signal based on the first set of audio signals, and the decorrelation adapter (203) is arranged to adapt the decorrelation coefficients in response to an update value determined based on the first set of audio signals. Example 10: The audio device according to Example 9, wherein the first spatial filter set (301) is arranged to filter the first audio signal set. Example 11: An audio device according to Example 9 or 10, wherein the first filter set (303) is arranged to filter the first audio signal set after it has been filtered by the first spatial filter set (601), and the beamforming adapter (205) is arranged to perform the comparison using the first audio signal set before it has been filtered by the first spatial filter set (601). Example 12: An audio device according to any one of Examples 8 or 9, wherein the first spatial filter set (701) is arranged to filter the second audio signal set. Example 13: An audio device according to any of the foregoing embodiments, wherein the decorrelation adapter (203) is arranged to determine the adaptive spatial decorrelation as a set of spatial decorrelation filters by performing the following steps, wherein each output audio signal of a filter in the set of spatial decorrelation filters is linked to an input audio signal in the first set of audio signals: The first audio signal set is segmented into time segments, and for at least some of the time segments, the following steps are performed: A frequency bin representation of the first audio signal set is generated, wherein each frequency bin of the frequency bin representation of the first audio signal set includes a frequency bin value for each audio signal in the first audio signal set; A frequency bin representation of a set of output signals is generated, each frequency bin in the frequency bin representation of the set of output signals including a frequency bin value for each output signal in the set of output signals, the frequency bin value for a given frequency bin for a given output signal in the set of output signals is generated as a weighted combination of the frequency bin values ​​for the given frequency bin of the first audio signal set, the weighted combination having the decorrelation coefficients as weights; and In response to a correlation metric between a first previous frequency bin value for a first frequency bin of a first output signal linked to a first input audio signal and a second previous frequency bin value for a second output signal linked to the first frequency bin, a first weight is updated for the contribution of the second frequency bin value of the first frequency bin to the first frequency bin value of the first frequency bin for the second input audio signal linked to the second output signal. Example 14: An audio device according to Example 13, wherein the decorrelation adapter (203) is arranged to update the first weight in response to the product of a first value and a second value, the first value being one of a first previous frequency bin value and a second previous frequency bin value, and the second value being the complex conjugate of the other of the first previous frequency bin value and the second previous frequency bin value. Example 15: An operating method for an audio device, the method comprising: A receiver (101) is arranged to receive a first set of audio signals, the first set of audio signals including audio signals of the scene captured from different locations; An audio beamformer (103) is arranged to generate a beamformed output audio signal based on the first set of audio signals, the audio beamformer (103) comprising: A signal processor (201) is arranged to generate the beamforming output audio signal based on signal processing of the first audio signal set, the signal processing including spatial beamforming and adaptive spatial decorrelation for the first audio signal set, the adaptive spatial decorrelation depending on a decorrelation parameter set, and the spatial beamforming depending on a beamforming parameter set. A decorrelation adapter (203) is configured to adapt the decorrelation parameter set according to the first set of audio signals; A beamforming adapter (205) is arranged to adapt the beamforming parameter set according to the beamforming output audio signal. An equalizer (105) is arranged to apply frequency equalization to the beamforming output audio signal to generate an output signal; and An equalization adapter (107) is arranged to adapt the frequency equalization according to the set of beamforming parameters. Example 16: A computer program product including computer program code units adapted to perform all the steps of Example 15 when the program is run on a computer.

Claims

1. An audio apparatus, comprising: a receiver (101) arranged to receive a first set of audio signals, the first set of audio signals comprising audio signals capturing audio of a scene from different positions; an audio beamformer (103) arranged to generate a beamformed output audio signal from the first set of audio signals, the audio beamformer (103) comprising: a signal processor (201) arranged to generate the beamformed output audio signal from a signal processing of the first set of audio signals, the signal processing comprising spatial beamforming for the first set of audio signals and spatial decorrelation, the spatial decorrelation depending on a set of decorrelation parameters and the spatial beamforming depending on a set of beamforming parameters, the beamforming parameters being parameters describing at least one of filtering and weighting of a signal of a combination of the beamforming; a beamforming adaptor (205) arranged to adapt the set of beamforming parameters from the beamformed output audio signal; an equalizer (105) arranged to apply a frequency equalization to the beamformed output audio signal to generate an output signal; and an equalization adaptor (107) arranged to adapt the frequency equalization from the set of beamforming parameters. The equalization adaptor (107) is arranged to adapt the frequency equalization from the set of decorrelation parameters.

2. The audio apparatus of claim 1, wherein, The signal processor (201) comprises:

3. The audio device of claim 1, wherein, a spatial decorrelator (301) arranged to receive the first set of audio signals and to generate a decorrelated first set of audio signals by performing the spatial decorrelation; and a beamforming circuit (303) arranged to perform the spatial beamforming by combining the decorrelated first set of audio signals, the combining depending on the set of beamforming parameters. The equalization adaptor (107) is arranged to adapt the frequency equalization to have a minimum phase frequency response.

4. The audio apparatus of any previous claim, wherein, The equalization adaptor (107) is arranged to adapt the frequency equalization to have a low-pass frequency response.

5. The audio apparatus of any previous claim, wherein, The equalization adaptor (107) is arranged to transition from a first frequency equalization to a second frequency equalization by determining a first intermediate sample of the output signal for the first frequency equalization and a second intermediate sample of the output signal for the second frequency equalization, and generating samples of the output signal as a weighted combination of the first and second intermediate samples, the equalization adaptor (107) being arranged to gradually change the relative weights of the first and second intermediate samples during a transition time interval.

6. The audio apparatus of any previous claim, wherein, The spatial decorrelation is an adaptive spatial decorrelation, and the audio apparatus further comprises:

7. The audio device of claim 1, wherein, a decorrelation adaptor (203) arranged to adapt the set of decorrelation parameters from the first set of audio signals. At least one of the adaptation of the set of beamforming parameters and the adaptation (203) of the set of decorrelation parameters is different for the set of active time intervals and the set of inactive time intervals.

8. The audio apparatus of claim 7, further comprising an audio detector (801) arranged to determine a set of active time intervals during which an audio source is active and a set of inactive time intervals during which the audio source is inactive; and wherein, ​ 9. The audio device of claim 7, wherein, The signal processor (201) comprises: a first filter set (303) arranged to filter the first set of audio signals, and a combiner (303) arranged to combine outputs of the first filter set (303) to generate the beamformed output audio signal; a feedback circuit (401) comprising a second filter set arranged to generate a second set of audio signals from a filtering of the beamformed output audio signal, each filter of the second filter set having a frequency response that is a complex conjugate of a frequency response of a filter of the first filter set (303); a first spatial filter set (301, 601, 701) arranged to apply a first spatial filtering to at least one of the first set of audio signals and the second set of audio signals, the first spatial filter set (301, 601, 701) having coefficients determined from decorrelation coefficients comprised in the decorrelation parameters; and wherein the beamforming adapter (205) is arranged to adapt the first filter set (303) and the second filter set in response to the comparison of the first set of audio signals and the second set of audio signals; and the decorrelation adapter (203) is arranged to determine decorrelation coefficients for a spatial decorrelation filter set to generate a decorrelated output signal from the first set of audio signals, the decorrelation adapter (203) being arranged to adapt the decorrelation coefficients in response to update values determined from the first set of audio signals.

10. The audio device of claim 9, wherein, The first spatial filter set (301) is arranged to filter the first set of audio signals.

11. The audio apparatus of claim 9 or 10, wherein, The first filter set (303) is arranged to filter the first set of audio signals after filtering by the first spatial filter set (601), and the beamforming adapter (205) is arranged to perform the comparison using the first set of audio signals before filtering by the first spatial filter set (601).

12. The audio device of claim 9, wherein, The first spatial filter set (701) is arranged to filter the second set of audio signals.

13. The audio apparatus of any previous claim, wherein, The decorrelation adapter (203) is arranged to determine the adaptive spatial decorrelation as a spatial decorrelation filter set by performing the following steps, each output audio signal of a filter of the spatial decorrelation filter set being linked to one input audio signal of the first set of audio signals: segmenting the first set of audio signals into time segments, and for at least some time segments, performing the following steps: generating a frequency bin representation of the first set of audio signals, each frequency bin of the frequency bin representation of the first set of audio signals comprising a frequency bin value for each of the audio signals in the first set of audio signals; generating a frequency bin representation of a set of output signals, each frequency bin in the frequency bin representation of the set of output signals comprising a frequency bin value for each of the output signals, the frequency bin value for a given frequency bin for a given output signal of the set of output signals being generated as a weighted combination of frequency bin values for the given frequency bin of the first set of audio signals, the weighted combination having as weights the decorrelation coefficients; and updating, in response to a measure of correlation between a first previous frequency bin value for a first frequency bin of a first output signal linked to a first input audio signal and a second previous frequency bin value for the first frequency bin of the second output signal, a first weight of a contribution of a second frequency bin value for the first frequency bin from a second input audio signal linked to the second output signal to a first frequency bin value for the first frequency bin of the first output signal.

14. The audio device of claim 13, wherein, The decorrelation adapter (203) is arranged to update the first weight in response to a product of a first value and a second value, the first value being one of the first previous frequency bin value and the second previous frequency bin value, and the second value being a complex conjugate of the other of the first previous frequency bin value and the second previous frequency bin value.

15. An operating method for an audio apparatus, the method comprising: receiving a first set of audio signals, the first set of audio signals comprising audio signals capturing audio of a scene from different positions; generating a beamformed output audio signal from the first set of audio signals, wherein generating the beamformed output audio signal comprises: generating the beamformed output audio signal from signal processing of the first set of audio signals, the signal processing comprising spatial beamforming for the first set of audio signals and adaptive spatial decorrelation, the adaptive spatial decorrelation depending on a set of decorrelation parameters, and the spatial beamforming depending on a set of beamforming parameters, the beamforming parameters being parameters describing at least one of filtering and weighting of input signals to a combination of the beamforming; adapting the set of decorrelation parameters from the first set of audio signals; adapting the set of beamforming parameters from the beamformed output audio signal; applying a frequency equalization to the beamformed output audio signal to generate an output signal; and adapting the frequency equalization from the set of beamforming parameters.

16. A computer program product comprising computer program code means adapted to perform all the steps of claim 15 when said program is run on a computer.

Citation Information

Patent Citations

  • Signal localization arrangement

    US6774934B1

  • Audio processing arrangement with multiple sources

    US7146012B1