Fitting of frequency equalization to beamforming dependence

The audio device uses spatial beamforming and decorrelation with adaptive parameters and frequency equalization to improve sound source isolation and audio quality, addressing the limitations of existing beamforming techniques by reducing complexity and noise sensitivity.

JP2026518306APending Publication Date: 2026-06-04KONINKLIJKE PHILIPS NV

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
KONINKLIJKE PHILIPS NV
Filing Date
2024-05-16
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing audio beamforming techniques struggle to optimally isolate sound sources, particularly in complex environments, often leading to suboptimal performance, high complexity, and resource usage, while failing to provide ideal sound characteristics for further processing.

Method used

An audio device employing spatial beamforming and spatial decorrelation with adaptive parameters, combined with frequency equalization, to generate an output audio signal that effectively isolates desired sound sources by adapting to acoustic and spatial characteristics, reducing sensitivity to noise, and minimizing frequency distortion.

Benefits of technology

The approach enhances sound source extraction, reduces complexity, and improves the quality of the captured audio by minimizing noise interference and frequency distortion, resulting in a more natural-sounding output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026518306000001_ABST
    Figure 2026518306000001_ABST
Patent Text Reader

Abstract

The audio device includes a receiver 101 configured to receive an audio signal from which an audio beamformer 103 generates a beamformed output audio signal. The beamformer 103 includes a signal processor 201 that applies signal processing to the audio signal. Signal processing includes spatial beamforming and spatial decorrelation, which is often adapted spatial decorrelation. A decorrelation adapter 203 adapts decorrelation parameters depending on the audio signal, and a beamform adapter 205 adapts beamform parameters depending on the beamformed output audio signal. An equalizer 105 applies frequency equalization to the beamformed output audio signal to generate an output signal, and an equalization adapter 107 adapts frequency equalization depending on the beamform parameters, typically also depending on the decorrelation parameters. This approach provides improved fidelity to the desired sound source and attenuates interfering sound sources.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an apparatus and method for generating an audio output signal, and more specifically, but not limited to, to extracting audio from a desired sound source, such as a desired speaker, using beamforming. [Background technology]

[0002] Capturing speech, specifically utterances, has become increasingly important in recent decades. For example, capturing speech or other sounds is becoming increasingly crucial for a variety of applications, including telecommunications, remote conferencing, gaming, and voice user interfaces. However, a problem in many scenarios and applications is that the desired sound source is not typically the only sound source in the environment. Rather, in a typical voice environment, there are many other sound sources / noise sources that are captured by the microphone. Speech processing is often used to improve speech capture, particularly by post-processing the captured speech time intervals to enhance the resulting audio signal.

[0003] In many embodiments, speech is represented by multiple different audio signals that reflect the same audio scene or environment. In particular, in many practical applications, speech is captured by multiple microphones at different locations. For example, a linear array of multiple microphones is often used to capture speech in an environment such as a room. The use of multiple microphones allows spatial information of the speech to be captured. Many different applications leverage such spatial information to enable improved and / or new services. [Overview of the Initiative] [Problems that the invention aims to solve]

[0004] One frequently used approach is to attempt to isolate sound sources by applying beamforming, which involves forming a beam oriented relative to the direction of arrival of sound from a particular source. However, while this offers favorable performance in many scenarios, it is not optimal in all cases. For example, it does not provide optimal source isolation in some cases, and in fact, in some applications, such spatial beamforming does not provide ideal sound characteristics for further processing to achieve a given effect.

[0005] Therefore, while spatial source separation, specifically separation based on audio beamforming, is highly advantageous in many scenarios and applications, there is a demand for improving the performance and operation of such approaches. However, there is also a demand for typically low complexity and / or resource usage (e.g., computational resources), and these preferences often conflict with each other.

[0006] Therefore, an improved approach is advantageous, particularly one that enables reduced complexity, increased flexibility, easier implementation, reduced cost, improved voice capture, improved source discrimination, improved source isolation, improved support for voice / speech applications, reduced reliance on known or static acoustic properties, improved flexibility and customization for different voice environments and scenarios, improved voice beamforming, a better trade-off between performance and complexity / resource usage, and / or improved performance.

[0007] Therefore, the present invention seeks to mitigate, mitigate, or eliminate one or more of the aforementioned disadvantages, either individually or in any combination. [Means for solving the problem]

[0008] According to an aspect of the present invention, an audio device is provided, comprising: a receiver configured to receive a first set of audio signals, the first set of audio signals including audio signals that capture the sound of a scene from different locations; and an audio beamformer configured to generate a beamform output audio signal from the first set of audio signals, the audio beamformer comprising: a signal processor configured to generate a beamform output audio signal from signal processing of the first set of audio signals, the signal processor including spatial beamforming and spatial decorrelation for the first set of audio signals, the spatial decorrelation depending on a set of decorrelation parameters, the spatial beamforming depending on a set of beamform parameters, the beamform parameters being parameters that indicate at least one of filtering and weighting of the input signal for a beamforming combination; a beamform adapter configured to adapt a set of beamform parameters depending on the beamform output audio signal; an equalizer configured to apply frequency equalization to the beamform output audio signal to generate an output signal; and an equalization adapter configured to adapt frequency equalization depending on the set of beamform parameters.

[0009] This approach provides improved operation and / or performance in many embodiments. This approach enables improved beamforming, particularly for focusing on specific sound sources in a scene. This approach enables improved extraction / separation of speech from specific sources in the presence of other sound sources and noise sources in the scene.

[0010] This approach enables the operation of the audio device to effectively adapt to current conditions, particularly the acoustic and spatial characteristics of the sound source and scene. This approach provides reduced sensitivity to noise and unwanted sound sources in the scene and captured audio.

[0011] This approach enables efficient operation and low complexity in many embodiments. Different fits work synergistically to provide improved separation of the desired sound source from other captured audio in the scene. Furthermore, the use of multiple fits provides improved operation while allowing the use of less complex fitting algorithms and criteria.

[0012] This approach provides a higher quality output audio signal. In many embodiments, this approach reduces or mitigates frequency distortion. This approach results in a more natural-sounding capture of the desired sound source. Higher fidelity of the captured audio is achieved in many embodiments.

[0013] The equalizer is configured to apply frequency equalization to the beamform output audio signal by applying a frequency response. The equalizer is configured to apply frequency equalization to the beamform output audio signal by filtering out beamform output audio signals that have a non-constant magnitude frequency response. The equalization adapter is configured to adapt the frequency response.

[0014] According to an optional feature of the present invention, the equalization adapter is configured to adapt frequency equalization depending on a set of decorrelation parameters.

[0015] This provides improved performance and / or operation in many embodiments. In particular, it generates an output audio signal that more accurately captures the desired sound source in an audio scene in many embodiments. This allows for improved frequency equalization and reduced frequency distortion in many scenarios.

[0016] According to an optional feature of the present invention, the signal processor comprises: a spatial decorrelator configured to receive a first set of audio signals and generate a first set of decorrelated audio signals by performing spatial decorrelation; and a beamforming circuit configured to perform spatial beamforming by combining the first set of decorrelated audio signals, the combination of which depends on a set of beamforming parameters.

[0017] This provides improved performance and / or operation in many embodiments. In particular, in many embodiments, it facilitates implementation and provides accurate voice capture while maintaining low complexity.

[0018] According to an optional feature of the present invention, the equalization adapter is configured to have a minimum phase frequency response by adapting frequency equalization.

[0019] This provides improved performance and / or operation in many embodiments. According to an optional feature of the present invention, the equalization adapter is configured to adapt frequency equalization to have a low-pass frequency response.

[0020] This provides improved performance and / or operation in many embodiments. In particular, it generates an output audio signal that more accurately captures the desired sound source in an audio scene in many embodiments. This allows for improved frequency equalization and reduced frequency distortion in many scenarios.

[0021] According to an optional feature of the present invention, the equalization adapter is configured to transition from a first frequency equalization to a second frequency equalization by determining a first intermediate sample of the output signal for a first frequency equalization and a second intermediate sample of the output signal for a second frequency equalization, and generating a sample of the output signal as a weighted combination of the first intermediate sample and the second intermediate sample, wherein the equalization adapter is configured to gradually change the relative weights of the first intermediate sample and the second intermediate sample during the transition time interval.

[0022] This provides improved performance and / or operation in many embodiments. It reduces audio artifacts in many scenarios.

[0023] According to an optional feature of the present invention, the spatial decorrelation is adaptive spatial decorrelation, and the audio device further comprises a decorrelation adapter configured to adapt a set of decorrelation parameters depending on a first set of audio signals.

[0024] This provides improved performance and / or operation in many embodiments.

[0025] According to an optional feature of the present invention, the audio device further comprises an audio detector configured to determine, among a set of active time intervals, a set of active time intervals in which the sound source is active, and among a set of inactive time intervals, a set of inactive time intervals in which the sound source is not active, wherein at least one of the fitting of a set of beamform parameters and the fitting of a set of decorrelation coefficients differs for the set of active time intervals and the set of inactive time intervals.

[0026] This provides improved performance and / or operation in many embodiments. Typically, it provides improved adaptation, which usually leads to improved extraction of the desired sound source.

[0027] The active time interval is the active time interval of the desired / preferred / speech source, and the inactive time interval is the inactive time interval of the desired / preferred / speech source.

[0028] In some embodiments, the beamform adapter is configured to fit a first set of filters and a second set of filters at a higher fitting rate during a set of active time intervals than during a set of inactive time intervals.

[0029] This provides improved performance and / or operation in many embodiments. Typically, it provides improved adaptation, which usually leads to improved extraction of the desired sound source.

[0030] In some embodiments, the beamform adapter is configured to adapt a first set of filters and a second set of filters for a period of time of only one set of time intervals from a set of active time intervals and a set of inactive time intervals.

[0031] This provides improved performance and / or operation in many embodiments. Typically, it provides improved adaptation, which usually leads to improved extraction of the desired sound source.

[0032] In some embodiments, the beamform adapter is configured to fit a first set of filters and a second set of filters during a set of active time intervals, but not during a set of inactive time intervals.

[0033] In some embodiments, the decorrelation adapter is configured to fit the decorrelation coefficient at a higher rate during periods of inactive time intervals than during periods of active time intervals.

[0034] This provides improved performance and / or operation in many embodiments. Typically, it provides improved adaptation, which usually leads to improved extraction of the desired sound source.

[0035] In some embodiments, the decorrelation adapter is configured to fit the decorrelation coefficient for a period of only one set of time intervals from a set of inactive time intervals and a set of active time intervals.

[0036] This provides improved performance and / or operation in many embodiments. Typically, it provides improved adaptation, which usually leads to improved extraction of the desired sound source.

[0037] In some embodiments, the fitting coefficient processor is configured to fit the uncorrelated coefficients over a set of inactive time intervals, but not over a set of active time intervals.

[0038] According to an optional feature of the present invention, the signal processor is a feedback circuit comprising: a first set of filters configured to filter a first set of audio signals; a combiner configured to combine the outputs of the first set of filters to generate a beamform output audio signal; a second set of filters configured to generate a second set of audio signals from the filtered beamform output audio signal, wherein each filter in the second set of filters has a frequency response that is the complex conjugate of the filters in the first set of filters; and a first spatial filter for at least one of the first set of audio signals and the second set of audio signals. A first set of spatial filters configured to apply taring, wherein the first set of spatial filters has coefficients determined from decorrelation coefficients included in the decorrelation parameters, the beamform adapter is configured to adapt the first set of filters and the second set of filters in response to a comparison of a first set of audio signals and a second set of audio signals, the decorrelation adapter is configured to determine decorrelation coefficients for a set of spatial decorrelation filters that generate a decorrelation output signal from the first set of audio signals, and the decorrelation adapter is configured to adapt the decorrelation coefficients in response to updated values ​​determined from the first set of audio signals.

[0039] This provides improved performance and / or operation in many embodiments. It offers particularly advantageous beamforming for specific approaches and works efficiently in conjunction with spatial decorrelation and equalization.

[0040] Each filter in the first set of filters is linked to a filter in the second set of filters, and this filter has a frequency response that is the complex conjugate of the frequency response of the filter in the first set of filters.

[0041] The first and second sets of filters each have an equal number of linked-pair filters having a complex conjugate frequency response. For each audio signal in the first set of audio signals, there exists one filter from the first set of filters, one linked-pair filter from the second set of filters (having a complex conjugate frequency response), and one audio signal in the second set of audio signals. The comparison is between the linked-pair audio signals in the first set of audio signals and the second set of audio signals. The fit of a given filter in the first set of filters and a given linked-pair filter in the second set of filters depends on a comparison (perhaps only) between the signal in the first set of audio signals filtered by the given filter in the first set of filters and the signal in the second set of audio signals produced by filtering the beamform output audio signal by the given linked-pair filter in the second set of filters. The fit is made such that the difference is reduced.

[0042] Each of the second set of audio signals is an estimate of the contribution of the first set of audio signals from the audio captured by the beamform output audio signal to the linked audio signal, and is therefore typically an estimate of the contribution from the desired source.

[0043] A set of spatial decorrelation filters contains one spatial filter for each signal in a first set of signals. Each filter in the set of spatial decorrelation filters produces a filtered / modified version of one of the audio signals in the first set of audio signals. Together, the set of spatial decorrelation filters produces a modified first set of audio signals with a higher degree of decorrelation. The spatial decorrelation filters perform filtering over the first set of audio signals. The output of the decorrelation filters (the corresponding modified audio signals in the first set of audio signals) depends on the first set of multiple (unmodified) audio signals. Specifically, the spatial decorrelation filters are frequency domain filters, and the decorrelation coefficients are frequency domain coefficients. For a given spatial decorrelation filter, the output value for a given frequency bin at a given time is a weighted combination of multiple values ​​from the first set of (unmodified) audio signals for a given frequency bin at a given time. The decorrelation output signal has the correlation of the reduced normalized cross-channel signal with respect to the first set of input signals to the set of spatial decorrelation filters.

[0044] A set of spatial filters includes one spatial filter for each signal in a first set / second set of signals. Each filter in the set of spatial filters produces a filtered / modified version of one of the first or second sets of audio signals. The set of spatial filters together produces a modified first or second set of audio signals. The spatial filters perform filtering across the first set of audio signals. The output of the spatial filter (the corresponding modified audio signal in the first or second set of audio signals) depends on multiple of the first or second sets of (unmodified) audio signals. The spatial filter is specifically a frequency domain filter, and its coefficients are frequency domain coefficients. For a given spatial filter, the output value for a given frequency bin at a given time is a weighted combination of multiple values ​​from the first or second set of (unmodified) audio signals for a given frequency bin at a given time.

[0045] According to an optional feature of the present invention, a first set of spatial filters is configured to filter a first set of audio signals.

[0046] This provides improved performance and / or operation in many embodiments. In many embodiments and scenarios, this provides particularly attractive performance and / or implementation.

[0047] In some embodiments, the first set of spatial filters is configured to have coefficients set to decorrelation coefficients determined for the set of spatial decorrelation filters.

[0048] This provides improved performance and / or operation in many embodiments. In many embodiments and scenarios, this provides particularly attractive performance and / or implementation.

[0049] According to an optional feature of the present invention, a first set of filters is configured to filter a first set of audio signals after filtering by a first set of spatial filters, and a beamform adapter is configured to perform a comparison using the first set of audio signals before filtering by the first set of spatial filters.

[0050] This provides improved performance and / or operation in many embodiments. In many embodiments and scenarios, this provides particularly attractive performance and / or implementation.

[0051] In some embodiments, the first set of spatial filters is configured to have coefficients that match the coefficients of the spatial filters, which are a cascade of two of the sets of uncorrelated filters.

[0052] This provides improved performance and / or operation in many embodiments. In many embodiments and scenarios, this provides particularly attractive performance and / or implementation.

[0053] According to an optional feature of the present invention, a first set of spatial filters is configured to filter a second set of audio signals.

[0054] This provides improved performance and / or operation in many embodiments. In many embodiments and scenarios, this provides particularly attractive performance and / or implementation.

[0055] In some embodiments, a first set of spatial filters is configured to have coefficients determined in accordance with a set of inverse spatial decorrelation filters, where the set of inverse spatial decorrelation filters is the inverse filter of the set of spatial decorrelation filters.

[0056] This provides improved performance and / or operation in many embodiments. In many embodiments and scenarios, this provides particularly attractive performance and / or implementation.

[0057] In some embodiments, a first set of spatial filters is configured to have coefficients that coincide with the coefficients of a spatial filter, which is a cascade of two sets of spatial inverse filters, each of which comprises an inverse filter of a set of spatial decorrelating filters.

[0058] This provides improved performance and / or operation in many embodiments. In many embodiments and scenarios, this provides particularly attractive performance and / or implementation.

[0059] According to an optional feature of the present invention, the decorrelation adapter is configured to determine the adapted spatial decorrelation as a set of spatial decorrelation filters, and each output audio signal of the set of spatial decorrelation filters is: a step of segmenting a first set of audio signals into time segments; and for at least some time segments: a step of generating a frequency bin representation of the first set of audio signals, wherein each frequency bin of the frequency bin representation of the first set of audio signals includes a frequency bin value for each of the audio signals in the first set of audio signals; and a step of generating a frequency bin representation of a set of output signals, wherein each frequency bin of the frequency bin representation of the set of output signals includes a frequency bin value for each of the output signals, for a given frequency bin of the set of output signals A frequency bin value for a given output signal is generated as a weighted combination of frequency bin values ​​for a first set of audio signals for a given frequency bin, wherein the weighted combination has a decorrelation coefficient as its weight, and the input audio signal is linked to one of the first set of audio signals by performing the steps of generating and updating a first weight for the contribution of the first frequency bin to the first frequency bin for the first output signal linked to the first input audio signal, from the second frequency bin value for the first frequency bin for the second input audio signal linked to the second output signal, according to a correlation measure between the first previous frequency bin value of the first output signal for the first frequency bin and the second previous frequency bin value of the second output signal for the first frequency bin.

[0060] This provides improved performance and / or operation in many embodiments. In many embodiments and scenarios, this provides particularly attractive performance and / or implementation.

[0061] This provides the advantageous generation of an output audio signal with typically increased decorrelation compared to the input signal. In many embodiments, this approach provides an efficient fitting of the operation that results in improved decorrelation. The fitting is typically performed with low complexity and / or resource usage. Specifically, this approach applies local fitting of individual weights to further achieve efficient fitting.

[0062] The generation of the output signal set is adapted in many embodiments and for many applications to provide enhanced audio processing, particularly beamforming, by providing increased decorrelation to the input signal.

[0063] The first and second output audio signals are typically different output audio signals.

[0064] According to an optional feature of the present invention, the decorrelated adapter is configured to update a first weight in accordance with the product of a first value and a second value, where the first value is one of a first previous frequency bin value and a second previous frequency bin value, and the second value is the complex conjugate of the other of the first previous frequency bin value and the second previous frequency bin value.

[0065] This provides improved performance and / or operation in many embodiments. Typically, it provides improved fit, leading to increased decorrelation of the output audio signal in many scenarios.

[0066] In some embodiments, the audio device is configured to update a second weight, which is the contribution of a third frequency bin value to the first frequency bin value, in accordance with the magnitude of a first previous frequency bin value.

[0067] This provides improved performance and / or operation in many embodiments. Typically, it provides improved fit, leading to increased decorrelation of the output audio signal in many scenarios. In particular, it provides improved fit of the generated output signal. In many embodiments, the weight update, which reflects the contribution of the linked input signal to the output signal, depends on the magnitude / amplitude of that linked input signal. For example, the update requires compensating the weights for the level of the input signal to produce a normalized output signal.

[0068] This approach allows for normalization / signal compensation / level compensation, for example, to provide the desired output level.

[0069] In some embodiments, the audio device is configured to set a predetermined weight for the contribution of a third frequency bin value, which is the frequency bin value of the first frequency bin for a first input audio signal, to the first frequency bin value.

[0070] This provides improved performance and / or operation in many embodiments. Typically, it provides improved fit, leading to increased decorrelation of the output audio signal in many scenarios. In many embodiments, it provides improved fit while ensuring convergence of the fit toward non-zero signal levels. It works very efficiently with weighted fit that is not for linked signal pairs.

[0071] In many embodiments, the adapter is configured to maintain a constant weight without weight adjustment or updating.

[0072] In some embodiments, the audio device is configured such that the weights for the contribution of a third frequency bin value, which is the frequency bin value of the first frequency bin for a first input audio signal, to the first frequency bin value are real values.

[0073] This provides improved performance and / or operation in many embodiments. Typically, it provides improved fit, leading to increased decorrelation of the output audio signal in many scenarios.

[0074] The weights between linked input / output signals are constrained to be determined / become, favorably, as real-valued weights. This leads to improved performance and fit, guaranteeing convergence to non-zero level solutions.

[0075] In some embodiments, the audio device is configured such that a second weight, which is the weight for the contribution of a first frequency bin to a fourth frequency bin value for a second output audio signal from a first input audio signal, is the complex conjugate of the first weight.

[0076] This provides improved performance and / or operation in many embodiments. Typically, it provides improved fit, leading to increased decorrelation of the output audio signal in many scenarios.

[0077] In many embodiments, the two weights for two pairs of input / output signals are their complex conjugates. In some embodiments, the weights of the weighted combination of input audio signals other than the first input audio signal are complex numerical weights.

[0078] This provides improved performance and / or operation in many embodiments. The use of complex values ​​for weights for unlinked input signals provides improved frequency-domain operation.

[0079] In some embodiments, the audio device is: y(ω)=W(ω)X(ω) It is configured to determine the output bin value for a given frequency bin ω from, where y(ω) is a vector containing the frequency bin values ​​for the output audio signal for a given frequency bin ω, x(ω) is a vector containing the frequency bin values ​​for the input audio signal for a given frequency bin ω, and W(ω) is a matrix having rows containing the weights of a weighted combination for the output audio signal.

[0080] This provides improved performance and / or operation in many embodiments. Typically, it provides improved fit, leading to increased decorrelation of the output audio signal in many scenarios.

[0081] The matrix W(ω) is favorably Hermitian. In many embodiments, the diagonals of the matrix W(ω) are constrained to be real-valued, set to predetermined values, and / or maintained as fixed values ​​but not updated / adapted. The weights / coefficients outside the diagonals are generally complex-valued.

[0082] In some embodiments, the audio device is: w ij (k+1,ω)=w ij (k,ω)-η(k,ω)[y i (k,ω)y j * (k,ω)] The weights of matrix W(ω) are determined accordingly. ij It is configured to fit, where i is the row index of matrix W(ω), j is the column index of matrix W(ω), k is the time segment index, ω represents the frequency bin, and η(k,ω) is the scaling parameter for fitting the fitting speed.

[0083] This provides improved performance and / or operation in many embodiments. Typically, it provides improved fit, leading to increased decorrelation of the output audio signal in many scenarios.

[0084] In some embodiments, the audio device is configured to compensate for correlation values ​​with respect to the signal levels of a first frequency bin.

[0085] This provides improved performance and / or operation in many embodiments. Typically, it provides improved fit, leading to increased decorrelation of the output audio signal in many scenarios. This allows for update rate compensation for signal fluctuations.

[0086] In some embodiments, the audio device is configured to initialize the weights such that the weighted combination includes at least one zero-value weight and one non-zero-value weight.

[0087] This provides improved performance and / or operation in many embodiments. This typically provides improved fit, leading to increased decorrelation of the output audio signal in many scenarios. This enables more efficient and / or faster fit and convergence toward favorable decorrelation. In many embodiments, the matrix W(ω) is initialized with zero values ​​for weights or coefficients for unlinked signals and fixed non-zero real values ​​for linked signals. Typically, the weights are set to, for example, 1 for the weights on the diagonals, and all other weights are initially set to zero.

[0088] In some embodiments, the weighted combination involves applying time-domain windowing to the frequency representation of the weights formed by the weights for a first input audio signal and a second input audio signal for different frequency bins.

[0089] This provides improved performance and / or operation in many embodiments. Typically, it provides improved fit, leading to increased decorrelation of the output audio signal in many scenarios.

[0090] Applying time-domain windowing to the frequency representation of weights involves: converting the frequency representation of weights to the time-domain representation of weights; applying a window to the time-domain representation to generate a modified time-domain representation; and converting the modified time-domain representation back to the frequency domain.

[0091] According to aspects of the present invention, a method of operation for an audio device is provided, the method comprising: receiving a first set of audio signals, the first set of audio signals including audio signals capturing sound of a scene from different locations; generating a beamform output audio signal from the first set of audio signals, wherein the step of generating the beamform output audio signal includes: generating a beamform output audio signal from signal processing of the first set of audio signals, the signal processing includes spatial beamforming and adaptive spatial decorrelation for the first set of audio signals, the adaptive spatial decorrelation depends on a set of decorrelation parameters, the spatial beamforming depends on a set of beamform parameters, and the beamform parameters are parameters indicating at least one of filtering and weighting of the input signal for a beamforming combination; fitting a set of decorrelation parameters depending on the first set of audio signals; fitting a set of beamform parameters depending on the beamform output audio signal; applying frequency equalization to the beamform output audio signal to generate an output signal; and fitting frequency equalization depending on the set of beamform parameters.

[0092] These and other aspects, features and advantages of the present invention will become clear and evident by referring to the embodiments described below.

[0093] Embodiments of the present invention are described with reference to the drawings, merely as examples. [Brief explanation of the drawing]

[0094] [Figure 1] Examples of elements of an audio device according to several embodiments of the present invention are shown. [Figure 2] Examples of elements of an audio beamformer for an audio device according to several embodiments of the present invention are shown. [Figure 3] Examples of elements of an audio beamformer for an audio device according to several embodiments of the present invention are shown. [Figure 4] Examples of elements of an audio beamformer for an audio device according to several embodiments of the present invention are shown. [Figure 5] Examples of elements of an audio beamformer for an audio device according to several embodiments of the present invention are shown. [Figure 6] Examples of elements of an audio beamformer for an audio device according to several embodiments of the present invention are shown. [Figure 7] Examples of elements of an audio beamformer for an audio device according to several embodiments of the present invention are shown. [Figure 8] Examples of elements of an audio beamformer for an audio device according to several embodiments of the present invention are shown. [Figure 9] Some elements of possible processor configurations for implementing elements of an audio device according to some embodiments of the present invention are shown. [Modes for carrying out the invention]

[0095] The following description focuses on embodiments of the present invention applicable to voice capture, such as speech capture for remote conferencing equipment. However, it is understood that this approach is applicable to many other voice signals, voice processing systems, and scenarios for capturing and / or processing voice.

[0096] Figure 1 shows an example of an audio device configured to generate an output audio signal from a first set of audio signals representing captured audio from scenes at different locations.

[0097] The audio device in Figure 1 specifically includes a receiver 101 that receives a first set of audio signals and, in a particular example, a set of microphone signals from a set of microphones that capture the audio scene from different locations. The microphones are configured, for example, relatively close to each other in a linear array. For example, the maximum distance between capture points for audio signals is not more than 1 meter, 50 cm, 25 cm, or in some cases even 10 cm, in many embodiments.

[0098] The input audio signal is received from different sources, including internal or external sources. The following describes an embodiment in which the receiver 101 is coupled to multiple microphones, such as a linear array of microphones, to provide a set of input audio signals in the form of microphone signals.

[0099] The receiver 101 is coupled to an audio beamformer 103 configured to generate a beamformed audio signal from a first set of audio signals. As will be described in more detail later, the audio beamformer 103 is configured to generate a beamformed audio signal by performing signal processing, including spatial beamforming and spatial decorrelation, of the first set of audio signals. Thus, both beamforming and decorrelation are performed by the audio beamformer 103 to generate a beamformed audio signal that is specifically instructed to extract or isolate desired sound sources in an audio scene. As will be described in more detail later, the audio beamformer 103 is an adaptive beamformer in which both the decorrelation function and the beamforming function are adapted, and the audio beamformer 103 is specifically configured to adapt its processing according to a first set of input audio signals and the generated beamformed audio signal.

[0100] The audio beamformer 103 therefore provides a process that achieves the effects of both spatial decorrelation and adaptive beamforming. Spatial decorrelation is controlled by a set of decorrelation parameters, and spatial beamforming is based on a set of beamform parameters. As will be described in more detail later, the decorrelation parameters are specifically decorrelation coefficients for a set of spatial filters configured to produce N output signals from N input signals, each output signal being produced as a weighted combination of samples from the N input signals. The decorrelation coefficients are specifically the weights for such weighted combinations.

[0101] Beamform parameters are specifically parameters that indicate or define a combination of signals to generate a beamformed output audio signal. Beamforming involves filtering multiple signals and combining the resulting filtered output signals. Beamform parameters specifically define the filter (specifically the filter coefficients) and / or the combination (e.g., the weights of each signal). It is understood that applying weights to a signal is considered equivalent to part of the filtering and / or part of the combination operation. For example, if a beamformed output audio signal is generated by weighting (e.g., complex weights) and then adding the input signals, the weights can be considered as filtering and / or the combination can be considered as including weighting (if beamforming does not involve any (other) filtering of the input signals). Specifically, the weights for the input signals are considered as filter coefficients.

[0102] The audio beamformer 103 adapts spatial decorrelation by adapting decorrelation parameters and adapts beamform behavior by adapting beamform parameters.

[0103] The audio beamformer 103 is coupled to an equalizer 105 configured to apply frequency equalization to the beamform output audio signal to generate the output signal of the audio device. Specifically, the equalizer applies filtering to the beamform output audio signal using a filter having a given frequency response. The given frequency response is fixed, constant, or undefined, but can be dynamically adapted during operation.

[0104] The audio device includes an equalization adapter 107 configured to adapt frequency equalization depending on the set of beamform parameters used by the audio beamformer 103. The equalization adapter 107 is coupled to the audio beamformer 103 and receives the beamform parameters currently applied to determine the frequency response of the equalization. Thus, the equalization adapter 107 is configured to adapt the frequency response of the equalizer 105 depending on the beamform parameters currently applied by the audio beamformer 103.

[0105] In many embodiments, the frequency response fitting of the equalizer 105 also depends on the decorrelation parameters determined for fitted decorrelation by the audio beamformer 103. Therefore, in many embodiments, the equalization adapter 107 also receives the decorrelation parameters and proceeds to evaluate them in order to determine / fit the current frequency response.

[0106] The inventors have found that the combination of spatial decorrelation and audio beamforming provides improved performance, particularly enabling improved extraction of the desired sound source in the presence of strong, and sometimes dominant, noise sources. The beamforming and decorrelation operations work synergistically to suppress unwanted sources by decorrelating them with respect to beamforming, while still allowing coherence to remain for effective beamforming towards the desired source.

[0107] However, the inventors also understood that while such approaches offer efficient performance and favorable operation in many scenarios, the combined effects also affect the overall frequency response, resulting in some frequency distortion. In particular, different frequency intervals have different correlations, and the combined effect of a decorrelator seeking to remove correlation and beamforming operation seeking to leverage correlation by combining it with the desired signal component results in a non-flat frequency response. The inventors further understood that the frequency shock is not constant or fixed, but is significantly dependent on the operation performed by the audio beamformer 103.

[0108] The audio device in Figure 1 reflects these understandings, and in particular, the equalizer 105 provides frequency compensation by applying a frequency response that compensates for the frequency distortion of the audio beamformer 103. The frequency response of the equalizer 105 is determined such that the combined frequency response of the audio beamformer 103 and the equalizer 105 is flatter than the frequency response of the audio beamformer 103 alone. The equalization adapter 107 adapts the equalizer 105 so that the resulting compensated frequency response is dynamically changed to reflect the current operation of the audio beamformer 103. Specifically, the compensated frequency response (the frequency response of the equalizer 105) is adapted based on the beamforming parameters, and therefore based on the parameters of the signal combination as part of the combination. Advantageously, in many embodiments, the compensated frequency response is also adapted based on the decorrelation parameters. Thus, in many embodiments, the compensated frequency response is adapted to reflect both the current operation of the decorrelation function and the beamforming function.

[0109] Therefore, the compensated frequency response is not determined by evaluating the signal components or the frequency content / distribution of the signal, or by frequency analysis or comparison of the signal components, but is controlled by parameters that determine the operation of the speech beamformer 103. These typically change much slower than the signal characteristics, making it possible to implement operation with very low complexity using significantly reduced computational resources. Nevertheless, this approach has been found to provide highly efficient and high-performance operation, from which an improved output signal can be generated. Specifically, the output signal sounds much more natural and provides speech or other sounds that are closer to actual speech in the speech scene. Frequency distortion can often be reduced, or even almost eliminated, from the output signal.

[0110] In many embodiments, the equalization adapter 107 is configured to adapt frequency equalization to have a low-pass frequency response. In many embodiments, the equalization adapter 107 is configured to generate a compensating frequency response that has a magnitude that decreases monotonically with increasing frequencies in frequency intervals of at least 50 Hz or 100 Hz to 500 Hz, 1 kHz, 2 kHz, or 4 kHz. The equalization adapter 107 generates a frequency response to provide a low-frequency boost to the beamform output audio signal.

[0111] The inventors understood that decorrelation has a frequency effect on the beamforming output because it results in a frequency-dependent decorrelation to the desired signal that beamforming is adapted to extract. For example, decorrelation of a noise source tends to affect lower frequencies of the desired sound source as well, because lower frequencies tend to have a higher correlation than higher frequencies due to longer wavelengths. Therefore, for different sound sources and signals, lower frequencies tend to have similar correlations across different microphones at different locations. Thus, decorrelation of one sound source (e.g., noise or dominant source) tends to decorrelate other sources as well. This is more pronounced at lower frequencies than at higher frequencies.

[0112] Such decorrelation tends to reduce beamforming efficiency and thus result in reduced amplitude because it seeks to leverage the correlation of desired signal components in different captured signals. Since the decorrelation of the desired signal is more pronounced at lower frequencies, the effect results in frequency-dependent attenuation of the audio captured by the beamform output audio signal. The equalization adapter 107 therefore controls the equalizer 105 to provide a low-frequency boost (low-pass characteristic) to compensate for this attenuation.

[0113] In many embodiments, the equalizer 105 is configured to determine a frequency equalization / compensation frequency response to have / become a minimum phase frequency response. This provides improved performance and perceived quality in many embodiments. This reduces, specifically minimizes, the delay of the equalizer 105, and therefore the delay of the capture operation and the audio device as a whole.

[0114] In many embodiments, the equalization adapter 107 is configured to first determine the magnitude-compensated frequency response based on beamform parameters (and optionally, decorrelation parameters), and then proceed to determine the minimum phase filter for this magnitude response. Such approaches will be described in more detail later.

[0115] It is understood that the specific operations performed by the audio beamformer 103, as well as specific decorrelation and beamforming combinations, will differ in different embodiments. The effect of the beamforming operation on the overall frequency response will therefore depend on the specific details of the implementation, and thus its dependence on the specific fitting and beamforming parameters of the compensated frequency response will also vary and be embodiment-specific.

[0116] The combination of adaptive beamformers and spatial decorrelators yields very good results, but it has been found to come at the cost of some (linear) frequency distortion. Due to decorrelation, the frequency response of the desired speech / utterance at the beamformer output is modified on a contextual basis without the decorrelation function. Typically, depending on the location of both the desired sound source / speaker and point noise source, and the size of the microphone array, decorrelation results in attenuation at low frequencies. At low frequencies, the wavelength of acoustic waves is large (e.g., over 3 meters at 100 Hz), and the input covariance matrices for both the point noise source and the desired sound source / utterance source do not differ significantly at low frequencies. The function of the decorrelator is to decorrelate the noise signal, and as a result the noise signal is attenuated, but as a side effect, the desired speech / utterance signal is also attenuated (partially) at low frequencies. The signal-to-noise ratio is typically increased considerably, but acceptable results can still be obtained for some applications such as speech recognition and wake word detection. However, for other applications, frequency distortion is highly undesirable. For example, in speech communication, while an improvement in the signal-to-noise ratio leads to better speech comprehension, the beamformer output tends to sound less natural due to the lack of low frequencies. Equalizers are configured to compensate for this by boosting low frequencies. Since the desired frequency characteristics of the equalizer cannot be known in advance because they depend on the beamform and decorrelation parameters, which vary with fitting, fitted equalizers are employed.

[0117] Figure 2 shows an example of the elements of the audio beamformer 103. The audio beamformer 103 includes a signal processor 201 configured to generate a beamformed audio signal by a process that includes both spatial beamforming and spatial decorrelation effects applied to a first set of audio signals.

[0118] In some embodiments, spatial decorrelation is non-adaptive spatial decorrelation (for example, parameters are set manually or once at initialization (based on microphone distance measurements or other measurements)). However, in many embodiments, spatial decorrelation is advantageously adapted spatial decorrelation, and the following description focuses on embodiments that employ adapted spatial decorrelation.

[0119] To perform adaptive spatial decorrelation, the audio beamformer 103 includes a decorrelation adapter 203 configured to adapt adaptive spatial decorrelation parameters based on the input signal to the audio beamformer 103, i.e., a first set of audio signals. The audio beamformer 103 further includes a beamform adapter 205 configured to adapt beamform parameters, specifically beamform / combination parameters that determine the beamform filter. The adaptation is based on the generated beamform output audio signal, and in many embodiments, the beamform parameters are adapted to maximize the signal level of the output signal. Many different techniques for performing adaptive spatial decorrelation and beamforming are known and can be used without prejudice to the present invention.

[0120] In many embodiments, the signal processor 201 is configured to first perform decorrelation on a first set of audio signals, i.e., on the input signals to the signal processor 201, to generate a first set of decorrelated audio signals, and then apply a beamform function to the first set of decorrelated audio signals to generate a beamform output audio signal. Such an approach is shown in Figure 3, where the signal processor 201 is shown to include a spatial decorrelator, which in a particular example is an adapted spatial decorrelator 301. The adapted spatial decorrelator 301 is configured to receive a first set of audio signals and generate a first set of decorrelated audio signals by performing adapted spatial decorrelation, specifically by filtering the first set of audio signals with a set of spatial decorrelation filters. The adapted spatial decorrelator 301 performs a decorrelation function based on decorrelation parameters, such as those determined by the decorrelation adapter 203, which specifically determines filter parameters, such as filter coefficients, for the spatial decorrelation filters.

[0121] The first set of uncorrelated audio signals is therefore the set of audio signals about which the correlation / coherence (of at least one source) is reduced with respect to the first set of audio signals.

[0122] The decorrelation adapter 203 specifically includes a set of spatial filters, which are to be fitted to generate an output signal (a first set of decorrelated audio signals) that specifically corresponds to the input signal for one (dominant) sound source, but with increased decorrelation of the signals. The output audio signal is generated to have increased spatial decorrelation, with lower inter-speech signal cross-correlation for the output audio signal than for the input audio signal. Specifically, the output signal is generated to have the same combined energy / power as the input signal (or a given scaling thereof), but with increased decorrelation (reduced correlation) between the signals. The output audio signal is generated to contain all the audio / signal components of the input signal in the output audio signal, but with redistribution to different signals to achieve increased decorrelation.

[0123] A decorrelation filter is specifically configured to produce an output signal with lower coherence / normalized correlation than the input signal. Therefore, the output signal of the decorrelation filter has lower coherence / normalized correlation than the coherence of the input signal to the decorrelation filter.

[0124] In some embodiments, the decorrelation adapter 203 is specifically configured to decorrelate / decoherenize noise / interference / unwanted sound sources from the captured signal. Specifically, it is configured to adapt decorrelation when noise / interference / unwanted sound sources are active. In many embodiments, it is configured not to adapt during the time when the desired sound sources are active. The beamforming adapter is configured to adapt when the desired sound sources are active, but not when noise / interference / unwanted sound sources are active. Thus, in many embodiments, the decorrelation adapter 203 is configured to adapt decorrelation for noise / interference / unwanted sound sources, and the beamforming adapter 205 is configured to adapt beamforming for the desired sound sources.

[0125] It is understood that in different embodiments, different algorithms and functions are used to decorrelate the input signal, and in particular, different approaches are used to determine appropriate filter coefficients for the decorrelation filter.

[0126] In some embodiments, the decorrelation adapter 203 includes, for example, a neural network configured to generate updated values ​​to modify the filter coefficients of the adaptive spatial decorrelation 301 based on input samples. The trained network is trained with training data that includes, for example, many different situations and weights, manually determined relevant updated values ​​to modify toward increased decorrelation.

[0127] As another example, the decorrelation adapter 203 includes a location processor configured to generate updated values ​​for correcting filter coefficients based on visual cues related to (changing) location.

[0128] As another example, adapter 203 determines the cross-correlation matrix of the input signal and calculates its eigenvalue decomposition. The eigenvectors and eigenvalues ​​can then be used to construct the decorrelation.

[0129] As another example, the decorrelation adapter 203 includes a microphone distance processor, which is configured to generate updated values ​​for correcting the filter coefficients in the case of diffuse noise only. For diffuse noise, the cross-correlation matrix of the input signals depends only on the distance between microphones and can then be calculated. In the specific case of diffuse noise only and the distance between microphones is fixed (static array), the calculation is required only during initialization. In such cases, therefore a static or non-adaptive spatial decorrelationizer 301 is employed.

[0130] A second set of output signals supplied to the beamforming circuit 303 is shown in Figure 3, and the beamforming circuit 303 is configured to perform spatial beamforming by combining a first set of uncorrelated audio signals, the combination of which depends on a set of beamforming parameters. Specifically, the combination involves filtering each of the first set of uncorrelated audio signals, and the filter outputs are then combined, for example, by (weighted) addition.

[0131] Many different beamformer algorithms are known, and it is understood that any suitable approach or algorithm may be used by the audio beamformer 103.

[0132] For example, the beamformer adapter 205 determines the cross-correlation matrix of its input signal, calculates its eigenvalue decomposition, and determines the eigenvector belonging to the largest eigenvalue (principal component analysis). The eigenvector can then be used to construct the beamform parameters of the beamform circuit 303.

[0133] The beamform circuit 303 is configured to receive a set of multiple audio signals from which a single signal is generated. The beamform circuit 303 seeks to combine the input signals so that a given sound source 0 contribution is constructively combined. Specifically, the beamform circuit 303 combines signals to perform beamforming for a set of spatial audio signals that capture sound in the environment, such as a set of signals from a microphone array. The beamform circuit 303 is configured to generate a beamformed output audio signal from a first set of audio signals.

[0134] Figure 4 shows an example of a possible beamform circuit 303 that is particularly suitable and advantageous for the audio device described.

[0135] The beamform circuit 303 in Figure 4 specifically consists of a first set of filters f1 that filter a first set of audio signals supplied to the beamform circuit 303 (specifically, a first set of uncorrelated audio signals or directly captured microphone signals, as will be described later). * (ω), f2 * (ω) is included. Each filter is configured to filter one of the audio signals. Furthermore, the filters are adaptive filters that are dynamically adapted to achieve adaptive beamforming. The first set of filters is also referred to as beamformer filters below.

[0136] The audio configuration further includes a feedback circuit 401 which contains a second set of filters that generate a second set of audio signals from the filtering of the beamform output audio signal. The second set of filters is also referred to as the feedback filter. The input signal to the beamform circuit 303 is also referred to as the beamformer input signal, and the signal generated by the feedback circuit is also referred to as the feedback signal. The feedback circuit 401 is used as part of the adaptation of beamform operation and can be considered as part of the beamform adapter 205. However, for clarity, it is shown separately from the beamform adapter 205 in Figure 4 (and subsequent figures).

[0137] The first and second sets of filters are specifically matched / linked such that for each filter in the first set (i.e., each beamformer filter), there exists a second set of filters (i.e., feedback filters) having a frequency response that is the complex conjugate of the corresponding beamformer filter. Thus, the filters of beamform circuit 303 and feedback circuit 401 are together fitted such that the filter coefficients provide a complex conjugate frequency response (equivalent to / corresponding to the time-inverted filter impulse response).

[0138] Therefore, the feedback circuit 401 generates a set of feedback signals from the beamform output audio signal, each feedback signal being generated from the beamform output audio signal by time-inverted / complex conjugate filtering with a matching filter.

[0139] The beamform adapter 205 is configured to fit a first set of filters and a second set of filters, namely beamformer filters and feedback filters. The beamform adapter 205 is configured to perform fitting based on a comparison between the first set and the second set of audio signals, specifically on a comparison between the beamformer input signal and the feedback signal. Specifically, a difference scale is determined for each matching / linked beamformer input signal and feedback signal (the beamformer filter and feedback filter having complex conjugate frequency responses), and the filters among the matching filters are fitted to reduce / minimize the difference scale.

[0140] The configuration of the beamform circuit 303, the feedback circuit 401, and the beamform adapter 205 interacts to provide highly efficient adaptive beamforming, which involves fitting filters to a set of spatial audio signals to form a beam directed towards a sound source, providing the captured audio in the beam formed by the beamformed audio signal. The filter fitting to minimize the difference scale allows the beamformer configuration to detect and track a given sound source. Specifically, in the ideal case, each of the beamformer filters concentrates all the energy from a given sound source captured in one of the beamformer input signals into a single signal value. In the ideal case, this is achieved by the beamformer filter having an impulse response that is the inverse of the time of the acoustic impulse response to the microphone capturing the signal from the sound source. In the ideal case, this is achieved for all signals / filters, resulting in all the captured energy from a particular sound source being combined into a single value by summing the outputs of the beamformer filters, thereby generating a sample value for the beamformed audio signal that maximizes the energy captured from the sound source.

[0141] Furthermore, since each of the feedback filters is the complex conjugate frequency response of the corresponding beamformer filter, this is identical to the acoustic impulse response (signal) from the sound source to the microphone in an ideal scenario. Therefore, a feedback signal is generated by filtering the beamformer output audio signal with the feedback filter, which reflects the signal captured by the microphone from the sound source as represented by the beamformer output audio signal. In an ideal scenario where only one sound source exists and the filter is ideally adapted, the difference between the generated feedback signal and the captured microphone signal is zero. In the presence of other sound sources and / or noise, there is a difference reflected in the difference measure. However, for uncorrelated noise / audio, the difference, and thus typically the difference measure, typically averages to zero and thus tends not to prevent efficient adaptation. Therefore, adaptation towards an optimal beamforming filter is typically achievable by adapting the filter for a given beamformer input signal to minimize the difference measure for the corresponding feedback signal.

[0142] Such a beamforming approach is described in U.S. Patent No. 7,146,012. In the following, a more detailed description and analysis of the configuration are provided. This description will be based on the example of FIG. 4, in which the configuration of the adaptive beamformer, where the microphone signal is filtered by two filters f1 * (ω) and f2 * (ω), consists of a filter section, and the output is added to obtain the output signal z(ω). In the update section, the output signal z(ω) is supplied to two adaptive filters f1 * (ω) and f2 * (ω) each having the respective microphone signal as a reference. <​​​​​2 =1. Specifically, this constraint is implied by the configuration because the conjugates of filters f1(ω) and f2(ω) in the update portion are replicated to the filter portion. The adaptive filter in the update portion is a "normal" unconstrained adaptive filter.

[0144] The constraints implied by the construction can be understood by looking at the optimal solution, i.e., the case where both residual signals are zero. Assume we have an utterance signal s(ω) and transfer functions h1(ω) and h2(ω) from the utterance source to the microphone. The microphone signal is then given by: u1(ω)=h1(ω)s(ω) and u2(ω)=h2(ω)s(ω).

[0145] The output signal z(ω) is then as follows: z(ω)=u1(ω)f1 * (ω)+u2(ω)f2 * (ω)= =s(ω)(h1(ω)f1 * (ω)+h2(ω)f2 * (ω)). In the update section, two equations are obtained after convergence, where ω is omitted for convenience: (h1f1 * +h2f2 * )f1=h1(1) (h1f1 * +h2f2 * )f2=h2(2) (3) follows (1) and (2):

number

number

number

number

number

[0146] More generally, f(ω) = [f1(ω).......f nmics (ω)] T Using matrix notation, Here, f m * (ω) is the beamformer filter corresponding to the m-th microphone, h(ω) = [h1(ω).......h nmics (ω)] T h m (ω) is the transfer function from source s(ω) to the m-th microphone, and for output z(ω): z(ω)=s(ω)h T (ω)f * (ω) (4) This can be obtained.

[0147] In the update section, the equation is given after convergence: z(ω)f(ω)=s(ω)h(ω) (5) It has, This is because (z(ω) and s(ω) are scalars) f(ω)=α(ω)h(ω) (6) Leading to, Here, α(ω) is a complex scaler.

[0148] By using (6) and substituting (4) into (5):

number

[0149] By substituting (7) into (6), the final solution after convergence (8) is given:

number

[0150] Regarding constraints, next: f H (ω)f(ω)=1 This can be obtained.

[0151] To understand that the solution given by Eq. 8 is also the optimal solution for maximizing output power, we look at the expected value of the update after convergence, which consists of the correlation between the residual signals (x1(ω) and (x2(ω)) in Figure 4) and the conjugate of the input signal of the fitted filter:

number

[0152] this is:

number

[0153] z(ω)=u T (ω)f * (ω), therefore z * (ω)=u H By using f(ω):

number

number

number

number

[0154] If the microphone signal specifically contains only desired signals such as speech, then u(ω)=s(ω)h(ω), and then f H (ω)R uu Maximizing (ω)f(ω) corresponds to maximizing the desired signal / utterance in the output.

[0155] Considering the case where u(ω)=s(ω)h(ω)+v(ω), where v(ω) is uncorrelated noise with equal dispersion across all microphones, that is,

number

number

number

[0156] Now it's v(ω)'s turn: R uu (ω)=R ss (ω)+R nn (ω) Assuming that the correlation noise consists of such noise, R nn (ω) is the noise covariance matrix with non-zero off-diagonal elements.

[0157] f H (ω)R uu The maximization of (ω)f(ω) is as follows: f H (ω)R uu (ω)f(ω)=f H (ω)R ss (ω)f(ω)+f H (ω)R nn (ω)f(ω) To maximize [something].

[0158] This is f H (ω)R uu This means that maximizing (ω)f(ω) does not necessarily lead to a better signal-to-noise ratio, because the choice of f determines not only the amount of speech in the output, but also the amount of noise in the output.

[0159] As described, the improved performance is achieved by including an adapted spatial decorrelation function, specifically by including an adapted spatial decorrelation 301 at the input to decorrelate a first set of audio signals. Figure 5 shows such an example in which the beamforming approach described with reference to Figure 4 is further enhanced by the introduction of an adapted spatial filter on the input signal.

[0160] Spatial filtering is performed across the signal and is based on decorrelation coefficients / weights determined based on the input audio signal to fit the current audio characteristics. As described in more detail below, spatial filtering is placed at different locations in the beamforming loop structure, including as inverse spatial decorrelation filtering in the feedback loop or as direct spatial decorrelation filtering on the input signal used by beamforming. This approach has been found to provide significantly improved performance in many real-world scenarios, such as greatly improved source separation. In relation to the above analysis, this approach is particularly useful for noise contribution f H (ω)R nn We achieve that (ω)f(ω) occurs independently of the choice of f.

[0161] Figure 5 shows an audio device that provides improved performance. In the example, the audio device has a beamforming structure described with reference to Figure 4. However, this approach is further enhanced by the fact that this structure operates on signals that have been spatially decorrelated by the adaptive spatial decorrelator 205.

[0162] The audio beamformer 103 in Figure 2 receives a first set of audio signals and, in certain examples, a set of microphone signals from a set of microphones capturing the audio scene from different locations. The microphones are configured, for example, relatively close to each other in a linear array. For example, the maximum distance between capture points for audio signals may not exceed 1 meter, 50 cm, 25 cm, or in some cases even 10 cm, in many embodiments.

[0163] The input audio signal is received from different sources, including internal or external sources. The following describes an embodiment in which the audio beamformer 103 is coupled to multiple microphones, such as linear array microphones, to provide a set of input audio signals in the form of microphone signals.

[0164] However, instead of directly performing the operation shown in Figure 4 on the microphone signal, the audio device in Figure 5 is configured to apply adaptive spatial decorrelation to the audio signal before beamforming and adaptation operations. Spatial decorrelation is adaptive decorrelation that is adapted based on the input audio signal, and therefore the decorrelation is continuously adapted to provide increased decorrelation.

[0165] Specifically, the audio device in Figure 5 includes an adapted spatial decorrelator 301 in the form of a set of spatial (decorrelating) filters that apply a spatial filter to a first set of audio signals. Thus, after filtering by the spatial filter, the first set of audio signals is modified to have increased decorrelation / decoherence (for at least one sound source) compared to before spatial decorrelation filtering.

[0166] The spatial filter is based on coefficients fitted by the decorrelation adapter 203, which in this example is a decorrelation adapter 203 that dynamically fits and updates the filter coefficients used by the set of filters of the adapted spatial decorrelationizer 301. The spatial filter is therefore set to have coefficients determined by the decorrelation / coefficient adapter 203.

[0167] Beamforming, decorrelation, and fitting are typically performed in the frequency domain. Receiver 101 includes a segmenter configured to segment a set of input audio signals into time segments. In many embodiments, segmentation is typically fixed segmentation into time segments of fixed and equal duration, such as division into time segments / intervals having fixed durations of, for example, 10–20 msec. In some embodiments, segmentation is adaptive so that the segments have variable durations. For example, the input audio signal has a variable sampling rate, and the segments are determined to have a fixed number of samples. Segmentation is typically performed into segments having a given fixed number of time-domain samples of the input signal. For example, in many embodiments, the segmenter is configured to divide the input signal into consecutive segments of, for example, 256 or 512 samples.

[0168] The receiver 101 is configured to generate a frequency bin representation of an input audio signal, and a first set of input signals to be further processed is typically represented in the frequency domain by the frequency bin representation. The audio device is configured to perform frequency domain processing of the frequency domain representation of the input audio signal. The signal representation and processing are based on frequency bins, and therefore the signal is represented by the values ​​of the frequency bins, and these values ​​are processed to generate frequency bin values ​​for the output signal. In many embodiments, the frequency bins have the same size and therefore cover frequency intervals of the same size. However, in other embodiments, the frequency bins have different bandwidths, and for example, perceptually weighted bin frequency intervals are used.

[0169] In some embodiments, the input audio signal is already provided in frequency representation and no further processing or operation is required. However, in some cases, reconstruction into a suitable segment representation is desired, which may include, for example, aligning the frequency representation into time segments using interpolation between frequency values.

[0170] In other embodiments, a filter bank, such as an orthogonal mirror filter QMF, is applied to the time-domain input signal to generate a frequency bin representation. However, in many embodiments, a discrete Fourier transform (DFT), specifically a fast Fourier transform (FFT), is applied to generate a frequency representation.

[0171] In the audio device of FIG. 5, the spatial filter of the adaptive spatial decorrelator 301 specifically processes the audio signal in the frequency domain. In the following description, the first set of audio signals is also referred to as the input audio signal (to the spatial filter) before filtering, and the resulting signal is also referred to as the output audio signal (from the spatial filter).

[0172] For each frequency bin, the output frequency bin value is generated from one or more input frequency bin values of one or more input signals, as described in more detail below. The output signal is generated, for example, at least for one sound source that is the dominant sound source, to reduce (typically / average) the correlation between signals with respect to the correlation of the input signals.

[0173] The set of spatial filters is configured to filter the input audio signal. The filtering is spatial filtering in that, for a given output signal, the output value is determined from a plurality, typically all, of the input audio signals (for the same time / segment and for the same frequency bin). The spatial filtering is specifically performed on a frequency bin basis such that the frequency bin value for a given frequency bin of the output signal is generated from the frequency bin values of the input signals for that frequency bin. The filtering / weighted combination is over the signals and is not typical time / frequency filtering.

[0174] Specifically, the frequency bin value for a given frequency bin is determined as a weighted combination of the frequency bin values of the input signals for that frequency bin. The combination is specifically addition, and the frequency bin value is determined as the weighted addition of the frequency bin values of the input signals for that frequency bin. The determination of the bin value for a given frequency bin is determined as the vector multiplication of a vector of weights / coefficients of the weighted addition and a vector containing the bin values of the input signals:

Number

[0175] Representing the output bin value for a given frequency bin ω as a vector y(ω), the determination of the output signal is: y(ω)=W(ω)x(ω) is determined, where the matrix W(ω) represents the weights / coefficients of the weighted addition for different output signals, and x(ω) is a vector containing the input signal values.

[0176] For example, for an example having only three input signals and output signals, the output bin value for the frequency bin ω is:

Number

[0177] The decorrelation adapter 203 seeks to adapt the spatial filter to become a spatial decorrelation filter that seeks to generate an output signal that corresponds to the input signal, but with increased signal decorrelation. The output audio signal is generated to have increased spatial decorrelation, with lower inter-speech signal cross-correlation for the output audio signal than for the input audio signal. Specifically, the output signal is generated to have the same combined energy / power as the input signal (or with a given scaling thereof), but with increased decorrelation (reduced correlation) between signals. The output audio signal is generated to contain all the audio / signal components of the input signal in the output audio signal, but with redistribution into different signals to achieve increased decorrelation.

[0178] A decorrelation filter is specifically configured to produce an output signal with lower coherence / normalized correlation than the input signal. Therefore, the output signal of the decorrelation filter has lower coherence / normalized correlation than the coherence of the input signal to the decorrelation filter.

[0179] The decorrelation adapter 203 is configured to determine update values ​​for the weighted combinations that form a set of decorrelation filters. Specifically, the update values ​​are determined for the matrix W(ω). The decorrelation adapter 203 then updates the weighted combinations based on the update values.

[0180] The decorrelation adapter 203 is configured to apply a fitting approach to determine the update value, which allows the output signal of the set of decorrelation filters 301 to represent the audio of the input signal to the set of decorrelation filters 301, but with the output signal being typically more decorrelated than the input signal.

[0181] The decorrelated adapter 203 is configured to use a specific approach to fitting weights based on the generated output signals. The operation is based on each output audio signal linked to a single input audio signal. Precise linking between the output and input signals is not required, and linking / pairing (including random in principle) of many different input signals for each output signal is used. However, the processing differs with respect to weights that reflect contributions from the output signals and linked / paired input signals, compared to weights that reflect contributions from the output signals and unlinked / paired input signals. For example, in some embodiments, the weights for linked signals (i.e., for input signals linked to an output signal generated by a weighted combination including the weights) are set to fixed values, are not updated, and / or the weights for linked signals are limited to real-valued weights, while other weights are generally complex-valued.

[0182] The decorrelated adapter 203 employs a fitting / update approach, where the updated values ​​are determined based on a correlation measure between the output bin values ​​for a given output signal and the output bin values ​​for a given (unlinked) input signal and a linked output signal, for a given weight representing the contribution of a given unlinked input signal to the bin values ​​for a given output signal. The updated values ​​are then applied to modify the given weights in subsequent segments, or the updated values ​​for the weights in a given segment are determined according to two output bin values ​​from the previous (typically immediately preceding) segment, where the two output values ​​represent the input and output signals to which the weights relate, respectively.

[0183] The described approach applies to multiple, typically all, weights used when determining output bin values ​​based on input signals that are not typically linked. Regarding weights related to input signals linked to output signals, other considerations are used, such as setting weights to fixed values, as will be explained in more detail later.

[0184] Specifically, the update value is determined by the product of the output bin value for the weight and the complex conjugate of the output bin value linked to the input signal for the weight, or equivalently, by the product of the complex conjugate of the output bin value for the weight and the output bin value linked to the input signal for the weight.

[0185] As a specific example, the updated value for segment k+1 for frequency bin ω is: [y i (k,ω)y j * (k,ω)] By, or equivalently: [yi * (k,ω)y j (k,ω)] It is determined depending on the correlation scale given by, where y i (k,ω) is the output bin value for the output signal i, which is determined based on the weights, and y j (k,ω) is the output bin value for output signal j linked to the input signal from which the contribution is determined (i.e., the input signal bin value multiplied by weights to determine the contribution to the output bin value for signal i).

[0186] [y i (k,ω)y j * The measure (k,ω) (or conjugate value) indicates the correlation of time-domain signals in a given segment. In a particular example, this value is then weighted w i,j (k+1,ω) is updated and used to fit the changes.

[0187] As mentioned above, a set of uncorrelated filters gives the output bin values ​​for the output signal for a given frequency bin ω: y(ω)=W(ω)X(ω) configured to be determined from, where y(ω) is a vector containing the frequency bin values for the output signal for a given frequency bin ω, x(ω) is a vector containing the frequency bin values for the input audio signal for a given frequency bin ω, and W(ω) is a matrix having a row containing the weights of the weighted combination for the output audio signal.

[0188] In an example, decorrelating adapter 203 specifically adapts at least some of the weights w ij of: w ij (k + 1, ω) = w ij (k, ω) - η(k, ω)[y i (k, ω)y j * (k, ω)] in accordance with, where i is the row index of the matrix W(ω), j is the column index of the matrix W(ω), k is the time segment index, ω represents the frequency bin, and η(k, ω) is a scaling parameter for adapting the adaptation speed. Typically, decorrelating adapter 203 is configured to adapt all weights that are not associated with a linked output signal (i.e., "cross-signal" weights) of the input signal.

[0189] In some embodiments, decorrelating adapter 203 is configured to adapt the update rate / speed of the weight adaptation. For example, in some embodiments, the adapter is configured to compensate a correlation measure for a given weight depending on the signal level of the output bin value whose contribution is determined by the weight.

[0190] As a specific example, the compensation value [y i (k, ω)y j * (k, ω)] is compensated by the signal level |y i (k, ω)| of the output bin value. The compensation is included, for example, to normalize such that the value of the update step depends less on the signal level of the decorrelated signal from which it was generated.

[0191] In many embodiments, such compensation or normalization is specifically performed on a frequency bin basis, i.e., compensation differs at different frequency bins. This improves performance in many scenarios and results in an improved fit of weights, typically producing a decorrelated signal.

[0192] The compensation is incorporated, for example, into the scaling parameter η(k,ω) of the previous update formula. Thus, in many embodiments, the decorrelation adapter 203 is configured to adapt / change the scaling parameter η(k,ω) differently at different frequency bins.

[0193] In many embodiments, the input signal vector x(ω) and output signal vector y(ω) are configured such that linked signals occupy the same position in their respective vectors; specifically, y1 is linked to x1, y2 to x2, y3 to x3, and so on. In this case, the weights for the linked signals lie on the diagonal of the weight matrix W(k,ω). In many embodiments, the diagonal values ​​are set to fixed real values, for example, a constant value of 1.

[0194] In many embodiments, the weights / space filters / weighted combinations are configured such that the weights for the contribution of the first input signal (not linked to the first output signal) to the first output signal are the complex conjugate of the contribution of the second input signal (linked to the first input signal) to the second output signal (linked to the first input signal). Thus, the two weights for the two pairs of linked input / output signals are complex conjugates.

[0195] In the example of weights for linked input and output signals arranged on the diagonal of the weight matrix W(ω), this results in a Hermitian matrix. In fact, in many embodiments, the weight matrix W(ω) is a Hermitian matrix. Specifically, the coefficients / weights of the weight matrix W(ω) are based on: w ij =wji * It satisfies the condition.

[0196] As mentioned above, the weights for the contribution of linked input signals to the output signal bin values ​​(corresponding to the diagonal values ​​of the weight matrix W(ω) in certain examples) are treated differently from the weights for unlinked input signals. The weights for linked input signals will also be referred to as linked weights for brevity below, and the weights for unlinked input signals will also be referred to as unlinked weights for brevity below. Therefore, in certain examples, the weight matrix W(ω) is a Hermitian matrix containing linked weights on the diagonal and unlinked weights outside the diagonal.

[0197] In many approaches, fitting unlinked weights is done to reduce the correlation metric. Specifically, each updated value is determined to reduce the correlation metric. As a whole, the fit therefore seeks to reduce the cross-correlation between output signals. However, linked weights are determined differently to ensure that the output signals maintain appropriate speech energy / power / levels. In fact, if linked weights are instead fitted to reduce the autocorrelation of the output signals with respect to the weights, there is a high risk that the fit will converge to a solution where all weights, and therefore the output signals, are essentially zero (which indeed results in the lowest correlation). Furthermore, speech devices are configured to produce signals with low cross-correlation, but not to reduce autocorrelation.

[0198] Therefore, in many embodiments, the linked weights are set to ensure that the output signal is generated to have the desired (combined) energy / power / level.

[0199] In some cases, the decorrelated adapter 203 is configured to fit the linked weights, while in other cases, the adapter is configured not to fit the linked weights.

[0200] For example, in some embodiments, the linked weights are simply set to fixed, constant values ​​that are not fitted. For example, in many embodiments, the linked weights are specifically set to a constant scalar value such as value 1 (i.e., a unit gain is applied to the linked input signal). For example, the weights on the diagonal of the weight matrix W(ω) are set to 1.

[0201] Therefore, in many embodiments, the weights for the contribution of linked input signal frequency bin values ​​to a given output signal frequency bin value are set to a predetermined value. In many embodiments, this value is kept constant without any modifications.

[0202] Such an approach offers highly efficient performance and provides a very accurate representation of the original sound of the input signal, but it has been found that it results in an overall fit that produces an output signal with a set of output signals that have increased decorrelation.

[0203] In some embodiments, linked weights are also fitted, but in a different way than unlinked weights. In particular, in many embodiments, linked weights are fitted based on the output signal.

[0204] Specifically, in many embodiments, the linked weights for the first input signal and the linked output signal are adapted based on the generated output bin value of the linked audio signal, specifically based on the magnitude of the output bin value.

[0205] Such an approach, for example, allows for the normalization and / or setting of the desired energy level for a signal.

[0206] In many embodiments, linked weights are constrained to be real-valued weights, while unlinked weights are generally complex-valued. In particular, in many embodiments, the weight matrix W(ω) is a Hermitian matrix with real values ​​on the diagonal and complex values ​​outside the diagonal.

[0207] Such an approach offers specific advantageous behavior and fit in many scenarios and embodiments. It has been found to provide relatively low complexity and computational resources while offering highly efficient spatial decorrelation.

[0208] The fitting gradually adjusts the weights to increase the decorrelation between signals. In many embodiments, the fitting is configured to converge toward a suitable weight matrix W(ω) regardless of the initial values, and in fact in some cases the fitting is initialized with random values ​​for the weights.

[0209] However, in many embodiments, the fitting is started with favorable initial values ​​that result in a fit that, for example, produces a faster fit and / or is more likely to converge toward more optimal weights for the uncorrelated signals.

[0210] In particular, in many embodiments, the weight matrix W(ω) is composed of several weights that are zero, but at least some weights are non-zero. In many embodiments, the number of nearly zero weights is two, three, five, or more times the number of weights that are set to non-zero values. This has been found to tend to provide improved fit in many scenarios.

[0211] In many embodiments, in particular, the decorrelation adapter 203 is configured to initialize the weights, with linked weights typically set to non-zero values, such as predetermined non-zero real numbers, while unlinked weights are set to approximately zero. Thus, in the above example where linked signals are located at the same position in the vector, this results in an initial weight matrix W(ω) having non-zero values ​​on the diagonal and (approximately) zero values ​​outside the diagonal.

[0212] Such initialization offers particularly advantageous performance in many embodiments and scenarios. This reflects the tendency for input signals to be somewhat uncorrelated, due to the fact that audio signals typically represent speech at different locations. Therefore, starting with the assumption that input signals are perfectly correlated is often advantageous and leads to faster and more frequent improved fitting.

[0213] Weights, specifically unlinked weights, are not necessarily exactly zero, but are set to low values ​​close to zero in some embodiments. However, the initial non-zero values ​​are at least 5, 10, 20, or 100 times higher than the initial near-zero values.

[0214] The described approach generates an output signal that represents the same speech as the input signal, but provides a highly efficient adapted spatial decorrelator with increased decorrelation. This approach has been found to provide highly efficient adaptation for a wide range of scenarios and many different acoustic environments, as well as for many different sound sources. For example, it has been found to provide highly efficient decorrelation of speaker signals in environments with multiple speakers.

[0215] The fitting approach is even more computationally efficient, allowing for local and individual fitting of individual weights based on only two signals closely related to the weights (specifically, only two frequency bin values), but this process still yields an efficient and often greatly optimized global optimization of the spatial filtering, specifically the weight matrix W(ω). Local fitting has been found to lead to a very favorable global optimization in many embodiments.

[0216] A specific advantage of this approach is that it is used to decorrelate convolutive mixtures, and is not limited to decorrelating only instantaneous mixtures. For convolutive mixtures, the full impulse response determines how signals from different sources combine in the microphone (i.e., delay / timing characteristics are prominent), whereas for instantaneous mixtures, the scalar representation is sufficient to determine how the sources combine in the microphone (i.e., delay / timing characteristics are not prominent). By converting convolutive mixtures to the frequency domain, the mixture can be considered as a complex-valued instantaneous mixture per frequency bin.

[0217] The decorrelation adapter 203 then determines coefficients for a set of spatial decorrelation filters of the adapted spatial decorrelation device 301, thereby modifying a first set of speech signals so that they represent the same speech, but with increased decorrelation / decoherence for at least one of the sound sources. Such decorrelation of speech signals, to which the beamforming approach described above is then applied, significantly improves overall performance in many scenarios. In fact, this results in improved separation and selection of specific sound sources, such as a particular speaker, in many scenarios. Thus, counterintuitively, decorrelation of signals provides improved beamforming, even though beamforming is essentially based on leveraging correlation between speech signals from different locations to extract / separate sound sources by spatially forming a beam toward the desired source. In fact, decorrelation essentially breaks the link between the speech signal and the specific location in the speech scene that is typically utilized by beamforming operation. However, the inventors have found that despite this, decorrelation provides very favorable effects and improved performance in many scenarios. For example, in the presence of a strong noise source, this approach facilitates and / or improves the extraction / isolation of a specific desired sound source, such as a speaker.

[0218] The specific fit described above offers a highly advantageous approach in many embodiments. This typically results in a spatial decorrelation filter that provides a very accurate fit with low complexity, producing a highly decorrelated signal. In particular, this allows local fitting of individual weights / filter coefficients to produce a highly efficient global decorrelation of a first set of speech signals.

[0219] However, in other embodiments, it is understood that other approaches are used to fit the spatial filter / decorrelation filter.

[0220] In what follows, the output z(ω) of the acoustic beamformer 103 of FIG. 5 is derived. In this example, also referred to as Configuration A below, the set of spatial filters is placed immediately after the microphones, fully outside the beamformer. This has the advantage that nothing needs to be changed to the beamformer algorithm, which still only has a different input constraint f H (ω)f(ω)=1.

[0221] A suitable decorrelator is as described above. This decorrelator transforms the noise covariance matrix R nn (ω) to: W(ω)R nn (ω)W H (ω)=Λ(ω) (9) where Λ(ω)=diag(λ1(ω).....λ Nmics (ω)) is a diagonal matrix and W(ω) is the nmics x nmics decorrelation matrix of the adaptive spatial decorrelator 301. f H (ω)R uu (ω)f(ω), instead of maximizing, the beamformer now maximizes f H (ω)W(ω)R uu (ω)W H (ω)f(ω), which is f H (ω)W(ω)R ss (ω)W H (ω)f(ω)+f H (ω)W(ω)R nn (ω)W H (ω)f(ω) =f H (ω)W(ω)R ss (ω)W H f(ω)+f H (ω)Λ(ω)f(ω) and can be expressed as.

[0222] If Λ(ω) is a scaled version of the identity matrix and can be expressed as β(ω)I, then f H (ω)Λ(ω)f(ω) is converted to β(ω), the noise contribution is independent of the choice of f(ω), and fH (ω)W(ω)R uu (ω)W H Maximizing (ω)f(ω) leads to maximizing the signal-to-noise ratio in the output.

[0223] The selection of β(ω) in the adaptive spatial decorrelator 301 does not affect the SNR, but it does affect the speech level at the output. If we select β(ω)=1 and assume that the noise level at the input doubles, all coefficients in W(ω) are scaled by 0.5, and therefore the speech level at the output of the adaptive spatial decorrelator 301 is also reduced. The adaptive spatial decorrelator 301 continues to update (per frequency bin) according to the input level. As a result, the beamformer also continues to update.

number

[0224] The optimal coefficient for the beamformer can be found by using equation (8) instead of h(ω) in the output of the spatial decorrelator: W(ω)h(ω).

[0225] W(ω) Hermit, therefore W H Using a decorrelator with (ω)=W(ω), for the optimal solution:

number

[0226] The output of the combination is:

number

number

[0227] Substituting (12) and (11) into (10), the output is:

number

[0228] Using homospersity and uncorrelated noise in the input, diag(W(ω))=I and therefore W 2 Note that using a decorrelator with (ω)=W(ω)=I yields the same solution as when no decorrelator is used or when only a beamformer is used.

[0229] The determination of the beamform output audio signal z(ω) can be used to determine the optimal compensation frequency response of the equalizer 105 for a given desired frequency response. in particular:

number

number

[0230] It is desirable that the speech has the same frequency characteristics, as in the case where correlated noise is present and when no noise or equivalent uncorrelated noise is present. this is:

number

number

[0231] in particular,

number

number

[0232] This demonstrates that this approach produces the desired compensated frequency response as given above.

[0233] The phase characteristics are determined to provide appropriate performance, and in many embodiments, this is determined to be the minimum phase characteristic in order to minimize delay. Different algorithms and approaches for producing the minimum phase characteristic from a given amplitude spectrum are known and are not further described herein for the sake of brevity. (See, for example, AVOppenheim and RWShafer, "Digital Signal Processing," Englewood Cliffs (New York, USA): Prentice-Hall, 1975, or AVOppenheim and RWShafer, "Discrete-time Signal Processing," Englewood Cliffs (New York, USA): Prentice-Hall, 1989).

[0234] In the example described above, the spatial decorrelation function and beamforming function are achieved by the adaptive spatial decorrelation 301, which applies decorrelation to a first set of audio signals and the resulting output signal from the adaptive spatial decorrelation 301, and then to the beamforming circuit 303 that operates on it. However, the same overall function can be achieved in other ways, including performing decorrelation or inverse decorrelation in the beamformer's feedback or in different parts of the adaptive circuit.

[0235] In the example in Figure 5, the decorrelation and spatial decorrelation filters are applied directly to a first set of audio signals before they are supplied to both the beamform circuit 303 and the beamform adapter 205. However, other approaches are used to adapt the operation based on applying spatial / cross-signal filtering using coefficients derived from the determined decorrelation coefficients.

[0236] Specifically, Figure 6 shows an example also referred to as configuration B, in which the audio beamformer 103 is configured to perform spatial filtering of a first set of audio signals before the first set of audio signals is filtered by a first set of filters; that is, beamforming is based on a first set of audio signals after the first set of audio signals has been filtered by a set of spatial filters 601. However, in this example, the beamform adapter 205 receives the first set of audio signals before any filtering by the set of filters 601; that is, the second set of filters is applied only to the beamformer path and not to the fitted path. The coefficients for this set of spatial filters 601 are determined from the decorrelation coefficients determined by the decorrelation adapter 203. In fact, in some embodiments, the approach described with reference to the configuration in Figure 5 is applied directly, and the resulting second set of filters is applied (only) to the signals in the beamformed path.

[0237] However, in many embodiments, the set of spatial filters 301 in the configuration example of Figure 5 differs from the set of spatial filters 601 in the configuration example of Figure 6. In particular, in many embodiments, the set of spatial filters is modified to correspond to a cascade of two sets of uncorrelated filters, as determined by the uncorrelated adapter 203. Thus, the filter coefficients of the set of spatial filters 601 have coefficients that match those of the spatial filters which are a cascade of two sets of uncorrelated filters. This can typically be considered equivalent to a double / repeated filtering of a first set of audio signals by the set of uncorrelated filters determined by the uncorrelated adapter 203.

[0238] In particular, the decorrelation adapter 203, as previously described: y(ω)=W(ω)X(ω) The weights / coefficients of the weight matrix W(ω) are determined for a decorrelation filter that can decorrelate the first set of audio signals according to the following.

[0239] In this case, the set of spatial filters 601 is generated to correspond to a cascaded application of two such filters, namely: y(ω)=W(ω)W(ω) Quant(ω)=W 2 (ω)x(ω) It is configured to correspond to this.

[0240] Therefore, in this example, the decorrelation adapter 203 proceeds to determine appropriate coefficients for a set of spatial decorrelation filters that decorrelate a first set of audio signals by performing a fitting as described for the example in Figure 5. In some cases, such filters are then applied by a set of spatial filters 601 in Figure 6. However, in many embodiments, the determined decorrelation filters are not used directly, but rather the coefficients for the set of spatial filters 601 are determined from the determined coefficients for the decorrelation filters. In such cases, the decorrelation adapter 203 performs the decorrelation filter as part of the coefficient determination / fitting and this (e.g., [y i (k,ω)y j * A first set of signals (from (k,ω)] is applied to determine the update value. However, in the example of a particular case, such filtered signals are used only for the fitting process and not further in beamforming / processing. Instead, the set of spatial filters 601 applied is derived from the decorrelation coefficient / weight matrix W(ω), specifically W 2 It is generated as (ω).

[0241] This approach provides improved overall performance, including an enhanced beamforming experience. In fact, this is the set of spatial filters 601 shown in Figure 6. 2 For the example where the coefficients are set to correspond to , it can then be shown that this produces the same performance, results, and output signals as the example in Figure 5. It can also be shown that these approaches produce the same optimal solution.

[0242] Another possible example of applying spatial filtering by a set of spatial filters determined from the decorrelation coefficients determined by the decorrelation adapter 203 is shown in Figure 7, also referred to as configuration C. In this example, a first set of audio signals is not filtered by the set of spatial filters, while a second set of audio signals generated by the feedback circuit 401 is filtered by the set of spatial filters 701. Thus, in this example, the feedback signal, rather than the input to the audio beamformer 103, is filtered by the set of spatial filters.

[0243] Furthermore, the set of spatial filters is determined as a set of inverse spatial decorrelation filters, where each of the inverse spatial decorrelation filters includes the inverse filtering of the determined decorrelation filters.

[0244] For example, if a spatially decorrelated filter is represented by a weight matrix W(ω), then the inverse spatially decorrelated filter is represented by the weight matrix W -1 (ω) is determined by the set of spatial filters applied to the second set of audio signals. -1 This includes filtering corresponding to (ω), where W(ω) is the spatial decorrelation determined by the decorrelation adapter 203.

[0245] In many embodiments, the set of spatial filters applied to a second set of audio signals is determined to have filter coefficients corresponding to the coefficients of a spatial filter, which is a cascade of two sets of spatial inverse filters, each of which is an inverse filter of a spatial decorrelation filter determined by the decorrelation adapter 203.

[0246] In particular, the decorrelation adapter 203 therefore: y(ω)=W(ω)X(ω) The weights / coefficients of the weight matrix W(ω) are determined for a decorrelation filter that can decorrelate the first set of audio signals according to the following.

[0247] In this case, the set of spatial filters 301 is generated to correspond to the cascading application of two such filters, i.e., this is: y(ω)=W -1 (ω)W -1 (ω)s(ω)=W -2 (ω)s(ω) It is configured to correspond to, where s(ω) represents a second set of audio signals generated by the feedback circuit 401.

[0248] Therefore, in the example, the decorrelation adapter 203 proceeds to determine appropriate coefficients for a spatial decorrelation filter that decorrelates a first set of audio signals by performing a fitting as described for the example in Figure 5. The determined decorrelation filter coefficients are not used directly, but rather the coefficients for a set of spatial filters 701 are determined from the determined coefficients. In such a case, the decorrelation adapter 203 performs the decorrelation filter as part of the coefficient determination / fitting, and this is (e.g., [y i (k,ω)y j * A first set of signals (from (k,ω)] is applied to determine the update value. However, in certain examples, such filtered signals are used only for the fitting process and not further for beamforming / processing.

[0249] This approach provides improved overall performance, including an enhanced beamforming experience. In fact, this is the set of spatial filters 701 shown in Figure 7. -2 For the example where the coefficient is set to the corresponding value, it is possible to show that this then produces the same performance and operation as the example in Figure 5, and specifically that these produce the same optimal solution.

[0250] The audio device employs a highly adaptive approach, particularly well-suited for adapting to extract specific sound sources, such as a desired speaker. This approach is especially advantageous, for example, when a strong noise source is present in the captured audio environment. The adaptive audio device has different adaptations for each of the spatial filter set (via decorrelation filter / decorrelation coefficient adaptation) and beamformer filter adaptations. The decorrelation adapter 203 and beamformer adapter 205 work synergistically to provide a highly advantageous and high-performance adaptation to the current audio characteristics of the scene.

[0251] For configuration B (shown in Figure 6), decorrelation is performed using the noise covariance matrix R nn (ω) : W(ω)R nn (ω)W H (ω)=Λ(ω) It should be noted that it is possible to convert W(ω) and R nn (ω) is positive definite, and its inverses exist: W -1 (ω)W(ω)R nn (ω)W H (ω)(W H (ω)) -1 =W -1 (ω)Λ(ω)(W H (ω)) -1 R nn (ω)=W -1 (ω)Λ(ω)(W H (ω)) -1 It can be written as follows: Λ(ω) is a scaled version of the identity matrix,

number

number

number

[0252] Next, we will prove that this combination maximizes the speech-to-noise ratio in the output. First, we derive the converged filter coefficients, as we did previously for the beamformer.

[0253] In the output of the filtering section: z(ω)=f H (ω)W 2 (ω)u(ω)=s(ω)f H (ω)W 2 (ω)h(ω) (Eq.13) It is possessed.

[0254] In the update section: z(ω)f(ω)=s(ω)h(ω) (Eq14) Since z(ω) and s(ω) are scalars: f(ω)=α(ω)h(ω) This can be expressed as follows: By substituting Eq13 into Eq14:

number

number

number

[0255] Next, the expected value of updates after convergence:

number

[0256] Next, regarding the update formula (Eq. 15):

number

number

number

[0257] Equation 16 shows that after convergence f(ω), the eigenvectors and matrix R uu (ω)W 2 The corresponding eigenvalue of (ω)

number

number

[0258] f H (ω)W 2 By multiplying by (ω):

number

[0259] Next time: R uu (ω)=R ss (ω)+R nn Considering (ω), here the adapted decorrelator is adapted to the noise source. Next, the output power

number

number

[0260] The noise contribution is independent of the selection of f(ω), and therefore maximizing the FSB output does not necessarily correspond to maximizing the speech-to-noise ratio in the output, but rather to maximizing the speech output itself.

[0261] The optimal solution after convergence for solutions A and B is: f A (ω)=W(ω)f B (ω) It is associated with.

[0262] This means that the outputs for both solutions will be equal.

[0263] For solution C in Figure 7, the matrix block is placed after the filter in the feedback circuit 401. Here the constraint is f H (ω)Γ(ω)f(ω) = 1, where Γ(ω) is the coherence matrix, which is equal to the normalized covariance matrix. Since it is desirable to use the decorrelator again,

number

[0264] Therefore, W -2 (ω) is equal to the normalized covariance matrix, W -2 (ω) can be used instead of Γ(ω), and the constraint is f H (ω)W -2 (ω)f(ω)=1.

[0265] As before, the optimal solution for the filter coefficient f(ω) is found by considering the signal flow in the filtering and updating sections.

[0266] In the filtering section: z(ω)=s(ω)f H (ω)h(ω) (Eq.17) It is provided, and in the updated section: z(ω)W -2 (ω)f(ω)=s(ω)h(ω) (Eq.18) f(ω)=α(ω)W 2 (ω)h(ω) (Eq.19) We have such that z(ω) and s(ω) are complex scalers.

[0267] By using (17) and substituting (19) into (18):

number

number

number

number

number

number

number

number

number

[0268] After the convergence f(ω), the eigenvectors and matrix W are shown. 2 (ω)R uu The corresponding eigenvalue of (ω)

number

number

[0269] The left and right sides of Eq. 20 are given by f H By multiplying by (ω):

number

[0270] Next time: R uu (ω)=R ss (ω)+R nn Considering (ω), here the adapted decorrelator is adapted to the noise source. Next, the output power

number

number

[0271] Maximizing the FSB output does not necessarily correspond to maximizing the speech-to-noise ratio in the output itself, but rather to maximizing the speech-to-noise ratio in the output.

[0272] The optimal solution after convergence for solutions C and B is f C (ω)=W 2 (ω)f B It is associated with (ω).

[0273] This means that the output for both solutions will be equal. Therefore, different constructions / solutions:

[0274] [Table 1] This is summarized by:

[0275] As can be seen, by appropriately selecting coefficients, it is possible to generate exactly the same (optimal) beamform output audio signal.

[0276] However, as can be seen, the coefficients selected for the set of spatial filters to achieve this output will differ.

[0277] The choice of approach used depends on the preferences and requirements of each individual embodiment.

[0278] The advantage of solution A is that the interaction between the decorrelator and the beamformer is reduced in the sense that synergistic effects can be achieved through reduced modifications of the beamformer, feedback circuit, and beamform adapter.

[0279] The advantage of solution B is that, if it is desired to calculate the direction of arrival (DOA) of the sound source using beamformer coefficients, this is easier to do for this configuration.

[0280] For example, an approach to calculating DOA is described in U.S. Patent No. 6,774,934 for a “normal” beamformer with a speech source where uncorrelated equivalent noise may be added. For this use case, the optimal coefficients are:

number

[0281] Part of the procedure is to calculate the cross-power spectrum for at least one pair of filter coefficients, e.g., f1(ω) and f2(ω). This allows:

number

[0282] Phase difference is a more important factor than amplitude. While it is possible to directly use the beamformer coefficient for solution B, for methods A and C, the coefficient is W -1 (ω) and W -2This means that each element is pre-multiplied by (ω).

[0283] Different configurations therefore result in the generation of exactly identical beamform output audio signals. Consequently, identical compensation frequency responses are also favorably applied by equalizer 105. A suitable approach for determining the appropriate compensation frequency response from beamform and decorrelation parameters is determined using the same considerations and approaches as in configuration A.

[0284] For configuration B, the magnitude of the compensated frequency response is determined by the equalization adapter 107:

number

[0285] in this case

number

[0286] |h eq Substituting (ω)| into the previous equation is as follows:

number

[0287] In this case, it should be noted that the compensated frequency response depends only on the beamform parameters and not on the decorrelation parameters.

[0288] For configuration C, the magnitude of the compensated frequency response is determined by the equalization adapter 107:

number

[0289] in this case

number

[0290] |h eq Substituting (ω)| into the previous equation is as follows:

number

[0291] In many embodiments, the audio device is configured to perform different adaptations at different times. Therefore, the decorrelation adapter 203 and the beamform adapter 205 are controlled to perform their respective adaptations over different time periods, and in many embodiments, these do not overlap. Thus, at a given time, either the decorrelation adapter 203 or the beamform adapter 105 is adapted, but not both (and for some periods neither is adapted).

[0292] For example, as shown in Figure 8, the audio device of Figure 5 (or the corresponding in Figure 6 or Figure 7) includes an audio detector 801 configured to detect when the sound source is active and when it is not. Specifically, the audio detector 801 is configured to divide time into (at least) an active set of time intervals in which the sound source is active, and an inactive set of time intervals in which the sound source is not active. Specifically, the sound source is a desired sound source, such as a desired speaker.

[0293] In some embodiments, low-complexity speech detection is used to determine whether a source is active or inactive. For example, a speech device is used to extract the dominant sound source in an environment, such as the most powerful speaker. For instance, in a typical teleconferencing application where only one person is speaking at the time, the speech device is used to extract the current speaker's speech from background or ambient noise.

[0294] In such applications, for example, the speech detector 801 simply detects sound source activity based on whether the level / energy exceeds a given threshold. If the signal level exceeds a given threshold (which is, for example, a dynamic threshold set based on a longer-term averaged signal level), the speech detector 801 considers the sound source / speaker to be active; otherwise, it considers the sound source / speaker to be inactive. Thus, the speech detector 801 divides time into active time intervals (due to the sound source being active) if the captured speech exceeds the threshold, and inactive time intervals (due to the sound source being inactive) if the signal level falls below the threshold.

[0295] In many embodiments, more complex detection is used, and the technique is actually used to separate different sound sources. For example, in many embodiments, detection involves considering whether a sound source has desired / expected characteristics that match a particular sound source. For example, the speech detector 801 distinguishes speech from other types of speech based on an evaluation of whether the captured speech has characteristics that match speech.

[0296] Many different techniques and algorithms are known for utterance / voice activity detection, more generally for detecting that a sound source is active, and it is understood that any suitable approach may be used without prejudice to the present invention.

[0297] In several embodiments, trained artificial neural networks have been found to be capable of performing effective speech / voice activity detection (and / or noise detection). The artificial neural networks are trained on speech and all types of non-speech, providing an indication for each frame whether it is noise or speech.

[0298] In many embodiments, detection is relatively fast, and the speech detector 801 is configured to designate relatively short time intervals as active or inactive time intervals. For example, in some embodiments, the speech detector 801 is configured to detect silent pauses / intervals during periods of normal speech and to designate such intervals as inactive time intervals. For example, time intervals shorter than, say, 5 msec, 10 msec, or 20 msec are identified / designated as active or inactive time intervals in some embodiments.

[0299] In many embodiments, the speech detector 801 directly detects activity about a sound source based on a first set of received microphone signals / speech signals / beamformer input signals / beamform output speech signals. For example, the sum of the energy of the captured speech is determined and compared to a threshold. However, in other embodiments, other signals, such as a second set of speech signals, are considered. In many embodiments, the speech detector 801 bases activity detection on the generated beamform output speech signal. For example, level detection or speech detection is applied directly to the beamform output speech signal. This provides improved performance in many embodiments, as the output signal is generated to focus on a desired signal, such as a desired speaker.

[0300] The sound detector 801 is configured to control the fit based on sound source activity detection, specifically the fit of the decorrelation coefficient and / or beamform filter coefficient differs between inactive and active time intervals.

[0301] In fact, in some embodiments, the fitting of the decorrelation coefficient is not performed during active time intervals, but only during inactive time intervals. The speech detector 801 specifically requests to detect whether a desired or preferred sound source / speaker is active. The speech detector 801 further controls the decorrelation adapter 203 to fit the decorrelation coefficient only if this sound source is not active. Thus, the decorrelation adapter 203 is configured to be fitted to provide a decorrelation filter that requests the decorrelation of unwanted speech. Thus, the goal is to be fitted to reduce the correlation of unwanted speech rather than updating the decorrelation filter to decorrelate the entire captured speech. Such an approach provides highly efficient and improved operation and performance in many situations. The decorrelation approach for unwanted speech allows beamforming to provide efficient and high-performance extraction of desired signals, while enabling improved beamforming with improved exclusion of unwanted speech by beamforming.

[0302] In some embodiments, the decorrelation adapter 203 is configured to fit the decorrelation coefficient in both active and inactive time intervals, but the fitting rate is higher during inactive time intervals than during active time intervals. Therefore, instead of fitting the decorrelation coefficient only during inactive time intervals, some fitting also exists during active time intervals, but this fitting is slower, for example, 2, 5, 10, or 100 times slower. This is advantageous in some scenarios, such as when the active time interval is much longer than the inactive time interval. In some embodiments, the update rate during active and / or inactive time intervals is dynamically adjusted, for example, depending on the characteristics of the time interval or input signal. For example, the decorrelation adapter 203 switches between updating only during inactive time intervals and updating (at a lower rate) during active time intervals as well if the duration of the active time interval exceeds a given duration.

[0303] In many embodiments, beamforming filter adaptation is not performed during inactive time intervals, but only during active time intervals. Specifically, the speech detector 801 detects whether a desired or preferred sound source / speaker is active and requests that the beamforming adapter 205 be controlled to adapt the beamforming filter only if this sound source is active. Thus, the beamforming adapter 205 is specifically configured to adapt the beamforming so that it specifically (effectively) forms a beam toward the desired sound source.

[0304] In some embodiments, the beamform adapter 205 is configured to adapt beamformer coefficients in both active and inactive time intervals, but the adaptation rate is higher during active time intervals than during inactive time intervals. Thus, instead of adapting the beamform filter only during active time intervals, some adaptation also exists during inactive time intervals, but this adaptation is slower, for example, 2, 5, 10, or 100 times slower. This is advantageous in some scenarios, such as when the audio detector 801 is configured to provide detection with a very low risk of false detection of an active sound source, and the risk of not detecting an active sound source is relatively high. In such cases, adaptation is still desirable during inactive time intervals, but the update rate is (typically) very low. Thus, a low update rate is used during time intervals when the desired source is active or inactive, and a high update rate is used during time intervals when the desired source is almost certainly present.

[0305] In many embodiments, the audio device is configured to adapt the decorrelation coefficient only during periods of inactive time intervals and the beamforming filter only during periods of active time intervals. This provides highly efficient performance and, in particular, significantly improved extraction / separation of the desired sound source / speaker in the presence of noise / unwanted speech, such as dominant and correlated noise.

[0306] In many practical applications, such as numerous speech capture and processing applications, both beamforming and decorrelation adaptation controls are crucial for the effective operation of speech devices. It is desirable that the beamformer adapts to decorrelation for speech that is not always active by its nature, and for portions consisting solely of noise. In use cases where noise is continuously present, the beamformer can only adapt if noise is also present. If the noise is uncorrelated, or if it is decorrelated by decorrelation, the noise does not affect the adaptation if it is equivariant. If the noise is not equivariant, the beamformer diverges during periods of noise only, but if the near end becomes active, the beamformer can often be quickly (almost imperceptibly) adjusted to the desired speaker. Depending on the application, further adaptation controls may not be used to allow the beamformer to quickly find new sources, or a voice activity detector may be used, which can range from simple energy-based detectors to more sophisticated detectors that also use pitch-based speech characteristics.

[0307] The above explanation focuses on a scenario where the active time interval is the active time of the desired / speech source, and the inactive time interval is the active time interval of the unwanted source / noise / interference. In this case, the fitting rate for the beamformer filter is higher during the active time interval than during the inactive time interval; specifically, the fitting rate for the beamformer filter is performed only during the active time interval and not during the inactive time interval. Similarly, the fitting rate for the decorrelation coefficient is higher during the inactive time interval than during the active time interval; specifically, the fitting rate for the filter is performed only during the inactive time interval and not during the active time interval.

[0308] Such time intervals are determined directly, for example, by the use of a speech detector that detects and determines the active time intervals of utterances that are considered as active time intervals. The remaining time intervals, i.e., active time intervals of non-utterances, are considered as inactive time intervals.

[0309] In some embodiments, detection of unwanted audio characteristics is used, specifically such as the detection of activity of interfering objects or noise sources. For example, detection of activity of unwanted sound sources, such as music detectors, silence detectors, or specific noise detectors, may be used. In such cases, the time intervals detected are considered as inactive time intervals. The remaining time intervals are considered as active time intervals.

[0310] Instead, if we consider that the active time interval corresponds to the active time interval where the unwanted sound source is not desired, and the inactive time interval corresponds to the inactive time interval where the unwanted sound source is not desired, then the aforementioned fitting is inverted (the filter fitting rate becomes higher during the inactive time interval, and the decorrelation coefficient fitting rate becomes higher during the active time interval; specifically, in many embodiments, the beamformer filter is fitted only during the inactive time interval, and the decorrelation coefficient is fitted only during the active time interval).

[0311] In this approach, the decorrelation function and beamforming function are dynamically updated, and accordingly, the equalizer 105 is dynamically fitted to have a variable compensated frequency response after the decorrelation and beamforming adjustments. In many embodiments, the operation is performed within a segment or block, and new updated parameters are determined for each block.

[0312] However, abruptly changing the compensated frequency response has been found to produce audio artifacts in many scenarios and applications. In many embodiments, the equalization adapter 107 is configured to change the compensated frequency response gradually, without abrupt and large changes in the compensated frequency response. Specifically, the compensated frequency response is not simply changed between different segments, but rather a gradual change is performed within a segment.

[0313] In many embodiments, the equalization adapter 107 is configured to interpolate between the current and new compensated frequency responses, with the interpolation gradually changing from the former to the latter. Specifically, the beamform output audio signal is subjected to both the current and new frequency equalizations, and specifically, samples of the beamform output audio signal are filtered by both the current and new compensated frequency responses. Thus, samples are generated for two signals corresponding to the current and new frequency equalizations, respectively. The equalizer 105 then generates the output signal as an interpolation between these, and specifically, samples of the output signal (in instants of the same sample) are generated as a weighted combination of samples from the two signals. The weighting is then gradually changed from a higher weighting of the sample for the current frequency equalization to a higher weighting of the sample for the new frequency equalization. The rate of change in combination / interpolation is selected to provide appropriate performance for a particular embodiment.

[0314] The equalization adapter 107 then controls the equalizer 105 to transition from one frequency equalization / compensation frequency response to another by determining intermediate samples of the output signal for two frequency equalizations and then generating samples of the output signal as a weighted combination of intermediate samples for different frequency equalizations, where the equalization adapter 107 is then configured to gradually change the relative weighting of the samples for the different frequency equalizations over the duration of the transition time interval.

[0315] Therefore h eq When (ω) is updated, it is preferable to perform a smooth transition. Good results are obtained when the old and new h for the transition frame (typically a block of 256 samples) eq This is obtained by calculating the output for both (ω) and then calculating the output with smooth transitions.

[0316] old h eq A sample of a frame processed with (ω) is called "inold", and the old h eq By indicating the sample frame processed by (ω) as "innew", the output is then determined by the following (pseudocode) approach with α=1, for example β=0.01:

number

[0317] The previous description focused on adaptive spatial decorrelation performed by the adaptive spatial decorrelationizer 301, but it is understood that in other embodiments, non-adaptive spatial decorrelation may be performed.

[0318] The voice device is specifically implemented by one or more appropriately programmed processors. Different functional blocks are implemented by different separate processors, and / or, for example, by the same processor. Examples of appropriate processors are provided below.

[0319] Figure 9 is a block diagram showing an exemplary processor 900 according to an embodiment of the present disclosure. The processor 900 is used to implement one or more processors that implement devices or elements thereof, including in particular one or more artificial neural networks, as described above. The processor 900 is any suitable type of processor, including, but not limited to, a microprocessor, a microcontroller, a digital signal processor (DSP), a field-programmable array (FPGA) programmed to constitute a processor, a graphics processing unit (GPU), an application-specific integrated circuit (ASIC) designed to constitute a processor, or a combination thereof.

[0320] The processor 900 includes one or more cores 902. Each core 902 includes one or more arithmetic logic units (ALUs) 904. In some embodiments, the core 902 includes a floating-point logic unit (FLPU) 906 and / or a digital signal processing unit (DSPU) 908 in addition to or instead of the ALUs 904.

[0321] The processor 900 includes one or more registers 912 that are communicatively coupled to the core 902. The registers 912 are implemented using dedicated logic gate circuits (e.g., flip-flops) and / or any memory technology. In some embodiments, the registers 912 are implemented using static memory. The registers provide the core 902 with data, instructions, and addresses.

[0322] In some embodiments, the processor 900 includes one or more levels of cache memory 910 communicably coupled to the core 902. The cache memory 910 provides computer-readable instructions for execution to the core 902. The cache memory 910 provides data for processing by the core 902. In some embodiments, computer-readable instructions are provided to the cache memory 910 by local memory, for example, local memory attached to an external bus 916. The cache memory 910 is implemented in any suitable type of cache memory, such as metal-oxide-semiconductor (MOS) memory such as static random access memory (SRAM), dynamic random access memory (DRAM), and / or any other suitable memory technology.

[0323] The processor 900 includes a controller 914, which controls inputs to the processor 900 from other processors and / or components included in the system, and / or outputs from the processor 900 to other processors and / or components included in the system. The controller 914 controls the data paths in the ALU 904, FPLU 906, and / or DSPU 908. The controller 914 is implemented as one or more state machines, data paths, and / or dedicated control logic. The gates of the controller 914 are implemented as standalone gates, FPGAs, ASICs, or any other suitable technology.

[0324] Registers 912 and cache 910 communicate with controller 914 and core 902 via internal connections 920A, 920B, 920C, and 920D. These internal connections are implemented as buses, multiplexers, crossbar switches, and / or any other suitable connection techniques.

[0325] Inputs and outputs for the processor 900 are provided via a bus 916 comprising one or more conductive wires. The bus 916 is communicatively coupled to one or more components of the processor 900, such as a controller 914, a cache 910, and / or registers 912. The bus 916 is coupled to one or more components of the system.

[0326] Bus 916 is coupled to one or more external memories. The external memory includes read-only memory (ROM) 932. ROM 932 is masked ROM, electronically programmable read-only memory (EPROM), or any other suitable technology. The external memory includes random access memory (RAM) 933. RAM 933 is static RAM, battery-backed static RAM, dynamic RAM (DRAM), or any other suitable technology. The external memory includes electrically erasable programmable read-only memory (EEPROM®) 935. The external memory includes flash memory 934. The external memory includes magnetic storage devices such as disk 936. In some embodiments, the external memory is contained within the system.

[0327] It is understood that the above description refers to embodiments of the invention with reference to different functional circuits, units, and processors for clarity. However, it will become clear that any appropriate distribution of functions between different functional circuits, units, or processors can be used without prejudice to the invention. For example, functions indicated to be performed by separate processors or controllers are performed by the same processor or controller. Therefore, references to specific functional units or circuits should be understood not as indicating a strict logical or physical structure or composition, but only as references to appropriate means for providing the described functions.

[0328] The present invention can be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. The present invention may optionally be implemented at least partially as computer software running on one or more data processors and / or digital signal processors. Elements and components of embodiments of the present invention may be implemented physically, functionally, and logically in any suitable manner. In fact, functionality may be implemented in a single unit, in multiple units, or as part of other functional units. Thus, the present invention may be implemented in a single unit or physically and functionally distributed across different units, circuits, and processors.

[0329] Although the present invention has been described in relation to several embodiments, it is not intended to be limited to any particular form described herein. Rather, the scope of the present invention is limited only by the appended claims. Furthermore, while some features may appear to be described in relation to a particular embodiment, those skilled in the art will recognize that various features of the described embodiments can be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.

[0330] Furthermore, even if described individually, steps of multiple means, elements, circuits, or methods may be performed, for example, by a single circuit, unit, or processor. Moreover, even if individual features are included in different claims, they may be advantageously combined in some cases, and inclusion in different claims does not imply that the combination of features is unfeasible and / or unfavorable. Also, inclusion of a feature within one category of claims does not imply limitation to that category, but rather indicates that the feature is equally applicable to other claim categories where applicable. Furthermore, the order of features in a claim does not imply any particular order in which the features must be operated, and the order of individual steps in a method claim, in particular, does not imply that the steps must be performed in a specific order. Rather, the steps are performed in any appropriate order. Furthermore, singular references do not exclude plural references. Thus, references to "a," "an," "first," "second," etc., do not exclude plural references. Reference numerals in the claims are provided merely as examples for clarity and should never be understood as limiting the scope of the claims.

[0331] Overall, examples of audio devices, methods of operation for audio devices, and computer programs for implementing the methods are shown by the following embodiments.

[0332] Embodiment Embodiment 1. A receiver (101) configured to receive a first set of audio signals, wherein the first set of audio signals includes audio signals that capture the sound of a scene from different locations, A voice beamformer (103) configured to generate a beamform output voice signal from a first set of voice signals, wherein the voice beamformer (103) is: A signal processor (201) configured to generate a beamformed output audio signal from signal processing of a first set of audio signals, wherein the signal processing includes spatial beamforming and spatial decorrelation of the first set of audio signals, the spatial decorrelation depending on a set of decorrelation parameters, and the spatial beamforming depending on a set of beamform parameters, A beamform adapter (205) configured to adapt a set of beamform parameters depending on the beamform output audio signal, An equalizer (105) configured to generate an output signal by applying frequency equalization to the beamform output audio signal, The system comprises an equalization adapter (107) configured to adapt frequency equalization depending on a set of beamform parameters, Audio beamformer (103) and A voice device equipped with the following features.

[0333] Embodiment 2. The audio device of Embodiment 1, wherein the equalization adapter (107) is configured to adapt frequency equalization depending on a set of uncorrelated parameters.

[0334] Embodiment 3. The signal processor (201) is: A spatial decorrelator (301) is configured to receive a first set of audio signals and generate a first set of decorrelated audio signals by performing spatial decorrelation, A beamforming circuit (303) configured to perform spatial beamforming by combining a first set of uncorrelated audio signals, wherein the combination depends on a set of beamforming parameters. An audio device according to Embodiment 1, comprising the following features.

[0335] Embodiment 4. An audio device according to any one of Embodiments 1 to 3, wherein the equalization adapter (107) is configured to adapt frequency equalization to have a minimum phase frequency response.

[0336] Embodiment 5. An audio device according to any one of Embodiments 1 to 4, wherein the equalization adapter (107) is configured to adapt frequency equalization and have a low-pass frequency response.

[0337] Embodiment 6. An audio device according to any one of Embodiments 1 to 5, wherein the equalization adapter (107) is configured to transition from the first frequency equalization to the second frequency equalization by determining a first intermediate sample of the output signal for the first frequency equalization and a second intermediate sample of the output signal for the second frequency equalization, and generating a sample of the output signal as a weighted combination of the first intermediate sample and the second intermediate sample, and the equalization adapter (107) is configured to gradually change the relative weights of the first intermediate sample and the second intermediate sample during the transition time interval.

[0338] Embodiment 7. Spatial decorrelation is adaptive spatial decorrelation, and the audio device is: The audio device of Embodiment 1 further comprises a decorrelation adapter (203) configured to adapt a set of decorrelation parameters depending on a first set of audio signals.

[0339] Embodiment 8. The speech device of Embodiment 7, further comprising a speech detector (801) configured to determine a set of active time intervals in which the sound source is active, and a set of inactive time intervals in which the sound source is not active, wherein at least one of the fittings of a first set of beamform parameters (303) and a second set of filters (401) and the fitting of a set of decorrelation parameters (203) differs for the set of active time intervals and the set of inactive time intervals.

[0340] Embodiment 9. The signal processor (201) is: A first set of filters (303) configured to filter a first set of audio signals, A combiner (303) configured to combine the outputs of a first set of filters (303) to generate a beamform output audio signal, A feedback circuit (401) includes a second set of filters configured to generate a second set of audio signals from filtering of a beamform output audio signal, wherein each filter in the second set of filters has a frequency response that is the complex conjugate of the filters in the first set of filters (303), A first set of spatial filters (301, 601, 701) configured to apply a first spatial filtering to at least one of a first set of audio signals and a second set of audio signals, wherein the first set of spatial filters (301, 601, 701) has coefficients determined from decorrelation coefficients included in the decorrelation parameter and Equipped with, The beamform adapter (205) is configured to adapt a first set of filters (303) and a second set of filters in accordance with a comparison between a first set of audio signals and a second set of audio signals. The audio device of Embodiment 7, wherein a decorrelation adapter (203) is configured to determine decorrelation coefficients for a set of spatial decorrelation filters that generate a decorrelation output signal from a first set of audio signals, and the decorrelation adapter (203) is configured to adapt the decorrelation coefficients according to updated values ​​determined from the first set of audio signals.

[0341] Embodiment 10. The audio device of Embodiment 9, wherein a first set of spatial filters (301) is configured to filter a first set of audio signals.

[0342] Embodiment 11. An audio device of Embodiment 9 or 10, wherein a first set of filters (303) is configured to filter a first set of audio signals after filtering by a first set of spatial filters (601), and a beamform adapter (205) is configured to perform a comparison using the first set of audio signals before filtering by the first set of spatial filters (601).

[0343] Embodiment 12. The audio device of Embodiment 8 or 9, wherein a first set of spatial filters (701) is configured to filter a second set of audio signals.

[0344] Embodiment 13. The decorrelation adapter (203) is configured to determine the adapted spatial decorrelation as a set of spatial decorrelation filters, and each output audio signal of the filters in the set of spatial decorrelation filters is: Steps include segmenting a first set of audio signals into time segments, and for at least some time segments: A step of generating a frequency bin representation of a first set of audio signals, wherein each frequency bin in the frequency bin representation of the first set of audio signals includes a frequency bin value for each of the audio signals in the first set of audio signals. A step of generating a frequency bin representation of a set of output signals, wherein each frequency bin of the frequency bin representation of the set of output signals contains a frequency bin value for each of the output signals, and the frequency bin values ​​for a given output signal in the set of output signals for a given frequency bin are generated as a weighted combination of the frequency bin values ​​of a first set of audio signals for a given frequency bin, and the weighted combination has a decorrelation coefficient as a weight. A step of updating a first weight for the contribution of the first frequency bin to the first frequency bin for the first output signal linked to the first input audio signal, from the second frequency bin for the first frequency bin for the second input audio signal linked to the second output signal, according to a correlation measure between the first previous frequency bin value of the first output signal for the first frequency bin and the second previous frequency bin value of the second output signal for the first frequency bin. One of the two audio devices from Embodiments 1 to 12, which is linked to one input audio signal from a first set of audio signals by performing the following:

[0345] Embodiment 14. The audio device of Embodiment 13, wherein the uncorrelated adapter (203) is configured to update the first weight according to the product of a first value and a second value, the first value being one of a first previous frequency bin value and a second previous frequency bin value, and the second value being the complex conjugate of the other of the first previous frequency bin value and the second previous frequency bin value.

[0346] Embodiment 15. A method of operation for an audio device, the method being: A receiver (101) configured to receive a first set of audio signals, wherein the first set of audio signals includes audio signals that capture the sound of a scene from different locations, A voice beamformer (103) configured to generate a beamform output voice signal from a first set of voice signals, wherein the voice beamformer (103) is: A signal processor (201) configured to generate a beamformed output audio signal from signal processing of a first set of audio signals, wherein the signal processing includes spatial beamforming and adaptive spatial decorrelation of the first set of audio signals, the adaptive spatial decorrelation depending on a set of decorrelation parameters, and the spatial beamforming depending on a set of beamform parameters, A decorrelation adapter (203) configured to adapt a set of decorrelation parameters depending on a first set of audio signals, A beamform adapter (205) configured to adapt a set of beamform parameters depending on the beamform output audio signal, An equalizer (105) configured to generate an output signal by applying frequency equalization to the beamform output audio signal, The system comprises an equalization adapter (107) configured to adapt frequency equalization depending on a set of beamform parameters, Audio beamformer (103) and A method having

[0347] Embodiment 16. A computer program product comprising computer program code means, wherein the computer program code means is adapted to perform all the steps of Embodiment 15 when the program is executed on a computer.

Claims

1. A receiver that receives a first set of audio signals, the first set of audio signals includes audio signals that capture the sound of a scene from different locations, A sound beamformer that generates a beamform output sound signal from a first set of the aforementioned sound signals, wherein the sound beamformer is A signal processor that generates the beamform output audio signal from signal processing of a first set of audio signals, wherein the signal processing includes spatial beamforming and spatial decorrelation of the first set of audio signals, the spatial decorrelation depends on a set of decorrelation parameters, the spatial beamforming depends on a set of beamform parameters, and the beamform parameters are parameters that indicate at least one of signal filtering and weighting of the beamforming combination, A beamform adapter that adapts the set of beamform parameters depending on the beamform output audio signal, An equalizer that applies frequency equalization to the beamform output audio signal to generate an output signal, An equalization adapter that adapts the frequency equalization depending on the set of beamform parameters, It is equipped with an audio beamformer and A voice device equipped with the following features.

2. The audio device according to claim 1, wherein the equalization adapter adapts the frequency equalization depending on the set of decorrelation parameters.

3. The aforementioned signal processor A spatial decorrelator that receives a first set of audio signals and generates a first set of decorrelated audio signals by performing spatial decorrelation, A beamforming circuit that performs spatial beamforming by combining a first set of the uncorrelated audio signals, wherein the combination depends on a set of beamforming parameters. The audio device according to claim 1, comprising:

4. The audio device according to any one of claims 1 to 3, wherein the equalization adapter has a minimum phase frequency response adapted to the frequency equalization.

5. The audio device according to any one of claims 1 to 4, wherein the equalization adapter adapts the frequency equalization to have a low-pass frequency response.

6. The audio device according to any one of claims 1 to 5, wherein the equalization adapter determines a first intermediate sample of the output signal for the first frequency equalization and a second intermediate sample of the output signal for the second frequency equalization, and transitions from the first frequency equalization to the second frequency equalization by generating a sample of the output signal as a weighted combination of the first intermediate sample and the second intermediate sample, and the equalization adapter gradually changes the relative weights of the first intermediate sample and the second intermediate sample during the transition time interval.

7. The spatial decorrelation is adaptive spatial decorrelation, and the audio device is The audio device according to claim 1, further comprising a decorrelation adapter that adapts the set of decorrelation parameters depending on a first set of audio signals.

8. The audio device according to claim 7, further comprising a sound detector that determines a set of active time intervals in which a sound source is active and a set of inactive time intervals in which the sound source is not active, wherein at least one of the fit of the set of beamform parameters and the fit of the set of decorrelation parameters differs for the set of active time intervals and the set of inactive time intervals.

9. The aforementioned signal processor A first set of filters for filtering the first set of audio signals, A combiner that generates the beamform output audio signal by combining the outputs of the first set of filters, A feedback circuit including a second set of filters that generates a second set of audio signals from the filtering of the beamform output audio signal, wherein each filter in the second set of filters has a frequency response that is the complex conjugate of the filters in the first set of filters, A first set of spatial filters to apply a first spatial filtering to at least one of the first set of audio signals and the second set of audio signals, wherein the first set of spatial filters has coefficients determined from the decorrelation coefficients included in the decorrelation parameters. Equipped with, The beamform adapter adjusts the first set of filters and the second set of filters according to a comparison between the first set of audio signals and the second set of audio signals. The audio device according to claim 7, wherein the decorrelation adapter determines a decorrelation coefficient for a set of spatial decorrelation filters that generate a decorrelation output signal from a first set of audio signals, and the decorrelation adapter adapts the decorrelation coefficient according to the updated value determined from the first set of audio signals.

10. The audio device according to claim 9, wherein the first set of spatial filters filters the first set of audio signals.

11. The audio device according to claim 9 or 10, wherein a first set of filters filters a first set of audio signals after filtering by a first set of spatial filters, and the beamform adapter performs the comparison using the first set of audio signals before filtering by the first set of spatial filters.

12. The audio device according to claim 9, wherein the first set of spatial filters filters the second set of audio signals.

13. The decorrelation adapter determines the adapted spatial decorrelation as a set of spatial decorrelation filters, and each output audio signal of the filters in the set of spatial decorrelation filters is The steps of segmenting the first set of audio signals into time segments, and for at least some time segments, A step of generating a frequency bin representation of a first set of audio signals, wherein each frequency bin in the frequency bin representation of the first set of audio signals includes a frequency bin value for each of the audio signals in the first set of audio signals. A step of generating a frequency bin representation of a set of output signals, wherein each frequency bin of the frequency bin representation of the set of output signals includes a frequency bin value for each of the output signals, and the frequency bin values ​​for a given output signal in the set of output signals for a given frequency bin are generated as a weighted combination of the frequency bin values ​​of a first set of audio signals for a given frequency bin, and the weighted combination has the decorrelation coefficient as a weight, A step of updating a first weight for the contribution of the first frequency bin of the first frequency bin to the first output signal linked to the first input audio signal, from the second frequency bin of the first frequency bin for the second input audio signal linked to the second output signal, according to a correlation scale between the first previous frequency bin value of the first output signal for the first frequency bin and the second previous frequency bin value of the second output signal for the first frequency bin. The audio device according to any one of claims 1 to 12, which is linked to one input audio signal from a first set of audio signals by performing the following:

14. The audio device according to claim 13, wherein the decorrelated adapter updates the first weight according to the product of a first value and a second value, the first value being one of the first previous frequency bin value and the second previous frequency bin value, and the second value being the complex conjugate of the other of the first previous frequency bin value and the second previous frequency bin value.

15. A method of operation for an audio device, wherein the method is A receiving step of receiving a first set of audio signals, wherein the first set of audio signals includes audio signals that capture the sound of a scene from different locations, A step of generating a beamform output audio signal from a first set of audio signals, wherein the step of generating the beamform output audio signal is A step of generating the beamform output audio signal from signal processing of a first set of audio signals, wherein the signal processing includes spatial beamforming and spatial decorrelation of the first set of audio signals, the spatial decorrelation depends on a set of decorrelation parameters, the spatial beamforming depends on a set of beamform parameters, and the beamform parameters are parameters that indicate at least one of filtering and weighting of the input signal for the beamforming combination; A step of adapting the set of decorrelation parameters depending on a first set of audio signals, A step of adapting the set of beamform parameters depending on the beamform output audio signal, A step of applying frequency equalization to the beamform output audio signal to generate an output signal, A step of adapting the frequency equalization depending on the set of beamform parameters. The steps include generating and A method having

16. A computer program comprising computer program code means, wherein when the computer program is executed on a computer, the computer program code means performs all the steps of the method according to claim 15.