Audio device and its operating method

The audio device uses complex conjugate filters and spatial decorrelation to enhance sound source isolation, addressing the limitations of traditional beamforming by reducing complexity and noise interference.

JP2026518307APending Publication Date: 2026-06-04KONINKLIJKE PHILIPS NV

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
KONINKLIJKE PHILIPS NV
Filing Date
2024-05-23
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing audio beamforming techniques struggle to optimally isolate sound sources, particularly in complex environments with multiple noise sources, often requiring high computational resources and complexity, which conflicts with the need for flexibility and efficient source separation.

Method used

An audio device employing a receiver with a first set of filters and a feedback circuit of complex conjugate filters, combined with spatial decorrelation filters, adaptively filters audio signals to enhance source separation and reduce noise interference.

Benefits of technology

The device achieves improved source isolation with reduced complexity, increased flexibility, and efficient noise reduction, adapting to acoustic and spatial characteristics of the environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026518307000001_ABST
    Figure 2026518307000001_ABST
Patent Text Reader

Abstract

The audio device includes a receiver 201 configured to receive a set of input audio signals. An audio beamformer 101 performs beamforming by combining the outputs of filters. A feedback circuit 103 includes a set of matching filters having a complex conjugate frequency response, and a beamform adapter 105 adapts the filters in accordance with a comparison between the input audio signal and the feedback audio signal. A fitting coefficient processor 207 determines decorrelation coefficients for a set of spatial decorrelation filters that produce a decorrelated output signal from the input audio signal, and the decorrelation coefficients are adapted based on the input signal. Sets of spatial filters 205, 301, and 401 are configured to apply spatial filtering to the input audio signal or the feedback signal, and the sets of spatial filters have coefficients determined from the decorrelation coefficients.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an apparatus and method for generating an audio output signal, and more specifically, but not limited to, to extracting audio from a desired sound source, such as a desired speaker. [Background technology]

[0002] Capturing speech, specifically utterances, has become increasingly important in recent decades. For example, capturing speech or other sounds is becoming increasingly crucial for a variety of applications, including telecommunications, remote conferencing, gaming, and voice user interfaces. However, a problem in many scenarios and applications is that the desired sound source is not typically the only sound source in the environment. Rather, in a typical voice environment, there are many other sound sources / noise sources that are captured by the microphone. Speech processing is often used to improve speech capture, particularly by post-processing the captured speech time intervals to enhance the resulting audio signal.

[0003] In many embodiments, speech is represented by multiple different audio signals that reflect the same audio scene or environment. In particular, in many practical applications, speech is captured by multiple microphones at different locations. For example, a linear array of multiple microphones is often used to capture speech in an environment such as a room. The use of multiple microphones allows spatial information of the speech to be captured. Many different applications leverage such spatial information to enable improved and / or new services. [Overview of the Initiative] [Problems that the invention aims to solve]

[0004] One frequently used approach is to attempt to isolate sound sources by applying beamforming, which involves forming a beam oriented relative to the direction of arrival of sound from a particular source. However, while this offers favorable performance in many scenarios, it is not optimal in all cases. For example, it does not provide optimal source isolation in some cases, and in fact, in some applications, such spatial beamforming does not provide ideal sound characteristics for further processing to achieve a given effect.

[0005] Therefore, while spatial source separation, specifically separation based on audio beamforming, is highly advantageous in many scenarios and applications, there is a demand for improving the performance and operation of such approaches. However, there is also a demand for typically low complexity and / or resource usage (e.g., computational resources), and these preferences often conflict with each other.

[0006] Therefore, an improved approach is advantageous, particularly one that enables reduced complexity, increased flexibility, easier implementation, reduced cost, improved speech capture, improved source discrimination, improved spatial perception / source separation, improved support for speech / utterance applications, reduced reliance on known or static acoustic properties, improved flexibility and customization for different speech environments and scenarios, improved speech beamforming, a better trade-off between performance and complexity / resource usage, and / or improved performance.

[0007] Therefore, the present invention seeks to mitigate, mitigate, or eliminate one or more of the aforementioned disadvantages, either individually or in any combination. [Means for solving the problem]

[0008] According to an aspect of the present invention: A receiver configured to receive a first set of audio signals, the first set of audio signals including audio signals that capture the sound of a scene from different locations; an audio beamformer comprising a first set of filters configured to filter the first set of audio signals and a combiner configured to combine the outputs of the first set of filters to produce a beamform output audio signal; a feedback circuit comprising a second set of filters configured to produce a second set of audio signals from the filtering of the beamform output audio signal, wherein each filter in the second set of filters has a frequency response that is the complex conjugate of a filter in the first set of filters; and the first set of filters in response to a comparison between the first set of audio signals and the second set of audio signals. A speech device is provided, comprising: a beamform adapter configured to adapt a second set of sets and filters; an adaptive coefficient processor configured to determine decorrelation coefficients for a set of spatial decorrelation filters that generate a decorrelated output signal from a first set of speech signals, wherein the adaptive coefficient processor is configured to adapt the decorrelation coefficients according to updated values ​​determined from the first set of speech signals; and a first set of spatial filters configured to modify at least one of the first set of speech signals and the second set of speech signals by applying a first spatial filtering to at least one of the first set of speech signals and the second set of speech signals, wherein the first set of spatial filters has coefficients determined from the decorrelation coefficients.

[0009] This approach provides improved operation and / or performance in many embodiments. This approach enables improved beamforming, particularly for focusing on specific sound sources in a scene. This approach enables improved extraction / separation of speech from specific sources in the presence of other sound sources and noise sources in the scene.

[0010] This approach enables the operation of the audio device to effectively adapt to current conditions, particularly the acoustic and spatial characteristics of the sound source and scene. This approach provides reduced sensitivity to noise and unwanted sound sources in the scene and captured audio.

[0011] This approach enables efficient operation and low complexity in many embodiments. Different fits work synergistically to provide improved separation of the desired sound source from other captured audio in the scene. Furthermore, the use of multiple fits provides improved operation while allowing the use of less complex fitting algorithms and criteria.

[0012] Each filter in the first set of filters is linked to a filter in the second set of filters, and this filter has a frequency response that is the complex conjugate of the frequency response of the filter in the first set of filters.

[0013] The first and second sets of filters each have an equal number of linked-pair filters having a complex conjugate frequency response. For each audio signal in the first set of audio signals, there exists one filter from the first set of filters, one linked-pair filter from the second set of filters (having a complex conjugate frequency response), and one audio signal in the second set of audio signals. The comparison is between the linked-pair audio signals in the first set of audio signals and the second set of audio signals. The fit of a given filter in the first set of filters and a given linked-pair filter in the second set of filters depends on a comparison (perhaps only) between the signal in the first set of audio signals filtered by the given filter in the first set of filters and the signal in the second set of audio signals produced by filtering the beamform audio signals by the given linked-pair filter in the second set of filters. The fit is made such that the difference is reduced.

[0014] Each of the second set of audio signals is an estimate of the contribution of the first set of audio signals from the audio captured by the beamform output audio signal to the linked audio signal, and is therefore typically an estimate of the contribution from the desired source.

[0015] A set of spatial decorrelation filters contains one spatial filter for each signal in a first set of signals. Each filter in the set of spatial decorrelation filters produces a filtered / modified version of one of the audio signals in the first set of audio signals. Together, the set of spatial decorrelation filters produces a modified first set of audio signals with a higher degree of decorrelation. The spatial decorrelation filters perform filtering over the first set of audio signals. The output of the decorrelation filters (the corresponding modified audio signals in the first set of audio signals) depends on the first set of multiple (unmodified) audio signals. Specifically, the spatial decorrelation filters are frequency domain filters, and the decorrelation coefficients are frequency domain coefficients. For a given spatial decorrelation filter, the output value for a given frequency bin at a given time is a weighted combination of multiple values ​​from the first set of (unmodified) audio signals for a given frequency bin at a given time. The decorrelation output signal has the correlation of the reduced normalized cross-channel signal with respect to the first set of input signals to the set of spatial decorrelation filters.

[0016] A set of spatial filters includes one spatial filter for each signal in a first set / second set of signals. Each filter in the set of spatial filters produces a filtered / modified version of one of the first or second sets of audio signals. The set of spatial filters together produces a modified first or second set of audio signals. The spatial filters perform filtering across the first set of audio signals. The output of the spatial filter (the corresponding modified audio signal in the first or second set of audio signals) depends on multiple of the first or second sets of (unmodified) audio signals. The spatial filter is specifically a frequency domain filter, and its coefficients are frequency domain coefficients. For a given spatial filter, the output value for a given frequency bin at a given time is a weighted combination of multiple values ​​from the first or second set of (unmodified) audio signals for a given frequency bin at a given time.

[0017] According to an optional feature of the present invention, the audio device further comprises an audio detector configured to determine, among a set of active time intervals, a set of active time intervals in which the sound source is active, and among a set of inactive time intervals, a set of inactive time intervals in which the sound source is not active, wherein at least one of the fittings of a first set of filters and a second set of filters and the fitting of the decorrelation coefficient differs for the set of active time intervals and the set of inactive time intervals.

[0018] This provides improved performance and / or operation in many embodiments. Typically, it provides improved adaptation, which usually leads to improved extraction of the desired sound source.

[0019] The active time interval is the active time interval of the desired / preferred / speech source, and the inactive time interval is the inactive time interval of the desired / preferred / speech source.

[0020] According to an optional feature of the present invention, the beamform adapter is configured to fit a first set of filters and a second set of filters at a higher fitting rate during a set of active time intervals than during a set of inactive time intervals.

[0021] This provides improved performance and / or operation in many embodiments. Typically, it provides improved adaptation, which usually leads to improved extraction of the desired sound source.

[0022] According to an optional feature of the present invention, the beamform adapter is configured to adapt a first set of filters and a second set of filters for a period of time of only one set of time intervals from among the set of active time intervals and the set of inactive time intervals.

[0023] This provides improved performance and / or operation in many embodiments. Typically, it provides improved adaptation, which usually leads to improved extraction of the desired sound source.

[0024] In some embodiments, the beamform adapter is configured to fit a first set of filters and a second set of filters during a set of active time intervals, but not during a set of inactive time intervals.

[0025] According to an optional feature of the present invention, the fitting coefficient processor is configured to fit the decorrelation coefficients at a higher fitting rate during a set of inactive time intervals than during a set of active time intervals.

[0026] This provides improved performance and / or operation in many embodiments. Typically, it provides improved adaptation, which usually leads to improved extraction of the desired sound source.

[0027] According to an optional feature of the present invention, the fitting coefficient processor is configured to fit the uncorrelated coefficients for a period of time of only one set of time intervals from a set of inactive time intervals and a set of active time intervals.

[0028] This provides improved performance and / or operation in many embodiments. Typically, it provides improved adaptation, which usually leads to improved extraction of the desired sound source.

[0029] In some embodiments, the fitting coefficient processor is configured to fit the uncorrelated coefficients over a set of inactive time intervals, but not over a set of active time intervals.

[0030] According to an optional feature of the present invention, a first set of spatial filters is configured to filter a first set of audio signals.

[0031] This provides improved performance and / or operation in many embodiments. In many embodiments and scenarios, this provides particularly attractive performance and / or implementation.

[0032] A first set of spatial filters generates a first modified audio signal by filtering a first set of audio signals. The first modified audio signal is fed to a first set of audio beamformers / filters (the first modified audio signal is the first set of audio signals filtered by the first set of filters). Alternatively or additionally, the first modified audio signal is fed to a beamform adapter (the first modified audio signal is the first set of audio signals compared to a second set of audio signals).

[0033] According to an optional feature of the present invention, the first set of spatial filters is configured to have coefficients set to the decorrelation coefficients determined for the set of spatial decorrelation filters.

[0034] This provides improved performance and / or operation in many embodiments. In many embodiments and scenarios, this provides particularly attractive performance and / or implementation.

[0035] According to an optional feature of the present invention, a first set of filters for filtering is configured to filter a first set of audio signals after filtering by a first set of spatial filters, and a beamform adapter is configured to perform a comparison using the first set of audio signals before filtering by the first set of spatial filters.

[0036] This provides improved performance and / or operation in many embodiments. In many embodiments and scenarios, this provides particularly attractive performance and / or implementation.

[0037] According to an optional feature of the present invention, the first set of spatial filters is configured to have coefficients that match the coefficients of two cascades of spatial filters from the set of uncorrelated filters.

[0038] This provides improved performance and / or operation in many embodiments. In many embodiments and scenarios, this provides particularly attractive performance and / or implementation.

[0039] According to an optional feature of the present invention, a first set of spatial filters is configured to filter a second set of audio signals.

[0040] This provides improved performance and / or operation in many embodiments. In many embodiments and scenarios, this provides particularly attractive performance and / or implementation.

[0041] The first set of spatial filters filters the second set of audio signals to produce a second modified set of audio signals. This second modified set of audio signals is fed to the beamform adapter (the second modified set of audio signals is the second set of audio signals compared to the first set of audio signals).

[0042] According to an optional feature of the present invention, a first set of spatial filters is configured to have coefficients determined according to a set of inverse spatial decorrelation filters, and the set of inverse spatial decorrelation filters is the inverse filter of the set of spatial decorrelation filters.

[0043] This provides improved performance and / or operation in many embodiments. In many embodiments and scenarios, this provides particularly attractive performance and / or implementation.

[0044] According to an optional feature of the present invention, the first set of spatial filters is configured to have coefficients that match the coefficients of a spatial filter which is a cascade of two sets of spatial inverse filters, each of which comprises an inverse filter of a set of spatial decorrelating filters.

[0045] This provides improved performance and / or operation in many embodiments. In many embodiments and scenarios, this provides particularly attractive performance and / or implementation.

[0046] According to an optional feature of the present invention, the adaptive coefficient processor is configured to determine a set of spatially decorrelated filters to generate a set of output audio signals, where each output audio signal in the set of output signals is: a step of segmenting a first set of audio signals into time segments; and for at least some time segments: a step of generating a frequency bin representation of the first set of audio signals, wherein each frequency bin in the frequency bin representation of the first set of audio signals includes a frequency bin value for each of the audio signals in the first set of audio signals; and a step of generating a frequency bin representation of the set of output signals, wherein each frequency bin in the frequency bin representation of the set of output signals includes a frequency bin value for each of the output signals, where for a given frequency bin, The frequency bin values ​​for the output signal are generated as a weighted combination of the frequency bin values ​​of a first set of audio signals for a given frequency bin, wherein the weighted combination has a decorrelation coefficient as its weight, and the input audio signal is linked to one of the first set of audio signals by performing the steps of generating and updating a first weight for the contribution of the first frequency bin to the first frequency bin for the first output signal linked to the first input audio signal, from the second frequency bin value of the first frequency bin for the second input audio signal linked to the second output signal, according to a correlation measure between the first previous frequency bin value of the first output signal for the first frequency bin and the second previous frequency bin value of the second output signal for the first frequency bin.

[0047] This provides improved performance and / or operation in many embodiments. In many embodiments and scenarios, this provides particularly attractive performance and / or implementation.

[0048] This provides the advantageous generation of an output audio signal with typically increased decorrelation compared to the input signal. In many embodiments, this approach provides an efficient fitting of the operation that results in improved decorrelation. The fitting is typically performed with low complexity and / or resource usage. Specifically, this approach applies local fitting of individual weights to further achieve efficient fitting.

[0049] The generation of the output signal set is adapted in many embodiments and for many applications to provide enhanced audio processing, particularly beamforming, by providing increased decorrelation to the input signal.

[0050] The first and second output audio signals are typically different output audio signals.

[0051] According to an optional feature of the present invention, the adaptive coefficient processor is configured to update a first weight in accordance with the product of a first value and a second value, where the first value is one of a first previous frequency bin value and a second previous frequency bin value, and the second value is the complex conjugate of the other of the first previous frequency bin value and the second previous frequency bin value.

[0052] This provides improved performance and / or operation in many embodiments. Typically, it provides improved fit, leading to increased decorrelation of the output audio signal in many scenarios.

[0053] In some embodiments, the audio device is configured to update a second weight, which is the contribution of a third frequency bin value to the first frequency bin value, in accordance with the magnitude of a first previous frequency bin value.

[0054] This provides improved performance and / or operation in many embodiments. Typically, it provides improved fit, leading to increased decorrelation of the output audio signal in many scenarios. In particular, it provides improved fit of the generated output signal. In many embodiments, the weight update, which reflects the contribution of the linked input signal to the output signal, depends on the magnitude / amplitude of that linked input signal. For example, the update requires compensating the weights for the level of the input signal to produce a normalized output signal.

[0055] This approach allows for normalization / signal compensation / level compensation, for example, to provide the desired output level.

[0056] In some embodiments, the audio device is configured to set a predetermined weight for the contribution of a third frequency bin value, which is the frequency bin value of the first frequency bin for a first input audio signal, to the first frequency bin value.

[0057] This provides improved performance and / or operation in many embodiments. Typically, it provides improved fit, leading to increased decorrelation of the output audio signal in many scenarios. In many embodiments, it provides improved fit while ensuring convergence of the fit toward non-zero signal levels. It works very efficiently with weighted fit that is not for linked signal pairs.

[0058] In many embodiments, the adapter is configured to maintain a constant weight without weight adjustment or updating.

[0059] In some embodiments, the audio device is configured such that the weights for the contribution of a third frequency bin value, which is the frequency bin value of the first frequency bin for a first input audio signal, to the first frequency bin value are real values.

[0060] This provides improved performance and / or operation in many embodiments. Typically, it provides improved fit, leading to increased decorrelation of the output audio signal in many scenarios.

[0061] The weights between linked input / output signals are constrained to be determined / become, favorably, as real-valued weights. This leads to improved performance and fit, guaranteeing convergence to non-zero level solutions.

[0062] In some embodiments, the audio device is configured such that a second weight, which is the weight for the contribution of a first frequency bin to a fourth frequency bin value for a second output audio signal from a first input audio signal, is the complex conjugate of the first weight.

[0063] This provides improved performance and / or operation in many embodiments. Typically, it provides improved fit, leading to increased decorrelation of the output audio signal in many scenarios.

[0064] In many embodiments, the two weights for two pairs of input / output signals are their complex conjugates.

[0065] In some embodiments, the weights of the weighted combination of input audio signals other than the first input audio signal are complex numerical weights.

[0066] This provides improved performance and / or operation in many embodiments. The use of complex values ​​for weights for unlinked input signals provides improved frequency-domain operation.

[0067] In some embodiments, the audio device is: y(ω)=W(ω)X(ω) It is configured to determine the output bin value for a given frequency bin ω from, where y(ω) is a vector containing the frequency bin values ​​for the output audio signal for a given frequency bin ω, x(ω) is a vector containing the frequency bin values ​​for the input audio signal for a given frequency bin ω, and W(ω) is a matrix having rows containing the weights of a weighted combination for the output audio signal.

[0068] This provides improved performance and / or operation in many embodiments. Typically, it provides improved fit, leading to increased decorrelation of the output audio signal in many scenarios.

[0069] The matrix W(ω) is favorably Hermitian. In many embodiments, the diagonals of the matrix W(ω) are constrained to be real-valued, set to predetermined values, and / or maintained as fixed values ​​but not updated / adapted. The weights / coefficients outside the diagonals are generally complex-valued.

[0070] In some embodiments, the audio device is: w ij (k+1,ω)=w ij (k,ω)-η(k,ω)[y i (k,ω)y j * (k,ω)] The weights of matrix W(ω) are determined accordingly. ij It is configured to fit, where i is the row index of matrix W(ω), j is the column index of matrix W(ω), k is the time segment index, ω represents the frequency bin, and η(k,ω) is the scaling parameter for fitting the fitting speed.

[0071] This provides improved performance and / or operation in many embodiments. Typically, it provides improved fit, leading to increased decorrelation of the output audio signal in many scenarios.

[0072] In some embodiments, the audio device is configured to compensate for correlation values ​​with respect to the signal levels of a first frequency bin.

[0073] This provides improved performance and / or operation in many embodiments. Typically, it provides improved fit, leading to increased decorrelation of the output audio signal in many scenarios. This allows for update rate compensation for signal fluctuations.

[0074] In some embodiments, the audio device is configured to initialize the weights such that the weighted combination includes at least one zero-value weight and one non-zero-value weight.

[0075] This provides improved performance and / or operation in many embodiments. This typically provides improved fit, leading to increased decorrelation of the output audio signal in many scenarios. This enables more efficient and / or faster fit and convergence toward favorable decorrelation. In many embodiments, the matrix W(ω) is initialized with zero values ​​for weights or coefficients for unlinked signals and fixed non-zero real values ​​for linked signals. Typically, the weights are set to, for example, 1 for the weights on the diagonals, and all other weights are initially set to zero.

[0076] In some embodiments, the weighted combination involves applying time-domain windowing to the frequency representation of the weights formed by the weights for a first input audio signal and a second input audio signal for different frequency bins.

[0077] This provides improved performance and / or operation in many embodiments. Typically, it provides improved fit, leading to increased decorrelation of the output audio signal in many scenarios.

[0078] Applying time-domain windowing to the frequency representation of weights involves: converting the frequency representation of weights to the time-domain representation of weights; applying a window to the time-domain representation to generate a modified time-domain representation; and converting the modified time-domain representation back to the frequency domain.

[0079] According to an aspect of the present invention, a method for generating an output audio signal, the method comprising: receiving a first set of audio signals, the first set of audio signals including audio signals capturing sound of a scene from different locations; generating an output signal by an audio beamformer: filtering the first set of audio signals with a first set of filters; and combining the outputs of the first set of filters to generate an output audio signal; and generating a second set of audio signals from the filtered output audio signals with a second set of filters, each filter in the second set of filters having a frequency response which is the complex conjugate of the filters in the first set of filters; and generating the audio signal A method is provided comprising the steps of: adapting a first set of filters and a second set of filters in accordance with a comparison of a set 1 with a second set of audio signals; determining decorrelation coefficients for a set of spatial decorrelation filters that generate a decorrelated output signal from a first set of audio signals, including adapting decorrelation coefficients in accordance with updated values ​​determined from the first set of audio signals; and modifying the first set of spatial filters by applying a first spatial filtering to at least one of the first set of audio signals and the second set of audio signals, wherein the first set of spatial filters has coefficients determined from the decorrelation coefficients.

[0080] These and other aspects, features and advantages of the present invention will become clear and evident by referring to the embodiments described below.

[0081] Embodiments of the present invention are described with reference to the drawings, merely as examples. [Brief explanation of the drawing]

[0082] [Figure 1] An example of a beamformer is shown. [Figure 2] Examples of audio devices according to several embodiments of the present invention are shown. [Figure 3] Examples of audio devices according to several embodiments of the present invention are shown. [Figure 4] Examples of audio devices according to several embodiments of the present invention are shown. [Figure 5] Examples of audio devices according to several embodiments of the present invention are shown. [Figure 6] Examples of audio devices according to several embodiments of the present invention are shown. [Figure 7] Some elements of possible processor configurations for implementing elements of an audio device according to some embodiments of the present invention are shown. [Modes for carrying out the invention]

[0083] The following description focuses on embodiments of the present invention applicable to voice capture, such as speech capture for remote conferencing equipment. However, it is understood that this approach is applicable to many other voice signals, voice processing systems, and scenarios for capturing and / or processing voice.

[0084] Figure 1 shows an example of a adaptive audio beamform configuration. For clarity, the explanatory example focuses on a case where the audio device is configured to process two audio input signals, but it will be understood that the example can be easily extended to cases where more input signals are processed.

[0085] The configuration is set up to generate a first set of input signals that include audio signals capturing the sound of the scene from different locations. Specifically, the first set of audio signals is an audio signal generated from a set of microphone signals from a set of microphones located at different locations in the sound environment, such as a microphone array, or more often specifically a linear set of microphones.

[0086] The configuration includes a beamformer 101 configured to receive a set of multiple audio signals from which a single signal is generated. The beamformer 101 requests that the input signals be combined so that a given sound source contribution is constructively combined. Specifically, the beamformer 101 combines signals to perform beamforming for a set of spatial audio signals that capture speech in an environment, such as a set of signals from a microphone array. The beamformer 101 is configured to generate a beamformed output audio signal from a first set of audio signals.

[0087] Specifically, the beamformer 101 is a first set of filters f1 that filters a first set of audio signals. * (ω), f2 * (ω) is included. Each filter is configured to filter one of the audio signals. Furthermore, the filters are adaptive filters that are dynamically adapted to achieve adaptive beamforming. The first set of filters is also referred to as beamformer filters below.

[0088] The audio configuration further includes a feedback circuit 103 which contains a second set of filters that generate a second set of audio signals from the filtering of the beamform output audio signal. The second set of filters is also called a feedback filter. The input signal to the beamformer is also called the beamformer input signal, and the signal generated by the feedback circuit is also called the feedback signal.

[0089] The first and second sets of filters are specifically matched / linked such that for each filter in the first set (i.e., each beamformer filter), there exists a second set of filters (i.e., feedback filters) having a frequency response that is the complex conjugate of the corresponding beamformer filter. Thus, the filters of beamformer 101 and feedback circuit 103 are together fitted such that the filter coefficients provide a complex conjugate frequency response (equivalent to / corresponding to the time-inverted filter impulse response).

[0090] Therefore, the feedback circuit 103 generates a set of feedback signals from the beamform output audio signal, each feedback signal being generated from the beamform output audio signal by time-inverted / complex conjugate filtering with a matching filter.

[0091] The audio configuration further comprises a beamform adapter 105 configured to fit a first set of filters and a second set of filters, namely beamformer filters and feedback filters. The beamform adapter 105 is configured to perform fitting based on a comparison of the first set and the second set of audio signals, specifically on a comparison of beamformer signals and feedback signals. Specifically, a difference scale is determined for each matching / linked beamformer signal and feedback signal (the beamformer filter and feedback filter having complex conjugate frequency responses), and the filters among the matching filters are fitted to reduce / minimize the difference scale.

[0092] The configuration of the beamformer 101, feedback circuit 103, and beamform adapter 105 interacts to provide highly efficient adaptive beamforming, which involves fitting filters to a set of spatial audio signals to form a beam directed towards a sound source, providing the captured audio in the beam formed by the beamformed audio signal. Filter fitting to minimize the difference scale allows the beamformer configuration to detect and track a given sound source. Specifically, in the ideal case, each of the beamformer filters concentrates all the energy from a given sound source captured in one of the beamformer signals into a single signal value. In the ideal case, this is achieved by the beamformer filter having an impulse response that is the inverse of the time of the acoustic impulse response to the microphone capturing the signal from the sound source. In the ideal case, this is achieved for all signals / filters, resulting in all the captured energy from a particular sound source being combined into a single value by summing the outputs of the beamformer filters, thereby generating a sample value for the beamformed audio signal that maximizes the energy captured from the sound source.

[0093] Furthermore, since each of the feedback filters is the complex conjugate frequency response of the corresponding beamformer filter, this is identical to the acoustic impulse response (signal) from the sound source to the microphone in an ideal case. Therefore, a feedback signal is generated by filtering the beamformer output audio signal with the feedback filter, which reflects the signal captured by the microphone from the sound source as represented by the beamformer output audio signal. In an ideal case where only one sound source exists and the filter is ideally adapted, the difference between the generated feedback signal and the captured microphone signal is zero. In the presence of other sound sources and / or noise, there is a difference reflected in the difference metric. However, for uncorrelated noise / audio, the difference, and thus typically the difference metric, is typically zero on average and thus tends not to prevent efficient adaptation. Therefore, by adapting the filter for a given beamformer signal to minimize the difference metric for the corresponding feedback signal, adaptation towards an optimal beamforming filter can typically be achieved.

[0094] Such a beamforming approach is described in U.S. Patent No. 7,146,012. In the following, a more detailed description and analysis of the configuration are provided. This description will be based on the example of FIG. 1, in which the configuration of the adaptive beamformer, where the microphone signal is filtered by two filters f1 * (ω) and f2 * (ω), consists of a filter section and the output is added to obtain the output signal z(ω). In the update section, the output signal z(ω) is supplied to two adaptive filters f1 * (ω) and f2 * (ω) each having the respective microphone signal as a reference.

[0095] With this configuration, an adaptive beamformer is provided that maximizes the output power under the following constraints: |f1(ω)| 2 +|f2(ω)|2 =1. Specifically, this constraint is implied by the configuration because the conjugates of filters f1(ω) and f2(ω) in the update portion are replicated to the filter portion. The adaptive filter in the update portion is a "normal" unconstrained adaptive filter.

[0096] The constraints implied by the construction can be understood by looking at the optimal solution, i.e., the case where both residual signals are zero. Assume we have an utterance signal s(ω) and transfer functions h1(ω) and h2(ω) from the utterance source to the microphone. The microphone signal is then given by: u1(ω)=h1(ω)s(ω) and u2(ω)=h2(ω)s(ω).

[0097] The output signal z(ω) is then as follows: z(ω)=u1(ω)f1 * (ω)+u2(ω)f2 * (ω)= =s(ω)(h1(ω)f1 * (ω)+h2(ω)f2 * (ω)). In the update section, two equations are obtained after convergence, where ω is omitted for convenience: (h1f1 * +h2f2 * )f1=h1(1) (h1f1 * +h2f2 * )f2=h2(2) (3) follows (1) and (2):

number

number

number

number

number

[0098] More generally, f(ω) = [f1(ω).......f nmics (ω)] T Using matrix notation, Here, f m * (ω) is the beamformer filter corresponding to the m-th microphone, h(ω) = [h1(ω).......h nmics (ω)] T h m (ω) is the transfer function from source s(ω) to the m-th microphone, and for output z(ω): z(ω)=s(ω)h T (ω)f * (ω) (4) This can be obtained.

[0099] In the update section, the equation is given after convergence: z(ω)f(ω)=s(ω)h(ω) (5) It has, This is because (z(ω) and s(ω) are scalars) f(ω)=α(ω)h(ω) (6) Leading to, Here, α(ω) is a complex scaler.

[0100] By using (6) and substituting (4) into (5):

number

[0101] By substituting (7) into (6), the final solution after convergence (8) is given:

number

[0102] Regarding constraints, next: f H (ω)f(ω)=1 This can be obtained.

[0103] To understand that the solution given by Eq. 8 is also the optimal solution for maximizing output power, we look at the expected value of the update after convergence, which consists of the correlation between the residual signals (x1(ω) and (x2(ω)) in Figure 1) and the conjugate of the input signal of the fitted filter:

number

[0104] this is:

number

[0105] z(ω)=u T (ω)f * (ω), therefore z * (ω)=u H By using f(ω):

number

number

number

number

[0106] If the microphone signal contains only speech, then u(ω)=s(ω)h(ω), and then f H (ω)R uu Maximizing (ω)f(ω) corresponds to maximizing the utterance in the output.

[0107] Considering the case where u(ω)=s(ω)h(ω)+v(ω), where v(ω) is uncorrelated noise with equal dispersion across all microphones, that is,

number

number

number

[0108] Now it's v(ω)'s turn: R uu (ω)=R ss (ω)+R nn (ω) Assuming that the correlation noise consists of such noise, R nn (ω) is the noise covariance matrix with non-zero off-diagonal elements.

[0109] f H (ω)R uu The maximization of (ω)f(ω) is as follows: f H (ω)R uu (ω)f(ω)=f H (ω)R ss (ω)f(ω)+f H (ω)R nn (ω)f(ω) To maximize [something].

[0110] This is f H (ω)R uu This means that maximizing (ω)f(ω) does not necessarily lead to a better signal-to-noise ratio, because the choice of f determines not only the amount of speech in the output, but also the amount of noise in the output.

[0111] An approach that provides improved performance in many scenarios is described below. In this approach, the beamforming approach, as described with reference to Figure 1, is further enhanced by the introduction of signal-adapted spatial filtering. Spatial filtering is performed across the signal and is based on decorrelation coefficients / weights determined based on the input audio signal to adapt to the current audio characteristics. As described in more detail below, spatial filtering is placed at different locations in the beamforming loop structure, including as inverse spatial decorrelation filtering in the feedback loop or as direct spatial decorrelation filtering on the input signal used by beamforming. This approach has been found to provide significantly improved performance in many practical scenarios, such as significantly improved source separation. In relation to the above analysis, this approach is particularly effective for noise contribution f H (ω)R nn We achieve that (ω)f(ω) occurs independently of the choice of f.

[0112] Figure 2 shows an audio device that provides improved performance. In the example, the audio device has a beamforming structure described with reference to Figure 1. However, this approach is further enhanced by the fact that this structure operates on signals that have been spatially decorrelated by an adapted spatial decorrelator.

[0113] The audio device in Figure 2 includes a receiver 201 that receives a first set of audio signals and, in certain examples, a set of microphone signals from a set of microphones that capture the audio scene from different locations. The microphones are configured, for example, relatively close to each other in a linear array. For example, the maximum distance between capture points for the audio signals is not more than 1 meter, 50 cm, 25 cm, or in some cases even 10 cm, in many embodiments.

[0114] The input audio signal is received from different sources, including internal or external sources. Below, an embodiment is described in which the receiver 201 is coupled to multiple microphones, such as a linear array microphone, to provide a set of input audio signals in the form of microphone signals.

[0115] However, instead of directly performing the operation shown in Figure 1 on the microphone signal, the audio device in Figure 2 is configured to apply adaptive spatial decorrelation to the audio signal before beamforming and adaptation operations. Spatial decorrelation is adaptive decorrelation that is adapted based on the input audio signal, and therefore the decorrelation is continuously adapted to provide increased decorrelation.

[0116] Specifically, the audio device in Figure 2 includes a spatial decorrelator 205 in the form of a set of spatial (decorrelating) filters that apply a spatial filter to a first set of audio signals. Thus, after filtering by the spatial filter, the first set of audio signals is modified to have increased decorrelation (for at least one sound source) compared to before spatial decorrelation filtering.

[0117] The spatial filter is based on coefficients fitted by a fitting coefficient processor 207, which dynamically fits and updates the filter coefficients used by the set of filters of the decorrelator 205. The spatial filter is therefore configured to have coefficients determined by the coefficient adapter 207.

[0118] Beamforming, decorrelation, and fitting are typically performed in the frequency domain. The receiver 201 includes a segmenter configured to segment a set of input audio signals into time segments. In many embodiments, segmentation is typically fixed segmentation into time segments of fixed and equal duration, such as division into time segments / intervals having fixed durations of, for example, 10 to 20 msec. In some embodiments, segmentation is adaptive so that the segments have variable durations. For example, the input audio signal has a variable sampling rate, and the segments are determined to have a fixed number of samples.

[0119] Segmentation is typically performed on a time-domain sample of the input signal into segments having a given fixed number of samples. For example, in many embodiments, the segmenter 203 is configured to divide the input signal into consecutive segments of, for example, 256 or 512 samples.

[0120] The receiver 201 is configured to generate a frequency bin representation of an input audio signal, and a first set of input signals to be further processed is typically represented in the frequency domain by the frequency bin representation. The audio device is configured to perform frequency domain processing of the frequency domain representation of the input audio signal. The signal representation and processing are based on frequency bins, and therefore the signal is represented by the values ​​of the frequency bins, and these values ​​are processed to generate frequency bin values ​​for the output signal. In many embodiments, the frequency bins have the same size and therefore cover frequency intervals of the same size. However, in other embodiments, the frequency bins have different bandwidths, and for example, perceptually weighted bin frequency intervals are used.

[0121] In some embodiments, the input audio signal is already provided in frequency representation and no further processing or operation is required. However, in some such cases, reconstruction into a suitable segment representation is desired, which may include, for example, aligning the frequency representation into time segments using interpolation between frequency values.

[0122] In other embodiments, a filter bank, such as a quadrature mirror filter (QMF), is applied to a time-domain input signal to generate a frequency bin representation. However, in many embodiments, a discrete Fourier transform (DFT), specifically a fast Fourier transform (FFT), is applied to generate the frequency representation.

[0123] In the audio device shown in Figure 2, the spatial filter 205 specifically processes the audio signal in the frequency domain. In the following description, the first set of audio signals is also referred to as the input audio signal (to the spatial filter) before filtering, and the resulting signal is also referred to as the output audio signal (from the spatial filter).

[0124] For each frequency bin, the output frequency bin value is generated from one or more input frequency bin values ​​of one or more input signals, as will be described in more detail below. The output signal is generated such that, for example, one dominant source, the inter-signal correlation with respect to the correlation of the input signals is (typically / on average) reduced.

[0125] A set of spatial filters is configured to filter an input audio signal. The filtering is spatial in that, for a given output signal, the output value is determined from multiple, typically all, input audio signals (for the same time / segment and for the same frequency bin). Spatial filtering is specifically performed on a frequency bin basis, such that the frequency bin values ​​for a given frequency bin of the output signal are generated from the frequency bin values ​​of the input signals for that frequency bin. The filtered / weighted combination is across the signal and not typical time / frequency filtering.

[0126] Specifically, the frequency bin value for a given frequency bin is determined as a weighted combination of the frequency bin values ​​of the input signal for that frequency bin. The combination is specifically an addition, and the frequency bin value is determined as the weighted sum of the frequency bin values ​​of the input signal for that frequency bin. The determination of the bin value for a given frequency bin is determined as a vector multiplication of the weight / coefficient vector of the weighted sum and a vector containing the bin values ​​of the input signal:

number

[0127] If we represent the output bin values ​​for a given frequency bin ω as a vector y(ω), then the determination of the output signal is: y(ω)=W(ω)X(ω) This is determined as follows, where the matrix W(ω) represents the weights / coefficients of the weighted sum for different output signals, and x(ω) is a vector containing the input signal values.

[0128] For example, in a case with only three input and output signals, the output bin value for frequency bin ω is:

number

[0129] The adaptive coefficient processor 207 seeks to adapt a spatial filter to become a spatial decorrelation filter that corresponds to the input signal but generates an output signal with increased signal decorrelation. The output audio signal is generated to have increased spatial decorrelation, with lower inter-speech signal cross-correlation for the output audio signal than for the input audio signal. Specifically, the output signal is generated to have the same combined energy / power as the input signal (or with a given scaling thereof), but with increased decorrelation (reduced correlation) between signals. The output audio signal is generated to contain all the audio / signal components of the input signal in the output audio signal, but with redistribution into different signals to achieve increased decorrelation.

[0130] A decorrelation filter is specifically configured to produce an output signal with lower coherence / normalized correlation than the input signal. Therefore, the output signal of the decorrelation filter has lower coherence / normalized correlation than the coherence of the input signal to the decorrelation filter.

[0131] Many different approaches and algorithms are known for decorrelating a set of signals, and it is understood that many different algorithms are known for adaptively fitting such decorrelation filters. The adaptive coefficient processor 207 provides an adaptive decorrelation operation that determines decorrelation coefficients for a set of adaptive filters applied to a set of audio signals to produce a decorrelated signal.

[0132] Many known adaptive decorrelation approaches are based on a feedback system where an audio output signal is generated and used to generate feedback update parameters for the filter coefficients. In such cases, the adaptive coefficient processor 207 implements the filter and feedback circuit to update and fit the decorrelation coefficients. However, it is understood that other approaches are possible, such as employing a trained artificial neural network that directly receives samples of the input audio signal and generates (or updates) coefficient values ​​without necessarily requiring that a decorrelated output signal is generated.

[0133] In the described example, adapter 207 is specifically configured to determine update values ​​for the weighted combinations that form a set of uncorrelated filters. Specifically, the update values ​​are determined for the matrix W(ω). Adapter 207 then updates the weighted combinations based on the update values.

[0134] The adapter 207 is configured to apply a fitting approach to determine the update value, which allows the output signals of the set of decorrelated filters 205 to represent the audio of the input signals to the set of decorrelated filters 205, but with the output signals typically being more decorrelated than the input signals.

[0135] Adapter 207 is configured to use a specific approach to fitting weights based on the generated output signals. The operation is based on each output audio signal linked to a single input audio signal. Precise linking between the output and input signals is not required, and linking / pairing (in principle, including random) of many different input signals for each output signal is used. However, the processing differs with respect to weights that reflect contributions from the output signal and linked / paired input signals, compared to weights that reflect contributions from the output signal and unlinked / paired input signals. For example, in some embodiments, the weights for linked signals (i.e., for input signals linked to an output signal generated by a weighted combination including the weights) are set to fixed values, are not updated, and / or the weights for linked signals are limited to real-valued weights, while other weights are generally complex-valued.

[0136] Adapter 207 employs a fitting / update approach, where the update value is determined based on a correlation measure between the output bin value for a given output signal and the output bin value for a given (unlinked) input signal and linked output signal, for a given weight representing the contribution of a given unlinked input signal to the bin value for a given output signal. The update value is then applied to modify the given weight in subsequent segments, or the update value for the weight in a given segment is determined according to two output bin values ​​from the previous (typically immediately preceding) segment, where the two output values ​​represent the input and output signals to which the weight is related, respectively.

[0137] The described approach applies to multiple, typically all, weights used when determining output bin values ​​based on input signals that are not typically linked. Regarding weights related to input signals linked to output signals, other considerations are used, such as setting weights to fixed values, as will be explained in more detail later.

[0138] Specifically, the update value is determined by the product of the output bin value for the weight and the complex conjugate of the output bin value linked to the input signal for the weight, or equivalently, by the product of the complex conjugate of the output bin value for the weight and the output bin value linked to the input signal for the weight.

[0139] As a specific example, the updated value for segment k+1 for frequency bin ω is: [y i (k,ω)y j * (k,ω)] By, or equivalently: [yi * (k,ω)y j (k,ω)] It is determined depending on the correlation scale given by, where y i (k,ω) is the output bin value for the output signal i, which is determined based on the weights, and y j (k,ω) is the output bin value for output signal j linked to the input signal from which the contribution is determined (i.e., the input signal bin value multiplied by weights to determine the contribution to the output bin value for signal i).

[0140] [y i (k,ω)y j * The measure (k,ω) (or conjugate value) indicates the correlation of time-domain signals in a given segment. In a particular example, this value is then weighted w i,j (k+1,ω) is updated and used to fit the new value.

[0141] As mentioned above, a set of uncorrelated filters gives the output bin values ​​for the output signal for a given frequency bin ω: y(ω)=W(ω)X(ω) configured to be determined from, where y(ω) is a vector containing the frequency bin values for the output signal for a given frequency bin ω, x(ω) is a vector containing the frequency bin values for the input audio signal for a given frequency bin ω, and W(ω) is a matrix having a row containing the weighted combination weights for the output audio signal.

[0142] In an example, the adapter 207 specifically adapts at least a part of the weights w ij of: w ij (k + 1, ω) = w ij (k, ω) - η(k, ω)[y i (k, ω)y j * (k, ω)] configured to adapt according to, where i is the row index of the matrix W(ω), j is the column index of the matrix W(ω), k is the time segment index, ω represents the frequency bin, and η(k, ω) is a scaling parameter for adapting the adaptation speed. Typically, the adapter 207 is configured to adapt all weights that are not associated with the input signal linked to the output signal (i.e., the "cross-signal" weights).

[0143] In some embodiments, the adapter 207 is configured to adapt the update rate / speed of the weight adaptation. For example, in some embodiments, the adapter is configured to compensate a correlation measure for a given weight depending on the signal level of the output bin value whose contribution is determined by the weight.

[0144] As a specific example, the compensation value [y i (k, ω)y j * (k, ω)] is compensated by the signal level |y i (k, ω)| of the output bin value. The compensation is included, for example, to normalize so that the value of the update step depends less on the signal level of the decorrelated signal generated.

[0145] In many embodiments, such compensation or normalization is specifically performed on a frequency bin basis, i.e., compensation differs at different frequency bins. This improves performance in many scenarios and results in an improved fit of weights, typically producing a decorrelated signal.

[0146] The compensation is incorporated, for example, into the scaling parameter η(k,ω) in the previous update formula. Thus, in many embodiments, the adapter 207 is configured to adapt / change the scaling parameter η(k,ω) differently at different frequency bins.

[0147] In many embodiments, the input signal vector x(ω) and output signal vector y(ω) are configured such that linked signals occupy the same position in their respective vectors; specifically, y1 is linked to x1, y2 to x2, y3 to x3, and so on. In this case, the weights for the linked signals lie on the diagonal of the weight matrix W(k,ω). In many embodiments, the diagonal values ​​are set to fixed real values, for example, a constant value of 1.

[0148] In many embodiments, the weights / space filters / weighted combinations are configured such that the weights for the contribution of the first input signal (not linked to the first output signal) to the first output signal are the complex conjugate of the contribution of the second input signal (linked to the first input signal) to the second output signal (linked to the first input signal). Thus, the two weights for the two pairs of linked input / output signals are complex conjugates.

[0149] In the example of weights for linked input and output signals arranged on the diagonal of the weight matrix W(ω), this results in a Hermitian matrix. In fact, in many embodiments, the weight matrix W(ω) is a Hermitian matrix. Specifically, the coefficients / weights of the weight matrix W(ω) are based on: w ij =w ji* It satisfies the condition.

[0150] As mentioned above, the weights for the contribution of linked input signals to the output signal bin values ​​(corresponding to the diagonal values ​​of the weight matrix W(ω) in certain examples) are treated differently from the weights for unlinked input signals. The weights for linked input signals will also be referred to as linked weights for brevity below, and the weights for unlinked input signals will also be referred to as unlinked weights for brevity below. Therefore, in certain examples, the weight matrix W(ω) is a Hermitian matrix containing linked weights on the diagonal and unlinked weights outside the diagonal.

[0151] In many approaches, fitting unlinked weights is done to reduce the correlation metric. Specifically, each updated value is determined to reduce the correlation metric. As a whole, the fit therefore seeks to reduce the cross-correlation between output signals. However, linked weights are determined differently to ensure that the output signals maintain appropriate speech energy / power / levels. In fact, if linked weights are instead fitted to reduce the autocorrelation of the output signals with respect to the weights, there is a high risk that the fit will converge to a solution where all weights, and therefore the output signals, are essentially zero (which indeed results in the lowest correlation). Furthermore, speech devices are configured to produce signals with low cross-correlation, but not to reduce autocorrelation.

[0152] Therefore, in many embodiments, the linked weights are set to ensure that the output signal is generated to have the desired (combined) energy / power / level.

[0153] In some cases, the adapter 207 is configured to fit the linked weights, and in other cases, the adapter is configured not to fit the linked weights.

[0154] For example, in some embodiments, the linked weights are simply set to fixed, constant values ​​that are not fitted. For example, in many embodiments, the linked weights are specifically set to a constant scalar value such as value 1 (i.e., a unit gain is applied to the linked input signal). For example, the weights on the diagonal of the weight matrix W(ω) are set to 1.

[0155] Therefore, in many embodiments, the weights for the contribution of linked input signal frequency bin values ​​to a given output signal frequency bin value are set to a predetermined value. In many embodiments, this value is kept constant without any modifications.

[0156] Such an approach offers highly efficient performance and provides a very accurate representation of the original sound of the input signal, but it has been found that it results in an overall fit that produces an output signal with a set of output signals that have increased decorrelation.

[0157] In some embodiments, linked weights are also fitted, but in a different way than unlinked weights. In particular, in many embodiments, linked weights are fitted based on the output signal.

[0158] Specifically, in many embodiments, the linked weights for the first input signal and the linked output signal are adapted based on the generated output bin value of the linked audio signal, specifically based on the magnitude of the output bin value.

[0159] Such an approach, for example, allows for the normalization and / or setting of the desired energy level for a signal.

[0160] In many embodiments, linked weights are constrained to be real-valued weights, while unlinked weights are generally complex-valued. In particular, in many embodiments, the weight matrix W(ω) is a Hermitian matrix with real values ​​on the diagonal and complex values ​​outside the diagonal.

[0161] Such an approach offers specific advantageous behavior and fit in many scenarios and embodiments. It has been found to provide relatively low complexity and computational resources while offering highly efficient spatial decorrelation.

[0162] The fitting gradually adjusts the weights to increase the decorrelation between signals. In many embodiments, the fitting is configured to converge toward a suitable weight matrix W(ω) regardless of the initial values, and in fact in some cases the fitting is initialized with random values ​​for the weights.

[0163] However, in many embodiments, the fitting is started with favorable initial values ​​that result in a fit that, for example, produces a faster fit and / or is more likely to converge toward more optimal weights for the uncorrelated signals.

[0164] In particular, in many embodiments, the weight matrix W(ω) is composed of several weights that are zero, but at least some weights are non-zero. In many embodiments, the number of nearly zero weights is two, three, five, or more times the number of weights that are set to non-zero values. This has been found to tend to provide improved fit in many scenarios.

[0165] In many embodiments, in particular, the adapter 207 is configured to initialize the weights, with linked weights typically set to non-zero values, such as predetermined non-zero real numbers, while unlinked weights are set to approximately zero. Thus, in the above example where linked signals are located at the same position in the vector, this results in an initial weight matrix W(ω) having non-zero values ​​on the diagonal and (approximately) zero values ​​outside the diagonal.

[0166] Such initialization offers particularly advantageous performance in many embodiments and scenarios. This reflects the tendency for input signals to be somewhat uncorrelated, due to the fact that audio signals typically represent speech at different locations. Therefore, starting with the assumption that input signals are perfectly correlated is often advantageous and leads to faster and more frequent improved fitting.

[0167] Weights, specifically unlinked weights, are not necessarily exactly zero, but are set to low values ​​close to zero in some embodiments. However, the initial non-zero values ​​are at least 5, 10, 20, or 100 times higher than the initial near-zero values.

[0168] The described approach generates an output signal that represents the same speech as the input signal, but provides a highly efficient adapted spatial decorrelator with increased decorrelation. This approach has been found to provide highly efficient adaptation for a wide range of scenarios and many different acoustic environments, as well as for many different sound sources. For example, it has been found to provide highly efficient decorrelation of speaker signals in environments with multiple speakers.

[0169] The fitting approach is even more computationally efficient, allowing for local and individual fitting of individual weights based on only two signals closely related to the weights (specifically, only two frequency bin values), but this process still yields an efficient and often greatly optimized global optimization of the spatial filtering, specifically the weight matrix W(ω). Local fitting has been found to lead to a very favorable global optimization in many embodiments.

[0170] A specific advantage of this approach is that it is used to decorrelate convolutive mixtures, and is not limited to decorrelating only instantaneous mixtures. For convolutive mixtures, the full impulse response determines how signals from different sources combine in the microphone (i.e., delay / timing characteristics are prominent), whereas for instantaneous mixtures, the scalar representation is sufficient to determine how the sources combine in the microphone (i.e., delay / timing characteristics are not prominent). By converting convolutive mixtures to the frequency domain, the mixture can be considered as a complex-valued instantaneous mixture per frequency bin.

[0171] The adaptive coefficient processor 207 thus determines coefficients for the set of spatial decorrelation filters of the decorrelator 205, thereby modifying the first set of audio signals such that they represent the same audio, but with an increase in decorrelation for at least one of the sound sources. Such decorrelation of the audio signals to which the beamforming approach described above is then applied significantly improves the overall performance in many scenarios. In fact, this results in improved separation and selection of specific sound sources, such as a particular speaker, in many scenarios. Thus, counterintuitively, decorrelation of the signals provides improved beamforming, although it is essentially based on exploiting and adapting the correlation between audio signals from different positions to spatially form a beam towards the source for which beamforming is desired, in order to extract / separate the sound source. In fact, decorrelation essentially breaks the link between the audio signal and a particular location in the audio scene typically exploited by the beamforming operation. However, the inventors have understood that, nevertheless, decorrelation provides highly advantageous effects and improved performance in many scenarios. For example, in the presence of a strong noise source, this approach facilitates and / or improves the extraction / isolation of specific desired sound sources, specifically speakers.

[0172] The above specific adaptation provides a highly advantageous approach in many embodiments. This typically results in a spatial decorrelation filter that provides a very accurate adaptation, albeit with low complexity, to generate a highly decorrelated signal. In particular, this enables the local adaptation of individual weights / filter coefficients to result in a very efficient global decorrelation of the first set of audio signals.

[0173] However, in other embodiments, it is understood that other approaches for adapting the spatial filter / decorrelating filter are used. For example, in some embodiments, the adaptive coefficient processor 207 includes a neural network based on input samples and is configured to generate updated values to modify the filter coefficients. For example, for each segment, the frequency bin values for all the audio signals are fed to a trained neural network that generates updated values as outputs for each weight. Each weight is then updated by this updated value. The trained network is trained, for example, by training data that includes a number of different frequency bin values and the associated updated values determined manually to modify towards increased decorrelation.

[0174] As another example, the adaptive coefficient processor 207 comprises a location processor configured to generate updated values to modify the filter coefficients based on visual cues related to (varying) locations. As another example, the adapter 207 determines the cross - correlation matrix of the input signal and calculates its eigenvalue decomposition. The eigenvectors and eigenvalues can be used to construct a decorrelating matrix.

[0175] In the example of FIG. 2, the decorrelating and spatial decorrelating filters are applied directly to the first set of audio signals before the first set of audio signals is fed to both the beamformer 101 and the beamformer adapter 105. However, other approaches are used that adapt the operation based on applying spatial / cross - signal filtering using coefficients derived from the determined decorrelating coefficients.

[0176] Specifically, Figure 3 shows an example in which the audio device is configured to perform spatial filtering of a first set of audio signals before the first set of audio signals is filtered by a first set of filters; that is, beamforming is based on a first set of audio signals after the first set of audio signals has been filtered by a set of spatial filters 301. However, in this example, the beamform adapter 105 receives the first set of audio signals before any filtering by a subset of filters 301; that is, a second set of filters is applied only to the beamformer path and not to the fitted path.

[0177] The coefficients for this set of spatial filters 301 are determined from the decorrelation coefficients determined by the coefficient processor 207. In fact, in some embodiments, the approach described with reference to the configuration in Figure 2 is applied directly, and the resulting second set of filters is applied (only) to the signal in the beamform path.

[0178] However, in many embodiments, the set of spatial filters 301 in the configuration example of Figure 3 differs from the set of spatial filters 201 in the configuration example of Figure 2. In particular, in many embodiments, the set of spatial filters is modified to correspond to a cascade of two sets of uncorrelated filters, as determined by the adaptive coefficient processor 207. Thus, the filter coefficients of the set of spatial filters 301 have coefficients that match the coefficients of the spatial filters, which are a cascade of two sets of uncorrelated filters. This can typically be considered equivalent to a double / repeated filtering of a first set of audio signals by a set of uncorrelated filters determined by the adaptive coefficient processor 207.

[0179] In particular, the adaptive coefficient processor 207, as described above: y(ω)=W(ω)X(ω) Determine the weights / coefficients of the weight matrix W(ω) for a decorrelating filter capable of decorrelating a first set of audio signals according to the following.

[0180] The set of spatial filters 301 is in this case generated to correspond to a cascaded application of two such filters, i.e.: y(ω) = W(ω)W(ω)x(ω) = W 2 (ω)x(ω) configured to correspond to.

[0181] Thus, in this example, the adaptive coefficient processor 207 proceeds to determine the appropriate coefficients for a set of spatial decorrelating filters that perform an adaptation as described for the example of FIG. 2 to decorrelate a first set of audio signals. In some cases, such filters are then applied by the set of spatial filters 301 in FIG. 3. However, in many embodiments, the determined decorrelating filter is not used directly, but rather the coefficients for the set of spatial filters 301 are determined from the determined coefficients for the decorrelating filter. In such cases, the adaptive coefficient processor 207 implements the decorrelating filter as part of the coefficient determination / adaptation and applies this to the first set of signals (e.g., to determine updated values from [y i (k,ω)y j * (k,ω)]). However, in the example of a particular example, such filtered signals are only used for the adaptation process and are not further used in beamforming / processing. Instead, the set of spatial filters 301 applied is generated from the decorrelating coefficients / weight matrix W(ω), specifically W 2 (ω).

[0182] This approach provides an improved overall performance including an improved beamforming experience. Indeed, this is such that the set of spatial filters 301 in FIG. 3 is W 2For an example where the coefficients are set to correspond to , it can then be shown that this produces the same performance, results, and output signals as the example in Figure 2. It can also be shown that these approaches produce the same optimal solution.

[0183] Another possible example of applying spatial filtering by a set of spatial filters determined from decorrelation coefficients determined by the adaptive coefficient processor 207 is shown in Figure 4. In this example, a first set of audio signals is not filtered by the set of spatial filters, while a second set of audio signals generated by the feedback circuit 103 is filtered by the set of spatial filters 401. Thus, in this example, the feedback signal, rather than the beamform signal, is filtered by the set of spatial filters.

[0184] Furthermore, the set of spatial filters is determined as a set of inverse spatial decorrelation filters, where each of the inverse spatial decorrelation filters includes the inverse filtering of the determined decorrelation filters.

[0185] For example, if a spatially decorrelated filter is represented by a weight matrix W(ω), then the inverse spatially decorrelated filter is represented by the weight matrix W -1 (ω) is determined by the set of spatial filters applied to the second set of audio signals. -1 This includes filtering corresponding to (ω), where W(ω) is the spatial decorrelation determined by the fitted coefficient processor 207.

[0186] In many embodiments, the set of spatial filters applied to a second set of audio signals is determined to have filter coefficients corresponding to the coefficients of a spatial filter, which is a cascade of two sets of spatial inverse filters, each of which is an inverse filter of a spatial decorrelation filter determined by the adaptive coefficient processor 207.

[0187] In particular, the adaptive coefficient processor 207 is therefore: y(ω)=W(ω)X(ω) The weights / coefficients of the weight matrix W(ω) are determined for a decorrelation filter that can decorrelate the first set of audio signals according to the following.

[0188] In this case, the set of spatial filters 301 is generated to correspond to the cascading application of two such filters, i.e., this is: y(ω)=W -1 (ω)W -1 (ω)s(ω)=W -2 (ω)s(ω) It is configured to correspond to, where s(ω) represents a second set of audio signals generated by the feedback circuit 103.

[0189] Therefore, in the example, the adaptive coefficient processor 207 proceeds to determine appropriate coefficients for a spatial decorrelation filter that decorrelates a first set of audio signals by performing a fit as described for the example in Figure 2. The determined decorrelation filter coefficients are not used directly, but rather the coefficients for a set of spatial filters 401 are determined from the determined coefficients. In such a case, the adaptive coefficient processor 207 performs the decorrelation filter as part of the coefficient determination / fitting, and this is (e.g., [y i (k,ω)y j * A first set of signals (from (k,ω)] is applied to determine the update value. However, in certain examples, such filtered signals are used only for the fitting process and not further for beamforming / processing.

[0190] This approach provides improved overall performance, including an enhanced beamforming experience. In fact, this is the set of spatial filters 401 shown in Figure 4. -2 For the example where the coefficient is set to the corresponding value, it is possible to show that this then produces the same performance and operation as the example in Figure 2, and specifically that these produce the same optimal solution.

[0191] The audio device implements a highly adaptive approach that is particularly suitable for adapting to extract a specific sound source, such as a desired speaker. This approach is particularly advantageous, for example, when there are strong noise sources in the captured audio environment. The adaptive audio device has different adaptations for each of a set of spatial filters (through adaptation of decorrelation filters / decorrelation coefficients) and adaptation of a beamformer filter. The adaptive coefficient processor 207 and the beamformer adapter 105 cooperate multiplicatively to provide a very advantageous and high-performance adaptation to the current audio characteristics of the scene.

[0192] In many embodiments, the audio device is configured to adapt different adaptations at different times. Thus, the adaptive coefficient processor 207 and the beamformer adapter 105 are controlled to be adapted during different time periods, and in many embodiments these do not overlap. Thus, at a given time, either the adaptive coefficient processor 207 or the beamformer adapter 105 is being adapted, but not both (although for some periods neither is being adapted).

[0193] For example, as shown in FIG. 5, the audio device of FIG. 2 (or the corresponding one of FIGS. 3 or 4) includes an audio detector 501 configured to detect when the sound source is active and when the sound source is not active. The audio detector 501 specifically divides time into a set of active time intervals during which, specifically, the sound source is active, and a set of non-active time intervals during which, specifically, the sound source is not active. Specifically, the sound source is a desired sound source, such as a desired speaker.

[0194] In some embodiments, low-complexity speech detection is used to determine whether a source is active or inactive. For example, a speech device is used to extract the dominant sound source in an environment, such as the most powerful speaker. For instance, in a typical teleconferencing application where only one person is speaking at the time, the speech device is used to extract the current speaker's speech from background or ambient noise.

[0195] In such applications, for example, the speech detector 501 simply detects sound source activity based on a level / energy exceeding a given threshold. If the signal level exceeds a given threshold (which is, for example, a dynamic threshold set based on a longer-term averaged signal level), the speech detector 501 considers the sound source / speaker to be active; otherwise, it considers the sound source / speaker to be inactive. Thus, the speech detector 501 divides time into active time intervals (due to the sound source being active) if the captured speech exceeds the threshold, and inactive time intervals (due to the sound source being inactive) if the signal level falls below the threshold.

[0196] In many embodiments, more complex detection is used, and the technique is actually used to separate different sound sources. For example, in many embodiments, detection involves considering whether a sound source has desired / expected characteristics that match a particular sound source. For example, the speech detector 501 distinguishes speech from other types of speech based on an evaluation of whether the captured speech has characteristics that match speech.

[0197] Many different techniques and algorithms are known for utterance / voice activity detection, more generally for detecting that a sound source is active, and it is understood that any suitable approach may be used without prejudice to the present invention.

[0198] In several embodiments, trained artificial neural networks have been found to be capable of performing effective speech / voice activity detection (and / or noise detection). The artificial neural networks are trained on speech and all types of non-speech, providing an indication for each frame whether it is noise or speech.

[0199] In many embodiments, detection is relatively fast, and the speech detector 501 is configured to designate relatively short time intervals as active or inactive time intervals. For example, in some embodiments, the speech detector 501 is configured to detect silent pauses / intervals during periods of normal speech and to designate such intervals as inactive time intervals. For example, time intervals shorter than, say, 5 msec, 10 msec, or 20 msec are identified / designated as active or inactive time intervals in some embodiments.

[0200] In many embodiments, the speech detector 501 directly detects activity about a sound source based on a first set of received microphone signals / speech signals / beamform signals. For example, the sum of the energy of the captured speech is determined and compared to a threshold. However, in other embodiments, other signals, such as a second set of speech signals, are considered. In many embodiments, the speech detector 501 bases activity detection on the generated beamform output speech signal. For example, level detection or speech detection is applied directly to the beamform output speech signal. This provides improved performance in many embodiments, as the output signal is generated to focus on a desired signal, such as a desired speaker.

[0201] The sound detector 501 is configured to control the fit based on sound source activity detection, specifically the fit of the decorrelation coefficient and / or beamform filter coefficient differs between inactive and active time intervals.

[0202] In fact, in some embodiments, the fitting of decorrelation coefficients is not performed during active time intervals, but only during inactive time intervals. The speech detector 501 specifically requests to detect whether a desired or preferred sound source / speaker is active. The speech detector 501 further controls the fitting coefficient processor 207 to fit decorrelation coefficients only if this sound source is not active. Thus, the fitting coefficient processor 207 is configured to fit to provide a decorrelation filter that requests the decorrelation of unwanted speech. Thus, the goal is to fit the decorrelation filter to reduce the decorrelation of unwanted speech, rather than updating it to decorrelate the entire captured speech. Such an approach provides highly efficient and improved operation and performance in many situations. The decorrelation approach for unwanted speech allows beamforming to provide efficient and high-performance extraction of desired signals, while enabling improved beamforming with improved exclusion of unwanted speech by beamforming.

[0203] In some embodiments, the adaptive coefficient processor 207 is configured to adapt the decorrelation coefficients in both active and inactive time intervals, but the adaptation rate is higher during inactive time intervals than during active time intervals. Therefore, instead of adapting the decorrelation coefficients only during inactive time intervals, some adaptation also exists during active time intervals, but this adaptation is slower, for example, 2, 5, 10, or 100 times slower. This is advantageous in some scenarios, such as when the active time interval is much longer than the inactive time interval. In some embodiments, the update rate during active and / or inactive time intervals is dynamically adapted, for example, depending on the characteristics of the time interval or input signal. For example, the adaptive coefficient processor 207 switches between updating only during inactive time intervals and updating (at a lower rate) during active time intervals as well if the duration of the active time interval exceeds a given duration.

[0204] In many embodiments, beamforming filter adaptation is not performed during inactive time intervals, but only during active time intervals. Specifically, the speech detector 501 detects whether a desired or preferred sound source / speaker is active and requests that the beamforming adapter 105 be controlled to adapt the beamforming filter only if this sound source is active. Thus, the beamforming adapter 105 is specifically configured to adapt the beamforming so that it specifically (effectively) forms a beam toward the desired sound source.

[0205] In some embodiments, the beamform adapter 105 is configured to adapt beamformer coefficients in both active and inactive time intervals, but the adaptation rate is higher during active time intervals than during inactive time intervals. Thus, instead of adapting the beamform filter only during active time intervals, some adaptation also exists during inactive time intervals, but this adaptation is slower, for example, 2, 5, 10, or 100 times slower. This is advantageous in some scenarios, such as when the audio detector 501 is configured to provide detection with a very low risk of false detection of an active sound source, and the risk of not detecting an active sound source is relatively high. In such cases, adaptation is still desirable during inactive time intervals, but the update rate is (typically) very low. Thus, a low update rate is used during time intervals when the desired source is active or inactive, and a high update rate is used during time intervals when the desired source is almost certainly present.

[0206] In many embodiments, the audio device is configured to adapt the decorrelation coefficient only during periods of inactive time intervals and the beamforming filter only during periods of active time intervals. This provides highly efficient performance and, in particular, significantly improved extraction / separation of the desired sound source / speaker in the presence of noise / unwanted speech, such as dominant and correlated noise.

[0207] In many practical applications, such as numerous speech capture and processing applications, both beamforming and decorrelation adaptation controls are crucial for the effective operation of speech devices. It is desirable that the beamformer adapts to decorrelation for speech that is not always active by its nature, and for portions consisting solely of noise. In use cases where noise is continuously present, the beamformer can only adapt if noise is also present. If the noise is uncorrelated, or if it is decorrelated by decorrelation, the noise does not affect the adaptation if it is equivariant. If the noise is not equivariant, the beamformer diverges during periods of noise only, but if the near end becomes active, the beamformer can often be quickly (almost imperceptibly) adjusted to the desired speaker. Depending on the application, further adaptation controls may not be used to allow the beamformer to quickly find new sources, or a voice activity detector may be used, which can range from simple energy-based detectors to more sophisticated detectors that also use pitch-based speech characteristics.

[0208] Detectors for noise are typically made more conservative, thereby remaining inactive when speech is also present, and initiating speech decorrelation. Depending on the type and level of noise, energy-based detectors that distinguish (static) noise from speech, or more advanced detectors such as neural network classifiers trained to distinguish noise from speech, can be used.

[0209] The above explanation focused on a scenario where the active time interval is the active time of the desired / speech source, and the inactive time interval is the active time interval of the unwanted source / noise / interference. In this case, the fitting rate for filters is higher during the active time interval than during the inactive time interval; specifically, the fitting rate for filters is performed only during the active time interval and not during the inactive time interval. Similarly, the fitting rate for decorrelation coefficients is higher during the inactive time interval than during the active time interval; specifically, the fitting rate for filters is performed only during the inactive time interval and not during the active time interval.

[0210] Such time intervals are determined directly, for example, by the use of a speech detector that detects and determines the active time intervals of utterances that are considered as active time intervals. The remaining time intervals, i.e., active time intervals of non-utterances, are considered as inactive time intervals.

[0211] In some embodiments, detection of unwanted audio characteristics is used, specifically such as the detection of activity of interfering objects or noise sources. For example, detection of activity of unwanted sound sources, such as music detectors, silence detectors, or specific noise detectors, may be used. In such cases, the time intervals detected are considered as inactive time intervals. The remaining time intervals are considered as active time intervals.

[0212] Instead, if we consider that the active time interval corresponds to the active time interval where the unwanted sound source is not desired, and the inactive time interval corresponds to the inactive time interval where the unwanted sound source is not desired, then the aforementioned fitting is inverted (the filter fitting rate becomes higher during the inactive time interval, and the decorrelation coefficient fitting rate becomes higher during the active time interval; specifically, in many embodiments, the filter fitting occurs only during the inactive time interval, and the decorrelation coefficient fitting occurs only during the active time interval).

[0213] In the following, it will be demonstrated that the specific configurations in Figures 2, 3, and 4 can provide the same solution. For brevity, these configurations and solutions will be referred to as A, B, and C, respectively. For the different solutions, the optimal filter coefficients and corresponding outputs will be derived.

[0214] In configuration A (Figure 2), the set of spatial filters is placed immediately after the microphone, completely outside the beamformer. This has the advantage that nothing needs to be changed to the beamformer algorithm, and this is still constrained by only the input being different f H (ω)f(ω)=1

[0215] A suitable decorrelator is one as described above. This decorrelator is the noise covariance matrix R nn (ω) : W(ω)R nn (ω)W H (ω)=Λ(ω) (Eq.9) Convert to this, where Λ(ω)=diag(λ1(ω)).....λ Nmics )(ω)) is a diagonal matrix, and W(ω) is the uncorrelated matrix of nmics x nmics in SDC. H (ω)R uu Instead of maximizing (ω)f(ω), the beamformer now uses f H (ω)W(ω)R uu (ω)W H Maximize (ω)f(ω), and this is fH (ω)W(ω)R ss (ω)W H (ω)f(ω)+f H (ω)W(ω)R nn (ω)W H (ω)f(ω) =f H (ω)W(ω)R ss (ω)W H f(ω)+f H (ω)Λ(ω)f(ω) It can be written as follows:

[0216] If Λ(ω) is a scaled version of the identity matrix and can be denoted as β(ω)I, then f H (ω)Λ(ω)f(ω) is converted to β(ω), and the noise contribution is independent of the choice of f(ω), f H (ω)W(ω)R uu (ω)W H Maximizing (ω)f(ω) leads to maximizing the signal-to-noise ratio in the output.

[0217] The choice of β(ω) in the decorrelator does not affect the SNR, but it does affect the speech level at the output. If we select β(ω)=1 and assume that the noise level at the input doubles, all coefficients in W(ω) are scaled by 0.5, and therefore the speech level at the output of the SDC is also reduced. The SDC continues to update (per frequency bin) according to the input level. As a result, the beamformer also continues to update.

number

[0218] The optimal coefficient for the beamformer can be found by using equation (8) instead of h(ω) in the output of the spatial decorrelator: W(ω)h(ω).

[0219] W(ω) Hermit, therefore W H Using a decorrelator with (ω)=W(ω), for the optimal solution:

number

[0220] The output of the combination is:

number

number

[0221] Substituting (12) and (11) into (10), the output is:

number

[0222] Using homospersity and uncorrelated noise in the input, diag(W(ω))=I and therefore W 2Note that by using a decorrelator with (ω)=W(ω)=I, the same solution can be obtained using only a beamformer as in the case without a decorrelator.

[0223] For configuration B (shown in Figure 3), decorrelation is performed using the noise covariance matrix R nn (ω) : W(ω)R nn (ω)W H (ω)=Λ(ω) It should be noted that it is possible to convert it to [a different format].

[0224] W(ω) and R nn (ω) is positive definite, and its inverses exist: W -1 (ω)W(ω)R nn (ω)W H (ω)(W H (ω)) -1 =W -1 (ω)Λ(ω)(W H (ω)) -1 R nn (ω)=W -1 (ω)Λ(ω)(W H (ω)) -1 It can be written as follows: Λ(ω) is a scaled version of the identity matrix,

number

number

number

[0225] Next, we will prove that this combination maximizes the speech-to-noise ratio in the output. First, we derive the converged filter coefficients, as we did previously for the beamformer.

[0226] In the output of the filtering section: z(ω)=f H (ω)W 2 (ω)u(ω)=s(ω)f H (ω)W 2 (ω)h(ω) (Eq.13) It is possessed.

[0227] In the update section: z(ω)f(ω)=s(ω)h(ω) (Eq14) It is possessed. Since z(ω) and s(ω) are scalars: f(ω)=α(ω)h(ω) It can be written as follows. By substituting Eq13 into Eq14:

number

number

number

[0228] Next, the expected value of updates after convergence:

number

[0229] Next, regarding the update formula (Eq. 15):

number

number

number

[0230] Equation 16 shows that after convergence f(ω), the eigenvectors and matrix R uu (ω)W 2 The corresponding eigenvalue of (ω)

number

number

[0231] f H (ω)W 2 By multiplying by (ω):

number

[0232] Next time: R uu (ω)=R ss (ω)+R nn Considering (ω), here the adapted decorrelator is adapted to the noise source. Next, the output power

number

number

[0233] The noise contribution is independent of the selection of f(ω), and therefore maximizing the FSB output does not necessarily correspond to maximizing the speech-to-noise ratio in the output, but rather to maximizing the speech output itself.

[0234] The optimal solution after convergence for solutions A and B is: f A (ω)=W(ω)f B (ω) It is associated with.

[0235] This means that the outputs for both solutions will be equal.

[0236] For solution C in Figure 4, the matrix block is placed after the filter in the feedback circuit 103. Here the constraint is f H(ω)Γ(ω)f(ω) = 1, where Γ(ω) is the coherence matrix, which is equal to the normalized covariance matrix. Since it is desirable to use the decorrelator again,

number

[0237] Therefore, W -2 (ω) is equal to the normalized covariance matrix, W -2 (ω) can be used instead of Γ(ω), and the constraint is f H (ω)W -2 (ω)f(ω)=1.

[0238] As before, the optimal solution for the filter coefficient f(ω) is found by considering the signal flow in the filtering and updating sections.

[0239] In the filtering section: z(ω)=s(ω)f H (ω)h(ω) (Eq.17) It is provided, and in the updated section: z(ω)W -2 (ω)f(ω)=s(ω)h(ω) (Eq.18) f(ω)=α(ω)W 2 (ω)h(ω) (Eq.19) We have such that z(ω) and s(ω) are complex scalers.

[0240] By using (17) and substituting (19) into (18):

number

number

number

number

number

number

number

number

[0241] After the convergence f(ω), the eigenvectors and matrix W are shown. 2 (ω)R uu The corresponding eigenvalue of (ω)

number

number

[0242] The left and right sides of Eq. 20 are given by f H By multiplying by (ω):

number

[0243] Next time: R uu (ω)=R ss (ω)+R nn Considering (ω), here the adapted decorrelator is adapted to the noise source. Next, the output power

number

number

[0244] Maximizing the FSB output does not necessarily correspond to maximizing the speech-to-noise ratio in the output itself, but rather to maximizing the speech-to-noise ratio in the output.

[0245] The optimal solution after convergence for solutions C and B is f C (ω)=W 2 (ω)f B It is associated with (ω).

[0246] This means that the output for both solutions will be equal. Therefore, different constructions / solutions:

[0247] [Table 1] This is summarized by:

[0248] As can be seen, by appropriately selecting coefficients, it is possible to generate exactly the same (optimal) beamform output audio signal.

[0249] However, as can be seen, the coefficients selected for the set of spatial filters to achieve this output will differ.

[0250] The choice of approach used depends on the preferences and requirements of each individual embodiment.

[0251] The advantage of solution A is that the interaction between the decorrelator and the beamformer is reduced in the sense that synergistic effects can be achieved through reduced modifications of the beamformer, feedback circuit, and beamform adapter.

[0252] The advantage of solution B is that, if it is desired to calculate the direction of arrival (DOA) of the sound source using beamformer coefficients, this is easier to do for this configuration.

[0253] For example, an approach to calculating DOA is described in U.S. Patent No. 6,774,934 for a “normal” beamformer with a speech source where uncorrelated equivalent noise may be added. For this use case, the optimal coefficients are:

number

[0254] Part of the procedure is to calculate the cross-power spectrum for at least one pair of filter coefficients, e.g., f1(ω) and f2(ω). This allows:

number

[0255] What is important is the phase difference, not the amplitude. For solution B, it is possible to use the beamformer coefficient directly, but for methods A and C, the coefficient is W -1 (ω) and W -2 This means that each must be multiplied by (ω) beforehand.

[0256] The voice device is specifically implemented by one or more appropriately programmed processors. Different functional blocks are implemented by different separate processors, and / or, for example, by the same processor. Examples of appropriate processors are provided below.

[0257] Figure 6 shows an example of an audio device that generates a beamform output audio signal. The audio device applies the principles described above and includes more detailed and specific embodiments shown in Figures 2 to 5, which are described with reference to them. Accordingly, corresponding reference numerals are used.

[0258] The audio device includes a receiver 201 that receives a first set of audio signals, including audio signals that capture the sound of a scene from different locations. The audio device further includes a combiner 603 that combines the output signals of an audio beamformer 101 and the first set of filters 603, each of which filters one of the first set of audio signals, to generate a beamformed output audio signal.

[0259] Furthermore, the audio device includes a feedback circuit 103 which includes a second set of filters. This second set of filters generates a second set of audio signals by filtering the beamform output audio signals generated by the audio beamformer 101. In addition, the first and second sets of filters are closely aligned by each filter from the second set of filters, which has a frequency response that is the complex conjugate of the filters in the first set of filters 601.

[0260] The audio device further comprises a beamform adapter 105 that dynamically updates / adapts a first set of filters 601 and therefore a second set of filters in response to a comparison between a first set of audio signals and a second set of audio signals.

[0261] As mentioned above, different approaches to updating and adapting filters are used in different embodiments, depending on the specific configuration and operation.

[0262] The audio device further comprises an adaptive coefficient processor 207 that determines decorrelation coefficients for a set of spatial decorrelation filters capable of generating a decorrelated output signal from a first set of audio signals. The adaptive coefficient processor 207 is configured to dynamically adapt the decorrelation filters based on updated values ​​determined from the first set of audio signals. Thus, the adaptive coefficient processor 207 dynamically determines the decorrelation coefficients, which results in the spatial decorrelation filters generating a decorrelated version of the first set of signals.

[0263] The audio device further includes a first set of spatial filters 205, 301, 401 that modify a first set and / or second set of audio signals (i.e., those received by receiver 201 or generated by feedback circuit 103) by applying spatial filtering to a first set and / or second set of audio signals. As previously stated, spatial filtering is applied to different signals in different embodiments, and embodiments in particular Figures 2 to 5 disclose different approaches to the application of filtering to different sets of signals, i.e., microphone signals and / or feedback signals. Thus, the first set of spatial filters 205, 301, 401 apply modification to the first set and / or second set of audio signals by applying filtering to these signals. A filter from the first set of spatial filters 205, 301, 401 is applied to each signal in the first set and / or second set of audio signals, and therefore modification to each signal in the first set and / or second set of audio signals is performed by one of the filters from the first set of spatial filters 205, 301, 401.

[0264] The coefficients of the first set of spatial filters 205, 301, and 401 are determined from the decorrelation coefficients determined by the fitted coefficient processor 207. Thus, the coefficients of the first set of spatial filters 205, 301, and 401 depend on the decorrelation of the input signals and are therefore fitted based on the first set of audio signals (input signals) and the correlations between them (because this determines the appropriate / necessary decorrelation coefficients for decorrelation of the first set of audio signals).

[0265] As mentioned above, the precise determination of the coefficients for the first set of filters from the uncorrelated filters depends on the specific preferences and requirements of a particular embodiment. The described approach is not limited to, and does not depend on, the specific approach used. In particular, it should be noted that the coefficient determination differs fundamentally depending on which exact signals are modified by being filtered by the first set of filters.

[0266] Figure 7 is a block diagram showing an exemplary processor 700 according to an embodiment of the present disclosure. The processor 700 is used to implement one or more processors that implement devices or elements thereof as described above (in particular, one or more artificial neural networks). The processor 700 is any suitable type of processor, including but not limited to a microprocessor, microcontroller, digital signal processor (DSP), field-programmable array (FPGA) programmed to constitute a processor, image processing unit (GPU), application-specific integrated circuit (ASIC) designed to constitute a processor, or a combination thereof.

[0267] The processor 700 includes one or more cores 702. Each core 702 includes one or more arithmetic logic units (ALUs) 704. In some embodiments, the core 702 includes a floating-point logic unit (FLPU) 706 and / or a digital signal processing unit (DSPU) 708 in addition to or instead of the ALU 704.

[0268] The processor 700 includes one or more registers 712 that are communicatively coupled to the core 702. The registers 712 are implemented using dedicated logic gate circuits (e.g., flip-flops) and / or any memory technology. In some embodiments, the registers 712 are implemented using static memory. The registers provide the core 702 with data, instructions, and addresses.

[0269] In some embodiments, the processor 700 includes one or more levels of cache memory 710 communicably coupled to the core 702. The cache memory 710 provides computer-readable instructions for execution to the core 702. The cache memory 710 provides data for processing by the core 702. In some embodiments, computer-readable instructions are provided to the cache memory 710 by local memory, for example, local memory attached to an external bus 716. The cache memory 710 is implemented in any suitable type of cache memory, such as metal-oxide-semiconductor (MOS) memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), and / or any other suitable memory technology.

[0270] The processor 700 includes a controller 714, which controls inputs to the processor 700 from other processors and / or components included in the system, and / or outputs from the processor 700 to other processors and / or components included in the system. The controller 714 controls the data paths in the ALU 704, FPLU 706, and / or DSPU 708. The controller 714 is implemented as one or more state machines, data paths, and / or dedicated control logic. The gates of the controller 714 are implemented as standalone gates, FPGAs, ASICs, or any other suitable technology.

[0271] Registers 712 and cache 710 communicate with controller 714 and core 702 via internal connections 720A, 720B, 720C, and 720D. These internal connections are implemented as buses, multiplexers, crossbar switches, and / or any other suitable connection techniques.

[0272] Inputs and outputs for the processor 700 are provided via a bus 716 comprising one or more conductive wires. The bus 716 is communicatively coupled to one or more components of the processor 700, such as a controller 714, a cache 710, and / or registers 712. The bus 716 is coupled to one or more components of the system.

[0273] Bus 716 is coupled to one or more external memories. The external memory includes read-only memory (ROM) 732. ROM 732 is masked ROM, electronically programmable read-only memory (EPROM), or any other suitable technology. The external memory includes random access memory (RAM) 733. RAM 733 is static RAM, battery-backed static RAM, dynamic RAM (DRAM), or any other suitable technology. The external memory includes electrically erasable programmable read-only memory (EEPROM®) 735. The external memory includes flash memory 734. The external memory includes magnetic storage devices such as disk 736. In some embodiments, the external memory is contained within the system.

[0274] It is understood that the above description refers to embodiments of the invention with reference to different functional circuits, units, and processors for clarity. However, it will become clear that any appropriate distribution of functions between different functional circuits, units, or processors can be used without prejudice to the invention. For example, functions indicated to be performed by separate processors or controllers are performed by the same processor or controller. Therefore, references to specific functional units or circuits should be understood not as indicating a strict logical or physical structure or composition, but only as references to appropriate means for providing the described functions.

[0275] The present invention can be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. The present invention may optionally be implemented at least partially as computer software running on one or more data processors and / or digital signal processors. Elements and components of embodiments of the present invention may be implemented physically, functionally, and logically in any suitable manner. In fact, functionality may be implemented in a single unit, in multiple units, or as part of other functional units. Thus, the present invention may be implemented in a single unit or physically and functionally distributed across different units, circuits, and processors.

[0276] Although the present invention has been described in relation to several embodiments, it is not intended to be limited to any particular form described herein. Rather, the scope of the present invention is limited only by the appended claims. Furthermore, while some features may appear to be described in relation to a particular embodiment, those skilled in the art will recognize that various features of the described embodiments can be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.

[0277] Furthermore, even if described individually, steps of multiple means, elements, circuits, or methods may be performed, for example, by a single circuit, unit, or processor. Moreover, even if individual features are included in different claims, they may be advantageously combined in some cases, and inclusion in different claims does not imply that the combination of features is unfeasible and / or unfavorable. Also, inclusion of a feature within one category of claims does not imply limitation to that category, but rather indicates that the feature is equally applicable to other claim categories where applicable. Furthermore, the order of features in a claim does not imply any particular order in which the features must be operated, and the order of individual steps in a method claim, in particular, does not imply that the steps must be performed in a specific order. Rather, the steps are performed in any appropriate order. Furthermore, singular references do not exclude plural references. Thus, references to "a," "an," "first," "second," etc., do not exclude plural references. Reference numerals in the claims are provided merely as examples for clarity and should never be understood as limiting the scope of the claims.

[0278] Embodiment A receiver (201) configured to receive a first set of audio signals, the first set of audio signals including audio signals that capture the sound of a scene from different locations, A first set of filters configured to filter a first set of audio signals and A combiner configured to combine the outputs of a first set of filters to generate a beamform output audio signal. A sound beamformer (101) equipped with, A feedback circuit (103) includes a second set of filters configured to generate a second set of audio signals from filtering of a beamform output audio signal, wherein each filter in the second set of filters has a frequency response that is the complex conjugate of a filter in the first set of filters, A beamform adapter (105) configured to adapt a first set of filters and a second set of filters in accordance with a comparison of a first set of audio signals and a second set of audio signals, A adaptive coefficient processor (207) configured to determine decorrelation coefficients for a set of spatial decorrelation filters that generate a decorrelated output signal from a first set of audio signals, wherein the adaptive coefficient processor (207) is configured to adapt the decorrelation coefficients according to updated values ​​determined from the first set of audio signals, A first set of spatial filters (205, 301, 401) configured to apply a first spatial filtering to at least one of a first set of audio signals and a second set of audio signals, wherein the first set of spatial filters (205, 301, 401) has coefficients determined from decorrelation coefficients and A voice device equipped with the following features.

[0279] A method for generating an output audio signal, the method being: A receiving step of receiving a first set of audio signals, wherein the first set of audio signals includes audio signals that capture the sound of a scene from different locations, Audio beamformer (101): The first set of filters filters a first set of audio signals, The combiner combines the outputs of a first set of filters to generate an output audio signal. The steps involve generating an output signal and A step of generating a second set of audio signals from filtering an output audio signal, wherein each filter in the second set of filters has a frequency response that is the complex conjugate of the filter in the first set of filters. A step of adapting a first set of filters and a second set of filters in accordance with a comparison of a first set of audio signals and a second set of audio signals, A step of determining decorrelation coefficients for a set of spatial decorrelation filters that generate a decorrelation output signal from a first set of audio signals, including fitting the decorrelation coefficients according to updated values ​​determined from a first set of audio signals, A method comprising the step of applying a first set of spatial filters (205, 301, 401) to at least one of a first set of audio signals and a second set of audio signals, wherein the first set of spatial filters (205, 301, 401) has coefficients determined from decorrelation coefficients.

[0280] The following subclaims apply equally to the embodiments described above.

Claims

1. A receiver that receives a first set of audio signals, wherein the first set of audio signals includes audio signals that capture the sound of a scene from different locations, A first set of filters for filtering the first set of audio signals and A combiner that generates a beamform output audio signal by combining the outputs of the first set of filters. An audio beamformer equipped with, A feedback circuit including a second set of filters that generates a second set of audio signals from the filtering of the beamform output audio signal, wherein each filter in the second set of filters has a frequency response that is the complex conjugate of the filter in the first set of filters, A beamform adapter that adjusts the first set of filters and the second set of filters according to a comparison between the first set of audio signals and the second set of audio signals, A adaptive coefficient processor that determines decorrelation coefficients for a set of spatial decorrelation filters that generate a decorrelated output signal from a first set of audio signals, wherein the adaptive coefficient processor adapts the decorrelation coefficients according to updated values ​​determined from the first set of audio signals. A first set of spatial filters that modifies at least one of the first set of audio signals and the second set of audio signals by applying a first spatial filtering to at least one of the first set of audio signals and the second set of audio signals, wherein the first set of spatial filters has coefficients determined from the decorrelation coefficients. A voice device equipped with the following features.

2. The audio device according to claim 1, further comprising a sound detector that determines a set of active time intervals in which the sound source is active and a set of inactive time intervals in which the sound source is not active, wherein at least one of the first set of filters and the second set of filters and the decorrelation coefficient differs for the set of active time intervals and the set of inactive time intervals.

3. The audio device according to claim 2, wherein the beamform adapter adapts the first set of filters and the second set of filters at a higher adaptation rate during the period of the active time interval than during the period of the inactive time interval.

4. The audio device according to claim 3, wherein the beamform adapter adapts the first set of filters and the second set of filters for a period of time of only one set of time intervals from the set of active time intervals and the set of inactive time intervals.

5. The audio device according to any one of claims 2 to 4, wherein the fitting coefficient processor fits the decorrelation coefficients at a higher fitting rate during the period of the inactive time interval than during the period of the active time interval.

6. The audio device according to claim 5, wherein the adapting coefficient processor adapts the decorrelation coefficients for a period of time only one of the set of time intervals from the set of inactive time intervals and the set of active time intervals.

7. The audio device according to any one of claims 1 to 6, wherein the first set of spatial filters filters the first set of audio signals.

8. The audio device according to claim 7, wherein the first set of spatial filters has coefficients set to the decorrelation coefficients determined for the set of spatial decorrelation filters.

9. The audio device according to any one of claims 1 to 6, wherein a first set of filters for filtering filters a first set of audio signals after filtering by the first set of spatial filters, and the beamform adapter performs the comparison using the first set of audio signals before filtering by the first set of spatial filters.

10. The audio device according to claim 9, wherein the first set of spatial filters has coefficients that match the coefficients of the spatial filters which are two cascades of the set of uncorrelated filters.

11. The audio device according to any one of claims 1 to 6, wherein the first set of spatial filters filters the second set of audio signals.

12. The audio device according to claim 11, wherein the first set of spatial filters has coefficients determined according to the set of inverse spatial decorrelation filters, and the set of inverse spatial decorrelation filters is the inverse filter of the set of spatial decorrelation filters.

13. The audio device according to claim 11, wherein the first set of spatial filters has coefficients that match the coefficients of a spatial filter which is a cascade of two sets of spatial inverse filters, each of which includes an inverse filter of the set of spatial decorrelating filters.

14. The adaptive coefficient processor determines the set of spatially uncorrelated filters to generate a set of output audio signals, and each of the output audio signals in the set of output signals is, The steps of segmenting the first set of audio signals into time segments, and for at least some time segments, A step of generating a frequency bin representation of a first set of audio signals, wherein each frequency bin in the frequency bin representation of the first set of audio signals includes a frequency bin value for each of the audio signals in the first set of audio signals. A step of generating a frequency bin representation of the set of output signals, wherein each frequency bin of the frequency bin representation of the set of output signals includes a frequency bin value for each of the output signals, and the frequency bin values ​​for a given output signal in the set of output signals for a given frequency bin are generated as a weighted combination of the frequency bin values ​​of a first set of audio signals for a given frequency bin, and the weighted combination has the decorrelation coefficient as a weight. A step of updating a first weight for the contribution of the first frequency bin to the first output signal linked to the first input audio signal, from the second frequency bin to the first frequency bin for the second input audio signal linked to the second output signal, according to a correlation scale between the first previous frequency bin value of the first output signal for the first frequency bin and the second previous frequency bin value of the second output signal for the first frequency bin. The audio device according to any one of claims 1 to 13, which is linked to one input audio signal from a first set of audio signals by performing the following:

15. The audio device according to claim 14, wherein the adaptive coefficient processor updates the first weight according to the product of a first value and a second value, the first value being one of the first previous frequency bin value and the second previous frequency bin value, and the second value being the complex conjugate of the other of the first previous frequency bin value and the second previous frequency bin value.

16. A method for generating an output audio signal, wherein the method is A receiving step of receiving a first set of audio signals, wherein the first set of audio signals includes audio signals that capture the sound of a scene from different locations, The audio beamformer, The first set of filters filters the first set of audio signals, The combiner combines the outputs of the first set of filters to generate an output audio signal. The steps involve generating an output signal and A step of generating a second set of audio signals from filtering the output audio signal, wherein each filter in the second set of filters has a frequency response that is the complex conjugate of the filter in the first set of filters, A step of adjusting the first set of filters and the second set of filters in accordance with a comparison of the first set of audio signals and the second set of audio signals, A step of determining the decorrelation coefficient for a set of spatial decorrelation filters that generate a decorrelation output signal from a first set of audio signals, which includes fitting the decorrelation coefficient according to an updated value determined from the first set of audio signals, A step of modifying at least one of the first set of audio signals and the second set of audio signals by applying a first spatial filtering to at least one of the first set of audio signals and the second set of audio signals, wherein the first set of spatial filters has coefficients determined from the decorrelation coefficients. A method having

17. A computer program comprising computer program code means, wherein when the computer program is executed on a computer, the computer program code means performs all the steps of the method according to claim 16.