Method and apparatus for capturing audio
The audio acquisition apparatus employs adaptive spatial uncorrelated filtering and multiple beamformers to enhance target audio capture in complex environments, addressing noise and reverberation challenges with improved performance and reduced complexity.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- KONINKLIJKE PHILIPS NV
- Filing Date
- 2024-05-20
- Publication Date
- 2026-06-24
AI Technical Summary
Existing audio capture systems face challenges in isolating target audio sources effectively, particularly in complex environments with noise and reverberation, leading to suboptimal performance and increased computational complexity.
An audio acquisition apparatus using adaptive spatial uncorrelated filtering and multiple beamformers, including a first unconstrained and constrained beamformers, to enhance audio source separation and reduce noise sensitivity, with adaptive parameters optimized for target audio capture.
Improves audio capture performance in challenging environments by enhancing target audio extraction, reducing noise interference, and adapting rapidly to new sources while maintaining low computational complexity.
Smart Images

Figure 2026520680000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to audio acquisition using beamforming, and more particularly, to speech acquisition using beamforming, although this is not exclusive. [Background technology]
[0002] Audio, particularly speech capture, has become increasingly important in recent decades. For example, capturing speech and other audio is becoming increasingly crucial in various applications such as communications, teleconferencing, gaming, and audio user interfaces. However, in many scenarios and applications, the problem is that the target audio source is not usually the only audio source in the environment. Rather, in a typical audio environment, there are many other audio / noise sources that are picked up by the microphone. Audio processing is often used to improve audio capture, especially by post-processing the captured audio to improve the resulting audio signal.
[0003] In many real-world applications, audio is captured by multiple microphones positioned at various locations. For example, a linear array of multiple microphones is often used to capture audio within an environment such as a room. By using multiple microphones, spatial information of the audio can be captured, and applications utilizing this spatial information have been developed, enabling improvements to services and the creation of new services.
[0004] One commonly used approach attempts to isolate audio sources by applying audio beamforming to create a beam in the direction from which audio is arriving from a specific audio source. While this can provide favorable performance in many scenarios, it is not optimal in all cases. For example, optimal source isolation may not be possible in some instances, and in fact, in some applications, spatial beamforming may not provide ideal audio characteristics for further processing to achieve a given effect. [Overview of the project] [Problems that the invention aims to solve]
[0005] Therefore, spatial audio source separation, particularly separation based on audio beamforming, is highly advantageous in many scenarios and applications, but there is a desire to improve the performance and usability of such approaches. However, there is also a desire to reduce complexity and resource usage (e.g., computational resource usage), and these priorities often conflict with each other.
[0006] Furthermore, in many applications, there is a strong desire for audio capture to adapt to specific audio scenes or sources, such as adapting to changes in who is currently speaking or the speaker's position. It is also desirable that speech capture be tolerant of other audio and noise present in the scene. For example, it is often desirable that the capture system can capture the target speaker even if another strong audio source is present.
[0007] Therefore, an improved approach is advantageous, particularly one that enables reduced complexity, increased flexibility, easier implementation, lower costs, improved audio capture, improved spatial awareness / distinguishing of audio sources, improved audio source isolation, improved support for audio / speech applications, reduced reliance on known or static acoustic properties, increased flexibility and customization for different audio environments and scenarios, improved audio beamforming, increased tolerance to noise and unwanted audio sources in audio scenes, improved trade-offs between performance and complexity / resource usage, and / or improved performance.
[0008] Therefore, the present invention preferably attempts to mitigate, alleviate, or eliminate one or more of the above-mentioned drawbacks, either individually or in any combination. [Means for solving the problem]
[0009] According to one aspect of the present invention, an apparatus for acquiring audio is provided. The apparatus includes a receiver that receives a set of first audio signals; an adaptive spatial uncorrelatedizer that applies spatial uncorrelated filtering to the first set of audio signals to generate a set of second audio signals; a first beamformer coupled to the adaptive spatial uncorrelatedizer that generates a first beamformed audio output signal from the second set of audio signals; a plurality of constrained beamformers coupled to the adaptive spatial uncorrelatedizer, each generating a constrained beamformed audio output signal from the second set of audio signals, each constrained beamformer having constrained beamform parameters, and the constrained beamform parameters of the plurality of constrained beamformers (309, 311) forming a set of beamform parameters; an output processor that generates an output audio signal from the constrained beamformed audio output signal; and an apparatus that adapts the beamform parameters of the first beamformer. The system comprises a first adapter, a difference processor that determines a difference scale for each of a plurality of constrained beamformers, wherein the difference scale for each of the plurality of constrained beamformers indicates the difference between the beam formed by the first beamformer and the beam formed by each of the plurality of constrained beamformers, a second adapter that applies constrained beamform parameters from a set of constrained beamform parameters for a plurality of constrained beamformers, with the constraint that the constrained beamform parameters from a set of constrained beamform parameters are applied only to constrained beamformers among the plurality of constrained beamformers for which the difference scale is determined to satisfy a similarity criterion, and a third adapter that applies coefficients for spatially uncorrelated filtering according to an update value determined from a first set of audio signals, wherein the update value is determined to reduce the correlation of the second set of audio signals.
[0010] The present invention improves audio capture in many embodiments. In particular, it often improves performance in reverberant environments and / or performance for audio sources. This approach improves speech capture, especially in many challenging audio environments. In many embodiments, this approach provides reliable and accurate beamforming while also providing rapid adaptation to new target audio sources. This approach provides an audio capture device with reduced sensitivity to noise, reverberation, and reflections. In particular, it often improves the capture of audio sources outside the reverberation radius.
[0011] This approach typically improves the capture of target audio in the presence of noise or potentially strong interference / noise sources. In many scenarios, this approach improves the extraction / separation / isolation of target audio or mitigates the effects of strong, undesirable audio sources. The synergistic effect of adaptive spatial decorrelation and beamforming facilitates the attenuation of noise audio sources, particularly improving the extraction of target audio sources, especially speech sources.
[0012] In some embodiments, the output audio signal from the audio acquisition device is generated according to a first beamformed audio output and / or a constrained beamformed audio output. In some embodiments, the output audio signal is generated as a combination of constrained beamformed audio outputs, specifically, for example, a selection combination is used that selects a single constrained beamformed audio output.
[0013] The difference scale reflects the difference between the beam formed by the first beamformer and the beam formed by the constrained beamformer from which the difference scale is generated. The difference scale is measured, for example, as a difference in beam direction. In many embodiments, the difference scale represents the difference between the beamformed audio output from the first beamformer and the beamformed audio output from the constrained beamformer. In some embodiments, the difference scale represents the difference between the beamform filter of the first beamformer and the beamform filter of the constrained beamformer. The difference scale may also be a distance scale, for example, a scale determined as the distance between the vectors of coefficients of the beamform filters of the first beamformer and the constrained beamformer.
[0014] It will be understood that similarity measures are equivalent to difference measures in that they provide information about the similarity between two features, and by doing so, they also provide information about the differences between those features (and vice versa).
[0015] Similarity criteria may include, for example, the requirement that a difference scale shows a difference below a given scale (for example, a difference scale with increasing values for increasing differences must be below a threshold).
[0016] A constrained beamformer is constrained in that its adaptation is limited to cases where the difference scale satisfies the similarity criterion. In contrast, the first beamformer is not subject to this requirement. In particular, the adaptation of the first beamformer is independent of any constrained beamformer, and especially independent of the beamforming of those beams.
[0017] An adaptation constraint requiring the difference scale to be below a threshold, for example, can be thought to correspond to the adaptation of only the constrained beamformer that is currently forming a beam corresponding to an audio source in a region close to the audio source to which the first beamformer is currently adapted.
[0018] The beamformer is adapted by adapting the filter parameters of the beamformer filter of the beamformer, particularly by adapting the filter coefficients. The adaptation attempts to optimize (maximize or minimize) a given adaptation parameter, such as maximizing the output signal level when an audio source is detected or minimizing the output signal level when only noise is detected. The adaptation attempts to modify the beamformer filter to optimize the measured parameters.
[0019] Spatial decorrelation filtering may be frequency-domain filtering in which, for each output signal, the output value for a given frequency bin is determined as a weighted combination of the values of the set of input signals for that frequency bin.
[0020] In many embodiments, the output processor further generates an output audio signal from the first beamformed audio output signal.
[0021] The second adapter can adapt the constrained beamformer parameters (of the set of constrained beamformer parameters) of a plurality of constrained beamformers with the constraint that only the constrained beamformer parameters (of the set of constrained beamformer parameters) of the constrained beamformer for which the difference measure is determined to satisfy the similarity criterion are adapted.
[0022] The set of beamformer parameters can be the combined / total / cumulative constrained beamformer parameters of each of the plurality of constrained beamformers (formed thereby / including them). The set of beamformer parameters can include all of the constrained beamformer parameters of the plurality of constrained beamformers.
[0023] According to an optional feature of the present invention, the third adapter designates, at least, the first beamformed audio output signal of the first constrained beamformer among the plurality of constrained beamformers as a speech audio signal or a noise audio signal, and the third adapter updates the coefficients of the spatial decorrelation filtering according to the difference scale of the first constrained beamformer only when the first beamformed audio output signal is designated as a noise audio signal.
[0024] Thereby, performance and / or operation are improved in many embodiments. Usually, the adaptation of the adaptive spatial decorrelator 303 is improved to clearly decorrelate the noise sources of the scene in the captured audio. Usually, the separation and extraction of the target audio (for example, the speech audio source) in the presence of strong interference / noise are improved.
[0025] The third adapter can, in particular, at least designate the first beamformed audio output signal of the first constrained beamformer as a speech audio signal or a noise audio signal. The third adapter can designate the first beamformed audio output signal according to the characteristics of the first beamformed audio output signal. The third adapter can designate the first beamformed audio output signal as a noise audio signal according to the fact that the speech detection applied to the first beamformed audio output signal fails to detect speech.
[0026] According to an optional feature of the present invention, when the first beamformed audio output signal is designated as a noise audio signal, the output processor generates an output audio signal that does not include the contribution from the first beamformed audio output signal.
[0027] This improves performance and / or operation in many embodiments. Typically, it improves the user experience.
[0028] In accordance with an optional feature of the present invention, the third adapter, upon detection that none of the multiple constrained beamformers have a difference scale that satisfies the similarity criterion, initializes a given constrained beamformer among the multiple constrained beamformers with beamform parameters that match those of the first beamformer.
[0029] This improves performance and / or operation in many embodiments. Typically, it improves the user experience.
[0030] Initializing a given constrained beamformer according to an optional feature of the present invention includes first designating a given constrained beamformed audio output signal of the given constrained beamformer as the spoken audio signal.
[0031] This improves performance and / or operation in many embodiments. Typically, it improves the user experience.
[0032] According to an optional feature of the present invention, the third adapter applies speech detection to a given constrained beamformed audio output signal and, upon detection that the given constrained beamformed audio output signal does not contain speech signal components at a level above a threshold, reclassifies the given constrained beamformed audio output signal from a speech audio signal to a noise audio signal.
[0033] This improves performance and / or operation in many embodiments. Typically, it improves the user experience.
[0034] According to an optional feature of the present invention, the third adapter initializes a given constrained beamformer only if the level of the first beamformed audio output signal exceeds the level of the constrained beamformed audio output signals from all constrained beamformers.
[0035] This improves performance and / or operation in many embodiments. Typically, it improves the user experience.
[0036] In accordance with an optional feature of the present invention, the second adapter applies only the constrained beamform parameters from a given set of constrained beamform parameters of a given constrained beamformer if a criterion is met that includes at least one requirement selected from the following groups: the requirement that the level of the constrained beamformed audio output signal of a given constrained beamformer is higher than that of any other constrained beamformer; the requirement that the level of the constrained beamformed audio output signal of a given constrained beamformer exceeds a given threshold; the requirement that the level of a point audio source in the constrained beamformed audio output signal of a given constrained beamformer is higher than that of any point audio source in any other constrained beamformed audio output signal; and the requirement that the signal-to-noise ratio of the constrained beamformed audio output signal of a given constrained beamformer exceeds a threshold.
[0037] According to an optional feature of the present invention, the maximum number of constrained beamformers that are updated simultaneously is 1.
[0038] In accordance with an optional feature of the present invention, the adaptive space uncorrelatedizer performs the steps of: linking each audio signal of a second set of audio signals with one audio signal of a first set of audio signals, segmenting the first set of audio signals into time segments, for at least some time segments, generating frequency bin representations of the first set of audio signals, wherein each frequency bin of the frequency bin representations of the first set of audio signals includes a frequency bin value for each of the audio signals of the first set; and generating frequency bin representations of a second set of audio signals, wherein each frequency bin of the frequency bin representations of the second set of audio signals includes a frequency bin value for each of the second set of audio signals, and the frequency bin values of the second output set of audio signals for a given audio signal for a given frequency bin are generated as a weighted combination of the frequency bin values of the first set of audio signals for a given frequency bin. A set of second audio signals is generated by performing the steps of weighting and updating the first weight of the contribution from the second frequency bin value of the first frequency bin of the second input audio signal linked to the second output signal to the first frequency bin, to the first frequency bin value of the first frequency bin of the first output signal linked to the first input audio signal, according to the correlation scale between the first previous frequency bin value of the first output signal and the second previous frequency bin value of the second output signal and the first frequency bin, respectively. The third adapter updates the weights of the weighted combination and updates the first weight of the contribution from the second frequency bin value of the first frequency bin of the second audio signal linked to the second output audio signal in the set of second audio signals, according to the correlation scale between the first previous frequency bin value of the first output audio signal and the first frequency bin,This includes updating the correlation scale between the second output audio signal and the second previous frequency bin value for the first frequency bin.
[0039] This improves performance and / or operation in many embodiments. In many embodiments and scenarios, this provides a particularly attractive performance and / or implementation form.
[0040] This favorably generates a second set of audio signals, which typically exhibits enhanced decorrelation / decoherence compared to the input signal. In many embodiments, this approach provides efficient adaptation of operations that improve decorrelation. These adaptations are typically performed with low complexity and / or resource usage. In particular, this approach applies local adaptations to individual weights while still achieving efficient global adaptation.
[0041] According to an optional feature of the present invention, the adaptive coefficient processor updates a first weight according to the product of a first value and a second value, where the first value is one of a first previous frequency bin value and a second previous frequency bin value, and the second value is the composite conjugate of the other of the first previous frequency bin value and the second previous frequency bin value.
[0042] According to an optional feature of the present invention, the adapter is y(ω)=W(ω)X(ω) The output bin values for a given frequency bin ω are determined from the following, where y(ω) is a vector containing the frequency bin values of the output audio signal for a given frequency bin ω, x(ω) is a vector containing the frequency bin values of the input audio signal for a given frequency bin ω, and W(ω) is a matrix with rows containing the weights of the weighted combination of the output audio signal.
[0043] This improves performance and / or operation in many embodiments. This typically provides an improved adaptation that facilitates the decorrelation / decoherence of a second set of audio signals in many scenarios.
[0044] According to an optional feature of the present invention, the adapter has weights w of matrix W(ω). ij of
number
[0045] This improves performance and / or operation in many embodiments. It typically provides an improved adaptation that facilitates the decorrelation of a second set of audio signals in many scenarios.
[0046] According to an aspect of the present invention, a method for acquiring audio is provided. The method comprises the steps of: receiving a set of first audio signals; applying spatially uncorrelated filtering to the set of first audio signals to generate a set of second audio signals; a first beamformer generating a first beamformed audio output signal from the set of second audio signals; each of a plurality of constrained beamformers generating a constrained beamformed audio output signal from the set of second audio signals, wherein each constrained beamformer has constrained beamforming parameters; generating an output audio signal from the constrained beamformed audio output signal; adapting the beamforming parameters of the first beamformer; and determining a difference scale for each of the plurality of constrained beamformers. The steps include: determining a difference scale which represents the difference between the beam formed by a first beamformer and the beam formed by each of a plurality of constrained beamformers; applying constrained beamform parameters from a set of constrained beamform parameters of a plurality of constrained beamformers, with the constraint that the constrained beamform parameters from the set of constrained beamform parameters are applied only to the constrained beamformers among the plurality of constrained beamformers which the difference scale is determined to satisfy a similarity criterion; and adapting coefficients for spatially uncorrelated filtering according to updated values determined from a first set of audio signals, wherein the updated values are determined to reduce the correlation of a second set of audio signals.
[0047] In some embodiments, the audio device updates a second weight for the contribution to the first frequency bin value from a third frequency bin value, which is the frequency bin value of the first frequency bin of the first input audio signal, depending on the magnitude of the first previous frequency bin value.
[0048] This improves performance and / or operation in many embodiments. It typically provides an improved adaptation that facilitates the decorrelation of the output audio signal in many scenarios. In particular, it provides an improved adaptation of the generated output signal. In many embodiments, the weight update, which reflects the contribution from the linked input signal to the output signal, depends on the magnitude / amplitude of the linked input signal. For example, the update attempts to compensate for the weight of the input signal level to produce a normalized output signal.
[0049] This approach provides, for example, normalization / signal correction / level correction to achieve the desired output level.
[0050] In some embodiments, the audio device sets a predetermined weight for the contribution of a third frequency bin value, which is the frequency bin value of the first frequency bin of the first input audio signal, to the first frequency bin value.
[0051] This improves performance and / or operation in many embodiments. It typically provides an improved adaptation that facilitates the decorrelation of the output audio signal in many scenarios. It provides an improved adaptation while ensuring the convergence of the adaptation to non-zero signal levels in many embodiments. It interacts very efficiently with weight adaptations that are not present in linked signal pairs.
[0052] In many embodiments, the adapter keeps the weights constant without making any adaptations or updating the weights.
[0053] In some embodiments, the audio device restricts the weights for the contribution of a third frequency bin value, which is the frequency bin value of the first frequency bin of the first input audio signal, to real values.
[0054] This improves performance and / or operation in many embodiments. It typically provides an improved adaptation that facilitates the decorrelation of the output audio signal in many scenarios.
[0055] The weights between linked input / output signals are advantageously determined as real-valued weights / constrained to real-valued weights. This leads to improved performance and adaptability, ensuring convergence of non-zero-level solutions.
[0056] In some embodiments, the audio device sets a second weight, which is the weight for the contribution of the first frequency bin of the second output audio signal from the first input audio signal to the fourth frequency bin value, to the complex conjugate of the first weight.
[0057] This improves performance and / or operation in many embodiments. It typically provides an improved adaptation that facilitates the decorrelation of the output audio signal in many scenarios.
[0058] In many embodiments, the two weights for two pairs of input / output signals are complex conjugates of each other.
[0059] In some embodiments, the weights of the weighted combination of input audio signals other than the first input audio signal are complex numerical weights.
[0060] This improves performance and / or operation in many embodiments. Using complex values for the weights of unlinked input signals improves operation in the frequency domain.
[0061] Advantageously, the matrix W(ω) is a Hermitian matrix. In many embodiments, the diagonals of the matrix W(ω) may be constrained to real values, set to predetermined values, and / or maintained as fixed values but not updated / adapted. The weights / coefficients outside the diagonals are generally complex values.
[0062] In some embodiments, the audio device corrects the correlation values of the signal levels of the first frequency bin.
[0063] This improves performance and / or operation in many embodiments. It typically provides an improved adaptation that facilitates the decorrelation of the output audio signal in many scenarios. This allows for correction of the update rate in response to signal changes.
[0064] In some embodiments, the audio device initializes the weights of the weighted combination to include at least one zero-value weight and one non-zero-value weight.
[0065] This improves performance and / or operation in many embodiments. It typically provides an improved adaptation that facilitates the decorrelation of output audio signals in many scenarios. This enables more efficient and / or rapid adaptation and convergence to favorable decorrelation. In many embodiments, the matrix W(ω) is initialized with zero values for the weights or coefficients of unlinked signals and fixed non-zero real values for linked signals. Typically, the weights are set to 1 for, for example, the diagonal weights, and all other weights are initially set to zero.
[0066] In some embodiments, the weighted combination involves applying a time-domain window to the frequency representation of the weights formed by the weights of the first and second input audio signals for various frequency bins.
[0067] This improves performance and / or operation in many embodiments. It typically provides an improved adaptation that facilitates the decorrelation of the output audio signal in many scenarios.
[0068] Applying a time-domain window to the frequency representation of a weight may include converting the frequency representation of the weight to a time-domain representation of the weight, applying a window to the time-domain representation to generate a modified time-domain representation, and converting the modified time-domain representation back to the frequency domain.
[0069] These and other aspects, features, and advantages of the present invention will become apparent from the embodiments described below and will be explained with reference to those embodiments. [Brief explanation of the drawing]
[0070] Embodiments of the present invention will be described below with reference to the drawings, as just one example.
[0071] [Figure 1] Figure 1 shows an example of the elements of a beamforming audio acquisition system. [Figure 2] Figure 2 shows an example of multiple beams formed by an audio acquisition system. [Figure 3] Figure 3 shows examples of elements of an audio acquisition device according to several embodiments of the present invention. [Figure 4] Figure 4 shows some elements of possible processor configurations for realizing elements of an audio device according to several embodiments of the present invention. [Modes for carrying out the invention]
[0072] The following description will focus on embodiments of the present invention applicable to beamforming-based speech capture audio systems, particularly for extracting one or more target speech sources from an audio scene.
[0073] Traditionally, the capture of speech and audio within an audio scene often includes echo cancellation, noise suppression, and more recently, beamforming. Figure 1 shows an example of an audio capture system based on beamforming. In this example, an array 101 consisting of multiple microphones is coupled to a beamformer 103. The beamformer 103 generates an audio source signal z(n) and one or more noise reference signals x(n).
[0074] In some embodiments, the microphone array 101 consists of only two microphones, but typically includes a larger number.
[0075] The beamformer 103 is a particularly adaptive beamformer, which can direct a single beam towards a speech source using an appropriate adaptive algorithm.
[0076] For example, U.S. Patents 7,146012 and 7,602926 disclose examples of adaptive beamformers that focus on processing speech but also provide a reference signal that often (almost) does not contain speech.
[0077] The beamformer generates an extended output signal z(n) by coherently summing the desired portion of the microphone signal by filtering the received signal with a forward-matching filter and summing the filtered outputs. The output signal is then filtered with a backward-adaptive filter having a conjugate filter response to the forward filter (in the frequency domain corresponding to the time-inverted impulse response in the time domain). An error signal is generated as the difference between the input signal and the output of the backward-adaptive filter, and the filter coefficients are adapted to minimize the error signal, thereby directing the audio beam toward the dominant signal. The generated error signal x(n) can be considered a noise reference signal particularly suitable for additional noise reduction on the extended output signal z(n).
[0078] Both the primary signal z(n) and the reference signal x(n) are typically contaminated with noise. If the noise in the two signals is coherent (for example, if there is an interference point noise source), the adaptive filter 105 can be used to reduce the coherent noise.
[0079] For this purpose, the noise reference signal x(n) is coupled to the input of the adaptive filter 105, and its output is subtracted from the audio source signal z(n) to produce a corrected signal r(n). The adaptive filter 105 is typically adapted to minimize the power of the corrected signal r(n) when the target audio source is inactive (e.g., when there is no speech), thereby suppressing coherent noise.
[0080] The corrected signal is fed to post-processor 107, where noise reduction is performed on the corrected signal r(n) based on the noise reference signal x(n). Specifically, post-processor 107 transforms the corrected signal r(n) and the noise reference signal x(n) into the frequency domain using a short-time Fourier transform. Next, for each frequency bin, the amplitude of R(ω) is corrected by subtracting a scaled version of the amplitude spectrum of X(ω). The resulting complex spectrum is returned to the time domain to obtain a noise-suppressed output signal q(n). This spectral subtraction technique was first described by SFBoll in "Suppression of Acoustic Noise in Speech Using Spectral Subtraction" (IEEE Trans.Acoustics, Speech and Signal Processing, vol.27, pp.113~120, Apr.1979).
[0081] While the system in Figure 1 offers highly efficient operation and favorable performance in many scenarios, it is not necessarily optimal in all scenarios. In fact, many conventional systems, including the example in Figure 1, tend to perform very well in applications where the target audio source / speaker is within the reverberation radius of the microphone array, i.e., where the direct energy of the target audio source is (preferably significantly) greater than the reflected energy of the target audio source, but otherwise tend to be less optimal. In typical environments, it has been found that the speaker should usually be within 1 to 1.5 meters of the microphone array.
[0082] However, there is a strong demand for audio-based hands-free solutions, applications, and systems where the user may be located further away from the microphone array. This is desired, for example, in many communication applications and many voice control systems and applications.
[0083] More specifically, when dealing with additional diffuse noise or target speakers outside the reverberation radius, the following problems may arise: Beamformers often have difficulty distinguishing between the echo of the target speech and diffuse background noise, resulting in speech distortion. Adaptive beamformers are slow to converge on the target speaker. While the adaptive beam is not yet converged, speech omissions occur in the reference signal, and if this reference signal is used for transient noise suppression and cancellation, speech distortion occurs. This problem becomes more pronounced when there are more target sources speaking after each other.
[0084] A solution to address the slow convergence of adaptive filters (due to background noise) is to compensate for this with several fixed beams directed in different directions, as shown in Figure 2. However, this approach was developed specifically for scenarios where the audio source of interest is within the reverberation radius. It is inefficient for audio sources outside the reverberation radius, and in such cases, especially when acoustic diffusion background noise is present, it often leads to less robust solutions.
[0085] Figure 3 shows examples of elements of an audio device according to several embodiments of the present invention. This audio device addresses or mitigates some of the shortcomings of conventional systems, particularly in many scenarios, such as some of the problems associated with using a single beamformer as shown in Figure 1.
[0086] The audio device includes a receiver 301 that receives a first set of audio signals. The receiver 301 specifically receives audio signals from a microphone array 301 consisting of multiple microphones that capture audio in the environment from various locations. In this example, the receiver 301 receives audio signals from a set of microphones that capture the audio scene from various locations. For example, the microphones are arranged in a linear array and are positioned relatively close to each other. For example, the maximum distance between audio signal acquisition points is, in many embodiments, 1 meter, 50 cm, 25 cm, or possibly 10 cm or less.
[0087] In some cases, the receiver may include an optional echo canceller that can cancel echoes originating from an acoustic source (for which a reference signal is available) that is linearly related to the echo of the microphone signal. This acoustic source may be, for example, a speaker. An echo-corrected signal is generated by applying the reference signal as input to an adaptive filter and subtracting its output from the microphone signal. This can be repeated for each microphone.
[0088] It will be understood that such echo cancellers are optional and may be omitted in many embodiments.
[0089] Receiver 301 is coupled to an adaptive spatial decorrelator 303. The adaptive spatial decorrelator 303 applies spatial filtering to a first set of audio signals to generate a second set of audio signals. The adaptive spatial decorrelator 303 comprises a set of filters, each filter generating one audio signal from the second set of audio signals based on multiple audio signals (usually all) from the first set of audio signals. Each audio signal in the second set of audio signals (also called the decorrel output signal) is generated as a weighted combination of the audio signals in the first set of audio signals (also called the decorrel input signals). The second set of audio signals is generated to correspond to the first set of audio signals (e.g., having the same total audio energy / power), but attempts to decorrelate at least one audio source in the audio scene captured by the microphone. This is achieved by adapting the filters, specifically by adapting the weights of the weighted sum of the decorrelation filters, as will be explained in more detail later.
[0090] Therefore, the output of the adaptive spatial decorrelator 303 corresponds to the input audio, but the decorrelation of at least one audio source is facilitated. That is, as will be described later, the adaptive spatial decorrelator 303 is adapted to attempt to decorrelate the dominant noise source.
[0091] Next, the output of the adaptive spatial uncorrelatedizer 303 is supplied to several different beamformers.
[0092] The adaptive space uncorrelatedizer 303 is specifically coupled to the first beamformer 305.
[0093] The first beamformer 305 combines the signals from the adaptive spatial uncorrelatedizer 303 so that an effective directional audio sensitivity is generated. Thus, the first beamformer 305 generates an output signal (referred to as the "first beamformed audio output") that corresponds to the selective capture of audio in the environment. The first beamformer 305 is an adaptive beamformer, and its directivity can be controlled by setting parameters for the beamform operation of the first beamformer 305 (referred to as the "first beamform parameters").
[0094] The first beamformer 305 is coupled with a first adapter 307 that adapts the first beamform parameters. Thus, the first adapter 307 adapts the parameters of the first beamformer 305 so that the beam can be steered. The first beamformer 305 is also, for brevity and clarity, referred to as an unconstrained or free-running beamformer, and the audio signal generated by the first beamformer 305 is also, for brevity and clarity, referred to as an unconstrained or free-running beamform signal.
[0095] The audio acquisition device comprises several constrained beamformers 309, 311. Each constrained beamformer combines signals from the adaptive spatial uncorrelatedizer 303 so as to generate the effective directional audio sensitivity of the microphone array 301. Thus, each of the constrained beamformers 309, 311 generates an audio output (referred to as the "constrained beamformed audio output") corresponding to the selective acquisition of audio in the environment. Similar to the first beamformer 305, the constrained beamformers 309, 311 are also adaptive beamformers, and the directivity of each constrained beamformer 309, 311 can be controlled by setting parameters of the constrained beamformers 309, 311 (referred to as the "constrained beamform parameters").
[0096] Accordingly, the audio acquisition device includes a second adapter 313 that adapts the constrained beamform parameters of multiple constrained beamformers, thereby adapting the beams formed by them.
[0097] Therefore, both the first beamformer 305 and the constrained beamformers 309, 311 are adaptive beamformers that can dynamically adapt to the beam actually formed. That is, the beamformers 305, 309, 311 are filter-and-combine beamformers (or, in most embodiments, filter-and-sum beamformers). A beamform filter is applied to each microphone signal, and the filtered outputs are combined, usually simply by summing them up.
[0098] In most embodiments, each beamform filter has a time-domain impulse response, which is not a simple Dirac pulse (corresponding to a simple delay and therefore to gain and phase offsets in the frequency domain), but rather typically an impulse response over time intervals of 2, 5, 10, or even 30 milliseconds or more.
[0099] The impulse response is often realized by a beamforming filter, which is a finite impulse response (FIR) filter with multiple coefficients. In such embodiments, the first adapter 307 and the second adapter 313 adapt the beamforming by adapting the filter coefficients. In many embodiments, the FIR filter has a coefficient corresponding to a fixed time offset (typically a sample time offset), and adapters 307 and 313 adapt this coefficient value. In other embodiments, the beamforming filter typically has significantly fewer coefficients (e.g., only two or three), but their timing is (also) adaptable.
[0100] A notable advantage of beamform filters with extended impulse response, rather than simple variable delay (or simple frequency-domain gain / phase adjustment), is that beamformers 305, 309, and 311 are not only adapted to the strongest, typically straight signal components. Rather, beamformers 305, 309, and 311 can be adapted to include further signal paths, typically corresponding to reflections. Therefore, this approach improves performance in most real-world environments, particularly in reflective and / or reverberant environments, and / or for audio sources far from the microphone array 301.
[0101] Various adaptive algorithms can be used in various embodiments, and it will be understood that various optimization parameters are known to those skilled in the art. For example, adapters 307 and 313 adapt beamform parameters to maximize the output signal value of the beamformer. As a specific example, consider a beamformer in which an received microphone signal is filtered by a forward-matching filter and the filtered outputs are summed. The output signal is filtered by a backward-adaptive filter having a conjugate filter response to the forward filter (in the frequency domain corresponding to the time-reversed impulse response in the time domain). An error signal is generated as the difference between the input signal and the output of the backward-adaptive filter, and the coefficients of the filter are adapted to minimize the error signal, thereby yielding maximum output power. Details of such approaches are described in U.S. Patents 7,146,012 and 7,602,926.
[0102] Furthermore, approaches such as those described in U.S. Patent Nos. 7,146012 and 7,602926 are based on adaptations that utilize both the audio source signal z(n) and the noise reference signal x(n) from the beamformer, and it will be understood that a similar approach can be used in the system shown in Figure 3.
[0103] The first beamformer 305 and the constrained beamformers 309 and 311 are shown in Figure 1 and correspond to those disclosed in U.S. Patent Nos. 7,146,012 and 7,602,926.
[0104] In many embodiments, the structure and implementation of the first beamformer 305 and the constrained beamformers 309 and 311 are identical. For example, the beamform filters have the same FIR filter structure with the same number of coefficients.
[0105] However, the operation and parameters of the first beamformer 305 and the constrained beamformers 309 and 311 differ, and in particular, the constrained beamformers 309 and 311 are constrained in a way that the first beamformer 305 is not. Specifically, the adaptation of the constrained beamformers 309 and 311 differs from the adaptation of the first beamformer 305 and is subject to several constraints.
[0106] In other words, while the constrained beamformers 309 and 311 are limited to adaptation (updating beamform filter parameters) only when certain criteria are met, the first beamformer 305 can adapt even when such criteria are not met. In fact, in many embodiments, the first adapter 307 can always adapt the beamform filter, and this is not constrained by any characteristics of the audio (or any of the constrained beamformers 309 and 311) captured by the first beamformer 305.
[0107] The criteria for applying the constrained beamformers 309 and 311 will be described later.
[0108] In many embodiments, the adaptability of the first beamformer 305 is higher than that of the constrained beamformers 309, 311. Therefore, in many embodiments, the first adapter 307 is configured to adapt to changes faster than the second adapter 313, and thus the first beamformer 305 is updated faster than the constrained beamformers 309, 311. This is achieved, for example, by low-pass filtering of a value to be maximized or minimized (e.g., signal level of the output signal or magnitude of the error signal) that has a higher cutoff frequency for the first beamformer 305 than for the constrained beamformers 309, 311. As another example, the maximum change per update of beamform parameters (particularly beamform filter coefficients) is greater for the first beamformer 305 than for the constrained beamformers 309, 311.
[0109] Therefore, in this system, multiple intensive (adaptively constrained) beamformers, which are slow and adapt only when certain criteria are met, are complemented by free-running, high-speed adaptive beamformers that are not subject to these constraints. Typically, slower and more intensive beamformers provide slower but more accurate and reliable adaptation to a particular audio environment than free-running beamformers. However, free-running beamformers can usually adapt quickly over larger parameter intervals.
[0110] In the system shown in Figure 3, these beamformers are used synergistically to improve performance. This is described in detail below.
[0111] The first beamformer 305 and the constrained beamformers 309 and 311 are coupled to an output processor 315. The output processor 315 receives beamformed audio output signals from the beamformers 305, 309, and 311. The exact output generated from the audio acquisition device depends on the specific preferences and requirements of each embodiment. In fact, in some embodiments, the output from the audio acquisition device consists simply of the audio output signals from the beamformers 305, 309, and 311.
[0112] In many embodiments, the output signal from the output processor 315 is generated as a combination of audio output signals from beamformers 305, 309, and 311. In fact, in some embodiments, a simple selection combination is made, for example, by selecting the audio output signal with the highest signal-to-noise ratio or simply the highest signal level.
[0113] Therefore, the output selection and post-processing of the output processor 315 are application-specific and vary in various implementations / embodiments. For example, all possible concentrated beam outputs may be provided, and selection can be made based on user-defined criteria (e.g., selecting the strongest speaker).
[0114] For example, in voice control applications, all output is forwarded to a voice trigger recognition device, which detects specific words or phrases to initialize voice control. In such an example, the audio output signal on which the trigger word or phrase is detected is then used by the voice recognition device, following the trigger phrase, to detect a specific command.
[0115] For communication applications, it is advantageous to select an audio output signal that, for example, has the strongest signal and also detects the presence of a specific point audio source.
[0116] In some embodiments, post-processing, such as noise suppression as shown in Figure 1, is applied to the output of the audio acquisition device (for example, by the output processor 315). This improves, for example, the performance of voice communication. Such post-processing may include nonlinear operations, but in some speech recognition devices, for example, it is more advantageous to restrict the processing to include only linear operations.
[0117] In the system shown in Figure 3, a particularly advantageous approach is taken for audio acquisition based on the synergistic interaction and interrelationship between the first beamformer 305 and the constrained beamformers 309 and 311.
[0118] For this purpose, the audio acquisition device includes a difference processor 317. The difference processor 317 determines the difference scale between one or more of the constrained beamformers 309, 311 and the first beamformer 305. The difference scale indicates the difference in the beams formed by the first beamformer 305 and the constrained beamformers 309, 311. Therefore, the difference scale of the first constrained beamformer 309 indicates the difference in the beams formed by the first beamformer 305 and the first constrained beamformer 309. In this way, the difference scale indicates how closely the two beamformers 305, 309 are adapted to the same audio source.
[0119] Different difference scales can be used in different embodiments and applications.
[0120] In some embodiments, the difference scale is determined based on the beamformed audio outputs generated from various beamformers 305, 309, and 311. As an example, a simple difference scale is generated by measuring the signal levels of the outputs of the first beamformer 305 and the first constrained beamformer 309 and comparing them to one another. The closer the signal levels are to each other, the lower the difference scale (typically, the difference scale also increases as a function of the actual signal level of, for example, the first beamformer 305).
[0121] In many embodiments, a more appropriate difference scale is generated by determining the correlation between the beamformed audio outputs from the first beamformer 305 and the first constrained beamformer 309. A higher correlation results in a lower difference scale.
[0122] Alternatively, the difference scale can also be determined based on a comparison between the beamform parameters of the first beamformer 305 and the beamform parameters of the first constrained beamformer 309. For example, the coefficients of the beamform filter of the first beamformer 305 and the beamform filter of the first constrained beamformer 309 for a given microphone are represented by two vectors. Next, the magnitude of the difference vector of these two vectors is calculated. This process can be repeated for all microphones, and the combined or average magnitude can be obtained and used as the difference scale. In this way, the generated difference scale reflects the difference in the beamform filter coefficients for the first beamformer 305 and the first constrained beamformer 309, and this is used as the beam difference scale.
[0123] Therefore, in the system shown in Figure 3, a difference scale is generated that reflects the difference between the beamform parameters of the first beamformer 305 and the beamform parameters of the first constrained beamformer 309, and / or the difference in these beamformed audio outputs.
[0124] It will be understood that generating, determining, and / or using a difference scale is directly equivalent to generating, determining, and / or using a similarity scale. In fact, one can usually be considered a monotonically decreasing function of the other, and therefore a difference scale is also a similarity scale (and vice versa), and one usually shows an increasing difference simply by increasing its value, while the other does this by decreasing its value.
[0125] The difference processor 317 is coupled to the second adapter 313 to provide a difference scale. The second adapter 313 adapts the constrained beamformers 309 and 311 according to the difference scale. Specifically, the second adapter 313 adapts the constrained beamform parameters only to constrained beamformers for which the difference scale is determined to satisfy the similarity criterion. Therefore, if the difference scale for a given constrained beamformer 309 or 311 has not been determined, or if the determined difference scale for a given constrained beamformer 309 or 311 indicates that the beams of the first beamformer 305 and the given constrained beamformers 309 or 311 are not sufficiently similar, no adaptation is performed.
[0126] Therefore, in the audio acquisition device of Figure 3, the beam adaptation of the constrained beamformers 309 and 311 is restricted. That is, the constrained beamformers 309 and 311 are constrained to adapt only when the current beam they form is close to the beam formed by the free-running first beamformer 305. In other words, each constrained beamformer 309 and 311 is adapted only when the first beamformer 305 is currently adapted to be sufficiently close to the individual constrained beamformers 309 and 311.
[0127] As a result, the adaptation of the constrained beamformers 309 and 311 is controlled by the operation of the first beamformer 305, and thus the beam formed by the first beamformer 305 effectively controls which constrained beamformers 309 and 311 are optimized / adapted to. This approach ensures that, in particular, the constrained beamformers 309 and 311 are adapted only when the target audio source is close to the current adaptation of the constrained beamformers 309 and 311.
[0128] Approaches that require beam similarity to enable adaptation have been shown to significantly improve performance when the audio source of interest, in this case the speaker of interest, is outside the reverberation radius. In fact, they have been shown to provide particularly desirable performance for weak audio sources in reverberation environments with non-dominant direct-path audio components.
[0129] In many embodiments, the constraints on adaptation have further requirements.
[0130] For example, in many embodiments, the adaptation is a requirement that the signal-to-noise ratio of the beamformed audio output exceeds a threshold. Thus, the adaptation of individual constrained beamformers 309, 311 is constrained to scenarios in which this is well adapted and the underlying signal of the adaptation reflects the audio signal of interest.
[0131] It will be understood that various approaches can be used to determine the signal-to-noise ratio in various embodiments. For example, the noise floor of a microphone signal can be determined by tracking the minimum value of a smoothed power estimate, with the instantaneous power for each frame or time interval being compared to this minimum. As another example, the noise floor of a beamformer output can be determined and compared to the instantaneous output power of the beamformed output.
[0132] In some embodiments, the adaptation of the constrained beamformers 309, 311 is restricted when speech components are detected at the output of the constrained beamformers 309, 311. This improves performance for speech acquisition applications. It will be understood that any suitable algorithm or approach can be used to detect speech in the audio signal.
[0133] It will be understood that the system typically operates using frame processing or block processing. Thus, consecutive time intervals or frames are defined, and within each time interval, the processing described is performed. For example, a microphone signal is divided into processing time intervals, and for each processing time interval, beamformers 305, 309, 311 generate the beamformed audio output signal for that time interval, determine the difference scale, select a constrained beamformer 309, 311, and update / adapt the constrained beamformer 309, 311. In many embodiments, it is advantageous for the processing time intervals to have a duration of 5 milliseconds to 50 milliseconds.
[0134] In some embodiments, it will be understood that different processing time intervals can be used for different aspects and functions of the audio acquisition device. For example, the difference scale and selection of the constrained beamformers 309, 311 for adaptation may be performed, for example, at a lower frequency than the beamforming processing time interval.
[0135] In many embodiments, adaptation relies on detecting point audio sources within the beamformed audio output. Therefore, in many embodiments, the audio device detects the audio source as part of the control of the beamformer. In many embodiments, the audio device specifically detects point audio sources within a second beamformed audio output.
[0136] In acoustics, a point audio source is a sound source that originates from a specific point in space. It will be understood that various algorithms or criteria can be used to estimate (detect) whether a point audio source exists within the beamformed audio output from given constrained beamformers 309, 311, and that those skilled in the art will be aware of such various approaches.
[0137] One approach, in particular, is based on identifying the characteristics of a single or dominant point source captured by the microphones of the microphone array 301. A single or dominant point source can be detected, for example, by observing the correlation between the microphone signals. High correlation suggests the presence of a dominant point source. Low correlation indicates the absence of a dominant point source, but rather that the captured signals originated from many uncorrelated sources. Therefore, in many embodiments, point audio sources are considered spatially correlated audio sources, and this spatial correlation is reflected by the correlation of the microphone signals.
[0138] In this example, the correlation is determined after filtering by the beamform filter. That is, the correlation of the beamform filter outputs of the constrained beamformers 309 and 311 is calculated, and if this exceeds a given threshold, it is considered that a point audio source has been detected.
[0139] In other embodiments, point sources are detected by evaluating the content of the beamformed audio output. For example, if the audio source detector 401 analyzes the beamformed audio output and detects a speech component of sufficient intensity within the beamformed audio output, this can be considered to correspond to a point audio source, and therefore, the detection of a strong speech component can be considered the detection of a point audio source.
[0140] The detection result is passed to the second adapter 313. The second adapter 313 adapts accordingly. That is, the second adapter 313 adapts only the constrained beamformers 309 and 311 that indicate that a point audio source has been detected by the audio source detector 401.
[0141] Therefore, the audio acquisition device restricts the application of the constrained beamformers 309, 311 so that only constrained beamformers 309, 311 are applied if a point audio source is present within the formed beam and the formed beam is close to the beam formed by the first beamformer 305. Thus, the application is typically restricted to constrained beamformers 309, 311 that are already close to the (target) point audio source. This approach enables very robust and accurate beamforming that works very well in environments where the target audio source is outside the reverberation radius. Furthermore, by operating and selectively updating multiple constrained beamformers 309, 311, this robustness and accuracy are complemented by relatively fast response times, allowing the entire system to quickly adapt to rapidly moving or newly generated audio sources.
[0142] In many embodiments, the audio acquisition device adapts only one constrained beamformer 309, 311 at a time. Therefore, the second adapter 313 selects one of the constrained beamformers 309, 311 at each adaptation time interval and adapts only this constrained beamformer by updating its beamform parameters.
[0143] The selection of a single constrained beamformer 309, 311 is usually performed automatically when selecting a constrained beamformer 309, 311 for adaptation, typically only when the currently formed beam is close to the beam formed by the first beamformer 305, and when a point audio source is detected within the beam.
[0144] However, in some embodiments, it is possible for multiple constrained beamformers 309, 311 to satisfy the criteria simultaneously. For example, if a point audio source is located near an area covered by two different constrained beamformers 309, 311 (or, for example, in an overlapping area), the point audio source will be detected within both beams, and these beams will be adapted to be close to each other as both are adapted to the point audio source.
[0145] Therefore, in such an embodiment, the second adapter 313 selects one of the constrained beamformers 309, 311 that satisfy two criteria and applies only this one. This reduces the risk of both beams being applied to the same point audio source, and thus reduces the risk of these operations interfering with each other.
[0146] In practice, the constrained beamformers 309 and 311 are adapted under the constraint that the corresponding difference scale must be sufficiently low, and by selecting only one constrained beamformer 309 or 311 for adaptation (e.g., for each processing time interval / frame), adaptation is distinguished between different constrained beamformers 309 and 311. This causes the constrained beamformers 309 and 311 to be adapted to cover different regions, and the nearest constrained beamformer 309 or 311 is automatically selected to adapt to / follow the audio source detected by the first beamformer 305. However, in contrast to the approach in Figure 2, for example, the regions are not fixed and predetermined, but are formed dynamically and automatically.
[0147] It should also be noted that the region depends on beamforming for multiple paths and is not usually limited to the angular direction of the destination region. For example, regions are distinguished based on the distance to the microphone array. Thus, the term "region" can be thought of as referring to the spatial location where the audio source yields an adaptation that satisfies the similarity requirement of the difference scale. Therefore, not only direct paths but also reflections are considered, for example, if reflections are taken into account in the beamform parameters and are determined based on both spatial and temporal aspects (especially if they depend on the complete impulse response of the beamform filter).
[0148] The selection of a single constrained beamformer 309, 311 is made particularly in accordance with the level of the captured audio. For example, the audio source detector 401 determines the audio level of each beamformed audio output from the constrained beamformers 309, 311 that meet the criteria and selects the constrained beamformer 309, 311 that yields the highest level. In some embodiments, the audio source detector 401 selects the constrained beamformer 309, 311 that has the highest value for point audio sources detected within the beamformed audio output. For example, the audio source detector 401 detects speech components in the beamformed audio outputs from two constrained beamformers 309, 311 and selects the one with the highest speech component level.
[0149] This approach involves highly selective adaptation of constrained beamformers 309 and 311, which are applied only in specific situations. This results in very robust beamforming by the constrained beamformers 309 and 311, improving the capture of the desired audio source. However, in many scenarios, the beamforming constraints also result in adaptation delays, and in fact, in many situations, new audio sources (e.g., new speakers) are not detected or are adapted only with a significant delay.
[0150] In this approach, the beamformers 305, 309, and 311 do not directly manipulate the microphone signal, but rather the modified signal resulting from the decorrelation operation by the adaptive spatial decorrelator 303. However, counterintuitively, even though the adaptive spatial decorrelator 303 removes the direct link between the spatial location of the captured audio and the audio signal being beamformed, the beamforming approach and algorithms offer very favorable performance, particularly providing highly efficient extraction / separation / selection of audio sources (especially speakers). In fact, in many scenarios, spatial decorrelation by the adaptive spatial decorrelator 303 significantly improves the performance of the beamforming algorithm, often improving the separation and isolation of the desired audio source. In particular, this approach often significantly reduces sensitivity to the presence of strong or dominant interfering or undesirable audio sources. Indeed, this approach has been shown to enable the separation and extraction of the desired audio source in situations where the audio scene and captured audio may be dominated by noise sources that may be captured at substantially higher levels than the desired audio source.
[0151] As a specific example, the beamforming approach and configuration described above, which uses a (typically fast) free-running beamformer 305 in combination with multiple constrained beamformers 309, 311, provides excellent performance in many scenarios, but may be less efficient in some scenarios where a continuous point noise interference source is stronger than the direct field contribution of the target speech source on the microphone. For example, when a television is playing relatively close to the microphone array and the commands of the target speaker must be recognized. Typically, the speaker signal from the television cannot be used by the audio device, so echo cancellation cannot be applied. Such scenarios are very important for many speech control / interface applications, such as voice-controlled personal devices and many communication applications. This not only degrades the individual beamforming experience but can also affect the interaction between unconstrained and constrained beamformers. In particular, unconstrained beamformers can often only detect and track strong interference sources and therefore cannot detect other speaker sources and transfer them to the constrained beamformers (even if these can, in principle, extract and track individual speaker sources). This means that a constrained beamformer will not be assigned to an active (speaking) speaker. However, by introducing the adaptive space decorrelator 303, this strong interference can be decorrelated, causing the interference source to appear as more decorrelated noise at the input to the beamformer. This improves beamforming, enhancing the isolation and separation of the target source, and in particular, allowing the unconstrained beamformer to detect audio sources other than the strong interference source.
[0152] Thus, in many scenarios, a strong synergistic effect is achieved by having the beamformer operate not on spatial audio signals corresponding to positions within the audio scene, but on signals that do not have a specific relationship to such positions (and whose spatial characteristics associated with the signals actually supplied to the beamformer differ between different parts and frequency ranges).
[0153] The audio device includes a third adapter 319, which dynamically adapts the adaptive spatial uncorrelatedizer 303 to provide appropriate uncorrelatedization. The adaptive spatial uncorrelatedizer 303 adapts the coefficients of spatial filtering depending on updated values determined from the input signal. For example, the processing is performed segment by segment, where new updated values for the filter coefficients are determined for each segment. The coefficients of subsequent segments are then modified based on the updated values. The adaptive spatial uncorrelatedizer 303 specifically includes a set of filters, each of which produces an output value as a weighted combination of the input signal, and the third adapter 319 adapts / updates the weights of the weighted combination.
[0154] The third adapter 319 determines the update value such that the update value reduces the correlation of the second set of audio signals, and decorrelates audio sources in particular, such as the dominant audio source.
[0155] Various approaches are known for adaptively decorrelating a set of (spatial) audio signals, and it will be understood that any suitable approach can be used without prejudice to the present invention. Below, we describe decorrelation and adaptive approaches that have been found to provide particularly advantageous performance for the audio apparatus in Figure 3, especially for processing by a beamformer.
[0156] Beamforming, decorrelation, and adaptation are typically performed in the frequency domain. The receiver 301 or adaptive space decorrelation 303 includes a segmenter that segments the set of input audio signals into time segments. In many embodiments, segmentation is typically fixed segmentation into time segments of fixed, equal duration, e.g., division into time segments / intervals with fixed durations of 10 to 20 milliseconds. In some embodiments, segmentation is adaptive so that the segments have varying durations. For example, the input audio signal has a variable sample rate, and the segments are determined to contain a fixed number of samples.
[0157] Segmentation is typically performed so that the input signal is divided into segments having a given fixed number of time-domain samples. For example, in many embodiments, the segmenter 103 divides the input signal into consecutive segments (e.g., 256 or 512 samples).
[0158] The receiver 301 or adaptive space uncorrelatedizer 303 generates a frequency bin representation of the input audio signal, and the first set of input signals to be further processed is typically represented in the frequency domain by the frequency bin representation. The audio device performs frequency domain processing on the frequency domain representation of the input audio signal. Since the representation and processing of the signal are based on frequency bins, the signal is represented by the values of the frequency bins, and these values are processed to generate the frequency bin values of the output signal. In many embodiments, the frequency bins have the same size and therefore cover frequency intervals of the same size. However, in other embodiments, the frequency bins have different bandwidths, and for example, perceptually weighted bin frequency intervals may be used.
[0159] In some embodiments, the input audio signal is already provided in frequency representation, and no further processing or manipulation may be required. However, in such cases, it may be desirable to rearrange the frequency representation into a suitable segment representation, for example, by using interpolation between frequency values to align the frequency representation into time segments.
[0160] In other embodiments, a filter bank, such as a quadrature mirror filter (QMF), may be applied to the time-domain input signal to generate the frequency bin representation. However, in many embodiments, the frequency representation is generated by applying the discrete Fourier transform (DFT), particularly the fast Fourier transform (FFT).
[0161] In the audio device shown in Figure 3, the adaptive spatial decorrelator 303 includes a spatial decorrelation filter that processes a first set of audio signals in the frequency domain. In the following description, the first set of audio signals may also be referred to as the input audio signals (to the adaptive spatial decorrelator 303), and the resulting signal (a second set of audio signals) may be referred to as the output audio signals (from the adaptive spatial decorrelator 303).
[0162] For each frequency bin, an output frequency bin value is generated from one or more input frequency bin values of one or more input signals, as will be explained in detail below. The output signal is generated to reduce the correlation between signals (usually / on average) in relation to the correlation of the input signals, for at least one audio source, for example, a dominant audio source.
[0163] A set of spatially uncorrelated filters filters an input audio signal. This filtering is spatial in that, for a given output signal (for the same time / segment and frequency bins), the output value is determined from multiple, usually all, input audio signals. Spatial filtering is particularly frequency bin-based, where the frequency bin value of the output signal for a given frequency bin is generated from the frequency bin value of the input signal for that frequency bin. The filtering / weighting combination spans the entire signal, rather than typical time / frequency filtering.
[0164] In other words, the frequency bin value of a given frequency bin is determined as a weighted combination of the frequency bin values of the input signal for that frequency bin. This combination may be a sum, and the frequency bin value is determined as the weighted sum of the frequency bin values of the input signal for that frequency bin. The determination of the bin value of a given frequency bin is determined by vector multiplication of the weight / coefficient vector of the weighted sum and the vector containing the bin values of the input signal.
number
[0165] Given a frequency bin ω, the output bin values are represented as a vector y(ω). The output signal is determined as follows: y(ω)=W(ω)X(ω) Here, the matrix W(ω) represents the weights / coefficients of the weighted sum of various output signals, and x(ω) is a vector containing the input signal values.
[0166] For example, in a case with only three input and output signals, the output bin values of the frequency bin ω are given by:
number
[0167] The third adapter 319 attempts to adapt a spatial filter to become a spatial decordation filter that corresponds to the input signal, but in particular promotes decordation of the signals for one (dominant) audio source. The output audio signal is generated with spatial decordation promoted, meaning that the cross-correlation between audio signals is lower for the output audio signal than for the input audio signal. That is, the output signal has the same combined energy / power as the input signal (or a given scaling thereof), but is generated with promoted decordation (reduced correlation) between signals. The output audio signal contains all the audio / signal components of the input signal, but is generated by redistributing them into different signals to promote decordation.
[0168] Correlationless filters, in particular, produce output signals with lower coherence / normalized correlation than the input signals. Therefore, the output signal of a correlationless filter has lower coherence / normalized correlation than the coherence of the input signal to the correlationless filter.
[0169] The third adapter 319 determines the updated values for the weight combinations that form the set of spatially uncorrelated filters. Specifically, the updated values are determined for the matrix W(ω). The third adapter 319 then updates the weight combinations based on the updated values.
[0170] The third adapter 319 applies an adaptive approach to the set of uncorrelated filters, determining update values that allow the output signals of the set of uncorrelated filters to represent the audio of the input signal, provided that the output signals are typically more uncorrelated than the input signals.
[0171] The third adapter 319 employs a specific approach to adapting weights based on the generated output signals. This operation is based on the assumption that each output audio signal is linked to one input audio signal. Precise linking between the output and input signals is not required, and many different (in principle, random) linking / pairings between each output and input signal are possible. However, the processing of a weight that reflects the contribution from an input signal linked / paired to the output signal is different from the processing of that weight that reflects the contribution from an input signal that is not linked / paired to the output signal. For example, in some embodiments, the weights of linked signals (i.e., input signals linked to output signals generated by weighted combinations including weights) are set to fixed values and are not updated, and / or the weights of linked signals are limited to real-valued weights, while other weights are generally complex-valued.
[0172] The third adapter 319 uses an adaptive / update approach in which updated values are determined for a given weight that represents the contribution of a given unlinked input signal to the bin value of a given output signal, based on a correlation measure between the output bin value of the given output signal and the output bin value of a given (unlinked) linked output signal. The updated values are then applied to modify the given weight in a subsequent segment, or the updated values of the weight in a given segment are determined according to the two output bin values of the (usually immediately preceding) segment, where these two output values represent the input and output signals to which the weight is associated, respectively.
[0173] The approach described typically applies to multiple, and usually all, weights used when determining output bin values based on unlinked input signals. For weights related to input signals linked to the output signals of the weights, other considerations are used, such as setting the weights to fixed values, as will be discussed later.
[0174] Specifically, the update value is determined by the product of the weight's output bin value and the complex conjugate of the output bin value linked to the weight's input signal, or equivalently, by the product of the weight's output bin value and the output bin value linked to the weight's input signal.
[0175] As a specific example, the update value of segment k+1 of the frequency bin ω is determined depending on the correlation measure given by:
number
number
[0176]
number
[0177] As mentioned above, the set of uncorrelated filters determines the output bin values of the output signals of a given frequency bin ω from the following equation: y(ω)=W(ω)X(ω) Here, y(ω) is a vector containing the frequency bin values of the output signal with a given frequency bin ω, x(ω) is a vector containing the frequency bin values of the input audio signal with a given frequency bin ω, and W(ω) is a matrix with rows containing the weights of the weighted combination of the output audio signal.
[0178] In this example, the third adapter 319 specifically determines the weights W of the matrix W(ω) according to the following equation. ij To adapt at least part of it:
number
[0179] In some embodiments, the third adapter 319 adapts the update rate / speed of the weight adaptation. For example, in some embodiments, the adapter corrects a correlation measure of a given weight depending on the signal level of the output bin value whose contribution is determined by the weight.
[0180] As a specific example, the correction value
Number
[0181] In many embodiments, such correction or normalization is performed particularly on a frequency bin basis. That is, the correction is different for different frequency bins. This improves the operation in many scenarios and usually improves the adaptation of the weights that generate a decorrelated signal.
[0182] For example, the correction is incorporated into the scaling parameter η(k,ω) of the aforementioned update equation. Thus, in many embodiments, the third adapter 319 adapts / modifies the scaling parameter η(k,ω) differently for different frequency bins.
[0183] In many embodiments, the arrangement of the input signal vector x(ω) and the output signal vector y(ω) is such that the linked signals are in the same positions within their respective vectors. That is, specifically, y1 is linked to x1, y2 is linked to x2, y3 is linked to x3, and so on. In this case, the weights of the linked signals are on the diagonal of the weight matrix W(k,ω). In many embodiments, the diagonal values are set to fixed real values, for example, specifically set to a constant value of 1.
[0184] In many embodiments, the weight / space filter / weight combination is such that the weight of the contribution from the first input signal (not linked to the first output signal) to the first output signal is the complex conjugate of the contribution from the second input signal (linked to the first input signal) to the second output signal (linked to the first input signal). Thus, the two weights for the two pairs of linked input / output signals are complex conjugates.
[0185] In examples where the weights of linked input and output signals lie on the diagonal of the weight matrix W(ω), this becomes a Hermitian matrix. In fact, in many embodiments, the weight matrix W(ω) is a Hermitian matrix. That is, the coefficients / weights of the weight matrix W(ω) satisfy the following criteria:
number
[0186] As mentioned above, the weights of the contributions of linked input signals to the output signal bin values (corresponding to the diagonal values of the weight matrix W(ω) in certain examples) are treated differently from the weights of unlinked input signals. For brevity, the weights of linked input signals will be referred to below as "linked weights," and the weights of unlinked input signals will be referred to below as "unlinked weights." Therefore, in certain examples, the weight matrix W(ω) is a Hermitian matrix containing linked weights on the diagonal and unlinked weights outside the diagonal.
[0187] In many approaches, the adaptation of unlinked weights is an adaptation that attempts to reduce the correlation measure. Specifically, each updated value is determined to reduce the correlation measure. Therefore, as a whole, the adaptation attempts to reduce the cross-correlation between output signals. Linked weights, however, are determined in a different way so that the output signals maintain appropriate audio energy / power / level. In fact, if linked weights are adapted to instead attempt to reduce the autocorrelation of the output signals with respect to the weights, there is a high risk that the adaptation will converge to a solution where all weights, and therefore the output signals, are virtually zero (which actually results in the lowest correlation). Furthermore, audio equipment attempts to produce signals with less cross-correlation, but not to reduce autocorrelation.
[0188] Therefore, in many embodiments, the linked weights are set to ensure that the output signal is generated to have the desired (combined) energy / power / level.
[0189] The third adapter 319 may or may not apply the linked weights.
[0190] For example, in some embodiments, the linked weights are simply set to fixed constant values that are not applied. For example, in many embodiments, the linked weights are set to constant scalar values, specifically the value 1 (i.e., the unitary gain applied to the linked input signal). For example, the weights on the diagonal of the weight matrix W(ω) are set to 1.
[0191] Therefore, in many embodiments, the weight of the contribution from the linked input signal frequency bin to a given output signal frequency bin is set to a predetermined value. In many embodiments, this value is kept constant without any adaptation.
[0192] Such an approach has been shown to provide highly efficient performance and a very accurate representation of the original audio of the input signal, but it has been shown to result in an overall adaptation where the output signal is generated in a set of output signals where decorrelation has been facilitated.
[0193] In some embodiments, linked weights may also be applied, but they are applied differently from unlinked weights. In particular, in many embodiments, linked weights are applied based on the output signal.
[0194] In other words, in many embodiments, the linked weights of the output signal linked to the first input signal are adapted based on the generated output bin value of the linked audio signal, and in particular based on the magnitude of the output bin value.
[0195] Such an approach allows, for example, the normalization and / or setting of the desired energy level of a signal.
[0196] In many embodiments, linked weights are constrained to real-valued weights, while unlinked weights are generally complex-valued. In particular, in many embodiments, the weight matrix W(ω) is a Hermitian matrix where real-valued values lie on the diagonal and complex-valued values lie outside the diagonal.
[0197] Such an approach offers particularly advantageous performance and adaptability in many scenarios and embodiments. It has been shown to provide highly efficient spatial decorrelation while maintaining relatively low complexity and computational resources.
[0198] Adaptation can gradually adjust the weights to facilitate the decorrelation between signals. In many embodiments, adaptation converges to a suitable weight matrix W(ω) regardless of the initial values, and in fact, in some cases, adaptation is initialized with random values for the weights.
[0199] However, in many embodiments, the adaptation is started with favorable initial values that, for example, result in faster adaptation or make it easier for the adaptation to converge to more optimal weights for decorrelating the signals.
[0200] In particular, in many embodiments, the weight matrix W(ω) has some weights that are zero, but at least some weights that are not zero. In many embodiments, the number of weights that are substantially zero is two, three, five, or more times the number of weights that are set to non-zero values. In many scenarios, this has been found to tend to improve adaptation.
[0201] In particular, in many embodiments, the third adapter 319 initializes the weights with linked weights set to non-zero values (usually predetermined non-zero real values), while the unlinked weights are set to virtually zero. Thus, in the above example where the linked signals are located at the same position in the vector, we obtain an initial weight matrix W(ω) that has non-zero values on the diagonal and (virtually) zero values outside the diagonal.
[0202] Such initialization offers particularly advantageous performance in many embodiments and scenarios. Typically, audio signals represent audio at various points in time, reflecting the tendency for input signals to be somewhat uncorrelated. Therefore, starting with the assumption that input signals are perfectly correlated is often advantageous, leading to faster and often improved adaptation.
[0203] It will be understood that weights, especially unlinked weights, are not necessarily strictly zero and may be set to low values close to zero in some embodiments. However, non-zero initial values may be at least 5, 10, 20, or 100 times higher than substantially zero initial values.
[0204] The described approach provides a highly efficient adaptive spatial uncorrelatedizer that generates an output signal representing the same audio as the input signal, but with enhanced uncorrelatedness. This approach has been shown to provide highly efficient adaptation in various scenarios and many different acoustic environments, as well as for many different audio sources. For example, it has been shown to provide highly efficient uncorrelatedization of speaker signals in environments with multiple speakers.
[0205] This adaptive approach is further computationally efficient, as it allows for local and individual adaptation of individual weights based on only two signals closely related to the weights (specifically, only two frequency bin values). However, this process also results in efficient and often significantly optimized global optimization of the weight matrix W(ω), particularly through spatial filtering. This local adaptation has been shown to lead to highly advantageous global adaptation in many embodiments.
[0206] A particular advantage of this approach is that it can be used to decorrelate convolutional mixes, and is not limited to decorrelating only instantaneous mixes. In convolutional mixes, the complete impulse response determines how signals from different audio sources combine at the microphone (i.e., delay / timing characteristics are important), whereas in instantaneous mixes, a scalar representation is sufficient to determine how the audio sources combine at the microphone (i.e., delay / timing characteristics are not important). By converting convolutional mixes to the frequency domain, the mix can be thought of as an instantaneous mix of complex values per frequency bin.
[0207] Therefore, the third adapter 319 determines the coefficients of the spatial decorrelation filter of the adaptive spatial decorrelationizer 303 so that these coefficients modify the first set of audio signals, which represent the same audio but are decorrelated. Such decorrelation of audio signals to which the aforementioned multi-beam beamforming approach is applied significantly improves overall performance in many scenarios. In fact, in many scenarios, the isolation and selection of specific audio sources, such as a particular speaker, is improved. Thus, counterintuitively, decorrelation of signals improves beamforming performance by spatially forming a beam toward the source of interest, even though beamforming is essentially based on extracting / separating audio sources by adapting to the correlation between audio signals from various locations. In fact, decorrelation essentially breaks the link between the audio signal and the specific location in the audio scene that is normally utilized by the beamforming operation. However, despite this, the inventors have recognized that decorrelation provides very favorable effects and performance improvements in many scenarios. For example, in the presence of a strong noise source, this approach facilitates and / or improves the extraction / isolation of a specific audio source (especially the speaker).
[0208] The specific adaptations described above offer a highly advantageous approach in many embodiments. Typically, they result in spatially decorrelated filters that produce highly decorrelated signals, offering high accuracy despite their low complexity. In particular, this approach can lead to highly efficient global decorrelatedization of a first set of audio signals by locally adapting individual weights / filter coefficients.
[0209] However, it will be understood that in other embodiments, other approaches may be used to adapt the spatial filter / decorrelation filter. For example, in some embodiments, the third adapter 319 includes a neural network that generates update values to modify the filter coefficients based on the input samples. For example, for each segment, all the frequency bin values of the audio signal are fed to the trained neural network, and as an output, an update value for each weight is generated. Each weight is then updated with this update value. The trained network is trained with training data that includes, for example, many different frequency bin values and the corresponding update values that are manually determined to modify the weights so that decorrelation is facilitated.
[0210] As another example, the third adapter 319 includes a position processor that generates updated values to correct the filter coefficients based on visual cues related to (changing) position. As yet another example, the third adapter 319 determines the cross-correlation matrix of the input signal and computes its eigenvalue decomposition. The eigenvectors and eigenvalues can be used to construct the uncorrelated matrix.
[0211] In certain situations, the audio device initializes the constrained beamformers 309, 311. Specifically, the audio device initializes the constrained beamformers 309, 311 in accordance with the first beamformer 305. More specifically, the audio device initializes one of the constrained beamformers 309, 311 to form a beam corresponding to the beam of the first beamformer 305. The audio device initializes the constrained beamformers 309, 311 with beamformer parameters that match the beamformer parameters of the first beamformer 305. In particular, the constrained beamformers 309, 311 are initialized with beamformers that result in audio beams formed in the direction of the audio beam formed by the first beamformer 305. In many embodiments, the audio device initializes the constrained beamformers 309, 311 by copying the filter parameters / values of the beam filter of the first beamformer 305 to the constrained beamformers 309, 311 being initialized. Therefore, the constrained beamformers 309 and 311 are initialized to form audio beams that match the audio beam of the first beamformer 305.
[0212] Initialization is performed, for example, when it is determined that there are no constrained beamformers whose difference scales satisfy the proximity / similarity criterion, such as when it is determined that there are no constrained beamformers whose difference scales satisfy the proximity / similarity criterion. Initialization requires the detection of audio sources present in the beam formed by the first beamformer 305. Initialization is performed when an audio source, such as a point audio source, is detected in the beamformed audio output signal of the first beamformer 305, and there are no constrained beamformers 309, 311 whose difference scales satisfy the similarity criterion. The detection of audio sources may be a less complex decision, such as determining that the level of the beamformed audio output signal of the first beamformer 305 exceeds a threshold, or it may be more complex, such as detecting specific characteristics, such as the characteristics of a particular point audio source.
[0213] In particular, the audio device sets the beamform parameter of one of the constrained beamformers 309, 311 according to the beamform parameter of the first beamformer 305 (hereinafter also referred to as the first beamform parameter). In some embodiments, the filters of the constrained beamformers 309, 311 and the first beamformer 305 are identical. For example, they have the same structure. In a specific example, the filters of the constrained beamformers 309, 311 and the first beamformer 305 are FIR filters of the same length (i.e., a given number of coefficients), and the currently adapted coefficient values from the filter of the first beamformer 305 are simply copied to the constrained beamformers 309, 311. That is, the coefficients of the constrained beamformers 309, 311 are set to the values of the first beamformer. In this way, the constrained beamformers 309, 311 are initialized with the same beam characteristics currently adapted by the first beamformer 305.
[0214] In some embodiments, the filter settings for the constrained beamformers 309, 311 are determined from the filter parameters of the first beamformer 305, but these can be adapted before application rather than used as is. For example, in some embodiments, the coefficients of the FIR filters are modified to initialize the beams of the constrained beamformers 309, 311 to be wider than (but formed in the same direction as) the beam of the first beamformer 305.
[0215] In many embodiments, the audio device, depending on the circumstances, initializes one of the constrained beamformers 309, 311 with an initial beam corresponding to the beam of the first beamformer 305. The system then processes the constrained beamformers 309, 311 as described above, and adapts them in particular when the aforementioned criteria are met.
[0216] The criteria for initializing the constrained beamformers 309 and 311 vary depending on the embodiment.
[0217] In many embodiments, the audio device initializes the constrained beamformers 309, 311 if the presence of a point audio source is detected at the first beamformed audio output but not at any constrained beamformed audio output.
[0218] Therefore, the audio device determines whether a point audio source is present in any of the beamformed audio outputs from either the constrained beamformers 309, 311, or the first beamformer 305. The detection / estimation result for each beamformed audio output is evaluated. If a point audio source is detected only for the first beamformer 305 and not for the constrained beamformers 309, 311, this reflects a situation where a point audio source, such as a speaker, exists and was detected by the first beamformer 305, but neither of the constrained beamformers 309, 311 detected or adapted to the point audio source. In this case, the constrained beamformers 309, 311 may never adapt to the point audio source (or adapt very slowly). Therefore, one of the constrained beamformers 309, 311 is initialized to form a beam corresponding to the point audio source. This beam is then likely to be close enough to the point audio source and will adapt to this new point audio source (usually slowly but surely).
[0219] Thus, this approach can combine the effects of both the high-speed first beamformer 305 and the more reliable constrained beamformers 309 and 311 to provide advantageous results.
[0220] In some embodiments, the audio device initializes the constrained beamformers 309 and 311 only if the difference scale of the constrained beamformers 309 and 311 exceeds a threshold. Specifically, if the determined minimum difference scale of the constrained beamformers 309 and 311 falls below the threshold, initialization is not performed. In such situations, the adaptation of the constrained beamformers 309 and 311 is close to the desired situation, but the less reliable adaptation of the first beamformer 305 is not very accurate and can be adapted to be closer to the first beamformer 305. Therefore, in such scenarios where the difference scale is sufficiently low, it is advantageous for the system to attempt to adapt automatically.
[0221] In some embodiments, the audio device initializes the constrained beamformers 309, 311 in particular when point audio sources are detected for both the first beamformer 305 and one of the constrained beamformers 309, 311, but the difference between them does not satisfy the similarity criterion. Specifically, the audio device sets the beamform parameters of the first constrained beamformers 309, 311 according to the beamform parameters of the first beamformer 305 when point audio sources are detected for both the beamformed audio output from the first beamformer 305 and the beamformed audio output from the constrained beamformers 309, 311, and the difference between them exceeds a threshold.
[0222] Such a scenario may reflect a situation where the constrained beamformers 309, 311 may have adapted to and captured a point audio source different from the point audio source captured by the first beamformer 305. Therefore, this scenario may particularly reflect a situation where the constrained beamformers 309, 311 may have captured the "wrong" point audio source. Consequently, the constrained beamformers 309, 311 are reinitialized to form a beam toward the desired point audio source.
[0223] In some embodiments, the number of active constrained beamformers 309, 311 can be changed. For example, an audio acquisition device includes the ability to form a relatively large number of constrained beamformers 309, 311. For example, up to eight simultaneous constrained beamformers 309, 311 can be implemented. However, not all of these are active at the same time, for example, to reduce power consumption or computational load.
[0224] In some embodiments, the audio device distinguishes between speech sources and noise sources, particularly non-speech sources. In particular, it detects when a strong captured audio source is likely to be a noise or interference source rather than speech. Specifically, in some embodiments, a third adapter 319 designates the constrained beamformed audio output signal of a given constrained beamformer as either a speech audio signal or a noise audio signal (particularly a non-speech audio signal). For example, when a constrained beamformer is initialized, the third adapter 319 performs speech detection on the beamformed audio output signal generated from the constrained beamformer. If it is detected to contain a sufficiently strong speech component, the output signal is designated as a speech audio signal; otherwise, it is designated as a noise audio signal. Correspondingly, the constrained beamformer is designated as either a speech-constrained beamformer or a noise-constrained beamformer.
[0225] In applications where speech extraction and isolation are sought, this approach thus separates whether the audio detected by the free-running beamformer 305 (and to which a new constrained beamformer has been assigned) is the desired speech audio source or the undesired noise (and typically non-speech) audio source. In many embodiments, therefore, designating an audio source / audio signal / beam / constrained beamformer as a noise source / signal / beam / constrained beamformer can also be referred to as designating a non-speech source / signal / beam / beamformer, and the following discussion will focus on the separation of speech audio signals from non-speech audio signals.
[0226] The audio device further processes and uses the generated beamformed output audio signal differently depending on how it is specified.
[0227] Specifically, the third adapter 319 controls the adaptation of the adaptive spatial uncorrelator 303 so as to depend on the beamformed output audio signal designated as a non-speech audio signal. That is, the adaptation and updating of the uncorrelation coefficient occurs only when the beamformed output audio signal designated as a non-speech audio signal has a difference scale that satisfies the proximity criterion. In other words, the third adapter 319 updates the adaptive spatial uncorrelator 303 only when the free-running beam acquires an audio signal that is sufficiently close to the beamformed output audio signal considered to be a noise audio signal.
[0228] Therefore, in some embodiments, the constrained beamformer is initialized to track / capture non-speech or noise audio signals, and the adaptive spatial uncorrelatedizer 303 is updated only when the free-running beam captures a matching audio signal. This updates the adaptive spatial uncorrelatedizer 303 and provides the effect of being adapted, in particular, to uncorrelated non-speech audio sources. This has been shown to significantly improve performance, in particular by greatly reducing sensitivity to strong or dominant noise or interfering audio sources in the audio scene.
[0229] In some embodiments, the audio device therefore assigns one (or more) of the constrained beamformers to capture a (typically strong or dominant) noise audio source and uses this to control the adaptation of the adaptive space uncorrelatedizer 303, in particular to uncorrelatedize the noise audio source.
[0230] It will be understood that initializing a constrained beamformer as a noise beamformer may require other requirements, similar to the initialization of the beamformer described above. In particular, it is necessary that the level of the audio signal captured by the first beamformer 305 is sufficiently high. In fact, in many embodiments, it is necessary that the level of the signal captured by the first beamformer 305 is higher than that of any of the currently active constrained beamformers (thus indicating that a new, strong audio source has been detected).
[0231] Similarly, the adaptation of the adaptive spatial uncorrelatedizer 303, and indeed the adaptation of the noise beamformer itself, may require various requirements as described above. In particular, criteria may include requirements such as the level of the constrained beamformed audio output signal being higher than any other constrained beamformer, the level of the constrained beamformed audio output signal being above a given threshold, the level of the point audio source in the constrained beamformed audio output signal being higher than the level of any point audio source in any other constrained beamformed audio output signal, and / or the signal-to-noise ratio of the constrained beamformed audio output signal being above a threshold.
[0232] In many embodiments, the output processor 315 generates an output audio signal that does not include a contribution from the beamformed output audio signal designated as a non-speech audio signal. Therefore, the output signal of the audio device does not include components from the beamformer output audio signal that are considered focused on the noise source. In many embodiments, the audio device therefore includes one or more constrained beamformers assigned purely to capture the noise source. One or more constrained beamformers are used to control the adaptation of the adaptive space uncorrelatedizer 303 so that the adaptive space uncorrelatedizer 303 is adapted to uncorrelatedize the noise source. The constrained beamformers and the beamformer output audio signals can be used solely for this purpose.
[0233] Many techniques and approaches are known for detecting audio signals and classifying them as either spoken or non-spoken audio signals, and it will be understood that an appropriate approach may be used without prejudice to the present invention.
[0234] However, a problem with many such suitable algorithms and techniques is that they tend to be relatively slow, requiring a certain amount of averaging over time before accurate results are obtained. Typically, accurate decisions take several seconds, and this delay before processing a new audio source is often detrimental, for example, in teleconferencing applications.
[0235] As described above, in many embodiments, when a new source is detected by the first beamformer 305, the audio device initializes a constrained beamformer, particularly in many embodiments, depending on the signal level and difference scale of the free-running beamformed output audio signal. Specifically, if there is no constrained beamformer that has a signal level above a threshold and possibly the highest level among any beamformed signals, and has a difference scale below a given threshold (indicating that a new audio source has been detected), a new constrained beamformer is initialized to track the new audio source.
[0236] In many embodiments, the third adapter 319 first designates the constrained beamformed audio output signal resulting from a given constrained beamformer as the spoken audio signal. Thus, whenever a new constrained beamformer is initialized, it is designated to generate the spoken audio signal. This designation is initially given without considering the characteristics of the generated audio signal, and in practice, in all initializations, the beamformer's output audio signal is designated as the spoken audio signal. Therefore, initially, the new signal from the constrained beamformer may be included in the output signal, and the adaptation of the adaptive space decorrelator 303 is not adapted to decorrelate this signal.
[0237] Next, the third adapter 319 applies speech detection to the new beamformed output audio signal to determine whether it actually corresponds to the capture of a speech audio signal. If it does, the third adapter 319 designates the beamformed output audio signal as a speech audio signal and includes it in the output signal. If speech detection does not detect a desired level of speech (based on any appropriate criteria), the third adapter 319 redesignates the beamformer output audio signal from a speech audio signal to a non-speech audio signal. Thus, the redesignation from a speech audio signal to a non-speech audio signal is performed in response to the detection that the beamformed output audio signal does not contain any speech signal components above a threshold level. Therefore, if no speech is detected (or if the detected speech is quiet, for example, compared to other audio components of the beamformer output audio signal), the third adapter 319 changes the designation of the beamformer output audio signal from a speech audio signal to a non-speech audio signal, specifically redesignating it from a speech audio signal to a noise audio signal. As a result, the beamformed output audio signal is removed from the output signal generated by the output processor 315 and instead used to adapt the adaptive space uncorrelator 303, which then adapts the adaptive space uncorrelator 303 to uncorrelate the audio noise source captured by the beamformed output audio signal.
[0238] In some embodiments, the audio device initially treats a new audio source as a spoken audio signal and includes it in the output signal, etc., when a new audio source is detected. However, if the audio source is then determined to be a non-spoken audio source, it is re-designated to be treated as a noise audio signal. Such an approach particularly mitigates the negative effects of delay when performing speech detection. Typically, significant delays, sometimes several seconds, are associated with reliable speech detection (primarily due to the temporal dynamics of speech), and delaying the processing of a new audio source until it is determined whether it provides desired audio or not (specifically, whether it is a spoken or non-spoken) results in a significant delay before the desired audio is heard at the remote end. In the described approach, this is mitigated by initializing the constrained beamformers 309, 311 to the new audio signal so that they are considered to form a candidate or temporary beam / beamformed output audio signal. During the initial time interval for determining whether a candidate beamformed output audio signal is a speech audio signal or a noise audio signal, the beamformed output audio signal is treated / designated as a speech audio signal and therefore included in the output signal. While this may sometimes add undesirable noise to the output signal, it can prevent delays in the start of speech until the speaker becomes audible at the remote end. Thus, a very favorable user experience is achieved.
[0239] Counterintuitively, the audio devices in these embodiments include a constrained beamformer that generates a beamformed output audio signal representing a noise source, and thus attempts to correlate the received signal component of the noise audio source with various signals. However, the adaptive space uncorrelator 303 is adapted (under the control of the generated beamformed output audio signal) to uncorrelate the noise audio source, thereby attempting to achieve, to some extent, the opposite of the noise constrained beamformer. However, it has been found that these two operations interact favorably and synergistically so that the uncorrelator can uncorrelate the noise audio signal to improve the extraction of the speech audio signal by the other constrained beamformer, while leaving sufficient correlation so that the noise constrained beamformer can extract the signal representing the noise audio source so that it can be used to control when the adaptive space uncorrelator 303 adapts.
[0240] As a specific example, the audio device can be based on block processing, which typically corresponds to 256 sample frames for an audio signal sampled at, for example, 16 kHz. For each frame, the outputs of the adaptive space uncorrelatedizer 303, beamformers 305, 309, 311, and adapters 303, 307, 313, 319 make a decision and determine the updated value. As an example, the audio device performs the following operations for each frame: 1) Calculate the output signals of the adaptive space uncorrelatedizer 303 and all beamformers 305, 309, and 311. 2) If a candidate beam exists, the third adapter 319 determines a new state for the candidate beam, which has three possible outcomes, based on the current state of speech detection: a) The beam is determined to be a speech beam, and the constrained beamformer and output signal are designated and treated as such. b) The beam is determined to be a noise beam, and the constrained beamformer and output signal are designated and treated as such. c) A sufficiently reliable decision has not yet been made, and the beam is still treated as a candidate beam that may also be designated as a speech beam / signal. 3) It is determined whether the free-running beamformer 305 has detected a new audio source (for example, there is a sufficiently strong signal and no constrained beamformers 309, 311 with a difference scale below the threshold). 4) If detected, the constrained beamformers 309 and 311 are initialized (for example, if the free-running beamformer 305 produces the strongest signal), and the generated beamformed output audio signal is classified as a candidate signal and temporarily designated as the speech audio signal. If no constrained beamformers 309 and 311 are available, the currently assigned constrained beamformers 309 and 311 are reassigned to the new signal. Initialization includes the coefficients of the free-running beamformer 305, which override the coefficients of the selected constrained beamformers 309 and 311. 5) Next, it is determined which of the constrained beamformers 309 and 311 is applicable (for example, only the constrained beamformers 309 and 311 with the lowest difference scale), and the corresponding beamform coefficients are updated / applied. 6) The third adapter 319 evaluates the beamformed output audio signal, designated as a non-speech audio signal, to determine whether the adaptive spatial uncorrelatedizer 303 can be updated. If it can be updated, the updated uncorrelated coefficient is determined. 7) The output processor 315 determines the output signal by combining it with the beamformer output audio signal designated as the speech audio signal.
[0241] Therefore, the audio device decorrelates the noise and adapts adapter 207 when a sufficiently strong (point) noise source is present. However, the decorrelation adaptation stops when speech is present / dominant, resulting in the decorrelation focusing more on decorrelating noise than speech.
[0242] For this purpose, noise-constrained beamformers 309 and 311 are used to control the adaptation of adapter 207, generating a noise beamform output audio signal that itself adapts to the noise. The noise beamform output audio signal is typically not used for any other purpose and is not included in the output signal. Despite the noise source signal being uncorrelated, it is usually found that there remains enough correlation for the free-running beamformer to focus on the noise when only the noise source is present. As soon as the speech source becomes active, the free-running beamformer 305 typically tracks the speech source very quickly. If the free-running beamformer is tracking towards the existing constrained beam, the distance to the noise source increases and the distance to the controlled beamformer decreases, which can further increase detection sensitivity.
[0243] Before using the noise-constrained beamformers 309 and 311 to update the adaptive spatial uncorrelatedizer (and the noise-constrained beamformers 309 and 311 themselves), it is necessary to determine whether the source detected by the free-running beamformer 305 is a speech source or a noise source. Depending on the characteristics of the noise, it may take several seconds to distinguish between speech and noise, so any new point noise source detected by the free-running beamformer 305 is considered a candidate source and is initially designated and treated as a normal speech source. However, if the source is subsequently detected to be a noise source, it is re-designated as such, usually removed from the output signal, and used to control the adaptation of the adaptive spatial uncorrelatedizer 303.
[0244] As described above, the third adapter 319 analyzes the beamformed output audio signal of the candidate beam to determine whether it is a speech audio signal or a noise audio signal. Therefore, the designation as a speech audio signal or a noise audio signal depends on the characteristics of the beamformed output audio signal.
[0245] In some cases, the third adapter 319 analyzes whether the beamformed output audio signal has certain characteristics known for a given audio source. For example, if the frequency spectrum is determined and it corresponds more to white noise than to, for example, a speech audio signal, then the beamformed output audio signal is determined to be a noise audio signal. Alternatively, the third adapter 319 can further detect whether the beamformed output audio signal is a speech audio signal, for example, by applying a known speech detection algorithm to the beamformed output audio signal.
[0246] Depending on the type of noise, various solutions exist for detecting speech. For example, if the source is known to be continuously active, asymmetric smoothing of the output power of a candidate speech beamformer can be applied. For example: P zz (ω,k+1)=αP zz (ω,k)+(1-α)z(ω,k)z * (ω,k) If (z(ω,k)z * (ω,k) <P zz If (ω,k+1), P zz (ω,k+1)=βP zz (ω,k)+(1-β)z(ω,k)z * (ω,k) This is the result. Here, z is the output signal of the beamformer, where β << α, for example, α = 0.95 and β = 0.1.
[0247] P across all or part of the frequency band zz If, after integrating (ω,k+1), the sum exceeds a certain threshold, it can be determined that the beamformed output audio signal contains noise rather than speech.
number
[0248] Here, lb and hb represent the lower and upper bandwidths, respectively. In this way, noise types that change much more slowly in amplitude and frequency compared to speech can be distinguished.
[0249] In some embodiments, a trained artificial neural network is used, and it has been found that this approach can indeed distinguish a greater number of noise types. The artificial neural network is trained on speech and all types of non-speech, and for each frame, it provides an indication of whether it is noise or speech. 0 indicates a high probability of it being speech, and 1 indicates a high probability of it being noise. If the characteristics of the noise and speech are similar, the artificial neural network can lower the probability of it being noise, for example to 0.6. To address this, the third adapter 319 does not make a decision on a frame-by-frame basis, but only after several frames. This can result in a delay. When the output of the neural network is close to 0 or 1, it can make a decision faster compared to a situation where the output is close to 0.5, for example. If the third adapter 319 has not yet made a decision, the beamformer state remains a candidate speech beam.
[0250] In many embodiments, the audio device includes multiple substantially identical constrained beamformers 309, 311, one of which is assigned as a noise-constrained beamformer 309, 311 that generates a beamformed output audio signal designated as a noise audio signal. In other embodiments, one constrained beamformer 309, 311 is specifically assigned as a noise-constrained beamformer 309, 311. In such cases, when a given beamformed output audio signal is designated as a noise audio signal, it is moved to the dedicated constrained beamformer 309, 311, for example, by copying the beamform coefficients.
[0251] Noise-constrained beamformers 309, 311 (dedicated or dynamically assigned) are adapted according to the beamformer difference scale. Specifically, the approach described for adapting constrained beamformers 309, 311 in general also applies to noise-constrained beamformers 309, 311. Specifically, noise-constrained beamformers 309, 311 are updated when their difference scale falls below a threshold (particularly when it is the lowest overall difference among the constrained beamformers 309, 311). In many embodiments, noise-constrained beamformers 309, 311 are adapted / updated simply when the adaptive space uncorrelatedizer 303 is adapted / updated. That is, noise-constrained beamformers 309, 311 can be adapted when the adaptive space uncorrelatedizer 303 is adapted.
[0252] Therefore, the third adapter 319 determines whether the adaptive space uncorrelatedizer 303 and the noise-constrained beamformers 309, 311 can be updated. A difference scale is used for this purpose, and this difference is calculated between the free-running beamformer and the noise beamformer. The difference scale is specifically bounded between 0.0 (far apart) and 1.0 (completely overlapping). When the adaptive space uncorrelatedizer 303 is at the beginning of convergence, the distance / difference is often close to 1.0. After convergence, the difference scale decreases due to uncorrelatedization, but is usually still high (0.8-0.9). To address this, a smoothed difference / difference is used separately from the difference / difference (called nsdist).
number
number
[0253] Therefore, the specific update rules for the adaptive spatial decorrelator 303 and the beamformers 309, 311 with noise constraints are as follows.
Number
[0254] When there is also a controlled speech beam, since the distance of the speech beam increases when the beam becomes active, the update rule can be further strengthened. The following equation is obtained:
Number
Number
Number
Number
[0255] In some embodiments, the audio device includes an audio source detector that detects a point audio source in the beamformed audio output, and the second adapter adapts the beamforming parameters with constraints only for the beamformer with constraints in which the presence of a point audio source is detected in the beamformed audio output with constraints.
[0256] In some embodiments, the audio source detector further detects a point audio source within a first beamformed audio output. The audio device further includes a controller that, if a point audio source is detected within the first beamformed audio output but not within any constrained beamformed audio output, sets the constrained beamform parameters of the first constrained beamformer according to the beamform parameters of the first beamformer.
[0257] In some embodiments, the controller sets the constrained beamform parameters of the first constrained beamformer according to the beamform parameters of the first beamformer, only if the difference scale of the first constrained beamformer exceeds a threshold.
[0258] In some embodiments, the audio source detector further detects an audio source within a first beamformed audio output. The audio device further includes a controller that sets the constrained beamform parameters of the first constrained beamformer according to the beamform parameters of the first beamformer when a point audio source is detected within a first beamformed audio output and a second beamformed audio output output from a first constrained beamformer, and a difference scale exceeding a threshold is determined relative to the first constrained beamformer.
[0259] In some embodiments, the constrained beamformers are an active subset of constrained beamformers selected from a pool of constrained beamformers, and the controller increases the number of active constrained beamformers to include the first constrained beamformer by initializing constrained beamformers from the pool of constrained beamformers using the beamform parameters of the first beamformer.
[0260] In some embodiments, the second adapter further adapts only the constrained beamform parameters of the first constrained beamformer if a criterion is met that includes at least one requirement selected from the following group: the requirement that the level of the second beamformed audio output from the first constrained beamformer is higher than that of any other second beamformed audio output; the requirement that the level of the point audio source in the second beamformed audio output from the first constrained beamformer is higher than that of any other second beamformed audio output; the requirement that the signal-to-noise ratio of the second beamformed audio output from the first constrained beamformer exceeds a threshold; and the requirement that the second beamformed audio output from the first constrained beamformer includes speech components.
[0261] In some embodiments, the difference processor determines the difference scale of the first constrained beamformer and reflects at least one of the following: The difference between the first parameter set and the constrained parameter set of the first constrained beamformer, and The difference between the first beamformed audio output from the first constrained beamformer and the constrained beamformed audio output.
[0262] In some embodiments, the adaptability of the first beamformer is higher than that of multiple constrained beamformers.
[0263] In some embodiments, the first beamformer and the multiple constrained beamformers are filter-and-combine beamformers.
[0264] In some embodiments, the first beamformer is a filter-and-combine beamformer comprising a plurality of first beamform filters, each having a first adaptive impulse response, and the second beamformer, which is one of a plurality of constrained beamformers, is a filter-and-combine beamformer comprising a plurality of second beamform filters, each having a second adaptive impulse response. The difference processor determines the difference scale between the beams of the first beamformer and the second beamformer in response to a comparison of the first adaptive impulse response and the second adaptive impulse response.
[0265] In some embodiments, the audio device further comprises a noise-reference beamformer that generates a beamformed audio output signal and at least one noise-reference signal, the noise-reference beamformer being one of a first beamformer and a plurality of constrained beamformers; a first transformer for generating a first frequency-domain signal from the frequency conversion of the beamformed audio output signal, the first frequency-domain signal being represented by time-frequency tile values; and a second transformer for generating a second frequency-domain signal from the frequency conversion of at least one noise-reference signal, the second frequency-domain signal being represented by time-frequency tile values. Lance, a difference processor for generating a time-frequency tile difference scale, wherein the time-frequency tile difference scale for a first frequency represents the difference between a first monotonic function of the norm of the time-frequency tile value of a first frequency domain signal for a first frequency and a second monotonic function of the norm of the time-frequency tile value of a second frequency domain signal for a first frequency, and a point audio source estimator for generating a point audio source estimate indicating whether a beamformed audio output signal contains a point audio source, wherein the point audio source estimator generates a point audio source estimate according to the combined difference values of the time-frequency tile difference scale for frequencies above a frequency threshold.
[0266] In some embodiments, the point audio source estimator detects the presence of a point audio source within the beamformed audio output when the combined difference value exceeds a threshold.
[0267] Figure 4 is a block diagram showing an exemplary processor 400 according to an embodiment of the present disclosure. Using the processor 400, one or more processors can be implemented to implement the aforementioned devices or elements thereof (in particular, including one or more artificial neural networks). The processor 400 may be any suitable processor type, including, but not limited to, a microprocessor, a microcontroller, a digital signal processor (DSP), a field-programmable array (FPGA) (the FPGA is programmed to form a processor), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC) (the ASIC is designed to form a processor), or a combination thereof.
[0268] The processor 400 may include one or more cores 402. A core 402 may include one or more arithmetic logic units (ALUs) 404. In some embodiments, the core 402 includes, in addition to or instead of, the ALUs 404, a floating-point logic unit (FPLU) 406 and / or a digital signal processing unit (DSPU) 408.
[0269] The processor 400 may include one or more registers 412 that are communicatively coupled to the core 402. The registers 412 may be implemented using dedicated logic gate circuits (e.g., flip-flops) and / or any memory technology. In some embodiments, the registers 412 are implemented using static memory. The registers can provide data, instructions, and addresses to the core 402.
[0270] In some embodiments, the processor 400 may include one or more levels of cache memory 410 communicably coupled to the core 402. The cache memory 410 may provide computer-readable instructions to the core 402 for execution. The cache memory 410 may provide data for the core 402 to process. In some embodiments, computer-readable instructions may be provided to the cache memory 410 by local memory (e.g., local memory attached to an external bus 416). The cache memory 410 may be implemented in any suitable cache memory type, such as metal oxide semiconductor (MOS) memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), and / or any other suitable memory technology.
[0271] The processor 400 may include a controller 414. The controller 414 may control inputs to the processor 400 from other processors and / or components included in the system, and / or outputs from the processor 400 to other processors and / or components included in the system. The controller 414 may control data paths in the ALU 404, FPLU 406, and / or DSPU 408. The controller 414 may be implemented as one or more state machines, data paths, and / or dedicated control logic. The gates of the controller 414 may be implemented as standalone gates, FPGAs, ASICs, or any other suitable technology.
[0272] The registers 412 and cache 410 can communicate with the controller 414 and core 402 via internal connections 420A, 420B, 420C, and 420D. The internal connections can be implemented as buses, multiplexers, crossbar switches, and / or any other suitable connection technology.
[0273] The input / output of the processor 400 is provided via the bus 416, which may include one or more conductive lines. The bus 416 may be communicatively coupled to one or more components of the processor 400, such as, for example, the controller 414, the cache 410, and / or the register 412. The bus 416 may be coupled to one or more components of the system.
[0274] The bus 416 may be coupled to one or more external memories. The external memory may include a read-only memory (ROM) 432. The ROM 432 may be a masked ROM, an electronically programmable read-only memory (EPROM), or any other suitable technology. The external memory may include a random access memory (RAM) 433. The RAM 433 may be a static RAM, a battery-backed static RAM, a dynamic RAM (DRAM), or any other suitable technology. The external memory may include an electrically erasable programmable read-only memory (EEPROM (registered trademark)) 435. The external memory may include a flash memory 434. The external memory may include a magnetic storage device such as a disk 436. In some embodiments, the external memory may be included in the system.
[0275] It will be understood that the foregoing description, for purposes of clarity, has described embodiments of the invention with reference to various functional circuits, units, and processors. However, it is clear that the functions can be appropriately distributed among the various functional circuits, units, or processors without detracting from the invention. For example, functions described as being performed by separate processors or controllers may be performed by the same processor or controller. Accordingly, references to specific functional units or circuits are to be regarded only as references to suitable means for providing the described functionality, and not as indicating any strict logical or physical structure or organization.
[0276] The present invention can be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. The present invention may optionally be implemented at least partially as computer software running on one or more data processors and / or digital signal processors. Elements and components of embodiments of the present invention can be implemented physically, functionally, and logically in any suitable way. In fact, functionality can be implemented as part of one unit, multiple units, or other functional units. Therefore, the present invention can be implemented in a single unit or physically and functionally distributed across different units, circuits, and processors.
[0277] While the present invention has been described in relation to several embodiments, it is not intended to be limited to any particular form described herein. Rather, the scope of the present invention is limited only by the appended claims. Furthermore, while some features may appear to be described in relation to a particular embodiment, those skilled in the art will recognize that various features of the described embodiments may be combined in accordance with the present invention. In the claims, the term “comprising” does not preclude the existence of other elements or steps.
[0278] Furthermore, although listed individually, multiple means, elements, circuits, or method steps can be implemented, for example, by a single circuit, unit, or processor. Additionally, individual features may be included in different claims, but these features can also be advantageously combined, and inclusion in various claims does not imply that such combinations are unfeasible and / or unfavorable. Furthermore, inclusion of a feature in one category of claims does not imply limitation to that category, but rather indicates that the feature can be similarly applied to other claim categories as needed. Moreover, the order of features in a claim does not imply a specific order in which the features must function, and in particular, the order of individual steps in a method claim does not imply that the steps must be performed in this order. Rather, the steps can be performed in any suitable order. Also, singular references do not preclude plural references; therefore, references such as "first," "second," etc., do not preclude plurals. Reference numerals in the claims are provided only as examples for clarity, and these examples should not be construed as limiting the claims in any way.
Claims
1. A receiver that receives a first set of audio signals, An adaptive spatial uncorrelatedizer that applies spatial uncorrelated filtering to the first set of audio signals to generate a second set of audio signals, A first beamformer coupled to the adaptive space uncorrelatedizer generates a first beamformed audio output signal from the second set of audio signals, A plurality of constrained beamformers coupled to the adaptive space uncorrelatedizer, each generating a constrained beamformed audio output signal from the second set of audio signals, wherein each constrained beamformer has constrained beamform parameters, and the constrained beamform parameters of the plurality of constrained beamformers form a set of beamform parameters, An output processor that generates an output audio signal from the constrained beamforming audio output signal, A first adapter for adapting the beamform parameters of the first beamformer, A difference processor that determines a difference scale for each of the plurality of constrained beamformers, wherein the difference scale for each of the plurality of constrained beamformers represents the difference between the beam formed by the first beamformer and the beam formed by each of the plurality of constrained beamformers, A second adapter for applying the constrained beamform parameters from the set of constrained beamform parameters to the plurality of constrained beamformers, with the constraint that the constrained beamform parameters from the set of constrained beamform parameters are applied only to the plurality of constrained beamformers whose difference scale is determined to satisfy the similarity criterion, A third adapter that adapts the coefficients of the spatially uncorrelated filtering according to an update value determined from the first set of audio signals, wherein the update value is determined to reduce the correlation of the second set of audio signals. An audio device equipped with the ability to capture audio.
2. The audio device according to claim 1, wherein the third adapter specifies at least the first constrained beamformed audio output signal of the first constrained beamformer among the plurality of constrained beamformers as a speech audio signal or a noise audio signal, and the third adapter updates the coefficient of the spatial uncorrelated filtering according to the difference scale of the first constrained beamformer when the first constrained beamformed audio output signal is specified as a noise audio signal.
3. The audio apparatus according to claim 2, wherein the output processor generates the output audio signal such that it does not include a contribution from the first constrained beamformed audio output signal when the first constrained beamformed audio output signal is designated as a noise audio signal.
4. The audio device according to claim 2 or 3, wherein the third adapter, in response to detection that none of the plurality of constrained beamformers has a difference scale that satisfies the similarity criterion, initializes a given constrained beamformer among the plurality of constrained beamformers with beamform parameters that match the beamform parameters of the first beamformer.
5. The audio device according to claim 4, wherein initializing a given constrained beamformer first includes designating a given constrained beamformed audio output signal of the given constrained beamformer as a spoken audio signal.
6. The audio device according to claim 5, wherein the third adapter applies speech detection to the given constrained beamformed audio output signal, and, upon detection that the given constrained beamformed audio output signal does not contain speech signal components at a level above a threshold, reclassifies the given constrained beamformed audio output signal from a speech audio signal to a noise audio signal.
7. The audio apparatus according to any one of claims 4 to 6, wherein the third adapter initializes the given constrained beamformer only if the level of the first beamformed audio output signal exceeds the level of the constrained beamformed audio output signals from all the constrained beamformers.
8. The second adapter described above is The requirement is that the level of the constrained beamformed audio output signal of the given constrained beamformer is higher than that of any other constrained beamformer. The requirement is that the level of the constrained beamformed audio output signal of the given constrained beamformer exceeds a given threshold, The requirement is that the level of a point audio source in the constrained beamformed audio output signal of the given constrained beamformer is higher than that of any point audio source in any other constrained beamformed audio output signal, and The requirement is that the signal-to-noise ratio of the constrained beamformed audio output signal of the given constrained beamformer exceeds a threshold. The audio device according to any one of claims 1 to 7, wherein only the constrained beamform parameters of a given set of constrained beamform parameters are applied if a criterion is met that includes at least one requirement selected from the group.
9. The audio device according to claim 2, wherein the maximum number of constrained beamformers updated simultaneously is 1.
10. The adaptive space uncorrelatedizer links each audio signal of the second set of audio signals with one audio signal of the first set of audio signals, and segments the first set of audio signals into time segments, for at least some time segments, A step of generating a frequency bin representation of the first set of audio signals, wherein each frequency bin of the frequency bin representation of the first set of audio signals includes a frequency bin value for each of the audio signals of the first set of audio signals. A segmentation step, comprising the steps of generating a frequency bin representation of the second set of audio signals, wherein each frequency bin of the frequency bin representation of the second set of audio signals includes a frequency bin value for each of the second set of audio signals, and the frequency bin values for a given audio signal of the second output set of audio signals for a given frequency bin are generated as a weighted combination of the frequency bin values of the first set of audio signals for the given frequency bin, A step of updating a first weight of the contribution of a second input audio signal linked to a second output signal from the second frequency bin value of the first frequency bin of a first output signal linked to a first input audio signal to the first frequency bin value of the first frequency bin, according to a correlation scale between the first previous frequency bin value of the first output signal with respect to the first frequency bin and the second previous frequency bin value of the second output signal with respect to the first frequency bin. By doing so, the second set of audio signals is generated, The audio apparatus according to any one of claims 1 to 9, wherein the third adapter updates the weights of the weight combination and updates the first weight of the contribution of the second frequency bin of the first frequency bin of the second audio signal in the set of second audio signals linked to the second output audio signal in the set of second audio signals, to the first frequency bin of the first frequency bin of the first output audio signal in the set of first audio signals, linked to the first audio signal in the set of first audio signals, according to a correlation scale between the first previous frequency bin of the first output audio signal and the second previous frequency bin of the second output audio signal with respect to the first frequency bin.
11. The audio apparatus according to claim 10, wherein the third adapter updates the first weight according to the product of a first value and a second value, the first value being one of the first previous frequency bin value and the second previous frequency bin value, and the second value being the composite conjugate of the other of the first previous frequency bin value and the second previous frequency bin value.
12. The third adapter described above is y(ω)=W(ω)x(ω) The audio device according to claim 10 or 11, wherein the output bin value of the given frequency bin ω is determined from, where y(ω) is a vector containing the frequency bin value of the output audio signal of the given frequency bin ω, x(ω) is a vector containing the frequency bin value of the input audio signal of the given frequency bin ω, and W(ω) is a matrix having rows containing the weights of the weighted combination of the output audio signal.
13. The third adapter is the weight w of the matrix W(ω). ij of [Number 18] The audio device according to claim 12, wherein the device is adapted according to the following, where i is the row index of the matrix W(ω), j is the column index of the matrix W(ω), k is the time segment index, ω represents the frequency bin, and η(k,ω) is a scaling parameter for adapting the adaptation speed.
14. The steps include receiving a first set of audio signals, The steps include applying spatially uncorrelated filtering to the first set of audio signals to generate a second set of audio signals, The first beamformer generates a first beamformed audio output signal from the second set of audio signals, A step of generating a constrained beamformed audio output signal from a set of second audio signals, wherein each of a plurality of constrained beamformers has a set of constrained beamform parameters, The steps include generating an output audio signal from the constrained beamforming audio output signal, The steps include: adapting the beamform parameters of the first beamformer; A step of determining a difference scale for each of the plurality of constrained beamformers, wherein the difference scale represents the difference between the beam formed by the first beamformer and the beam formed by each of the plurality of constrained beamformers. A step of applying the constrained beamform parameters from the set of constrained beamform parameters to the plurality of constrained beamformers, with the constraint that the constrained beamform parameters from the set of constrained beamform parameters are applied only to the plurality of constrained beamformers whose difference scale is determined to satisfy the similarity criterion, A step of adapting the coefficients of the spatial uncorrelated filtering according to an update value determined from the first set of audio signals, wherein the update value is determined to reduce the correlation of the second set of audio signals. How to import audio, including [specific audio formatting].
15. A computer program comprising computer program code means, wherein the computer program code means is adapted to perform all the steps of the method according to claim 14 when the computer program is executed on a computer.