Multi-source tracking and voice activity detection for planar microphone arrays

By employing multidimensional TDOA tracking and VAD mechanisms on a planar microphone array, the problem of separating the target audio signal from the noise source in a multi-stream audio environment is solved, achieving efficient enhancement and noise suppression of the target audio signal, applicable to any array geometry.

CN113113034BActive Publication Date: 2026-04-10SYNAPTICS INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SYNAPTICS INC
Filing Date
2021-01-08
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In existing multi-stream audio environments, separating the target audio signal from the noise source is difficult to perform efficiently. Existing methods are insufficient in terms of computational complexity and accuracy, especially for microphone arrays with non-specific array geometries.

Method used

By employing multi-dimensional TDOA tracking and VAD mechanisms, a microphone array is configured on a plane to utilize TDOA trajectory information for multi-source tracking and noise suppression, thereby avoiding ghosting and reducing computational complexity.

Benefits of technology

It achieves efficient enhancement and noise suppression of target audio signals, is applicable to arbitrary array geometries, reduces computational complexity, and improves processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113113034B_ABST
    Figure CN113113034B_ABST
Patent Text Reader

Abstract

Embodiments described herein provide a combined multi-source time difference of arrival (TDOA) tracking and voice activity detection (VAD) mechanism that can be applicable to general array geometries, e.g., microphone arrays located on a plane. The combined multi-source TDOA tracking and VAD mechanism scans azimuth and elevation angles of the microphone array in microphone pairs based on which a planar locus of physically allowable TDOAs can be formed in a multi-dimensional TDOA space of multiple microphone pairs. In this way, the multi-dimensional TDOA tracking reduces the number of computations typically involved in conventional TDOA by performing a TDOA search separately for each dimension.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] In accordance with one or more embodiments, the present disclosure relates generally to audio signal processing, and more particularly, for example, to systems and methods for multi-source tracking and multi-stream voice activity detection for general planar microphone arrays. BACKGROUND

[0002] In recent years, smart speakers and other voice-controlled devices and apparatuses have gained popularity. A smart speaker typically includes a microphone array for receiving audio input (e.g., a user’s spoken command) from an environment. When a target audio (e.g., a spoken command) is detected in the audio input, the smart speaker can translate the detected target audio into one or more commands and perform different tasks based on the commands.

[0003] One challenge with these smart speakers is to efficiently and effectively isolate the target audio (e.g., a spoken command) from noise or other active speakers in an operating environment. For example, in the presence of one or more noise sources, one or more speakers can be active. When the goal is to enhance a particular speaker, the speaker is referred to as a target speaker, while the rest of the speakers can be considered as interference sources. Existing voice enhancement algorithms mostly use multiple input channels (microphones) to exploit the spatial information of sources, such as blind source separation (BSS) methods related to independent component analysis (ICA) and spatial filtering or beamforming methods.

[0004] However, BSS methods are mainly designed for batch processing, which can be generally undesirable or even inapplicable in real applications due to large response latency. On the other hand, spatial filtering or beamforming methods usually require supervision under voice activity detection (VAD) as a cost function to be minimized, which can be overly dependent on the estimation of the covariance matrix related to noise / interference-only segments.

[0005] Accordingly, there is a need for improved systems and methods for detecting and processing target audio signal(s) in a multi-stream audio environment. BRIEF DESCRIPTION OF DRAWINGS

[0006] Aspects of the disclosure, together with its advantages, can be better understood by reference to the following figures and detailed description. It will be appreciated that same reference numerals are used to designate the same elements throughout the figures and the detailed description, in which the drawings are for purposes of illustrating the embodiments of the present disclosure and not for purposes of limiting the same. The components in the figures are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the present disclosure.

[0007] Figure 1FIG. illustrates an example operating environment for an audio processing device, in accordance with one or more embodiments of the disclosure.

[0008] Figure 2 is a block diagram of an example audio processing device, in accordance with one or more embodiments of the disclosure.

[0009] Figure 3 is a block diagram of an example audio signal processor for multi-track audio enhancement, in accordance with one or more embodiments of the disclosure.

[0010] Figure 4 is a block diagram of an example multi-track activity detection engine for processing multiple audio signals from a general microphone array, in accordance with various embodiments of the disclosure.

[0011] Figure 5A is a diagram illustrating example geometries of pairs of microphones, in accordance with one or more embodiments of the disclosure.

[0012] Figure 5B is a diagram illustrating an example grid of time-difference-of-arrival (TDOA) trajectory information corresponding to different microphone array geometries in a multi-dimensional space, in accordance with one or more embodiments of the disclosure.

[0013] Figure 6 is a logic flow diagram of an example method for enhancing multi-source audio signals through multi-source tracking and activity detection, in accordance with various embodiments of the disclosure.

[0014] Figure 7 is a logic flow diagram of an example process for computing TDOA trajectory information in a multi-dimensional space using a pair of microphones, in accordance with various embodiments of the disclosure. DETAILED DESCRIPTION

[0015] The present disclosure provides improved systems and methods for detecting and processing target audio signal(s) in a multi-stream audio environment.

[0016] Voice activity detection (VAD) can be used to supervise speech enhancement of a target audio in a process that utilizes spatial information of sources from multiple input channels. VAD can allow spatial statistics of interfering / noisy sources to be induced during silent periods of a desired speaker, such that when the desired speaker becomes active, the influence of noise / interference can then be eliminated. For example, VAD for each source can be derived to track spatial information of sources in the form of time-difference-of-arrival (TDOA) or direction-of-arrival (DOA) by constructing VAD and utilizing history of detections by determining when a detection occurs near an existing track. This process is often referred to as measurement-to-track (M2T) assignment. In this way, multiple VADs can be derived for all sources of interest.

[0017] In particular, existing DOA methods typically construct a single steering vector for the entire microphone array based on a closed-form mapping of azimuth and elevation angles, which can be used to exploit the special geometry of linear or circular arrays. Such DOA methods cannot be extended to general or arbitrary geometry of the microphone array. In addition, these closed-form mapping based DOA methods typically require extensive search in multi-dimensional space. For arbitrary geometry, existing TDOA based methods can be used, which can not be limited to a particular array geometry and can construct multiple steering vectors for each microphone pair to form a multi-dimensional TDOA vector (one dimension per pair). However, these existing methods risk introducing TDOA ghosting formed by the cross-intersection of peaks from the spectra of each TDOA pair. Thus, further post-processing involving the special array geometry is typically required to remove the TDOA ghosting.

[0018] In view of the need for multi-stream VAD that is not constrained by a particular array geometry, embodiments described herein provide a combined multi-source TDOA tracking and VAD mechanism that can be applicable to general array geometries (e.g., microphone arrays lying on a plane). By performing the TDOA search separately for each dimension, the combined multi-source TDOA tracking and VAD mechanism can reduce the number of computations typically involved in conventional TDOA.

[0019] In some embodiments, a multi-dimensional TDOA method for general array geometry lying on a plane is employed that avoids unwanted ghost TDOAs. In one embodiment, Cartesian coordinates of the generally configured microphones are obtained, and one of the microphones can be selected as a reference microphone. Azimuth and elevation angles of the microphones can be scanned, based on which a planar locus of physically allowable TDOAs can be formed in the multi-dimensional TDOA space of multiple microphone pairs. In this way, the formed planar locus avoids the formation of ghost TDOAs, and thus further post-processing to remove the ghost TDOAs is not required. Moreover, compared to the full DOA scanning method, the multi-dimensional TDOA method disclosed herein reduces computational complexity by performing the search separately in the paired TDOA domain related to each dimension, rather than in the full multi-dimensional space.

[0020] Figure 1 An example operating environment 100 in which an audio processing system according to various embodiments of the present disclosure can operate is illustrated. The operating environment 100 includes an audio processing device 105, a target audio source 110, and one or more noise sources 135-145. In Figure 1In the example illustrated in the middle, the operating environment 100 is illustrated as a room, but it is contemplated that the operating environment can include other areas, such as a vehicle interior, an office conference room, a home room, an outdoor stadium, or an airport. According to various embodiments of the present disclosure, the audio processing device 105 can include two or more audio sensing components (e.g., microphones) 115a-115d and, optionally, one or more audio output components (e.g., speakers) 120a-120b.

[0021] The audio processing device 105 can be configured to sense sound via the audio sensing components 115a-115d and generate a multi-channel audio input signal including two or more audio input signals. The audio processing device 105 can process the audio input signals using the audio processing techniques disclosed herein to enhance the audio signals received from the target audio source 110. For example, the processed audio signals can be transmitted to other components within the audio processing device 105, such as a speech recognition engine or a voice command processor, or to an external device. Thus, the audio processing device 105 can be a standalone device that processes audio signals, or can be a device that converts the processed audio signals into other signals (e.g., commands, instructions, etc.) for interacting with or controlling an external device. In other embodiments, the audio processing device 105 can be a communication device, such as a mobile phone or a Voice over IP (VoIP) enabled device, and the processed audio signals can be transmitted over a network to another device for output to a remote user. The communication device can also receive processed audio signals from a remote device and output the processed audio signals via the audio output components 120a-120b.

[0022] The target audio source 110 can be any source that produces sound that can be detected by the audio processing device 105. The target audio to be detected by the system can be defined based on criteria specified by a user or system requirements. For example, the target audio can be defined as human speech, sounds emitted by a particular animal or machine. In the illustrated example, the target audio is defined as human speech, and the target audio source 110 is a person. In addition to the target audio source 110, the operating environment 100 can include one or more noise sources 135-145. In various embodiments, sounds that are not target audio can be processed as noise. In the illustrated example, the noise sources 135-145 can include a speaker 135 playing music, a television 140 playing a television program, a movie, or a sporting event, and background conversations between non-target speakers 145. It will be appreciated that different noise sources can be present in various operating environments.

[0023] It is noted that the target audio and noise can arrive at the audio sensing components 115a-d of the audio processing device 105 from different directions and at different times. For example, the noise sources 135-145 can produce noise at different locations within the operating environment 100, and the target audio source (person) 110 can speak while moving between locations within the operating environment 100. In addition, the target audio and / or noise can reflect off of fixtures within the room 100 (e.g., walls). For example, consider the paths that the target audio can take from the target audio source 110 to reach each of the audio sensing components 115a-d. As indicated by arrows 125a-d, the target audio can propagate directly from the target audio source 110 to the audio sensing components 115a-d, respectively. In addition, the target audio can reflect off of walls 150a and 150b and reach the audio sensing components 115a-d indirectly from the target audio source 110, as indicated by arrows 130a-b. In various embodiments, the audio processing device 105 can use one or more audio processing techniques to estimate and apply a room impulse response to further enhance the target audio and suppress the noise.

[0024] Figure 2 An example audio processing device 200 is illustrated in accordance with various embodiments of the present disclosure. In some embodiments, the audio processing device 200 can be implemented as the audio processing device 105 of Figure 1 FIG. 1. The audio processing device 200 includes an audio sensor array 205, an audio signal processor 220, and a host system component 250.

[0025] The audio sensor array 205 includes two or more sensors, each of which can be implemented as a transducer that converts audio input having the form of a sound wave into an audio signal. In the illustrated environment, the audio sensor array 205 includes a plurality of microphones 205a-n, each of which generates an audio input signal that is provided to an audio input circuit 222 of the audio signal processor 220. In one embodiment, the audio sensor array 205 generates a multi-channel audio signal, where each channel corresponds to an audio input signal from one of the microphones 205a-n.

[0026] The audio signal processor 220 includes an audio input circuit 222, a digital signal processor 224, and optional audio output circuit 226. In various embodiments, the audio signal processor 220 can be implemented as an integrated circuit that includes analog circuitry, digital circuitry, and the digital signal processor 224, which is operable to execute program instructions stored in firmware. The audio input circuit 222, for example, can include an interface to the audio sensor array 205, anti-aliasing filters, analog-to-digital converter circuitry, echo cancellation circuitry, and other audio processing circuitry and components. The digital signal processor 224 is operable to process multi-channel digital audio signals to generate enhanced audio signals that are output to one or more host system components 250. In various embodiments, the digital signal processor 224 can be operable to perform echo cancellation, noise cancellation, target signal enhancement, post-filtering, and other audio signal processing functions.

[0027] The optional audio output circuit 226 processes audio signals received from the digital signal processor 224 for output to at least one loudspeaker, such as the loudspeakers 210a and 210b. In various embodiments, the audio output circuit 226 can include digital-to-analog converters to convert one or more digital audio signals to analog audio signals and one or more amplifiers to drive the loudspeakers 210a-210b.

[0028] The audio processing device 200 can be implemented as any device operable to receive and enhance target audio data, such as, for example, a mobile phone, a smart speaker, a tablet, a laptop computer, a desktop computer, a voice-controlled appliance, or a car. The host system components 250 can include various hardware and software components for operating the audio processing device 200. In the illustrated embodiment, the host system components 250 include a processor 252, user interface components 254, a communication interface 256 for communicating with external devices and networks, such as a network 280 (e.g., the Internet, a cloud, a local area network, or a cellular network) and a mobile device 284, and a memory 258.

[0029] The processor 252 and the digital signal processor 224 can include one or more of a processor, a microprocessor, a single-core processor, a multi-core processor, a microcontroller, a programmable logic device (PLD) (e.g., a field-programmable gate array (FPGA)), a digital signal processing (DSP) device, or other logic device (which can be configured by hardwiring, executing software instructions, or a combination of the two) to perform the various operations discussed herein for embodiments of the present disclosure. The host system components 250 are configured to interface and communicate with the audio signal processor 220 and other host system components 250, such as through a bus or other electronic communication interface.

[0030] It will be appreciated that, although the audio signal processor 220 and host system components 250 are shown as incorporating a combination of hardware components, circuitry, and software, in some embodiments, at least some or all of the functionality that the hardware components and circuitry are operable to perform can be implemented as software modules that are executed by the processor 252 and / or the digital signal processor 224 in response to software instructions and / or configuration data stored in the firmware or memory 258 of the digital signal processor 224.

[0031] The memory 258 can be implemented as one or more storage devices operable to store data and information including audio data and program instructions. The memory 258 can include one or more various types of storage devices including volatile and non-volatile storage devices such as RAM (random access memory), ROM (read only memory), EEPROM (electrically erasable programmable read only memory), flash memory, hard drives, and / or other types of memory.

[0032] The processor 252 can be operable to execute software instructions stored in the memory 258. In various embodiments, the speech recognition engine 260 is operable to process the enhanced audio signals received from the audio signal processor 220 including recognizing and executing voice commands. The voice communication component 262 can be operable to facilitate voice communications with one or more external devices such as the mobile device 284 or the user device 286, such as through a voice call over a mobile or cellular telephone network or a VoIP call over an IP network. In various embodiments, the voice communications include transmitting the enhanced audio signals to the external communication devices.

[0033] The user interface component 254 can include a display, a touchpad display, a keypad, one or more buttons, and / or other input / output components operable to enable a user to directly interact with the audio processing device 200.

[0034] The communication interface 256 facilitates communications between the audio processing device 200 and external devices. For example, the communication interface 256 can enable Wi-Fi (e.g., 802.11) or Bluetooth connectivity between the audio processing device 200 and one or more local devices such as the mobile device 284 or a wireless router providing network access to a remote server 282 such as through the network 280. In various embodiments, the communication interface 256 can include other wired and wireless communication components that facilitate direct or indirect communications between the audio processing device 200 and one or more other devices.

[0035] Figure 3An example audio signal processor 300 is illustrated in accordance with various embodiments of the present disclosure. In some embodiments, the audio signal processor 300 is embodied as one or more integrated circuits that include analog and digital circuitry and firmware logic implemented by a digital signal processor (such as the audio signal processor 220 of FIG. 4) or other processing device. As illustrated, the audio signal processor 300 includes an audio input circuit 315, a sub-band frequency analyzer 320, a multi-track VAD engine 325, an audio enhancement engine 330, and a synthesizer 335. Figure 2

[0036] The audio signal processor 300 receives a multi-channel audio input from a plurality of audio sensors, such as a sensor array 305 that includes at least two audio sensors 305a-n. The audio sensors 305a-305n can include microphones integrated with an audio processing device, such as the audio processing device 200 of FIG. 4, or external components connected thereto. According to various embodiments of the present disclosure, the audio signal processor 300 can or can not be aware of the arrangement of the audio sensors 305a-305n. Figure 2

[0037] The audio signal can be initially processed by the audio input circuit 315, which can include anti-aliasing filters, analog-to-digital converters, and / or other audio input circuitry. In various embodiments, the audio input circuit 315 outputs a digital, multi-channel, time-domain audio signal, where M is the number of sensor (e.g., microphone) inputs. The multi-channel audio signal is input to the sub-band frequency analyzer 320, which divides the multi-channel audio signal into successive frames and decomposes each frame of each channel into a plurality of frequency sub-bands. In various embodiments, the sub-band frequency analyzer 320 includes a Fourier transform process and outputs a plurality of frequency windows. The decomposed audio signal is then provided to the multi-track VAD engine 325 and the audio enhancement engine 330.

[0038] The multi-track VAD engine 325 is operable to analyze frames of one or more audio tracks and generate a VAD output indicating whether target audio activity is present in the current frame. As discussed above, the target audio can be any audio to be recognized by the audio system. When the target audio is human speech, the multi-track VAD engine 325 can be embodied to detect speech activity. In various embodiments, the multi-track VAD engine 325 is operable to receive a frame of audio data and generate, for each audio track, a VAD indication output regarding the presence or absence of the target audio on the corresponding audio track corresponding to the frame of audio data. Regarding the multi-track VAD engine 325, further details of the detailed components and operations of the multi-track VAD engine 325 are illustrated in FIG. 4. Figure 4

[0039] ​​​The audio enhancement engine 330 receives subband frames from the subband frequency analyzer 320 and receives VAD indications from the multi-track VAD engine 325. According to various embodiments of the present disclosure, the audio enhancement engine 330 is configured to process the subband frames based on the received multi-track VAD indications to enhance the multi-track audio signals. For example, the audio enhancement engine 330 can enhance portions of the audio signals that are determined to be from the direction of the target audio source and suppress other portions of the audio signals that are determined to be noise.

[0040] After enhancing the target audio signals, the audio enhancement engine 330 can pass the processed audio signals to the synthesizer 335. In various embodiments, the synthesizer 335 reconstructs one or more multi-channel audio signals on a frame-by-frame basis by combining the subbands to form an enhanced time-domain audio signal. The enhanced audio signal can then be transformed back to the time domain and sent to system components or external devices for further processing.

[0041] Figure 4 FIG. illustrates an example multi-track VAD engine 400 for processing multiple audio signals from a general microphone array, according to various embodiments of the present disclosure. The multi-track VAD engine 400 can be implemented as a combination of digital circuitry and logic executed by a digital signal processor. In some embodiments, the multi-track VAD engine 400 can be installed in an audio signal processor such as the 300 in Figure 3 The multi-track VAD engine 400 can provide further structural and functional details to the multi-track VAD engine 325 in Figure 3 The multi-track VAD engine 400 can provide further structural and functional details to the multi-track VAD engine 325 in

[0042] According to various embodiments of the present disclosure, the multi-track VAD engine 400 includes a subband analysis module 405, a block-based TDOA estimation module 410, a TDOA track computation module 420, and a multi-source tracking and multi-stream VAD estimation module 430.

[0043] The subband analysis module 405 receives multiple audio signals 402 represented by , i.e., time-domain audio signals of samples recorded at the mth microphone for a total of M microphones (e.g., similar to the audio sensors 305a-n in Figure 3 The audio signals 402 can be received via the audio input circuitry 315 in Figure 3 The audio signals 402 can be received via the audio input circuitry 315 in , .

[0044] The subband analysis module 405 is configured to obtain the audio signals 402 and transform them into a time-frequency domain representation 404, which is represented as corresponding to the original time-domain audio signals , where indicating a subband time index, and indicating a frequency band index. For example, the subband analysis module 405 can be similar to the subband frequency analyzer 320 in Figure 3 which performs a Fourier transform to convert an input time-domain audio signal into a frequency-domain representation. The subband analysis module 405 can then send the generated time-frequency domain representation 404 to the block-based TDOA estimation module 410 and the multi-source tracking and multi-stream VAD estimation module 430.

[0045] The TDOA trajectory computation module 420 is configured to scan a general microphone array (e.g., the audio sensors 305a-n forming a general array geometry). For example, for a given arbitrary microphone array geometry on a plane, a trajectory of admissible TDOA positions is computed once at system startup. This trajectory of points can avoid ghost formation.

[0046] For an array of M microphones, a first microphone can be chosen as a reference microphone, which in turn gives M-1 microphone pairs all relative to the first microphone. For example, Figure 5A An example microphone pair is illustrated. The microphone pair with index the (i-1)th pair includes microphone i 502 and the reference microphone 1 501 for an incident sound ray 505 having an azimuth angle of θ and an elevation angle of zero, emitted from a distant source (assuming a far-field model). The distance between the microphone pair of 501 and 502 and the angle between the two microphones are denoted as and which can be computed given the Cartesian coordinates of the ith microphone 502. For the general case when the incident sound ray 505 is angled at an azimuth angle θ and an elevation angle ϕ, the TDOA of the (i-1)th microphone pair can be computed as

[0047]

[0048] where c is the propagation speed.

[0049] After scanning different elevation and azimuth angles, the TDOA trajectory computation module 420 can construct a grid of admissible TDOAs. When all M microphones are located on a plane, the resulting TDOA trajectory (for all scanned θ and ϕ) also lies on a plane in (M-1) -dimensional space. Different layouts of M microphones can result in different planes in (M-1) -dimensional space.

[0050] For example, Figure 5BTwo different example microphone layouts are illustrated along with their corresponding TDOA grids. At 510, a set of M=4 microphones is shown, with the distance between the first and third microphones being 8 cm, and the grid for permissible TDOA generation is shown in M-1=3 dimensional space as shown at 515. When the distance between the first and third microphones increases to 16 cm, the grid for permissible TDOA generation is shown at 520, and at 525, it is shown.

[0051] Return to reference Figure 4 The TDOA trajectory calculation module 420 can then send the (M-1) dimensional TDOA 403 to the block-based TDOA estimation module 410. The block-based TDOA estimation module 410 receives the TDOA 403 and a time-frequency domain representation 404 of the multi-source audio. Based on the time-frequency domain representation 404, the TDOA estimation module 410 uses data obtained from consecutive frames to extract the source microphones (e.g., ...). Figure 3 The TDOA information of the audio sensor 305a-n shown is shown.

[0052] In one embodiment, the block-based TDOA estimation module 410 employs a guided minimum variance (STMV) beamformer to obtain TDOA information from the time-frequency domain representation 404 of the multi-source audio. Specifically, the block-based TDOA estimation module 410 can select a microphone as a reference microphone and then specify a total of M-1 microphone pairs by pairing the remaining M1 microphones with the reference microphone. The microphone pairs are... index.

[0053] For example, the first microphone can be chosen as the reference microphone, and therefore, This represents the time-frequency representation of the audio from the reference microphone. For the p-th microphone pair, the block-based TDOA estimation module 410 calculates the frequency representation of the p-th pair in matrix form as follows: ,in() T This represents the transpose. Then, the block-based TDOA estimation module 410 calculates the covariance matrix of the p-th input signal pair for each frequency band k:

[0054]

[0055] in() H This represents the transpose of Hermitian.

[0056] In some implementations, computation is performed on blocks of a certain number of consecutive frames. Summation within the block. For simplicity, block indices are ignored here.

[0057] Then, the block-based TDOA estimation module 410 can construct a steering matrix for each pair and frequency band as follows:

[0058]

[0059] where, is the TDOA of the p-th pair obtained from the TDOA trajectory computation module 420 after different scans of 0 and (omitted for brevity); is the frequency at frequency band k; and diag([a, b]) denotes a 2x2 diagonal matrix with diagonal elements a and b.

[0060] For each microphone pair p, the block-based TDOA estimation module 410 constructs a directional covariance matrix that is coherently aligned across all frequency bands by:

[0061] .

[0062] The directional covariance matrix is computed over all microphone pairs p and all scans of the azimuth / elevation of For reducing the computation over all scans, the TDOA space of each dimension p corresponding to the p-th microphone pair is linearly quantized into q segments at the beginning of the processing (at system startup). The TDOA trajectory points obtained by scanning each azimuth and elevation angle are mapped to the closest quantized point for each dimension. For each azimuth / elevation , the mapping is saved in memory, where is the quantized TDOA index of dimension p related to the scan angles 0 and.

[0063] For example, if there are M = 4 microphones and the azimuth and elevation scans are The number of different computations of that need to be performed is When the TDOA trajectory points are quantized, because some of the TDOA dimensions can be quantized to the same segment among q quantized segments, not all computations need to be performed. Thus, for example, if q = 50, the maximum number of different computations needed to compute is reduced to The pseudo code for performing the computation of with TDOA quantization can be shown in Algorithm 1 as follows:

[0064] .

[0065] Next, for each pair p , the direction that minimizes the beam power subject to the distortionless criterion (with its equivalent pseudo-likelihood solution) is computed as follows:

[0066]

[0067] where . Then, the block-based TDOA estimation module 410 can compute the STMV joint pseudo-likelihood for all M-1 pairs of microphones as:

[0068] .

[0069] Then, the azimuth and elevation angles that produce the maximum STMV joint pseudo-likelihood for all M-1 pairs are identified, denoted by the following equation

[0070]

[0071] The azimuth and elevation angle pair , can then be used for multi-source tracking and multi-stream VAD estimation. One possible solution can include directly tracking the angle between the two microphones of each microphone pair. However, due to the wrap-around effect of the azimuth in 360 degrees, if the angle between the paired microphones is directly tracked, track loss can occur when the microphone source crosses 0 o towards 360 o and vice versa. Therefore, to avoid such confusion, a polar transformation is used to compute the detection z in a circular manner based on the angle between the paired microphones as follows:

[0072]

[0073] where is a scaling constant that can expand the measurement space, thus allowing tracking with parameters related to meaningful concepts such as angles.

[0074] Then, the block-based TDOA estimation module 410 sends the computed detection z to the multi-source tracking and multi-stream VAD estimation module 430. If there is a maximum number of tracks, t , the TDOA obtained from the block-based TDOA estimation module 410 is to be tracked by updating the tracks obtained from the previous step in a recursive manner. Specifically, if the detection obtained at block (time-step) n-1 is denoted by z n-1 , and there are tn-1 a new detection z n 406, the multi-source tracking and multi-stream VAD estimation module 430 updates the existing tracks based on the gates of the existing tracks for the new detection z n The processing is done as follows:

[0075] If z n falls into the gate of only one of the previous t n-l tracks, then the particular track is updated to incorporate the detection z n .

[0076] If z n falls into the gates of multiple previous t n-l tracks, then the track closest to the detection z n is updated to incorporate the detection z n .

[0077] If z n does not fall into the gate of any previous t n-l track, and the maximum number of tracks has not been reached (e.g., t n-l < , then a new track is launched to incorporate the detection z n and the number of existing tracks is updated at time-step n, e.g., .

[0078] If z n does not fall into the gate of any previous t n-1 track, and the maximum number of tracks has been reached (e.g., t n-1 = , then the track with the lowest power among the existing t tracks is stopped and replaced with a new track to incorporate the detection z n .

[0079] For all other tracks that are not updated, launched, or replaced (as in the previous steps), then these tracks are updated with the same mean value, but the variance of each respective track is increased to account for uncertainty (e.g., based on a random walk model). The power of each respective track is also decayed so that future appearing sources have a chance to be launched. In this way, the tracking results 408 incorporating the latest detections 406 at time-step n can be output from the module 430, which is denoted by .

[0080] When all audio tracks have been updated, the module 430 generates the multi-stream VAD 412 using the most recent neighbors M2T assignment. Specifically, at time-step n, the multi-stream VAD 412 can be assigned by assigning 1 to the track closest to the detection zn The orbit and the allocation of 0 to other orbits to perform M2T allocation. In some implementations, a hangover can be applied to the VAD to have an intermediate value, such as -1, before it is fully allocated to zero after being 1 in a previous time step. In this way, from module 430, for example, to Figure 3 The audio enhancement engine 330 output is from The multi-stream VAD412 is used for audio enhancement, with each multi-stream VAD412 representing whether any speech activity was detected in the corresponding track.

[0081] Figure 6 An example method 600 for enhancing a multi-source audio signal through multi-source tracking and VAD according to various embodiments of the present disclosure is illustrated. In some embodiments, method 600 may be performed by one or more components of an audio signal processor 300 and / or one or more components of a multi-track VAD engine 400.

[0082] Method 600 begins with step 602, in which TDOA trajectory information can be calculated based on the spatial information of the microphone array. For example, TDOA trajectory information can be calculated once at system startup by scanning the microphone array using incident sound rays at varying azimuth and elevation angles. (See also: Regarding...) Figure 7 As further described, computations can be performed with reduced complexity in a multidimensional space constructed by pairing microphones from a microphone array.

[0083] refer to Figure 7 This provides further detailed steps for step 602, where in step 702, a first microphone can be selected from the microphone array as a reference microphone. In step 704, each remaining microphone from the microphone array can be paired with the reference microphone. In step 706, for each microphone pair, the selection can be based on the distance and angle between the two microphones in the corresponding pair (e.g., according to...). Figure 4 The equation (1) is described to calculate the TDOA position corresponding to a specific azimuth angle and a specific elevation angle of the incident sound ray. Figure 5A The paper also shows example microphone pairs with specific azimuth and elevation angles for the incident sound rays.

[0084] In step 708, if there are more microphone pairs to process, the method retrieves the next microphone pair in step 710 and repeats in step 706 until the TDOA positions of all microphone pairs have been calculated.

[0085] In step 712, if there are more scans of azimuth and elevation angles, the method retrieves the next scan of azimuth and elevation angles in step 714 and repeats in step 706 until the TDOA position of all scans of azimuth and elevation angles is calculated.

[0086] In step 712, when there are no more scans of azimuth and elevation angles to process (e.g., the TDOA positions have already been calculated for all microphone pairs on all scans of azimuth and elevation angles), a grid of TDOA position points can be formed in step 716. Figure 5B The image shows an example grid of TDOA location points corresponding to different geometries of the microphone array.

[0087] Return to reference Figure 6 When calculating TDOA trajectory information during system startup, method 600 proceeds to step 604. In step 604, one or more multi-source audio signals can be received from the microphone array. For example, this can be achieved via... Figure 3 The audio input circuit 315 in the middle is used to receive Figure 4 The time-domain sampling of the multi-source audio signal 402 in the middle.

[0088] In step 606, one or more multi-source audio signals can be transformed from a time-domain representation to a time-frequency representation. For example, sub-band analysis module 405 can transform a time-domain signal into a time-frequency domain representation, as per [reference to...]. Figure 4 As described.

[0089] In step 608, TDOA detection data can be calculated for one or more multi-source audio signals based on the calculated TDOA trajectory and according to the STMV beamformer. For example, for each microphone pair, time-frequency representations of one or more multi-source audio signals from the corresponding microphone pair can be used for each frequency band (e.g., based on information about...). Figure 4 The covariance matrix can then be calculated using the equation (2) described above. Then, the covariance matrix can be calculated for each frequency band based on the TDOA location (e.g., according to the information about the frequency band). Figure 4 The steering matrix is ​​constructed using the equation (3) described above, where the TDOA position is for different scans corresponding to the azimuth and elevation angles of the respective microphone pairs. The steering matrix can be based on the constructed steering matrix and the calculated covariance matrix (e.g., according to the information provided). Figure 4 The described equation (4) is used to construct the directional covariance matrix aligned across all frequency bands. The constructed directional covariance matrix can be based on (e.g., according to the equation regarding...) Figure 4 The described equation (5) determines the pseudo-likelihood solution that minimizes the beam power. This can then be determined by taking the product of all determined pseudo-likelihood solutions across all microphone pairs (e.g., based on the information provided). Figure 4 The equation (6) described above is used to calculate the joint pseudo-likelihood of STMV. Then, it can be (e.g., based on the information about...)Figure 4 Equation (7) described determines an azimuth and elevation angle that maximizes the STMV joint pseudo-likelihood. The determined azimuth and elevation angle can then be transformed into a polar coordinate representation representing the TDOA detection data (e.g., according to Equation (8) described Figure 4 Equation (8) described transforms the determined azimuth and elevation angle into a polar coordinate representation representing the TDOA detection data.

[0090] At step 610, the updated plurality of audio tracks and constructed VAD data can be used to generate one or more enhanced multi-source audio signals. For example, the enhanced multi-source audio signals can then be transmitted to various devices or components. For another example, the enhanced multi-source audio signals can be packaged and transmitted over a network to another audio output device (e.g., a smartphone, a computer, etc.). The enhanced multi-source audio signals can also be transmitted to speech processing circuitry such as an automatic speech recognition component for further processing. Figure 4 As described with respect to module 430 in FIG. 4, for another example, the method 600 can assign a first value to the VAD of the respective audio track when the respective audio track is closest to the TDOA detection, and a second value to the VAD of the other audio tracks. Figure 4 Figure 4 As described with respect to module 430 in FIG. 4, for another example, the method 600 can assign a first value to the VAD of the respective audio track when the respective audio track is closest to the TDOA detection, and a second value to the VAD of the other audio tracks.

[0091] At step 612, the updated plurality of audio tracks and constructed VAD data can be used to generate one or more enhanced multi-source audio signals. For example, the enhanced multi-source audio signals can then be transmitted to various devices or components. For another example, the enhanced multi-source audio signals can be packaged and transmitted over a network to another audio output device (e.g., a smartphone, a computer, etc.). The enhanced multi-source audio signals can also be transmitted to speech processing circuitry such as an automatic speech recognition component for further processing.

[0092] The foregoing disclosure is not intended to limit the application to the precise forms or specific uses disclosed. Thus, various alternatives and / or modifications of the embodiments described herein that are apparent to one of ordinary skill in the art are intended to be within the scope of the present disclosure. For example, the embodiments described herein can be used to provide locations of multiple sound sources in an environment in order to supervise human-machine interaction tasks (e.g., in applications incorporating additional information from other modalities such as video streams, 3D cameras, lidar, etc.). Having thus described embodiments of the disclosure, what is claimed as new and desired to be protected by Letters Patent is set forth in the following claims.

Claims

1. A method for enhancing multi-source audio by multi-source tracking and voice activity detection, comprising: receiving one or more multi-source audio signals from a microphone array via audio input circuitry; computing TDOA detection data for the one or more multi-source audio signals based on TDOA trajectory information constructed in a multi-dimensional space defined by pairs of microphones from the microphone array according to a steered minimum variance (STMV) beamformer; updating a plurality of audio tracks based on computed TDOA detection data up to a current time-step, wherein updating a plurality of audio tracks based on computed TDOA detection data up to a current time-step comprises: identifying a TDOA detection corresponding to the current time-step and a set of existing audio tracks that have been established up to the current time-step, and determining whether to incorporate the TDOA detection into one of the existing audio tracks or establish a new audio track based on a comparison between the TDOA detection and gates of the existing audio tracks; constructing voice activity detection (VAD) data for each of the plurality of audio tracks based on the computed TDOA detection data; and generating one or more enhanced multi-source audio signals using the updated plurality of audio tracks and constructed VAD data.

2. The method of claim 1, wherein the multi-dimensional space defined by pairs of microphones from the microphone array is formed by: selecting a first microphone from the microphone array as a reference microphone; and pairing each remaining microphone from the microphone array with the reference microphone.

3. The method of claim 2, wherein the TDOA trajectory information is computed once at a start-up phase based on spatial information of the pairs of microphones by: computing, for each microphone pair, a TDOA position corresponding to an azimuth angle and an elevation angle of an incident sound ray based on a distance and an angle between the two microphones in the respective pair; and forming a grid of TDOA position points by varying the azimuth angle and the elevation angle of the incident sound ray over all microphone pairs.

4. The method of claim 3, wherein the grid of TDOA position points lies on a first plane in the multi-dimensional space when the microphone array is physically located on a second plane in reality, the multi-dimensional space having a number of dimensions equal to a total number of microphone pairs.

5. The method of claim 2, wherein computing TDOA detection data for the one or more multi-source audio signals further comprises: for each microphone pair: computing a covariance matrix for each frequency band using a time-frequency representation of the one or more multi-source audio signals from the respective microphone pair; constructing a steering matrix for each frequency band based on TDOA positions for different scans corresponding to azimuth and elevation angles of the respective microphone pair; constructing a directional covariance matrix aligned across all frequency bands based on the constructed steering matrices and the computed covariance matrices; and ​ determine a pseudo-likelihood solution that minimizes beam power based on the constructed directional covariance matrix.

6. The method of claim 5, further comprising: computing an STMV joint pseudo-likelihood based on taking a product of all determined pseudo-likelihood solutions across all microphone pairs; determining a single azimuth and elevation angle that maximizes the STMV joint pseudo-likelihood; and converting the determined single azimuth and elevation angle to a polar coordinate representation that represents the TDOA detection data.

7. The method of claim 6, wherein constructing the directional covariance matrix that is aligned across all frequency bands based on the constructed steering matrix and the computed covariance matrix is repeated over all microphone pairs and all scans of azimuth and elevation angles.

8. The method of claim 6, wherein constructing the directional covariance matrix that is aligned across all frequency bands based on the constructed steering matrix and the computed covariance matrix is performed in a manner that reduces the number of repeated computations by: dividing the multi-dimensional space into a number of segments, where the number of segments is less than the total number of dimensions of the multi-dimensional space; mapping each TDOA location point from a grid of TDOA location points to the closest segment; and computing the directional covariance matrix using the number of segments and the mapping between the grid of TDOA location points and the number of segments instead of the grid of TDOA location points established according to all scans of azimuth and elevation angles.

9. The method of claim 1, wherein constructing VAD data for each of the plurality of audio tracks based on the computed TDOA detection data further comprises: assigning a first value to a respective audio track when the respective audio track is closest to the TDOA detection; and assigning a second value to other audio tracks.

10. An audio processing device for enhancing multi-source audio through multi-source tracking and voice activity detection, comprising: an audio input circuit configured to receive one or more multi-source audio signals from a microphone array; a time-difference-of-arrival (TDOA) estimator configured to compute TDOA detection data for the one or more multi-source audio signals based on steering minimum variance (STMV) beamformer, based on TDOA track information constructed in a multi-dimensional space defined by a plurality of microphone pairs from the microphone array; a multi-source audio tracker configured to update a plurality of audio tracks based on computed TDOA detection data up to a current time-step, and to construct voice activity detection (VAD) data for each of the plurality of audio tracks based on the computed TDOA detection data; and an audio enhancement engine configured to generate one or more enhanced multi-source audio signals using the updated plurality of audio tracks and constructed VAD data; wherein the multi-source audio tracker is configured to update the plurality of audio tracks based on the computed TDOA detection data up to the current time-step by: identifying a TDOA detection corresponding to a current time-step and a set of existing audio tracks that have been previously established up to the current time-step; and determining whether to incorporate the TDOA detection into one of the existing audio tracks or to establish a new audio track based on a comparison between the TDOA detection and the gates of the existing audio tracks.

11. The audio processing device of claim 10, wherein the multi-dimensional space defined by a plurality of microphone pairs from the microphone array is formed by: selecting a first microphone from the microphone array as a reference microphone; and pairing each remaining microphone from the microphone array with the reference microphone.

12. The audio processing device of claim 11, wherein the TDOA track information is computed once at a start-up phase based on spatial information of the plurality of microphone pairs by: for each microphone pair, computing a TDOA position corresponding to an azimuth angle and an elevation angle of an incident sound ray based on a distance and an angle between the two microphones in the respective pair; and forming a grid of TDOA position points by varying the azimuth angle and the elevation angle of the incident sound ray over all microphone pairs.

13. The audio processing device of claim 12, wherein when the microphone array is physically located on a second plane in reality, the grid of TDOA position points is located on a first plane in the multi-dimensional space, the multi-dimensional space having a number of dimensions equal to a total number of microphone pairs.

14. The audio processing device of claim 11, wherein the TDOA estimator is configured to compute the TDOA detection data by: for each microphone pair: computing a covariance matrix using a time-frequency representation of the one or more multi-source audio signals from the respective microphone pair for each frequency band; constructing a steering matrix for each frequency band based on TDOA positions for different scans of azimuth and elevation angles corresponding to the respective microphone pair; constructing a directional covariance matrix aligned across all frequency bands based on the constructed steering matrices and the computed covariance matrices; and determining a pseudo-likelihood solution that minimizes beam power based on the constructed directional covariance matrix.

15. The audio processing device of claim 14, wherein the TDOA estimator is further configured to compute the TDOA detection data by: computing an STMV joint pseudo-likelihood by taking a product of all determined pseudo-likelihood solutions across all microphone pairs; determining a pair of azimuth and elevation angles that maximizes the STMV joint pseudo-likelihood; and converting the determined pair of azimuth and elevation angles to a polar coordinate representation representative of the TDOA detection data. ​ 16. The audio processing device of claim 15, wherein the TDOA estimator is further configured to repeatedly construct the directional covariance matrix aligned across all frequency bands based on the constructed steering matrix and the computed covariance matrix over all microphone pairs and all scans of azimuth and elevation angles.

17. The audio processing device of claim 15, wherein the TDOA estimator is further configured to construct the directional covariance matrix aligned across all frequency bands based on the constructed steering matrix and the computed covariance matrix in a manner that reduces the number of repeated computations by: dividing the multi-dimensional space into a number of segments, wherein the number of segments is less than the total number of dimensions of the multi-dimensional space; mapping each TDOA location point from a grid of TDOA location points to the closest segment; and computing the directional covariance matrix using the number of segments and the mapping between the grid of TDOA location points and the number of segments instead of the grid of TDOA location points established according to all scans of azimuth and elevation angles.

18. The audio processing device of claim 10, wherein the multi-source audio tracker is configured to construct VAD data for each of the plurality of audio tracks based on the computed TDOA detections by: assigning a first value to a respective audio track when the respective audio track is closest to the TDOA detection; and assigning a second value to other audio tracks.

Citation Information

Patent Citations

  • Robust broadband steered minimum variance beam forming method suitable for any formation

    CN104035064A

  • 360-degree multi-source location detection, tracking and enhancement

    US20190355373A1