Multiple sound source tracking and speech activity detection for planar microphone arrays

The combination of multi-source TDOA tracking and voice activity detection addresses the inefficiencies in existing speaker enhancement algorithms by reducing computational complexity and enhancing target speech in multi-stream audio environments.

JP7742703B2Active Publication Date: 2025-09-22SYNAPTICS INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2020212089
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-01-10
Filing Date
2020-12-22
Publication Date
2025-09-22
Estimated Expiration
2040-12-22

AI Technical Summary

Technical Problem

Existing speaker enhancement algorithms for smart speakers face challenges in efficiently separating target speech from noise and other active speakers due to long response delays in batch processing and reliance on supervision under voice activity detection, which is not applicable for practical applications.

Method used

A combination of multi-source TDOA tracking and voice activity detection mechanism applicable to general array geometries, reducing computational complexity by performing TDOA searches separately for each dimension, avoiding ghost TDOA, and enhancing target audio signals while suppressing noise.

Benefits of technology

The solution effectively separates target speech from noise in multi-stream audio environments, reducing computational complexity and eliminating the need for extensive multidimensional searches, thereby improving the efficiency and accuracy of speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007742703000030
    Figure 0007742703000030
  • Figure 0007742703000031
    Figure 0007742703000031
  • Figure 0007742703000032
    Figure 0007742703000032
Patent Text Reader

Abstract

To provide a system and a method for multi-source tracking and multi-stream speech section detection with reduced computational complexity by a microphone array to detect and process a target audio signal in a multi-stream audio environment.SOLUTION: A method: updates an audio track for multiple sound source audio signals, on the basis of TDOA detection data calculated for multiple sound source audio signals, by a steered minimal dispersion (STMV) beamformer based on TDOA trajectory information constructed in a multidimensional space defined by pairs of microphones from a microphone array; constructs speech interval detection (VAD) data for each of multiple audio tracks on the basis of TDOA detection data; and generates one or more emphasized multiple sound source audio signals using the multiple updated audio tracks and the constructed VAD data.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure, according to one or more embodiments, relates generally to audio processing, and more particularly to systems and methods for multi-source tracking and multi-stream voice activity detection, for example for general planar microphone arrays. [Background technology]

[0002] Smart speakers and other voice-controlled devices and electronic devices have gained popularity in recent years. Smart speakers often include an array of microphones to receive voice input (e.g., a user's verbal commands) from the environment. When a target voice (e.g., a verbal command) is detected in the voice input, the smart speaker can convert the detected target voice into one or more commands and perform different tasks based on the commands.

[0003] One of the challenges for these smart speakers is efficiently and effectively separating target speech (e.g., verbal commands) from noise in the operating environment and other active speakers. For example, one or more speakers may be active in the presence of one or more noise sources. When the goal is to emphasize a specific speaker, that speaker is referred to as the target speaker, while the remaining speakers can be considered as interference sources. Existing speaker enhancement algorithms mainly exploit the spatial information of the sound sources using multiple input channels (microphones), such as blind source separation (BSS) methods related to independent component analysis (ICA), spatial filtering, or beamforming methods.

[0004] BSS methods, however, were primarily designed for batch processing and often have long response delays that make them undesirable or even inapplicable for practical applications.Spatial filtering or beamforming methods, on the other hand, often require supervision under voice activity detection (VAD) as a cost function to be minimized, which may rely too heavily on estimating the covariance matrix belonging to the noise / interference-only partition.

[0005] Therefore, there is a need for improved systems and methods for detecting and processing target audio signals in a multi-stream audio environment. [Brief explanation of the drawings]

[0006] Aspects of the present disclosure and its advantages may be better understood with reference to the following drawings and detailed description below. While like reference numerals are used to identify like elements shown in one or more of the drawings, it should be understood that the illustrations are for the purpose of illustrating embodiments of the present disclosure and not for the purpose of limiting the same. The components in the drawings are not necessarily to scale, emphasis instead being placed upon clearly illustrating the principles of the present disclosure.

[0007] [Figure 1] FIG. 1 illustrates an exemplary operating environment for an audio processing device in accordance with one or more embodiments of the present disclosure.

[0008] [Figure 2] FIG. 2 is a block diagram of an exemplary audio processing device in accordance with one or more embodiments of the present disclosure.

[0009] [Figure 3] FIG. 3 is a block diagram of an example audio processor for multi-track audio enhancement in accordance with one or more embodiments of the present disclosure.

[0010] [Figure 4]FIG. 4 is a block diagram of an exemplary multi-track activity detection engine for processing multiple audio signals from a typical microphone array, according to various embodiments of the present disclosure.

[0011] [Figure 5A] FIG. 5A is a diagram illustrating an exemplary geometry of microphone pairs, according to one or more embodiments of the present disclosure.

[0012] [Figure 5B] FIG. 5B illustrates a mesh of exemplary time difference of arrival (TDOA) trajectory information in a multi-dimensional space corresponding to different microphone array geometries, in accordance with one or more embodiments of the present disclosure.

[0013] [Figure 6] FIG. 6 is a logic flow diagram of an exemplary method for enhancing a multi-source audio signal through multi-source tracking and activity detection, according to various embodiments of the present disclosure.

[0014] [Figure 7] FIG. 7 is a logic flow diagram of an exemplary process for calculating TDOA trajectory information in a multi-dimensional space using microphone pairs, according to various embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0015] The present disclosure provides improved systems and methods for detecting and processing target audio signals in a multi-stream audio environment.

[0016] Voice activity detection (VAD) can be used to monitor speech enhancement of a target voice in a process that utilizes spatial information of sound sources obtained from multiple input channels. The VAD may provide spatial statistics of interference / noise sources during periods when the desired speaker is silent, so that the impact of noise / interference can be substantially nullified when the desired speaker becomes active. For example, the VAD for each sound source can be inferred to track spatial information in the form of the sound source's time difference of arrival (TDOA) or direction of arrival (DOA) by utilizing the VAD's detection history by determining when detections occurred in the vicinity of existing tracks. This process is commonly known as the Measurement-to-Track (M2T) problem. In this way, multiple VADs can be estimated for all of the sound sources of interest.

[0017] Specifically, existing DOA methods typically construct a single steering vector for the entire microphone array based on a closed-form mapping of azimuth and elevation angles. This method can be used to exploit specific configurations of linear or circular arrays. However, such DOA methods cannot be extended to general or arbitrary configurations of microphone arrays. Furthermore, these closed-form mapping-based DOA methods often require extensive searches in multidimensional space. For arbitrary configurations, existing TDOA-based methods can be used. These methods may not be limited to specific array configurations and may construct multiple steering vectors for each microphone pair, forming a multidimensional TDOA vector (one dimension for each pair). However, these existing methods risk inducing TDOA ghosts, which are formed by the intersection of peaks in the spectrum of each TDOA pair. As a result, additional post-processing is often required to remove TDOA ghosts, including those specific to the array configuration.

[0018] In view of the need for a multi-stream VAD that is not constrained to a particular array geometry, the embodiments described herein provide a combination of multi-source TDOA tracking and a VAD mechanism that is applicable to general array geometries (e.g., microphone arrays arranged on a plane). The combination of multi-source TDOA tracking and a VAD mechanism may reduce the number of calculations typically involved in conventional TDOA by performing a TDOA search separately for each dimension.

[0019] In some embodiments, a multidimensional TDOA method is employed for the placement of a general array arranged on a plane, avoiding unwanted ghost TDOA. In one embodiment, Cartesian coordinates of the microphones in a general configuration are obtained. One of the microphones may be selected as a reference microphone. The azimuth and elevation angles of the microphones may be scanned, based on which a physically possible planar trajectory of the TDOA can be formed in the multidimensional TDOA space of multiple microphone pairs. In this way, the formed planar trajectory avoids ghost TDOA, so further post-processing to remove the ghost TDOA is unnecessary. Furthermore, compared to the full DOA scanning method, the multidimensional TDOA method disclosed herein reduces computational complexity by performing a search for each dimension separately on the TDOA region of the pairs, rather than searching over the complete multidimensional space.

[0020] FIG. 1 illustrates an exemplary operating environment 100 in which a sound processing system according to various embodiments of the present disclosure may operate. The operating environment 100 includes a sound processing device 105, a target sound source 110, and one or more noise sources 135-145. In the example illustrated in FIG. 1, the operating environment 100 is shown as a room. However, it is contemplated that the operating environment may include other locations, such as the interior of a vehicle, an office conference room, a home room, an outdoor stadium, or an airport. In various embodiments of the present disclosure, the sound processing device 105 may include two or more sound sensing components (e.g., microphones) 115a-115d and, optionally, one or more sound output components (e.g., speakers) 120a-120b.

[0021] The audio processing device 105 may be configured to sense sound using audio sensing components 115a-115d and generate a multi-channel audio input signal including two or more audio input signals. The audio processing device 105 may process the audio input signal using audio processing techniques disclosed herein to enhance the audio signal received from the target audio source 110. For example, the processed audio signal may be communicated to other components within the audio processing device 105, such as a speech recognition engine or a voice command processor, or to an external device. Thus, the audio processing device 105 may be a standalone device that processes audio signals, or a device that converts the processed audio signals into other signals (e.g., commands, instructions, etc.) for communicating with or controlling external devices. In other embodiments, the audio processing device 105 may be a communications device, such as a cellular phone or a voice-over-IP (VoIP)-enabled device. The processed audio signal may then be communicated over a network to other devices for output to a remote user. The communications device may further receive processed audio signals from remote devices and output the processed audio signals using audio output components 120a-120b.

[0022] The target sound source 110 may be any sound source that produces a sound detectable by the sound processing device 105. The target sound to be detected by the system may be defined based on criteria specified by a user or system requirements. For example, the target sound may be defined as human speech, a sound made by a particular animal, or a machine. In the illustrated example, the target sound is defined as human speech, and the target sound source 110 is a human. In addition to the target sound source 110, the operating environment 100 may include one or more noise sources 135-145. In various embodiments, sounds that are not target sounds may be treated as noise. In the illustrated example, the noise sources 135-145 may include a loudspeaker 135 playing music, a television 140 playing a television program, movie, or sporting event, and background conversation between non-target speakers 145. It will be appreciated that other noise sources may be present in various operating environments.

[0023] It should be noted that the target sound and noise may arrive at the sound sensing components 115a-115d of the sound processing device 105 from different directions and at different times. For example, noise sources 135-145 may generate noise at different locations within the operating environment 100. And, the target sound source (human) 110 may speak while moving between multiple locations within the operating environment 100. Furthermore, the target sound and / or noise may reflect off fixtures (e.g., walls) within the operating environment 100. For example, consider the path that the target sound may take from the target sound source 110 to each of the sound sensing components 115a-115d. As indicated by arrows 125a-125d, the target sound may travel directly from the target sound source 110 to each of the sound sensing components 115a-115d. Additionally, the target sound may reach the sound sensing components 115a-115d indirectly from the target sound source 110 by reflecting off walls 150a and 150b, as indicated by arrows 130a-130b. In various embodiments, the sound processing device 105 may estimate and apply the room impulse response and further use one or more sound processing techniques to enhance the target sound and suppress noise.

[0024] 2 illustrates an exemplary audio processing device 200 according to various embodiments of the present disclosure. In some embodiments, audio processing device 200 may be implemented as audio processing device 105 of FIG. 1. Audio processing device 200 includes an audio sensor array 205, an audio signal processor 220, and a host system component 250.

[0025] The audio sensor array 205 comprises two or more sensors, each of which may be implemented as a transducer that converts audio input in the form of sound waves into an audio signal. In the illustrated environment, the audio sensor array 205 comprises multiple microphones 205a-205n, each of which generates an audio input signal that is provided to audio input circuitry 222 of the audio signal processor 220. In one embodiment, the audio sensor array 205 generates a multi-channel audio signal, with each channel corresponding to the audio input signal from one of the microphones 205a-n.

[0026] The audio signal processor 220 includes audio input circuitry 222, a digital signal processor 224, and optionally, audio output circuitry 226. In various embodiments, the audio signal processor 220 may be implemented as an integrated circuit including analog circuitry, digital circuitry, and the digital signal processor 224 operable to execute program instructions stored in firmware. The audio input circuitry 222 may include, for example, an interface to the audio sensor array 205, anti-aliasing filters, analog-to-digital converter circuitry, echo cancellation circuitry, and other audio processing circuitry and components. The digital signal processor 224 is operable to process the multi-channel digital audio signals to generate enhanced audio signals that are output to one or more host system components 250. In various embodiments, the digital signal processor 224 may be operable to perform echo cancellation, noise cancellation, target signal enhancement, post-filtering, and other audio signal processing functions.

[0027] Optional audio output circuitry 226 processes audio signals received from digital signal processor 224 for output to at least one speaker, such as speakers 210a and 210b. In various embodiments, audio output circuitry 226 may include a digital-to-analog converter to convert one or more digital audio signals to analog and one or more amplifiers to drive speakers 210a-210b.

[0028] The voice processing device 200 may be implemented as any device operable to receive and enhance target voice data, such as, for example, a mobile phone, a smart speaker, a tablet, a laptop computer, a desktop computer, a voice-controlled appliance, or an automobile. The host system component 250 may include various hardware and software components for operating the voice processing device 200. In the illustrated embodiment, the host system component 250 includes a processor 252, a user interface component 254, a communication interface 256 for communicating with external devices and a network such as a network 280 (e.g., the Internet, the cloud, a local area network, or a telephone network), a mobile device 284, and memory 258.

[0029] The processor 252 and the digital signal processor 224 may comprise one or more of a processor, a microprocessor, a single-core processor, a multi-core processor, a microcontroller, a programmable logic device (PLD) (e.g., a field programmable gate array (FPGA)), a digital signal processing (DSP) device, or other logic devices, which may be configured to perform various processes discussed herein in embodiments of the present disclosure through hardware, by executing software, or a combination of both. The host system component 250 is configured to connect to and communicate with the audio signal processor 220 and other host system components 250, for example, through a bus or other electronic communication interface.

[0030] Although audio signal processor 220 and host system components 250 are shown as incorporating a combination of hardware components, circuitry, and software, it will be understood that in some embodiments, at least some, or all, of the functions of the hardware components and circuitry operable to perform may be implemented as software modules executable by processor 252 and / or digital signal processor 224 in response to software instructions and / or configuration data stored in memory 258 or in firmware of digital signal processor 224.

[0031] Memory 258 may be implemented as one or more memory devices operable to store data and information, including audio data and program instructions. Memory 258 may comprise one or more of various types of memory devices, including volatile and non-volatile memory devices, such as Random Access Memory (RAM), Read-Only Memory (ROM), Electrically Erasable Read-Only Memory (EEPROM), flash memory, hard disk drives, and / or other types of memory.

[0032] The processor 252 may be operable to execute software instructions stored in the memory 258. In various embodiments, the speech recognition engine 260 may be operable to process the enhanced voice signal received from the voice signal processor 220. This processing includes identifying and executing voice commands. The voice communication component 262 may be operable to facilitate voice communication with one or more external devices, such as a mobile device 284 or a user device 286, such as calls using a mobile or cellular phone network or a VoIP call across an IP network. In various embodiments, the voice communication includes communicating the enhanced voice signal to the external communication device.

[0033] The user interface component 254 may include a display, a touchpad display, a keypad, one or more buttons, and / or other input / output components operable to allow a user to interact directly with the audio processing device 200.

[0034] Communications interface 256 facilitates communication between audio processing device 200 and external devices. For example, communications interface 256 may enable a Wi-Fi (e.g., 802.11), or Bluetooth connection between audio processing device 200 and one or more local devices, such as a mobile device 284 or a wireless router providing network access (such as via network 280) to a remote server 282. In various embodiments, communications interface 256 may include other wired or wireless communications components that facilitate direct or indirect communications between audio processing device 200 and one or more other devices.

[0035] 3 illustrates an exemplary audio signal processor 300 according to various embodiments of the present disclosure. In some embodiments, audio signal processor 300 is embodied as one or more integrated circuits including analog and digital circuitry and firmware logic implemented by a digital signal processor, such as audio signal processor 220 of FIG. 2. As shown, audio signal processor 300 includes audio input circuitry 315, a subband frequency analyzer 320, a multi-track VAD engine 325, a speech enhancement engine 330, and a synthesizer 335.

[0036] The audio signal processor 300 receives multi-channel audio input from multiple audio sensors, such as a sensor array 305 comprising at least two audio sensors 305a-n. The audio sensors 305a-305n may include, for example, multiple microphones integrated with an audio processing device, such as the audio processing device 200 of FIG. 2, or external components connected thereto. The arrangement of the audio sensors 305a-305n may be known or unknown to the audio signal processor 300, according to various embodiments of the present disclosure.

[0037] The audio signal may first be processed by audio input circuitry 315, which may include anti-aliasing filters, analog-to-digital converters, and / or other audio input circuitry. In various embodiments, audio input circuitry 315 outputs a digital, multi-channel, time-domain audio signal, where M is the number of sensor (e.g., microphone) inputs. The multi-channel audio signal is input to a subband frequency analyzer 320, which divides the multi-channel audio signal into a number of consecutive frames and decomposes each frame for each channel into a number of frequency subbands. In various embodiments, the subband frequency analyzer 320 includes Fourier transform processing and outputs a number of frequency bins. The decomposed audio signal is then provided to a multi-track VAD engine 325 and a speech enhancement engine 330.

[0038] The multi-track VAD engine 325 is operable to analyze frames of one or more audio tracks and generate a VAD output indicating whether target audio activity is present in the current frame. As discussed above, the target audio may be any audio to be recognized by an audio system. When the target audio is human speech, the multi-track VAD engine 325 may be implemented specifically for speech activity detection. In various embodiments, the multi-track VAD engine 325 is operable to receive frames of audio data and generate a VAD indicator output for each audio track regarding the presence or absence of target audio in the respective audio track corresponding to the frame of audio data. Detailed components and processing of the multi-track VAD engine 325 are further illustrated in connection with 400 in FIG. 4 .

[0039] The speech enhancement engine 330 receives the subband frames from the subband frequency analyzer 320 and the VAD indicators from the multi-track VAD engine 325. In various embodiments of the present disclosure, the speech enhancement engine 330 is configured to process the subband frames based on the received multi-track VAD indicators to enhance the multi-track speech signal. For example, the speech enhancement engine 330 may enhance portions of the speech signal determined to be from the direction of a target sound source and suppress other portions of the speech signal determined to be noise.

[0040] After enhancing the target speech signal, the speech enhancement engine 330 may pass the processed speech signal to a synthesizer 335. In various embodiments, the synthesizer 335 reconstructs one or more multi-channel speech signals frame by frame by combining subbands to form a time-domain enhanced speech signal, which is then converted back to the time domain and sent to a system component or external device for further processing.

[0041] FIG. 4 illustrates an exemplary multi-track VAD engine 400 for processing multiple audio signals from a common microphone array, according to various embodiments of the present disclosure. The multi-track VAD engine 400 may be implemented as a combination of digital circuitry and logic executed by a digital signal processor. In some embodiments, the multi-track VAD engine 400 may be installed in an audio processor, such as 300 of FIG. 3. The multi-track VAD engine 400 may provide further structural and functional details to the multi-track VAD engine 325 of FIG. 3.

[0042] In various embodiments of the present disclosure, the multi-track VAD engine 400 comprises a sub-band analysis module 405, a block-based TDOA estimation module 410, a TDOA trajectory calculation module 420, and a multi-source tracking and multi-stream VAD estimation module 430.

[0043] The subband analysis module 405 receives a plurality of audio signals 402. The audio signals 402 are divided into x m (t), m=1,...,M, is the audio signal recorded by the mth microphone of a total of M microphones (e.g., similar to audio sensors 305a-n in FIG. 3) sampled in the time domain. m (t), m=1, . . . , M may be received via audio input circuitry 315 of FIG.

[0044] The subband analysis module 405 is configured to obtain the audio signal 402 and convert the audio signal 402 into a time-frequency domain representation 404. The time-frequency domain representation 404 is a representation of the original time-domain audio signal x m Corresponding to (t), X m The time-frequency domain representation 404 may be represented as (l, k), where l denotes a subband time index and k denotes a frequency band index. For example, the subband analysis module 405 may be similar to the subband frequency analyzer 320 of FIG. 3, which performs a Fourier transform to convert the input time-domain audio signal into a frequency-domain representation. The subband analysis module 405 may then send the generated time-frequency domain representation 404 to the block-based TDOA estimation module 410 and the multi-source tracking and multi-stream VAD estimation module 430.

[0045] The TDOA trajectory calculation module 420 is configured to scan a general microphone array (e.g., the audio sensors 305a-n forming a general array configuration). For example, for any given microphone array configuration on a plane, a trajectory of allowable TDOA positions is calculated once at system startup. This trajectory of points allows for the avoidance of ghosting.

[0046] For an array of M microphones, a first microphone may be selected as the reference microphone. This sequentially results in M−1 microphone pairs, all associated with the first microphone. For example, FIG. 5A shows an exemplary microphone pair. A microphone pair is indexed as the i−1th pair, which includes the i-th microphone 502 and the first reference microphone 501 for an incident ray 505 emitted from a distant sound source (assuming a far-field model) with an azimuth angle θ and an elevation angle of zero. The distance between the microphone pair 501 and 502, along with the angle between the two microphones, is d i-1 and ψ i-1 and , respectively. These can be calculated given the Cartesian coordinates of the i-th microphone 502. In the general case where the incident ray 505 has angles of azimuth θ and elevation φ, the TDOA of the (i-1)-th microphone pair is given by

number

[0047] After scanning different azimuth and elevation angles, the TDOA trajectory computation module 420 may construct a mesh of acceptable TDOAs. If all M microphones are located on a plane, the resulting TDOA trajectory (for all scans of θ and φ) is

number

[0048] For example, in Figure 5B, two different exemplary microphone placements are shown along with their respective TDOA meshes. A set of M = 4 microphones is shown at 510, where the distance between the first and third microphones is 8 cm, and the resulting acceptable TDOA mesh is shown at 515 in an M - 1 = 3-dimensional space. If the distance between the first and third microphones is increased to 16 cm as shown at 520, the resulting acceptable TDOA mesh is shown at 525.

[0049] 4, the TDOA trajectory computation module 420 may then send the (M-1)-dimensional TDOA 403 to the block-based TDOA estimation module 410. The block-based TDOA estimation module 410 receives the time-frequency representation of the multi-source audio 404 and the TDOA 403. Based on the time-frequency representation of the multi-source audio 404 and the TDOA 403, the TDOA estimation module 410 extracts TDOA information for the source microphones (e.g., the audio sensors 305a-n in FIG. 3) using data obtained from successive frames.

[0050] In one embodiment, the block-based TDOA estimation module 410 uses a steered minimum variance (STMV) beamformer to obtain TDOA information from the time-frequency domain representation of the multi-source audio 404. More specifically, the block-based TDOA estimation module 410 may select a microphone as a reference microphone and then pair the remaining M-1 microphones with the reference microphone, thereby designating a total of M-1 microphone pairs. The microphone pairs are indexed by p=1, ..., M-1.

[0051] For example, the first microphone may be selected as the reference microphone, and accordingly, X1(l,k) may represent the time-frequency representation of the sound from the reference microphone. For the p-th microphone pair, the block-based TDOA estimation module 410 may calculate the frequency representation of the p-th pair in the matrix form:

number

number

[0052] In some implementations, R p The summation in computing (k) is done over a specific number of consecutive blocks of frames, where the block indices are omitted for simplicity.

[0053] The block-based TDOA estimation module 410 may then construct a steering matrix for each pair and frequency band as follows:

number

[0054] For each microphone pair p, the block-based TDOA estimation module 410 constructs an azimuth covariance matrix that is coherently aligned across all frequency bands as follows:

number

[0055] Orientation covariance matrix C p (τ p ) is calculated for all microphone pairs p and τ p This is repeated over all scans of azimuth / elevation angles (θ,φ) for the P microphone pair. To reduce the amount of calculations over all scans, each p-dimensional TDOA space corresponding to the p-th microphone pair is linearly quantized into q segments. At the start of processing (system startup), the TDOA trajectory points obtained from each scan of azimuth and elevation angles (θ,φ) are

number

number

number

[0056] For example, if there are M=4 microphones and the azimuth and elevation scans are

number

number

number

number

[0057] Then, for each pair p, the direction that minimizes the beam power in its equivalent pseudo-likelihood solution according to the distortion-free criterion is calculated as follows:

number

number

number

[0058] The azimuth and elevation angles that yield the maximum STMV joint pseudo-likelihood for the M-1 pairs are then identified as follows:

number

number

number

number

[0059] The block-based TDOA estimation module 410 then sends the calculated detection z to the multi-source tracking and multi-stream VAD estimation module 430.

number

[0060] z n But the leading t n-1 If one of the tracks is included in the gate, that particular track is considered a detection n Updated to incorporate

[0061] z n but preceded by (multiple) t n-1 If tracks are included in the overlapping gates, the detection n To incorporate this, detection z n The track closest to the

[0062] z n But the leading t n-1 is not included in any of the gates of the tracks, and is the maximum number of tracks

number

number

[0063] z n But the leading t n-1 is not included in any of the gates of the tracks, and is the maximum number of tracks

number

number

number

[0064] Since all other tracks have not been updated, initiated, or replaced (as in the previous step), these tracks are then updated with the same mean value. However, to account for uncertainty, the respective variance of each track is increased, for example based on a random walk model. The power of each track is also attenuated so that a sound source appearing in the future has a chance to be initiated. In this way, a tracking result 408 incorporating the latest detection 406 at time step n can be output to module 430. The tracking result 408 can be

number

[0065] When all audio tracks have been updated, the module 430 uses the nearest M2T allocation to generate the multi-stream VAD 412. Specifically, at time step n, the M2T allocation is n This may be done by assigning a value of 1 to the track closest to , and a value of 0 to the other tracks. In some implementations, a hangover may be applied to the VAD, so that it takes an intermediate value (e.g., -1) after being 1 in the previous time step but before being fully assigned zero. In this way, each track indicates whether speech activity was detected or not.

number

[0066] 6 shows an exemplary method 600 for enhancing a multi-source audio signal with multi-source tracking and VAD, according to various embodiments of the present disclosure. In some embodiments, method 600 may be performed by one or more components of audio signal processor 300 and / or by one or more components of multi-track VAD engine 400.

[0067] Method 600 begins at step 602, where TDOA trajectory information may be calculated based on spatial information of the microphone array. For example, the TDOA trajectory information may be calculated once at system startup by scanning the microphone array with incident rays having various azimuth and incidence angles. The calculation may be performed with reduced computational complexity in a multi-dimensional space constructed by pairing microphones from the microphone array, as further described with reference to FIG. 7.

[0068] Referring to FIG. 7, which provides further detailed steps for step 602, in step 702, a first microphone from the microphone array may be selected as a reference microphone. In step 704, the remaining microphones in the microphone array may each be paired with the reference microphone. In step 706, for each microphone pair, a TDOA position corresponding to a particular azimuth angle and a particular elevation angle of the incident ray may be calculated based on the distance and angle between the two microphones in each pair (e.g., according to equation (1) described with reference to FIG. 4). An exemplary microphone pair having a particular azimuth angle and a particular elevation angle of the incident ray is also shown in FIG. 5A.

[0069] If there are more microphone pairs to be processed in step 708, the method extracts the next microphone pair in step 710 and repeats step 706 until the TDOA positions for all microphone pairs have been calculated.

[0070] In step 712, if there are more scans in azimuth and elevation, the method extracts the next scan in azimuth and elevation in step 714 and repeats step 706 until TDOA positions have been calculated for all scans in azimuth and elevation.

[0071] If there are no more azimuth / elevation scans to process in step 712 (e.g., TDOA locations have been calculated across azimuth and elevation scans for all microphone pairs), a mesh of TDOA location points may be formed in step 716. An exemplary mesh of TDOA location points corresponding to different placements of the microphone array is shown in FIG. 5B.

[0072] Returning to Figure 6, once TDOA location information is determined at system startup, method 600 proceeds to step 604. In step 604, one or more multi-source audio signals may be received from a microphone array. For example, time-domain samples of multi-source audio 402 of Figure 4 may be received via audio input circuitry 315 of Figure 3.

[0073] In step 606, the one or more multi-source audio signals may be transformed from the time domain to a time-frequency representation. For example, as described in relation to Figure 4, the subband analysis module 405 may transform the time-domain signals into a time-frequency representation.

[0074] In step 608, TDOA detection data for one or more multi-source audio signals may be calculated by the STMV beamformer based on the calculated TDOA trajectories. For example, for each microphone pair, a covariance matrix for all frequency bands may be calculated (e.g., by equation (2) described in connection with FIG. 4) using the time-frequency representation of one or more multi-source audio signals from each microphone pair. Then, a steering matrix may be constructed for all frequency bands (e.g., by equation (3) described in connection with FIG. 4) based on the TDOA positions for different azimuth and elevation scans corresponding to each microphone pair. An azimuth covariance matrix may be constructed (e.g., by equation (4) described in connection with FIG. 4) aligned across all frequency bands based on the constructed steering matrix and the calculated covariance matrix. A pseudo-likelihood solution that minimizes the power of the beam may be determined (e.g., by equation (5) described with reference to FIG. 4) based on the constructed azimuth covariance matrix. The STMV joint pseudo-likelihood may then be calculated by taking the product of all pseudo-likelihood solutions determined across all microphone pairs (e.g., according to equation (6) described with reference to FIG. 4). The azimuth and elevation angle pair that maximizes the STMV joint pseudo-likelihood may be determined (e.g., according to equation (7) described with reference to FIG. 4). The determined azimuth and elevation angle pair may be converted into a polar coordinate representation representing the TDOA detection data (e.g., according to equation (8) described with reference to FIG. 4).

[0075] In step 610, multiple audio tracks may be updated, and VAD data may be constructed based on TDOA detection data calculated up to the current time step. For example, a TDOA detection corresponding to the current time step and a set of existing audio tracks established prior to the current time step may be identified. Method 600 may then determine whether to incorporate the TDOA detection into one of the existing audio tracks or construct a new audio track based on a comparison of the TDOA detection with the gate of the existing audio track (as described in connection with module 430 of FIG. 4). As another example, method 600 may assign a first value to the VAD of each audio track when the respective audio track is closest to the TDOA detection, and a second value to the VAD of the other audio track (as described in connection with module 430 of FIG. 4).

[0076] In step 612, one or more enhanced multi-source audio signals may be generated using the updated audio tracks and the constructed VAD data. For example, the enhanced multi-source signals may then be communicated to various devices or components. For example, the enhanced multi-source signals may be packetized and communicated over a network to other audio output devices (e.g., smartphones, computers, etc.). The enhanced multi-source signals may also be communicated to voice processing circuitry, such as an automated speech recognition component, for further processing.

[0077] The foregoing disclosure is not intended to limit the invention to the precise form or particular field of use disclosed. Accordingly, various alternative embodiments and / or modifications of the present disclosure, whether expressly described or implied herein, are contemplated in light of this disclosure. For example, the embodiments described herein may be used to provide the location of multiple sound sources within an environment (e.g., in applications combined with additional information from other modalities, such as video streams, 3D cameras, lidar, etc.) for the purpose of managing human-machine interaction tasks. Having described embodiments of the present disclosure, those skilled in the art will recognize advantages over conventional approaches and will recognize that changes in form and detail may be made without departing from the scope of the present disclosure. Accordingly, the present disclosure is limited only by the claims.

Claims

1. A method for enhancing multi-source speech by multi-source tracking and voice activity detection, receiving one or more multi-source audio signals from a microphone array via audio input circuitry; calculating TDOA detection data for the one or more multi-source audio signals by a steered minimum variance (STMV) beamformer based on TDOA trajectory information constructed within a multi-dimensional space defined by several microphone pairs from the microphone array; updating a plurality of audio tracks based on the TDOA detection data calculated up to a current time step; constructing voice activity detection (VAD) data for each of the plurality of audio tracks based on the calculated TDOA detection data; generating one or more enhanced multi-source audio signals using the updated audio tracks and the configured VAD data; and Including, the multi-dimensional space defined by a number of microphone pairs from the microphone array being: selecting a first microphone from the microphone array as a reference microphone; pairing each of the remaining microphones of the microphone array with the reference microphone; is formed by The TDOA trajectory information is For each microphone pair, determining a TDOA position corresponding to a particular azimuth angle and a particular elevation angle of the incident ray based on the distance and angle between the two microphones in each pair; forming a mesh of TDOA location points by varying the particular azimuth angle and the particular elevation angle of the incident ray across all of the microphone pairs; is calculated once in the startup stage based on the spatial information of the several microphone pairs by method.

2. When the microphone array is actually physically located on a second plane, the mesh of TDOA location points lies on a first plane in the multidimensional space having the same number of dimensions as the total number of microphone pairs.

10. The method of claim 1.

3. Calculating the TDOA detection data for the one or more multi-source audio signals includes, for each microphone pair: calculating a covariance matrix for all frequency bands using time-frequency representations of the one or more multi-source audio signals from each microphone pair; constructing a steering matrix for all frequency bands based on TDOA positions for different azimuth and elevation scans corresponding to each microphone pair; constructing an orientation covariance matrix aligned across all frequency bands based on the constructed steering matrix and the calculated covariance matrix; determining a pseudo-likelihood solution that minimizes the power of the beam based on the constructed orientation covariance matrix; Further comprising:

10. The method of claim 1.

4. A voice processing device for enhancing multi-source voices by multi-source tracking and voice activity detection, audio input circuitry configured to receive one or more multi-source audio signals from a microphone array; a time difference of arrival (TDOA) estimator configured to calculate TDOA detection data for the one or more multi-source audio signals by a steered minimum variance (STMV) beamformer based on TDOA trajectory information constructed in a multi-dimensional space defined by several microphone pairs from a microphone array; and a multi-source voice tracker configured to update a plurality of voice tracks based on the calculated TDOA detection data up to a current time step, and to construct voice activity detection (VAD) data for each of the plurality of voice tracks based on the calculated TDOA detection data; a speech enhancement engine configured to generate one or more enhanced multi-source speech signals using the updated plurality of speech tracks and the constructed VAD data; Equipped with the multi-dimensional space defined by a number of microphone pairs from the microphone array being: selecting a first microphone from the microphone array as a reference microphone; pairing each of the remaining microphones of the microphone array with the reference microphone; is formed by The TDOA trajectory information is For each microphone pair, determining a TDOA position corresponding to a particular azimuth angle and a particular elevation angle of the incident ray based on the distance and angle between the two microphones in each pair; forming a mesh of TDOA location points by varying the particular azimuth angle and the particular elevation angle of the incident ray across all of the microphone pairs; The audio processing device is calculated once in a startup stage based on spatial information of the several microphone pairs by

5. when the microphone array is actually physically located on a second plane, the mesh of TDOA location points lies on a first plane in the multidimensional space having as many dimensions as the total number of microphone pairs; The audio processing device of claim 4.

6. The TDOA estimator, for each microphone pair, calculating a covariance matrix for all frequency bands using time-frequency representations of the one or more multi-source audio signals from each microphone pair; constructing a steering matrix for all frequency bands based on TDOA positions for different scans in azimuth and elevation corresponding to each microphone pair; constructing an orientation covariance matrix aligned across all frequency bands based on the constructed steering matrix and the calculated covariance matrix; determining a pseudo-likelihood solution that minimizes the power of the beam based on the constructed orientation covariance matrix; configured to calculate the TDOA detection data by The audio processing device of claim 5.

7. The TDOA estimator calculating the STMV joint pseudo-likelihood by taking the product of all pseudo-likelihood solutions determined across all microphone pairs; determining a pair of azimuth and elevation angles that maximizes the STMV joint pseudo-likelihood; converting the determined azimuth and elevation angle pairs into a polar coordinate representation representing the TDOA detection data; and further configured to calculate the TDOA detection data by The audio processing device of claim 6.

8. the TDOA estimator is further configured to construct an orientation covariance matrix aligned across all of the frequency bands based on the constructed steering matrix and the calculated covariance matrix, and the calculation of the orientation covariance matrix is ​​repeated across all of the microphone pairs and all of the azimuth and elevation scans. The audio processing device of claim 7.

9. the multi-source audio tracker: Identifying a TDOA detection corresponding to a current time step and a set of existing audio tracks pre-established up to said current time step; determining whether to incorporate the TDOA detection into one of the existing audio tracks or establish a new audio track based on a comparison of the TDOA detection with the gating of the existing audio tracks; and updating the plurality of audio tracks based on the TDOA detection data calculated up to a current time step by The audio processing device of claim 4.

10. the multi-source audio tracker: assigning a first value to each audio track when the respective audio track is closest to the TDOA detection; assigning a second value to the other audio track; and constructing VAD data for the plurality of audio tracks based on the TDOA detection calculated by The audio processing device of claim 9.

Citation Information

Patent Citations

  • Voice interval detecting method, speech recognition method, voice interval detector, speech recognition device, and program and storage method therefor

    JP2012048119A

  • System and method for voice activity detection

    US20190385635A1

  • Robust speaker localization in presence of strong noise interference systems and methods

    US20210390952A1