Multi-source tracking and voice activity detection for planar microphone arrays

By employing multidimensional TDOA tracking and VAD mechanisms on a planar microphone array, the problem of separating the target audio signal from the noise source in a multi-stream audio environment is solved, achieving efficient signal separation and noise suppression while reducing computational complexity.

CN122177146APending Publication Date: 2026-06-09SYNAPTICS INC

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SYNAPTICS INC
Filing Date
2021-01-08
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

In existing multi-stream audio environments, it is difficult to efficiently separate the target audio signal from the noise source. Existing methods such as BSS and spatial filtering have response delays or rely on noise covariance matrix estimation, resulting in poor performance in practical applications.

Method used

Employing multidimensional TDOA tracking and VAD mechanisms, a microphone array with a general array geometry on a plane is used to perform multi-source tracking and noise suppression using TDOA trajectory information, avoiding ghosting and reducing computational complexity.

Benefits of technology

It achieves efficient separation of target audio signals and noise in multi-stream audio environments, reduces computational complexity, and improves the efficiency and accuracy of signal processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122177146A_ABST
    Figure CN122177146A_ABST
Patent Text Reader

Abstract

Embodiments described herein provide a combined multi-source time difference of arrival (TDOA) tracking and voice activity detection (VAD) mechanism that can be applicable to general array geometries, e.g., microphone arrays located on a plane. The combined multi-source TDOA tracking and VAD mechanism scans azimuth and elevation angles of the microphone array in microphone pairs based on which a planar locus of physically allowable TDOAs can be formed in a multi-dimensional TDOA space of multiple microphone pairs. In this way, the multi-dimensional TDOA tracking reduces the number of computations typically involved in conventional TDOA by performing a TDOA search separately for each dimension.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese patent application No. 202110023469.7, filed on January 8, 2021, entitled "Multi-source tracking and speech activity detection for planar microphone array". Technical Field

[0002] According to one or more embodiments, this disclosure relates generally to audio signal processing, and more particularly, for example, to systems and methods for multi-source tracking and multi-stream speech activity detection for general planar microphone arrays. Background Technology

[0003] In recent years, smart speakers and other voice-controlled devices and apparatuses have become increasingly popular. Smart speakers typically include a microphone array for receiving audio input from the environment (e.g., a user's verbal commands). When a target audio (e.g., a verbal command) is detected in the audio input, the smart speaker can translate the detected target audio into one or more commands and perform different tasks based on those commands.

[0004] One challenge for these smart speakers is to efficiently and effectively isolate target audio (e.g., verbal commands) from noise or other active speakers in an operational environment. For example, one or more speakers may be active in the presence of one or more noise sources. When the goal is to enhance a specific speaker, that speaker is referred to as the target speaker, while the remaining speakers can be considered sources of interference. Most existing speech enhancement algorithms utilize multiple input channels (microphones) to leverage the spatial information of the sources, such as blind source separation (BSS) methods associated with independent component analysis (ICA) and spatial filtering or beamforming methods.

[0005] However, the BSS method is primarily designed for batch processing, and its large response latency may often make it undesirable or even unsuitable for real-world applications. On the other hand, spatial filtering or beamforming methods typically require supervision under speech activity detection (VAD) as a cost function to be minimized, which may rely excessively on the estimation of the covariance matrix that is only related to noise / interference segments.

[0006] Therefore, there is a need for improved systems and methods for detecting and processing (one or more) target audio signals in multi-stream audio environments. Attached Figure Description

[0007] A better understanding of the aspects and advantages of this disclosure can be achieved by referring to the following accompanying drawings and the detailed description that follows. It should be understood that the same reference numerals are used to identify the same elements illustrated in one or more of the drawings, wherein the illustrations in the drawings are for the purpose of illustrating embodiments of the disclosure and not for the purpose of limiting the embodiments of the disclosure. The components in the drawings are not necessarily to scale, but rather the emphasis is on clearly illustrating the principles of the disclosure.

[0008] Figure 1 An example operating environment for an audio processing device according to one or more embodiments of the present disclosure is illustrated.

[0009] Figure 2 This is a block diagram of an example audio processing device according to one or more embodiments of the present disclosure.

[0010] Figure 3 This is a block diagram of an example audio signal processor for multitrack audio enhancement according to one or more embodiments of the present disclosure.

[0011] Figure 4 This is a block diagram of an example multitrack activity detection engine for processing multiple audio signals from a general-purpose microphone array, according to various embodiments of the present disclosure.

[0012] Figure 5A This is a diagram illustrating example geometry of a microphone pair according to one or more embodiments of the present disclosure.

[0013] Figure 5B This is a diagram illustrating example grids of time difference of arrival (TDOA) trajectory information corresponding to different microphone array geometries in a multidimensional space according to one or more embodiments of the present disclosure.

[0014] Figure 6 This is a logic flowchart of an example method for enhancing multi-source audio signals through multi-source tracking and activity detection, according to various embodiments of the present disclosure.

[0015] Figure 7 This is a logic flowchart of an example process for calculating TDOA trajectory information in a multidimensional space using a microphone, according to various embodiments of the present disclosure. Detailed Implementation

[0016] This disclosure provides improved systems and methods for detecting and processing (one or more) target audio signals in a multi-stream audio environment.

[0017] Speech Activity Detection (VAD) can be used to supervise speech enhancement of a target audio source while utilizing spatial information from sources obtained from multiple input channels. VADs allow for the induction of spatial statistics of interfering / noise sources during periods of silence from the desired speaker, so that the effects of noise / interference can subsequently be eliminated when the desired speaker becomes active. For example, the VAD for each source can be derived to track the spatial information of the source in the form of Time Difference of Arrival (TDOA) or Direction of Arrival (DOA) by constructing the VAD by determining when the detection occurs near an existing track and utilizing the history of the detection. This process is often referred to as measurement-to-track (M2T) assignment. In this way, multiple VADs can be derived for all sources of interest.

[0018] Specifically, existing DOA methods typically construct a single steering vector for the entire microphone array based on closed-form mappings of azimuth and elevation angles, which can be used to utilize the specific geometry of linear or circular arrays. Such DOA methods cannot be extended to general or arbitrary geometries of microphone arrays. Furthermore, these closed-form mapping-based DOA methods often require extensive searches in multidimensional space. For arbitrary geometries, existing TDOA-based methods can be used, which may not be limited to a specific array geometry and can construct multiple steering vectors for each microphone pair to form a multidimensional TDOA vector (one dimension per pair). However, these existing methods risk introducing TDOA ghosting caused by the cross-intersection of peaks from the spectrum of each TDOA pair. Therefore, further post-processing involving specific array geometries is usually required to remove TDOA ghosting.

[0019] Given the need for multi-stream VADs that are not constrained by particular array geometries, the embodiments described herein provide a combined multi-source TDOA tracking and VAD mechanism applicable to general array geometries, such as microphone arrays located in a plane. By performing TDOA searches separately for each dimension, the combined multi-source TDOA tracking and VAD mechanism can reduce the amount of computation typically involved in conventional TDOA.

[0020] In some embodiments, a multidimensional TDOA method for a general array geometry located on a plane is employed, which avoids unwanted ghosting TDOA. In one embodiment, the Cartesian coordinates of a generally configured microphone are obtained, and one of the microphones can be selected as a reference microphone. The azimuth and elevation angles of the microphones can be scanned, based on which a physically permissible planar trajectory of TDOA can be formed in the multidimensional TDOA space of multiple microphone pairs. In this way, the formed planar trajectory avoids the formation of ghosting TDOA, thus eliminating the need for further post-processing to remove ghosting TDOA. Moreover, compared to the full DOA scan method, the multidimensional TDOA method disclosed herein reduces computational complexity by separately performing searches in the pairwise TDOA domains associated with each dimension, rather than searching in the full multidimensional space.

[0021] Figure 1 An example operating environment 100 is illustrated in which an audio processing system according to various embodiments of the present disclosure may operate. Operating environment 100 includes an audio processing device 105, a target audio source 110, and one or more noise sources 135-145. Figure 1 In the example illustrated, the operating environment 100 is depicted as a room; however, it is conceivable that the operating environment may include other areas, such as a vehicle interior, an office conference room, a family room, an outdoor stadium, or an airport. According to various embodiments of this disclosure, the audio processing device 105 may include two or more audio sensing components (e.g., microphones) 115a-115d and, optionally, one or more audio output components (e.g., speakers) 120a-120b.

[0022] Audio processing device 105 can be configured to sense sound via audio sensing components 115a-115d and generate a multi-channel audio input signal comprising two or more audio input signals. Audio processing device 105 can use audio processing techniques disclosed herein to process the audio input signal to enhance the audio signal received from target audio source 110. For example, the processed audio signal can be transmitted to other components within audio processing device 105 (such as a voice recognition engine or voice command processor) or to an external device. Therefore, audio processing device 105 can be a standalone device for processing audio signals, or it can be a device that converts the processed audio signal into other signals (e.g., commands, instructions, etc.) for interacting with or controlling external devices. In other embodiments, audio processing device 105 can be a communication device, such as a mobile phone or a device implementing Voice over IP (VoIP), and the processed audio signal can be transmitted over a network to another device for output to a remote user. The communication device can also receive the processed audio signal from a remote device and output the processed audio signal via audio output components 120a-120b.

[0023] The target audio source 110 can be any source that produces sound detectable by the audio processing device 105. The target audio to be detected by the system can be defined based on criteria specified by the user or system requirements. For example, the target audio can be defined as human speech, a sound made by a particular animal, or a machine. In the illustrated example, the target audio is defined as human speech, and the target audio source 110 is a person. In addition to the target audio source 110, the operating environment 100 can include one or more noise sources 135-145. In various embodiments, sounds that are not the target audio can be processed as noise. In the illustrated example, noise sources 135-145 can include: a speaker 135 playing music; a television 140 playing a television program, movie, or sporting event; and background dialogue between non-target speakers 145. It will be understood that different noise sources can exist in a variety of operating environments.

[0024] Note that the target audio and noise may arrive at the audio sensing components 115a-115d of the audio processing device 105 from different directions and at different times. For example, noise sources 135-145 may generate noise at different locations within the operating environment 100, and the target audio source (person) 110 may speak while moving between locations within the operating environment 100. Furthermore, the target audio and / or noise may be reflected from fixed objects (e.g., walls) within the room 100. For example, consider the path that the target audio may take from the target audio source 110 to reach each of the audio sensing components 115a-115d. As indicated by arrows 125a-125d, the target audio may propagate directly from the target audio source 110 to the audio sensing components 115a-115d respectively. Additionally, the target audio may be reflected from walls 150a and 150b, and indirectly from the target audio source 110 to the audio sensing components 115a-115d, as indicated by arrows 130a-130b. In various embodiments, the audio processing device 105 may use one or more audio processing techniques to estimate and apply room impulse response to further enhance the target audio and suppress noise.

[0025] Figure 2 An example audio processing device 200 according to various embodiments of the present disclosure is illustrated. In some embodiments, the audio processing device 200 may be implemented as Figure 1 The audio processing device 105. The audio processing device 200 includes an audio sensor array 205, an audio signal processor 220, and a host system component 250.

[0026] The audio sensor array 205 includes two or more sensors, each of which can be implemented as a transducer that converts an audio input in the form of a sound wave into an audio signal. In the illustrated setting, the audio sensor array 205 includes a plurality of microphones 205a-205n, each microphone generating an audio input signal that is provided to the audio input circuitry 222 of the audio signal processor 220. In one embodiment, the audio sensor array 205 generates a multi-channel audio signal, wherein each channel corresponds to an audio input signal from one of the microphones 205a-n.

[0027] Audio signal processor 220 includes audio input circuitry 222, digital signal processor 224, and optional audio output circuitry 226. In various embodiments, audio signal processor 220 may be implemented as an integrated circuit including analog circuitry, digital circuitry, and digital signal processor 224, operable to execute program instructions stored in firmware. Audio input circuitry 222 may, for example, include an interface to audio sensor array 205, anti-aliasing filters, analog-to-digital converter circuitry, echo cancellation circuitry, and other audio processing circuitry and components. Digital signal processor 224 is operable to process multi-channel digital audio signals to generate enhanced audio signals, which are output to one or more host system components 250. In various embodiments, digital signal processor 224 may be operable to perform echo cancellation, noise cancellation, target signal enhancement, post-filtering, and other audio signal processing functions.

[0028] Optional audio output circuitry 226 processes the audio signals received from digital signal processor 224 for output to at least one speaker, such as speakers 210a and 210b. In various embodiments, audio output circuitry 226 may include a digital-to-analog converter that converts one or more digital audio signals into analog audio signals and one or more amplifiers for driving speakers 210a-210b.

[0029] The audio processing device 200 can be implemented as any device operable to receive and enhance target audio data, such as, for example, a mobile phone, smart speaker, tablet computer, laptop computer, desktop computer, voice-controlled device, or automobile. The host system component 250 may include various hardware and software components for operating the audio processing device 200. In the illustrated embodiment, the host system component 250 includes a processor 252, a user interface component 254, a communication interface 256 for communicating with external devices and networks, such as network 280 (e.g., the Internet, cloud, local area network, or cellular network) and mobile device 284, and memory 258.

[0030] Processor 252 and digital signal processor 224 may include one or more of a processor, microprocessor, single-core processor, multi-core processor, microcontroller, programmable logic device (PLD) (e.g., field-programmable gate array (FPGA)), digital signal processing (DSP) device, or other logic device (which may be configured by hardwiring, executing software instructions, or a combination of both) to perform the various operations discussed herein with respect to embodiments of this disclosure. Host system component 250 is configured to interface and communicate with audio signal processor 220 and other host system components 250, such as via a bus or other electronic communication interface.

[0031] It will be understood that although the audio signal processor 220 and the host system component 250 are shown as a combination of hardware components, circuitry, and software, in some embodiments, at least some or all of the functionality that the hardware components and circuitry are operable to perform may be implemented as software modules executed by the processor 252 and / or the digital signal processor 224 in response to software instructions and / or configuration data stored in the firmware or memory 258 of the digital signal processor 224.

[0032] The memory 258 can be implemented as one or more storage devices operable to store data and information including audio data and program instructions. The memory 258 may include one or more storage devices of various types, including volatile and non-volatile storage devices, such as RAM (random access memory), ROM (read-only memory), EEPROM (electrically erasable read-only memory), flash memory, hard disk drive and / or other types of memory.

[0033] Processor 252 may be operable to execute software instructions stored in memory 258. In various embodiments, voice recognition engine 260 is operable to process enhanced audio signals received from audio signal processor 220, including recognizing and executing voice commands. Voice communication component 262 may be operable to facilitate voice communication with one or more external devices, such as mobile device 284 or user equipment 286, such as via voice calls over mobile or cellular telephone networks or VoIP calls over IP networks. In various embodiments, voice communication includes transmitting enhanced audio signals to external communication devices.

[0034] User interface component 254 may include a display, touchpad display, keypad, one or more buttons and / or other input / output components operable to enable a user to interact directly with audio processing device 200.

[0035] Communication interface 256 facilitates communication between audio processing device 200 and external devices. For example, communication interface 256 may enable Wi-Fi (e.g., 802.11) or Bluetooth connectivity between audio processing device 200 and one or more local devices, such as mobile device 284 or a wireless router that provides network access to remote server 282 via network 280. In various embodiments, communication interface 256 may include other wired and wireless communication components that facilitate direct or indirect communication between audio processing device 200 and one or more other devices.

[0036] Figure 3 An example audio signal processor 300 according to various embodiments of the present disclosure is illustrated. In some embodiments, the audio signal processor 300 is embodied as one or more integrated circuits, which include analog and digital circuitry and a digital signal processor (such as...) Figure 2 The firmware logic is implemented in the audio signal processor 320. As illustrated, the audio signal processor 300 includes an audio input circuit 315, a sub-band frequency analyzer 320, a multi-track VAD engine 325, an audio enhancement engine 330, and a synthesizer 335.

[0037] The audio signal processor 300 receives multi-channel audio input from multiple audio sensors, such as a sensor array 305 including at least two audio sensors 305a-n. The audio sensors 305a-305n may include, for example, […]. Figure 2 The audio signal processor 300 may have a microphone integrated into an audio processing device such as audio processing device 200 or integrated with external components connected thereto. According to various embodiments of this disclosure, the audio signal processor 300 may or may not know the arrangement of the audio sensors 305a-305n.

[0038] The audio signal may initially be processed by audio input circuitry 315, which may include an anti-aliasing filter, an analog-to-digital converter, and / or other audio input circuitry. In various embodiments, audio input circuitry 315 outputs a digital, multi-channel, time-domain audio signal, where M is the number of sensor (e.g., microphone) inputs. The multi-channel audio signal is input to sub-band frequency analyzer 320, which divides the multi-channel audio signal into consecutive frames and decomposes each frame of each channel into multiple frequency sub-bands. In various embodiments, sub-band frequency analyzer 320 includes a Fourier transform process and outputs multiple frequency windows. The decomposed audio signal is then provided to multi-track VAD engine 325 and audio enhancement engine 330.

[0039] The multi-track VAD engine 325 is operable to analyze frames of one or more audio tracks and generate a VAD output indicating the presence or absence of target audio activity in the current frame. As discussed above, the target audio can be any audio to be recognized by the audio system. When the target audio is human speech, the multi-track VAD engine 325 can be specifically implemented for detecting speech activity. In various embodiments, the multi-track VAD engine 325 is operable to receive frames of audio data and generate a VAD indication output for each audio track regarding the presence or absence of target audio on the corresponding audio track corresponding to the frame of audio data. Figure 4 The diagram in section 400 further illustrates the detailed components and operation of the multi-track VAD engine 325.

[0040] Audio enhancement engine 330 receives subband frames from subband frequency analyzer 320 and VAD indications from multitrack VAD engine 325. According to various embodiments of this disclosure, audio enhancement engine 330 is configured to process subband frames based on the received multitrack VAD indications to enhance the multitrack audio signal. For example, audio enhancement engine 330 may enhance portions of the audio signal determined to originate from a target audio source and suppress other portions of the audio signal determined to be noise.

[0041] After enhancing the target audio signal, the audio enhancement engine 330 can pass the processed audio signal to the synthesizer 335. In various embodiments, the synthesizer 335 reconstructs one or more multi-channel audio signals on a frame-by-frame basis by combining subbands to form an enhanced time-domain audio signal. The enhanced audio signal can then be transformed back to the time domain and sent to system components or external devices for further processing.

[0042] Figure 4 An example multitrack VAD engine 400 for processing multiple audio signals from a general-purpose microphone array is illustrated according to various embodiments of the present disclosure. The multitrack VAD engine 400 can be implemented as a combination of digital circuitry and logic executed by a digital signal processor. In some embodiments, the multitrack VAD engine 400 can be installed in a... Figure 3 In audio signal processors like the 300, the multi-track VAD engine 400 can... Figure 3 Further structural and functional details are provided for the multi-track VAD engine 325.

[0043] According to various embodiments of the present disclosure, the multi-track VAD engine 400 includes a sub-band analysis module 405, a block-based TDOA estimation module 410, a TDOA trajectory calculation module 420, and a multi-source tracking and multi-stream VAD estimation module 430.

[0044] Subband analysis module 405 receives data from... The multiple audio signals 402 represent the signals from a total of M microphones (e.g., similar to...). Figure 3 The sampled time-domain audio signal recorded at the m-th microphone of the audio sensor 305a-n in the image. It can be transmitted via... Figure 3 The audio input circuit 315 in the middle is used to receive audio signals. , .

[0045] Subband analysis module 405 is configured to acquire audio signal 402 and transform it into time-frequency domain representation 404, which is represented as the original time-domain audio signal. corresponding ,in Indicates the sub-band time index, and Indicates the frequency band index. For example, the subband analysis module 405 can be similar to... Figure 3 The subband frequency analyzer 320 performs a Fourier transform to convert the input time-domain audio signal into a frequency-domain representation. The subband analysis module 405 can then send the generated time-frequency domain representation 404 to the block-based TDOA estimation module 410 and the multi-source tracking and multi-stream VAD estimation module 430.

[0046] The TDOA trajectory calculation module 420 is configured to scan a general microphone array (e.g., audio sensors 305a-n forming a general array geometry). For example, for any microphone array geometry given on a plane, the trajectory of the permissible TDOA position is calculated once at system startup. This trajectory of the point avoids ghosting.

[0047] For an array of M microphones, the first microphone can be chosen as the reference microphone, which in turn provides all M-1 microphone pairs relative to the first microphone. For example, Figure 5A An example microphone pair is illustrated. The microphone pair indexed as i-1 includes microphone i 502 and a reference microphone l 501 for incident sound ray 505, which has an azimuth angle θ emitted from a distant source (assuming a far-field model) and an elevation angle of zero. The distance between microphone pairs 501 and 502 and the angle between the two microphones are expressed as follows: and It can be calculated given the Cartesian coordinates of the i-th microphone 502. For the general case where the incident sound ray 505 forms an angle with azimuth θ and elevation ϕ, the TDOA of the (i-1)-th microphone pair can be calculated as follows: Where c is the propagation speed.

[0048] After scanning at different elevation and azimuth angles, the TDOA trajectory calculation module 420 can construct a permissible TDOA mesh. When all M microphones are located on the plane, the generated TDOA trajectory (for all scans of θ and ϕ) is... The M microphones also lie on a plane in (M-1) dimensional space. Different arrangements of the M microphones may result in different planes in (M-1) dimensional space.

[0049] For example, Figure 5B Two different example microphone layouts are illustrated along with their corresponding TDOA grids. At 510, a set of M=4 microphones is shown, with the distance between the first and third microphones being 8 cm, and the grid for permissible TDOA generation is shown in M-1=3 dimensional space as shown at 515. When the distance between the first and third microphones increases to 16 cm, the grid for permissible TDOA generation is shown at 520, and at 525, it is shown.

[0050] Return to reference Figure 4 The TDOA trajectory calculation module 420 can then send the (M-1) dimensional TDOA 403 to the block-based TDOA estimation module 410. The block-based TDOA estimation module 410 receives the TDOA 403 and a time-frequency domain representation 404 of the multi-source audio. Based on the time-frequency domain representation 404, the TDOA estimation module 410 uses data obtained from consecutive frames to extract the source microphones (e.g., ...). Figure 3 The TDOA information of the audio sensor 305a-n shown is shown.

[0051] In one embodiment, the block-based TDOA estimation module 410 employs a guided minimum variance (STMV) beamformer to obtain TDOA information from the time-frequency domain representation 404 of the multi-source audio. Specifically, the block-based TDOA estimation module 410 can select a microphone as a reference microphone and then specify a total of M-1 microphone pairs by pairing the remaining M1 microphones with the reference microphone. The microphone pairs are... index.

[0052] For example, the first microphone can be chosen as the reference microphone, and therefore, This represents the time-frequency representation of the audio from the reference microphone. For the p-th microphone pair, the block-based TDOA estimation module 410 calculates the frequency representation of the p-th pair in matrix form as follows: ,in() T This represents the transpose. Then, the block-based TDOA estimation module 410 calculates the covariance matrix of the p-th input signal pair for each frequency band k: in() H This represents the transpose of Hermitian.

[0053] In some implementations, computation is performed on blocks of a certain number of consecutive frames. Summation within the block. For simplicity, block indices are ignored here.

[0054] Then, the block-based TDOA estimation module 410 can construct a steering matrix for each pair and frequency band as follows: in, The p-th pair of TDOA is obtained from the TDOA trajectory calculation module 420 after different scans of θ and ϕ (omitted for brevity); It is the frequency at frequency band k; and diag([a, b]) represents a 2×2 diagonal matrix with diagonal elements a and b.

[0055] For each microphone pair p, the block-based TDOA estimation module 410 constructs a direction covariance matrix coherently aligned across all frequency bands in the following manner: .

[0056] Directional covariance matrix The calculation of all microphones for p and for Azimuth / Elevation The computation is repeated across all scans. To reduce computation across all scans, the TDOA space corresponding to each dimension p of the p-th microphone pair is linearly quantized into... q Segment. At the start of processing (when the system boots up), by scanning each azimuth and elevation angle. Obtained TDOA trajectory points It is mapped to the closest quantization point for each dimension. For each azimuth / elevation angle. , The mapping is stored in memory, where It is a TDOA index that quantizes the dimension p related to the scanning angles θ and ϕ.

[0057] For example, if there are M=4 microphones, and the azimuth and elevation scans are respectively... Required to be executed The number of different calculations is When TDOA trajectory points When quantized, because some of the TDOA dimensions can be quantized as q The same segment within each quantization segment, therefore not all calculations need to be performed. Thus, for example, if q = 50, then calculate The maximum number of different calculations required is reduced to Utilizing TDOA quantization for execution The pseudocode for the calculation can be shown in Algorithm 1 as follows: .

[0058] Next, for each pair p The direction for minimizing beam power, which follows the distortion-free criterion (with its equivalent pseudo-likelihood solution), is calculated as follows: in Then, the block-based TDOA estimation module 410 can calculate the joint pseudo-likelihood of STMV for all M-1 pairs of microphones as follows: .

[0059] Then, the azimuth and elevation angles that generate the maximum STMV joint pseudo-likelihood for all M-1 pairs are identified, expressed by the following equation. Then you can align the azimuth and elevation angles. , Used for multi-source tracking and multi-stream VAD estimation. One possible solution could be to directly track the angle between the two microphones in each microphone pair. However, due to the wrap-around effect of the azimuth angle in 360 degrees, if the angle between the paired microphones is tracked directly, when the microphone source crosses 0... o Oriented to 360 o Conversely, if the polarity changes, track loss may occur. Therefore, to avoid such disorder, the detected z is calculated in a circular manner based on the angle between paired microphones using a polarity transformation, as follows: in It is a scaling constant that can expand the measurement space, thereby allowing tracking to be performed using parameters that are relevant to meaningful concepts such as angles.

[0060] Then, the block-based TDOA estimation module 410 sends the calculated detection z to the multi-source tracking and multi-stream VAD estimation module 430. If the orbital is... The maximum number of TDOA values ​​obtained from the block-based TDOA estimation module 410 is tracked by recursively updating the trajectory obtained from previous steps. Specifically, if the detection obtained at block (time-step) n-1 is determined by z... n-1 This indicates that t existed until then. n-1For each orbit, a new detection z appearing at time step n... n 406, Multi-source tracking and multi-stream VAD estimation module 430 Based on existing track gate pair for new detection z n The following steps will be taken: If z n Only fall into the previous t n-l In one of the gates of the orbitals, the special orbital is updated to incorporate the detection z. n .

[0061] If z n Falling into multiple previous t n-l Among the overlapping gates of the orbits, the update is the one closest to the detection z. n The orbit to be incorporated into the detection z n .

[0062] If z n Not falling into any previous t n-l In the gate of the track, and the maximum number of tracks has not been reached. (For example, t) n-l < If so, a new orbit will be initiated to merge with the detection z. n And update the number of existing orbits at time step n, for example, .

[0063] If z n Not falling into any previous t n-1 The gate has a track, and the maximum number of tracks has been reached. (For example, t) n-1 = ), then in the existing The track with the lowest power among the tracks is stopped, and a new track is used to replace it to incorporate the detection z. n .

[0064] For all other orbits that are not updated, activated, or replaced (as in previous steps), these orbits are updated with the same average value, but the variance of each corresponding orbit is increased to account for uncertainty (e.g., based on a random walk model). The power of each corresponding orbit is also decayed, allowing future sources to be activated. In this way, the tracking result 408, which incorporates the latest detection 406 at time step n, can be output from module 430. express.

[0065] Once all audio tracks have been updated, module 430 generates multistream VAD 412 using the nearest neighbor M2T assignment. Specifically, at time step n, this can be achieved by assigning 1 to the nearest detection z. nThe orbit and the allocation of 0 to other orbits to perform M2T allocation. In some implementations, a hangover can be applied to the VAD to have an intermediate value, such as -1, before it is fully allocated to zero after being 1 in a previous time step. In this way, from module 430, for example, to Figure 3 The audio enhancement engine 330 output is from The multi-stream VAD412 is used for audio enhancement, with each multi-stream VAD412 representing whether any speech activity was detected in the corresponding track.

[0066] Figure 6 An example method 600 for enhancing a multi-source audio signal through multi-source tracking and VAD according to various embodiments of the present disclosure is illustrated. In some embodiments, method 600 may be performed by one or more components of an audio signal processor 300 and / or one or more components of a multi-track VAD engine 400.

[0067] Method 600 begins with step 602, in which TDOA trajectory information can be calculated based on the spatial information of the microphone array. For example, TDOA trajectory information can be calculated once at system startup by scanning the microphone array using incident sound rays at varying azimuth and elevation angles. (See also: Regarding...) Figure 7 As further described, computations can be performed with reduced complexity in a multidimensional space constructed by pairing microphones from a microphone array.

[0068] refer to Figure 7 This provides further detailed steps for step 602, where in step 702, a first microphone can be selected from the microphone array as a reference microphone. In step 704, each remaining microphone from the microphone array can be paired with the reference microphone. In step 706, for each microphone pair, the selection can be based on the distance and angle between the two microphones in the corresponding pair (e.g., according to...). Figure 4 The equation (1) is described to calculate the TDOA position corresponding to a specific azimuth angle and a specific elevation angle of the incident sound ray. Figure 5A The paper also shows example microphone pairs with specific azimuth and elevation angles for the incident sound rays.

[0069] In step 708, if there are more microphone pairs to process, the method retrieves the next microphone pair in step 710 and repeats in step 706 until the TDOA positions of all microphone pairs have been calculated.

[0070] In step 712, if there are more scans of azimuth and elevation angles, the method retrieves the next scan of azimuth and elevation angles in step 714 and repeats in step 706 until the TDOA position of all scans of azimuth and elevation angles is calculated.

[0071] In step 712, when there are no more scans of azimuth and elevation angles to process (e.g., the TDOA positions have already been calculated for all microphone pairs on all scans of azimuth and elevation angles), a grid of TDOA position points can be formed in step 716. Figure 5B The image shows an example grid of TDOA location points corresponding to different geometries of the microphone array.

[0072] Return to reference Figure 6 When calculating TDOA trajectory information during system startup, method 600 proceeds to step 604. In step 604, one or more multi-source audio signals can be received from the microphone array. For example, this can be achieved via... Figure 3 The audio input circuit 315 in the middle is used to receive Figure 4 The time-domain sampling of the multi-source audio signal 402 in the middle.

[0073] In step 606, one or more multi-source audio signals can be transformed from a time-domain representation to a time-frequency representation. For example, sub-band analysis module 405 can transform a time-domain signal into a time-frequency domain representation, as per [reference to...]. Figure 4 As described.

[0074] In step 608, TDOA detection data can be calculated for one or more multi-source audio signals based on the calculated TDOA trajectory and according to the STMV beamformer. For example, for each microphone pair, time-frequency representations of one or more multi-source audio signals from the corresponding microphone pair can be used for each frequency band (e.g., based on information about...). Figure 4 The covariance matrix can then be calculated using the equation (2) described above. Then, the covariance matrix can be calculated for each frequency band based on the TDOA location (e.g., according to the information about the frequency band). Figure 4 The steering matrix is ​​constructed using the equation (3) described above, where the TDOA position is for different scans corresponding to the azimuth and elevation angles of the respective microphone pairs. The steering matrix can be based on the constructed steering matrix and the calculated covariance matrix (e.g., according to the information provided). Figure 4 The described equation (4) is used to construct the directional covariance matrix aligned across all frequency bands. The constructed directional covariance matrix can be based on (e.g., according to the equation regarding...) Figure 4 The described equation (5) determines the pseudo-likelihood solution that minimizes the beam power. This can then be determined by taking the product of all determined pseudo-likelihood solutions across all microphone pairs (e.g., based on the information provided). Figure 4 The equation (6) described above is used to calculate the joint pseudo-likelihood of STMV. Then, it can be (e.g., based on the information about...) Figure 4 The equation (7) described determines a pair of azimuth and elevation angles that maximize the joint pseudo-likelihood of STMV. Then it can be (e.g., based on the information about...) Figure 4 The equation (8) described transforms a determined azimuth and elevation angle into polar coordinates representing the TDOA detection data.

[0075] In step 610, the TDOA detection data calculated up to the current time step can update multiple audio tracks and construct VAD data. For example, TDOA detections corresponding to the current time step and a set of existing audio tracks previously established up to the current time step can be identified. Then, method 600 can determine whether to incorporate the TDOA detection into one of the existing audio tracks or to create a new audio track based on a comparison between the gates of the existing audio tracks and the TDOA detections. Figure 4 As described in module 430. For another example, method 600 may assign a first value to the VAD of the corresponding audio track when the corresponding audio track is closest to the TDOA detection, and assign a second value to the VAD of other audio tracks, as per [reference to...]. Figure 4 As described in module 430.

[0076] In step 612, one or more enhanced multi-source audio signals can be generated using updated multiple audio tracks and constructed VAD data. For example, the enhanced multi-source audio signals can then be transmitted to various devices or components. As another example, the enhanced multi-source audio signals can be packaged and transmitted over a network to another audio output device (e.g., a smartphone, computer, etc.). The enhanced multi-source audio signals can also be transmitted to speech processing circuitry, such as an automatic speech recognition component, for further processing.

[0077] The foregoing disclosure is not intended to limit the invention to the precise forms disclosed or any particular field of use. Therefore, it is contemplated that various alternative embodiments and / or modifications to this disclosure are possible, whether explicitly described or implied herein. For example, the embodiments described herein can be used to provide the location of multiple sound sources in an environment for supervising human-computer interaction tasks, e.g., in applications incorporating additional information from other modalities such as video streams, 3D cameras, LiDAR, etc. Embodiments of this disclosure have been described thus, and those skilled in the art will recognize the advantages over conventional methods and that changes in form and detail may be made without departing from the scope of this disclosure. Therefore, this disclosure is limited only by the claims.

Claims

1. A method for enhancing multi-source audio through multi-source tracking and speech activity detection, comprising: Receive one or more multi-source audio signals from a microphone array via an audio input circuit; Construct a multidimensional space defined by multiple microphone pairs from the microphone array; For each microphone pair, based on the distance and angle between the two microphones in the corresponding microphone pair, calculate the location point corresponding to the azimuth angle and elevation angle of the incident sound ray; A grid of TDOA location points is formed by iterating through all microphone pairs and all scans of azimuth and elevation angles. as well as Multiple audio tracks are updated based on the grid of the TDOA location points.

2. The method of claim 1, wherein iterating through all microphone pairs and all scans of azimuth and elevation angles comprises: The changes apply to the azimuth and elevation angles of the incident sound rays for all microphone pairs; as well as For each variation and each microphone pair, calculate the corresponding TDOA position.

3. The method according to claim 1, further comprising: Based on the directional minimum variance (STMV) beamformer and the TDOA trajectory information constructed in the multidimensional space, TDOA detection data is calculated for the one or more multi-source audio signals. Multiple audio tracks are updated based on TDOA detection data calculated up to the current time step. Based on the calculated TDOA detection data, construct speech activity detection (VAD) data for each of the plurality of audio tracks; as well as One or more enhanced multi-source audio signals are generated using updated multiple audio tracks and constructed VAD data.

4. The method of claim 1, wherein the multidimensional space defined by a plurality of microphone pairs from the microphone array is formed by the following steps: Select a first microphone from the microphone array as a reference microphone; and Each remaining microphone from the microphone array is paired with the reference microphone.

5. The method of claim 1, wherein when the microphone array is physically located on a second plane in reality, the grid of the TDOA location points is located on a first plane in the multidimensional space, the multidimensional space having a number of dimensions equal to the total number of microphone pairs.

6. The method according to claim 1, further comprising: For each microphone pair: The covariance matrix is ​​calculated for each frequency band using the time-frequency representation of the one or more multi-source audio signals from the corresponding microphone pair; A steering matrix is ​​constructed for each frequency band based on the TDOA position, wherein the TDOA position is for different scans corresponding to the azimuth and elevation angles of the corresponding microphone pair; Based on the constructed steering matrix and the calculated covariance matrix, a directional covariance matrix aligned across all frequency bands is constructed. as well as Based on the constructed directional covariance matrix, a pseudo-likelihood solution that minimizes beam power is determined.

7. The method according to claim 6, further comprising: The STMV joint pseudo-likelihood is calculated by taking the product of all determined pseudo-likelihood solutions across all microphone pairs. Determine an azimuth and elevation angle that maximizes the joint pseudo-likelihood of the STMV; as well as The determined azimuth and elevation angles are converted into polar coordinates representing the TDOA detection data.

8. The method of claim 6, wherein the directional covariance matrix, which is constructed based on the constructed steering matrix and the calculated covariance matrix, is repeated across all microphone pairs and all scans of azimuth and elevation angles.

9. The method of claim 6, wherein constructing the orientation covariance matrix aligned across all frequency bands based on the constructed steering matrix and the calculated covariance matrix to reduce repetition is performed by the following steps: The multidimensional space is divided into multiple segments, wherein the number of segments is less than the total number of dimensions of the multidimensional space; Map each TDOA location point from the TDOA location point grid to the nearest segment; as well as The direction covariance matrix is ​​calculated using the number of segments and the mapping between the grid of the TDOA location points and the number of segments, rather than using the grid of the TDOA location points built from all scans of the azimuth and elevation angles.

10. The method of claim 1, wherein updating the plurality of audio tracks further comprises: Identify the TDOA detection corresponding to the current time step and a set of existing audio tracks that have been established up to the current time step; as well as The determination is based on a comparison between the TDOA detection and the gates of the existing audio tracks to determine whether to incorporate the TDOA detection into one of the existing audio tracks or to create a new audio track.

11. An audio processing apparatus for enhancing multi-source audio through multi-source tracking and speech activity detection, comprising: An audio input circuit configured to receive one or more multi-source audio signals from a microphone array; as well as One or more processors, which are configured to: Construct a multidimensional space defined by multiple microphone pairs from the microphone array; For each microphone pair, based on the distance and angle between the two microphones in the corresponding microphone pair, calculate the location point corresponding to the azimuth angle and elevation angle of the incident sound ray; A grid of TDOA location points is formed by iterating through all microphone pairs and all scans of azimuth and elevation angles. as well as Multiple audio tracks are updated based on the grid of the TDOA location points.

12. The audio processing apparatus of claim 11, wherein the multidimensional space is formed by the following steps: Select a first microphone from the microphone array as a reference microphone; and Each remaining microphone from the microphone array is paired with the reference microphone.

13. The audio processing apparatus of claim 11, wherein the one or more processors are further configured to: Based on the directional minimum variance (STMV) beamformer and the TDOA trajectory information constructed in the multidimensional space, TDOA detection data is calculated for the one or more multi-source audio signals. Multiple audio tracks are updated based on the calculated TDOA detection data up to the current time step; Based on the calculated TDOA detection data, construct speech activity detection (VAD) data for each of the plurality of audio tracks; as well as One or more enhanced multi-source audio signals are generated using the updated multiple audio tracks and the constructed VAD data.

14. The audio processing apparatus of claim 13, wherein when the microphone array is physically located on a second plane in reality, the grid of the TDOA location points is located on a first plane in the multidimensional space, the multidimensional space having a number of dimensions equal to the total number of microphone pairs.

15. The audio processing apparatus of claim 12, wherein the one or more processors are further configured to: For each microphone pair: The covariance matrix is ​​calculated for each frequency band using the time-frequency representation of the one or more multi-source audio signals from the corresponding microphone pair; A steering matrix is ​​constructed for each frequency band based on the TDOA position, wherein the TDOA position is for different scans corresponding to the azimuth and elevation angles of the corresponding microphone pair; Based on the constructed steering matrix and the calculated covariance matrix, a directional covariance matrix aligned across all frequency bands is constructed. as well as Based on the constructed directional covariance matrix, a pseudo-likelihood solution that minimizes beam power is determined.

16. The audio processing apparatus of claim 15, wherein the one or more processors are further configured to: The STMV joint pseudo-likelihood is calculated by taking the product of all determined pseudo-likelihood solutions across all microphone pairs. Determine an azimuth and elevation angle that maximizes the joint pseudo-likelihood of the STMV; and The determined azimuth and elevation angles are converted into polar coordinates representing the TDOA detection data.

17. The audio processing apparatus of claim 11, wherein the plurality of audio tracks are updated by the following steps: Identify the TDOA detection corresponding to the current time step and a set of existing audio tracks previously established up to the current time step; and The determination is based on a comparison between the TDOA detection and the gates of the existing audio tracks to determine whether to incorporate the TDOA detection into one of the existing audio tracks or to create a new audio track.

18. A non-transitory processor-readable medium storing a plurality of processor-executable instructions for enhancing multi-source audio through multi-source tracking and speech activity detection, the processor-executable instructions being executed by one or more processors to perform operations including: Receive one or more multi-source audio signals from a microphone array via an audio input circuit; Construct a multidimensional space defined by multiple microphone pairs from the microphone array; For each microphone pair, based on the distance and angle between the two microphones in the corresponding microphone pair, calculate the location point corresponding to the azimuth angle and elevation angle of the incident sound ray; A grid of TDOA location points is formed by iterating through all microphone pairs and all scans of azimuth and elevation angles. as well as Multiple audio tracks are updated based on the grid of the TDOA location points.

19. The non-transitory processor-readable medium of claim 18, wherein the multidimensional space is formed by the following steps: Select a first microphone from the microphone array as a reference microphone; and Each remaining microphone from the microphone array is paired with the reference microphone.

20. The non-transitory processor-readable medium of claim 18, wherein the operation further comprises: Based on the directional minimum variance (STMV) beamformer and the TDOA trajectory information constructed in the multidimensional space, TDOA detection data is calculated for the one or more multi-source audio signals. Multiple audio tracks are updated based on the calculated TDOA detection data up to the current time step; Based on the calculated TDOA detection data, construct speech activity detection (VAD) data for each of the plurality of audio tracks; as well as One or more enhanced multi-source audio signals are generated using the updated multiple audio tracks and the constructed VAD data.