Adaptive Spatial VAD and Time-Frequency Mask Estimation for Highly Unstable Noise Sources

By using a subband analysis module, a constrained minimum variance adaptive filter, a mask estimator and a spatial VAD in the speech activity detection system, the problem of difficult to distinguish between target audio signals and interfering noise signals in the noise environment is solved, and higher detection accuracy and performance are achieved.

CN111415686BActive Publication Date: 2025-05-30SYNAPTICS INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010013763.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-01-07
Filing Date
2020-01-07
Publication Date
2025-05-30
Estimated Expiration
2040-01-07

AI Technical Summary

Technical Problem

Existing Voice Activity Detection (VAD) technology is difficult to effectively distinguish between target audio signals and interfering noise signals in noisy environments, resulting in false alarms and reduced performance.

Method used

The subband analysis module, a constrained minimum variance adaptive filter, a mask estimator and a spatial VAD are used to perform multi-channel audio signal processing through these modules and methods to distinguish the target audio signal from the interference noise signal.

Benefits of technology

It improves the accuracy and performance of speech activity detection in noisy environments, reduces false alarm rates, and achieves a significant improvement in keyword recognition performance in difficult noise scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111415686B_ABST
    Figure CN111415686B_ABST
Patent Text Reader

Abstract

The system and method include: a first voice activity detector operable to detect speech in frames of a multi-channel audio input signal and output a speech determination; a constrained minimum variance adaptive filter operable to receive the multi-channel audio input signal and the speech determination and minimize the signal variance at the output of the filter, thereby producing an equalized target speech signal; a mask estimator operable to receive the equalized target speech signal and the speech determination and generate a spectral-temporal mask to distinguish the target speech from noise and interfering speech; and a second active speech detector operable to detect speech in frames of the speech discrimination signal. The audio input sensor array includes a plurality of microphones, each microphone generating a channel of the multi-channel audio input signal. A subband analysis module is operable to decompose each of the channels into a plurality of frequency subbands.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] According to one or more embodiments, the present application generally relates to systems and methods for audio signal detection and processing, and more particularly to, for example, voice activity detection systems and methods. Background Art

[0002] Voice activity detection (VAD) is used in various voice communication systems such as voice recognition systems, noise reduction systems, and sound source localization systems. In many applications, an audio signal is received through one or more microphones that sense acoustic activity in a noisy environment. The sensed audio signal may include speech to be detected and various noise signals (including non-target speech) that degrade speech intelligibility and / or reduce VAD performance. Conventional VAD techniques may also require relatively large processing or memory resources that are not practical for real-time voice activity detection in low-power, low-cost devices such as mobile phones, smart speakers, and laptop computers. In view of the foregoing, there is still a need in the art for improved VAD systems and methods. Summary of the Invention

[0003] Improved systems and methods for detecting a target audio signal (such as speech of a target person) in a noisy audio signal are disclosed herein. In one or more embodiments, the system includes a subband analysis module, an input voice activity detector, a constrained minimum variance adaptive filter, a mask estimator, and a spatial VAD.

[0004] The scope of the present disclosure is defined by the claims incorporated by reference into this section. A more complete understanding of the embodiments of the present invention and the recognition of additional advantages thereof will be provided to those skilled in the art by considering the following detailed description of one or more embodiments. Reference will be made to the accompanying drawings that will be briefly described first. Brief Description of the Drawings

[0005] Aspects of the present disclosure and its advantages can be better understood with reference to the following drawings and the subsequent detailed description. It should be appreciated that like reference numerals are used to identify like elements illustrated in one or more of the figures, in which the display is for the purpose of illustrating embodiments of the present disclosure and not for the purpose of limiting embodiments of the present disclosure. The components in the drawings are not necessarily drawn to scale, but emphasis is placed on clearly illustrating the principles of the present disclosure.

[0006] Figure 1 Illustrate an example system architecture of an adaptive spatial voice activity detection system according to one or more embodiments of the present disclosure.

[0007] Figure 2 Describe an example audio signal generated by components of an adaptive spatial voice activity detection system according to one or more embodiments of the present disclosure.

[0008] Figure 3 Describe an example target voice processing including direction of arrival according to one or more embodiments of the present disclosure.

[0009] Figure 4 Describe an example system including an implementation of adaptive spatial voice detection according to one or more embodiments of the present disclosure.

[0010] Figure 5 Describe an example audio signal processing system implementing adaptive spatial voice detection according to one or more embodiments of the present disclosure.

[0011] Figure 6 Describe an example voice activity detection method according to one or more embodiments of the present disclosure. DETAILED DESCRIPTION

[0012] Improved systems and methods for detecting voice activity in a noisy environment are disclosed herein.

[0013] Despite recent progress, speech recognition performed under noisy conditions remains a challenging task. In a multi-microphone setup, several multi-channel speech enhancement algorithms have been proposed, including algorithms that include the following: adaptive and non-adaptive beamforming, blind source separation based on independent component analysis or independent vector analysis, and multi-channel non-negative matrix factorization. A promising approach in the field of automatic speech recognition is the maximum signal-to-noise ratio (SNR) beamformer (which is also referred to as the generalized eigenvalue (GEV) beamformer) that aims to optimize a multi-channel filter to maximize the output SNR. One component for implementing an online maximum SNR beamformer algorithm is an estimator of the noise and input covariance matrices. The estimation is generally supervised by voice activity detection or by a deep neural network (DNN) that predicts a spectral-temporal mask that is related to voice activity. The VAD (or DNN-mask) has the goal of identifying parts of the signal where there is a high confidence of separately observed noise to update the noise covariance matrix. It is also required to identify parts of the signal where the noise overlaps with the target voice so that the input noise covariance matrix can be updated.

[0014] One drawback of existing systems is that the VAD and DNN-mask estimators are designed to distinguish speech from “non-speech” noise. However, in many real-world scenarios, noise sources (e.g., a television or radio) may also emit audio that contains speech components, which will produce false positives and ultimately degrade the overall performance of noise reduction. In the present disclosure, improved systems and methods are disclosed for generating multi-channel VAD predictions and spectro-temporal masks to distinguish between target speech and interfering speech emitted by noise sources. For example, interfering noise may be generated by a TV playing a movie, a show, or other media with audio content. The noise in this scenario will typically include a mixture of speech and non-speech sounds (such as music or other audio effects).

[0015] In various embodiments, a method for voice activity detection includes estimating a constrained adaptive filter that minimizes output variance without explicitly defining a target speech direction. The filter is trained when there is high confidence that the audio does not belong to the “speech” class. This supervision can be obtained through a deep neural network-based voice activity detector that is trained to distinguish speech from non-speech audio. Multi-channel filter estimation can be equivalent to the estimation of the relative transfer function (RTF) of the noise source. Since the filter output is minimized for audio emitted by the same noise source, the filter output will also be minimized when there is speech in the noise. Thus, it is possible to distinguish between the target and interfering speech. In some embodiments, the method includes running a power-based VAD at the output of the adaptive filter. The output of the filter can also be used to estimate a sub-band mask that identifies time-frequency points and can further be used to supervise noise reduction methods.

[0016] The methods disclosed herein have been successfully applied to supervise two-channel speech enhancement (SSP) in challenging noise scenarios such as a speaker uttering a trigger word at -10 dB SNR in a noisy TV where the TV noise is playing a movie that includes some speech. An improvement in keyword recognition performance has been measured, with the average hit rate score moving from approximately 30% (without using spatial VAD) to over 80% (using spatial VAD). Additionally, the methods disclosed herein have been successfully used to supervise direction-of-arrival (DOA) estimation, allowing for the position tracking of a target speaker under -10 dB SNR conditions with highly unstable noise.

[0017] The technical differences and advantages compared to other solutions will now be described. Existing single-channel based methods rely on the nature of the sound itself in the audio signal to produce a prediction as to whether the input frame contains speech or only non-speech noise. These methods cannot distinguish between target speech and interfering speech because both belong to the same sound category. Any detected speech (whether from the target user providing a voice command or from interfering speech) can be classified as speech in these systems.

[0018] Existing multi-channel based methods typically rely on strong geometric assumptions about the position of the target speaker. For example, the target speaker can be assumed to be (i) closer to one of the microphones, (ii) located in a predefined spatial region, and / or (iii) producing more coherent speech. These assumptions are not practical in many applications such as smart speaker applications that account for environments with coherent noise (e.g., speech from a television or radio) or 360-degree far-field voice control.

[0019] In contrast to existing voice activity detectors, the systems and methods disclosed herein utilize both the nature of the sound and the unique spatial signature of the sound in 3D space to produce a high speech / noise discrimination. Additionally, the systems and methods of the present disclosure do not require prior assumptions about the geometry or speaker position, thus providing greater flexibility for far-field applications compared to existing systems. In various embodiments, a supervised adaptive spatial voice activity detector is used and specifically adapted to remove false positives caused by speech sounds emitted from noise sources.

[0020] Reference Figure 1 , an example system 100 will now be described in accordance with various embodiments. The system 100 receives a multi-channel audio input signal 110, which is processed by a sub-band analysis module 120. In some embodiments, the multi-channel audio input signal 110 is generated by an audio input component that includes a plurality of audio sensors (e.g., a microphone array) and audio input processing circuitry. The multi-channel audio input signal 110 includes a plurality of audio channels M that are divided into a series of frames l. The sub-band analysis module 120 (e.g., using a Fourier transform process) divides the spectrum of the audio channels into a plurality of frequency sub-bands X i (k, l). The system 100 further includes an input voice activity detector (VAD) 130, a constrained minimum variance adaptive filter 140, a time-frequency (TF)-mask estimator 152, and a spatial voice activity detector VAD 154.

[0021] The input VAD 130 receives the output X of the sub-band analysis module 120 i(k, l) and identify the instances (e.g., audio frames) in which non-speech (such as noise) is detected separately (e.g., there is no speech). In some embodiments, the input VAD 130 is tuned to produce more false alarms than false rejections of speech activity. In other words, the goal of the input VAD 130 is to identify the frames in which a determination of the absence of speech is made with a high degree of confidence. In various embodiments, the input VAD 130 may include power-based speech detection techniques, which may include classifiers based on machine learning data, such as deep neural networks, support vector machines, and / or Gaussian mixture models that are trained to distinguish between speech and non-speech audio. In one embodiment, the input VAD 130 may implement an embodiment of the method set forth in co-pending application Ser. No. 15 / 832,709, entitled "VOICE ACTIVITY DETECTION SYSTEMS AND METHODS", which is incorporated herein by reference in its entirety.

[0022] The input VAD 130 outputs a variable v(l), and the output variable v(l) defines the state of the input VAD 130 for the observed frame l. In one embodiment, a value equal to "1" indicates that the observed frame is determined to include speech, and a value equal to "0" indicates the absence of speech in the observed frame. In other embodiments, the input VAD 130 may include other conventional VAD systems and methods operable to produce time-based speech activity determinations, which include VADs that analyze and produce speech activity determinations based on one or more channels, subbands, and / or frames of a multi-channel signal.

[0023] The constrained minimum variance adaptive filter 140 receives the multi-channel subband signal X i (k, l) and the speech determination v(l), and is operable to estimate an adaptive filter to minimize the signal variance at its output. For simplicity and effectiveness, a frequency domain implementation is disclosed herein, but the present disclosure is not limited to this approach. In the illustrated embodiment, for each channel i, the time domain signal x i (t) of this embodiment is transformed into an undersampled time-frequency domain representation by the subband-analysis module 120. This can be obtained by applying subband analysis or a short-time Fourier transform:

[0024]

[0025] X(k, l) = [X 1 (k, l),..., X M (k, l)] T

[0026] Among them, M indicates the number of input channels (M > 1). For sub-band k, the output of the filter can be defined as:

[0027] Y(k, l) == G(k) H X(k, l)

[0028] Among them, under the constraint |G(k) H e 1 | = 1 (where e 1 = [1...0] T (in some embodiments, this constraint is used to prevent from becoming an all-zero vector)), when only the noise source is active (for example, when v(l) indicates that no voice is detected), G(k) is optimized to minimize the output variance E[|Y(k)| 2 :

[0029] Subject to |G(k) H e 1 | = 1.

[0030] The closed-form solution of this optimization is:

[0031]

[0032] Among them, R n (k) is the covariance of the noise, which is calculated as follows:

[0033]

[0034] In an online implementation, the covariance matrix is updated for frame l and can be estimated using first-order recursive smoothing as follows:

[0035] R n (k, l + 1) = α(l)R n (k, l) + (1 - α(l))X(k, l) H X(k, l)

[0036] where α(l) = max(α, v(l)), and α is the smoothing constant (<1).

[0037] In some embodiments, an alternative way to estimate the filter G(k) is to impose the following constrained filter structure:

[0038] G(k) = [1|H(k)] T

[0039] H(k) = -[G 2 (k),..., G M (k)]

[0040] And, optimize without imposing any constraints in the adaptive. The adaptive solution to this optimization problem can be obtained by using the normalized least mean square (NLMS), which can be formulated as:

[0041]

[0042]

[0043] where μ is the adaptive step size, Z(k, l) = [X 2 (k, l),..., X M (k, l)] T , and, an additional term β|Y(k, l)| 2 (β > 1) is added to stabilize the learning and avoid numerical divergence.

[0044] The output variance of the constrained minimum variance adaptive filter |Y(k, l)| 2 is minimized for frames that contain audio emitted by a noise source. The attenuation of the filter is independent of the nature of the sound, but only depends on the spatial covariance matrix, and thus, for the noise part that contains interfering speech, the output will also be small. On the other hand, audio emitted from different points in space will have different spatial covariance matrices and thus will not be attenuated to the same extent as the noise source. Following the NLMS formula, for the case of M = 2 and a coherent noise source, the estimated filter G i (k)(i > 2) can be considered as the relative transfer function between the first microphone and the i-th microphone.

[0045] In the disclosed embodiments, the noise with covariance R n (k) is attenuated at the output Y(k, l), but this signal is not directly used as an enhanced version of the target speech. In various embodiments, the "distortionless" constraint is not imposed as is typically done in a minimum variance distortionless response (MVDR) beamformer because the target speaker direction or its RTF is not known in advance. Thus, in the illustrated embodiments, Y(k, l) will contain an equalized version of the target speech, where the spectral distortion depends on the similarity between the spatial covariance of the target speech and the spatial covariance of the noise. The SNR improvement at the output Y(k, l) is large enough to allow the estimation of the TF activity mask related to the speech by the TF-mask estimator 152 without explicitly estimating the true target speech variance.

[0046] First, for each subband k, a reference feature signal is calculated from |X 1 (k, l)| and |Y(k, l)| as follows:

[0047] F(k, l) = f(|X 1 (k, l)|, |Y(k, l)|)

[0048] In various embodiments, a possible formula for F(k, l) can be:

[0049]

[0050] It is actually the output amplitude weighted by the magnitude transfer function of the filter. However, alternative formulas are also possible.

[0051] For each sub-band k, the activity of the target speech can be determined by tracking the power level of the signal F(k, l) and detecting the non-stationary signal part. A single-channel power-based VAD can then be applied to each signal F(k, l) to produce a mask:

[0052] If speech is detected, then V(k, l) = 1,

[0053] otherwise, V(k, l) = 0.

[0054] In this embodiment, an example sub-band VAD is shown, but since many alternative algorithms are available, the present disclosure should not be considered limited to this formula.

[0055] For each sub-band k, the noise floor can be estimated by dual-rate smoothing as follows:

[0056] N(k, l + 1) = γN(k, l) + (1 - γ)F(k, l)

[0057] where,

[0058] If |F(k, l)| > N(k, l + 1), then γ = γ up ,

[0059] If |F(k, l)| < N(k, l + 1), then γ = γ down ,

[0060] where the smoothing constant γ up >> γ down .

[0061] Then, the target speech mask can be calculated as follows:

[0062]

[0063] If then V(k, l) = 0,,

[0064] Among them, SNR_threshold is a tunable parameter. In the illustrated embodiment, it is assumed that the adaptive filter can reduce the noise output variance under the noise basis, thereby generating a stable noise residue. This is possible when the noise is coherent and the subband signal representation has a resolution high enough to accurately model acoustic reverberation. In another embodiment, this assumption is not strict, and a method based on tracking the distribution of relative power levels is adopted, such as the method described in Ying, Dongwen et al.'s "Voice activity detection based on an unsupervised learning framework." (IEEE Transactions on Audio, Speech, and Language Processing 19.8 (2011): 2624-2633), which is incorporated herein by reference.

[0065] A frame-based spatial VAD can be calculated by integrating the feature signal F(k, l) (e.g., from the TF-mask estimator 152) into a single signal F(l):

[0066] where k ∈ K

[0067] where K is a subset of frequencies and the single-channel VAD criterion is applied to F(l) to obtain the binary-frame-based decision V(l). In some embodiments, V(k, l) can also be directly applied for each subband as:

[0068] V(l) = ∑ k V(k, l) > threshold.

[0069] In another embodiment, the full signal F(k, l) can be used to generate the prediction V(l), for example, by using strictly designed features (hard-engineered features) extracted from F(k, l) or using data-based maximum likelihood methods (e.g., deep neural networks, Gaussian mixture models, support vector machines, etc.).

[0070] Reference Figure 2, an example audio signal 200 generated by components of an adaptive spatial voice activity detection system in accordance with one or more embodiments of the present disclosure will now be described. In operation, a multi-channel audio signal is received via a plurality of input sensors. A first channel of the input audio signal 210 is illustrated and may include target speech and noise (both non-target speech and non-speech noise). An input voice activity detector (e.g., input VAD 130) detects frames in which there is a high likelihood of the absence of speech, and as illustrated (e.g., in signal 220), outputs a "0" for non-speech frames and a "1" for speech frames. Audio processing then proceeds to detect target speech activity from non-target speech activity and outputs an indication of "0" for the absence of target speech and an indication of "1" for the detected target speech, as illustrated in signal 230. In some embodiments, as previously discussed herein, the audio signal may include a noisy non-stationary noise source (e.g., a TV signal) that can be identified as non-target speech by the spatial VAD. The input audio signal 210 is then processed using the detection information from the spatial VAD (e.g., signal 230) to generate an enhanced target speech signal 240.

[0071] Figure 3 An example target speech processing including direction-of-arrival processing in accordance with one or more embodiments of the present disclosure is illustrated. Diagram 300 illustrates an example estimated direction of arrival for a speech source, which is estimated for each frame using a neural network-based voice activity detector. The direction of arrival of the speech is illustrated in diagram 310 and shows both the target speech (e.g., the person providing the voice command) and other speech generated by noise sources (e.g., speech detected from a television). The VAD output corresponds to the speech activity decision as illustrated in diagram 320, e.g., diagram 320 shows speech detected in all time frames, including the target speech and / or speech generated by the TV. The bottom diagram 350 illustrates an example of applying a spatial voice activity detector to the task of estimating the direction of arrival (DOA) of the target speech when there is noisy noise (e.g., TV noise). In this case, the target speech in diagram 360 is detected and the non-target speech (e.g., TV noise) is ignored, thus providing an improved speech activity detection as illustrated, for example, in diagram 370.

[0072] Figure 4 An audio processing apparatus 400 including spatial voice activity detection in accordance with various embodiments of the present disclosure is illustrated. The audio processing apparatus 400 includes an input of an audio sensor array 405, an audio signal processor 420, and a host system component 450.

[0073] The audio sensor array 405 includes one or more sensors, each of which can convert sound waves into an audio signal. In the illustrated environment, the audio sensor array 405 includes a plurality of microphones 405a - 405n, each of which generates an audio channel of a multi-channel audio signal.

[0074] The audio signal processor 420 includes an audio input circuitry 422, a digital signal processor 424, and an optional audio output circuitry 426. In various embodiments, the audio signal processor 420 may be implemented as an integrated circuit including analog circuitry, digital circuitry, and the digital signal processor 424, which is operable to execute program instructions stored in firmware. The audio input circuitry 422 may include, for example, an interface to the audio sensor array 405, an anti-aliasing filter, analog-to-digital converter circuitry, echo cancellation circuitry, and other audio processing circuitry and components as disclosed herein. The digital signal processor 424 is operable to process the multi-channel digital audio signal to generate an enhanced audio signal, which is output to one or more host system components 450. In various embodiments, the multi-channel audio signal includes a mixture of a noise signal and at least one desired target audio signal (e.g., human speech), and the digital signal processor 424 is operable to isolate or enhance the desired target signal while reducing the undesired noise signal. The digital signal processor 424 may be operable to perform echo cancellation, noise cancellation, target signal enhancement, post-filtering, and other audio signal processing functions. The digital signal processor 424 may further include an adaptive spatial target activity detector and mask estimation module 430, which is operable to implement Figures 1-3 and Figures 5-6 one or more embodiments of the systems and methods disclosed herein in

[0075] The digital signal processor 424 may include one or more of a processor, a microprocessor, a single-core processor, a multi-core processor, a microcontroller, a programmable logic device (PLD) (e.g., a field-programmable gate array (FPGA)), a digital signal processing (DSP) device, or other logic devices that can be configured to perform the various operations discussed herein for embodiments of the present disclosure by being hard-wired, executing software instructions, or a combination of both. The digital signal processor 424 is operable to interface and communicate with the host system components 450, such as via a bus or other electronic communication interface.

[0076] The optional audio output circuitry 426 processes the audio signals received from the digital signal processor 424 for output to at least one speaker, such as speakers 410a and 410b. In various embodiments, the audio output circuitry 426 may include a digital-to-analog converter that converts one or more digital audio signals into corresponding analog signals and one or more amplifiers for driving speakers 410a - 410b.

[0077] The audio processing device 400 may be implemented as any device operable to receive and detect target audio data, such as, for example, a mobile phone, a smart speaker, a tablet computer, a laptop computer, a desktop computer, a voice-controlled appliance, or an automobile. The host system components 450 may include various hardware and software components for operating the audio processing device 400. In the illustrated embodiment, the system components 450 include a processor 452, a user interface component 454, a communication interface 456 for communicating with external devices and networks (such as network 480 (e.g., the Internet, the cloud, a local area network, or a cellular network) and mobile device 484), and a memory 458.

[0078] The processor 452 may include one or more of a processor, a microprocessor, a single-core processor, a multi-core processor, a microcontroller, a programmable logic device (PLD) (e.g., a field-programmable gate array (FPGA)), a digital signal processing (DSP) device, or other logic devices that can be configured to perform the various operations discussed herein for embodiments of the present disclosure by being hard-wired, executing software instructions, or a combination of both. The host system components 450 are operable to interface and communicate with the audio signal processor 420 and other system components 450, such as via a bus or other electronic communication interface.

[0079] It will be appreciated that although the audio signal processor 420 and the host system components 450 are shown as a combination of integrated hardware components, circuitry, and software, in some embodiments, some or all of the functionality that the hardware components and circuitry are operable to perform may be implemented as software modules executed by the processor 452 and / or the digital signal processor 424 in response to software instructions and / or configuration data stored in the firmware of the memory 458 or the digital signal processor 424.

[0080] The memory 458 may be implemented as one or more memory devices operable to store data and information, including audio data and program instructions. The memory 458 may include one or more various types of memory devices, including volatile and non-volatile memory devices (such as RAM (random access memory), ROM (read-only memory), EEPROM (electrically erasable read-only memory)), flash memory, a hard disk drive, and / or other types of memory.

[0081] The processor 452 can be operable to execute software instructions stored in the memory 458. In various embodiments, the voice recognition engine 460 can be operable to process the enhanced audio signal received from the audio signal processor 420, including identifying and executing voice commands. The voice communication component 462 can be operable to facilitate voice communication with one or more external devices (such as the mobile device 484 or the user device 486), such as via a voice call over a mobile or cellular telephone network or a VoIP call over an IP (Internet Protocol) network. In various embodiments, the voice communication device includes transmitting the enhanced audio signal to an external communication device.

[0082] The user interface component 454 can include a display, a touchpad display, a keypad, one or more buttons, and / or other input / output components that are operable to enable a user to directly interact with the audio processing device 400.

[0083] The communication interface 456 facilitates communication between the audio processing device 400 and external devices. For example, the communication interface 456 can implement a Wi-Fi (e.g., 802.11) or Bluetooth connection between the audio processing device 400 and one or more local devices (such as the mobile device 484) or a wireless router that provides network access to a remote server 482 via the network 480. In various embodiments, the communication interface 456 can include other wired and wireless communication components that facilitate direct or indirect communication between the audio processing device 400 and one or more other devices.

[0084] Figure 5 Describe an audio signal processor 500 according to various embodiments of the present disclosure. In some embodiments, the audio signal processor 500 is embodied as one or more integrated circuits including analog and digital circuitry and firmware logic implemented by a digital signal processor (such as, Figure 4 the digital signal processor 424). As illustrated, the audio signal processor 500 includes an audio input circuitry 515, a subband frequency analyzer 520, an adaptive spatial target activity detector and mask estimation module 530, and a synthesizer 535.

[0085] The audio signal processor 500 receives multi-channel audio input from a plurality of audio sensors (such as a sensor array 505 including at least one audio sensor 505a-n). The audio sensors 505a - 505n can include microphones integrated with or external components connected to an audio processing device (such as, Figure 4 the audio processing device 400).

[0086] An audio signal may initially be processed by an audio input circuitry 515, which may include an anti-aliasing filter, an analog-to-digital converter, and / or other audio input circuitry. In various embodiments, the audio input circuitry 515 outputs a digital, multi-channel, time-domain audio signal having M channels, where M is the number of sensor (e.g., microphone) inputs. The multi-channel audio signal is input to a subband frequency analyzer 520, which divides the multi-channel audio signal into successive frames and decomposes each frame of each channel into a plurality of frequency subbands. In various embodiments, the subband frequency analyzer 520 includes a Fourier transform process. The decomposed audio signal is then provided to an adaptive spatial target activity detector and mask estimation module 530.

[0087] The adaptive spatial target activity detector and mask estimation module 530 is operable to analyze frames of one or more of the audio channels and generate a signal indicating whether a target audio is present in the current frame. As discussed herein, the target audio may be human speech (e.g., for voice command processing), and the adaptive spatial target activity detector and mask estimation module 530 may be operable to detect the target speech in a noisy environment (which includes non-target speech) and generate an enhanced target audio signal for further processing by, for example, a host system. In some embodiments, the enhanced target audio signal is reconstructed on a frame-by-frame basis by combining subbands of one or more channels to form an enhanced time-domain audio signal that is sent to a host system, another system component, or an external device for further processing, such as voice command processing.

[0088] Reference Figure 6, embodiments of method 600 for detecting target voice activity using the systems disclosed herein will now be described. In step 610, the system receives a multi-channel audio signal and decomposes the multi-channel audio signal into a plurality of sub-bands. The multi-channel input signal may be generated, for example, by a corresponding plurality of audio sensors (e.g., a microphone array), and the corresponding plurality of audio sensors generate sensor signals that are processed by an audio input circuitry. In some embodiments, each channel is decomposed into a plurality of frequency sub-bands. In step 620, the multi-channel audio signal is analyzed frame by frame to detect voice activity and generate a voice determination for each frame indicating the detection of speech or the absence of speech. In step 630, the multi-channel audio signal and the corresponding voice determination are used as inputs to estimate a constrained minimum variance adaptive filter. In various embodiments, in step 640, the minimum variance adaptive filter estimates an adaptive filter to minimize the signal variance at its output and generate an equalized target voice signal. In step 650, a feature signal and a noise floor are calculated based on the equalized target voice signal and the channels of the multi-channel audio signal. In step 660, a target voice mask is calculated using the feature signal and the noise floor.

[0089] In applicable cases, the various embodiments provided by the present disclosure may be implemented using hardware, software, or a combination of hardware and software. Additionally, in applicable cases, without departing from the spirit of the present disclosure, the various hardware components and / or software components set forth herein may be combined into composite components including software, hardware, and / or both. In applicable cases, without departing from the scope of the present disclosure, the various hardware components and / or software components set forth herein may be separated into sub-components including software, hardware, or both. Additionally, in applicable cases, it is contemplated that software components may be implemented as hardware components, and vice versa.

[0090] Software according to the present disclosure (such as program code and / or data) may be stored on one or more computer-readable media. It is also contemplated that the software identified herein may be implemented, networked, and / or otherwise using one or more general-purpose or special-purpose computers and / or computer systems. In applicable cases, the ordering of the various steps described herein may be changed, combined into composite steps, and / or separated into sub-steps to provide the features described herein.

[0091] The foregoing disclosure is not intended to limit the present disclosure to the precise forms or particular fields of use disclosed. Similarly, various alternative embodiments and / or modifications of the present disclosure are contemplated. Having thus described embodiments of the present disclosure, those of ordinary skill in the art will recognize that changes may be made in form and detail without departing from the scope of the present disclosure. Accordingly, the present disclosure is limited only by the claims.

Claims

1. A voice activity detection system, comprising: A first voice activity detector operable to detect speech in frames of a multi-channel audio input signal and output a speech determination indicating whether speech is detected or no speech is present in the frame; A constrained minimum variance adaptive filter operable to receive the multi-channel audio input signal and the speech determination and minimize the signal variance at the output of the filter, thereby producing an equalized target speech signal; A mask estimator operable to receive the equalized target speech signal and the speech determination and generate a spectral-temporal mask to distinguish the target speech from noise and interfering speech; and A second voice activity detector operable to detect speech in the frame of the multi-channel audio input signal at least in part based on the equalized target speech signal.

2. The system according to claim 1, further comprising an audio input sensor array including a plurality of microphones, each microphone generating a channel of the multi-channel audio input signal.

3. The system according to claim 2, further comprising a sub-band analysis module operable to decompose each of the channels into a plurality of frequency sub-bands.

4. The system according to claim 1, wherein, The first voice activity detector includes a neural network trained to identify speech in the frame of the multi-channel audio input signal.

5. The system according to claim 1, wherein, The constrained minimum variance adaptive filter is operable to minimize the output variance when the speech determination indicates no speech is present in the frame.

6. The system according to claim 1, wherein, The constrained minimum variance adaptive filter includes a normalized least mean square process.

7. The system according to claim 1, wherein, The mask estimator is further operable to generate a reference feature signal for each sub-band and frame of a selected channel of the multi-channel audio input signal.

8. The system according to claim 1, wherein, The second voice activity detector includes a single-channel power-based voice activity detector applied to each signal to produce a target speech mask.

9. The system according to claim 1, wherein, The system includes a speaker, a tablet computer, a mobile phone, and / or a laptop computer.

10. A voice activity detection method, comprising: Receiving a multi-channel audio input signal; Using a first voice activity detector to detect voice activity in frames of the multi-channel audio input signal and generate a speech determination indicating whether speech is detected or no speech is present in the frame; Applying a constrained minimum variance adaptive filter to the multi-channel audio input signal and the speech determination and minimizing the signal variance at the output of the filter, thereby producing an equalized target speech signal; Using the filtered multi-channel audio input signal and the speech determination to estimate a spectral mask to distinguish the target speech from noise and interfering speech; and Use a second voice activity detector to detect voice activity in the frame of the multi-channel audio input signal based at least in part on the equalized target speech signal.

11. The method according to claim 10, wherein, receiving the multi-channel audio input signal includes using a plurality of microphones to generate the multi-channel audio input signal, each microphone generating a corresponding channel of the multi-channel audio input signal.

12. The method according to claim 11, further comprising using a subband analysis module to decompose each of the channels into a plurality of frequency subbands.

13. The method according to claim 10, wherein, using the first voice activity detector to detect voice activity includes processing the frame of the multi-channel audio input signal through a neural network, the neural network being trained to identify the speech in the frame.

14. The method according to claim 10, wherein, applying the constrained minimum variance adaptive filter further includes minimizing the output variance when the voice determination indicates that there is no voice in the frame.

15. The method according to claim 10, wherein, applying the constrained minimum variance adaptive filter includes performing a normalized least mean squares process.

16. The method according to claim 10, further comprising generating a reference feature signal for each subband and frame of the selected channel of the multi-channel audio signal.

17. The method according to claim 10, wherein, the second voice activity detector includes a single-channel power-based voice activity detector, the single-channel power-based voice activity detector being applied to each signal to produce a target speech mask.

18. The method according to claim 10, wherein, the method is implemented by a speaker, a tablet computer, a mobile phone, and / or a laptop computer.

Citation Information

Patent Citations

  • Voice activity detection systems and methods

    US10504539B2

  • Systems, methods, and apparatus for voice activity detection

    CN103180900A

  • Systems, methods, apparatus, and computer-readable media for phase-based processing of multichannel signal

    US20100323652A1