User Voice Activity Detection Using Dynamic Classifiers

The dynamic classifier in the system addresses the inefficiency of conventional SVAD by adaptively clustering audio data from multiple microphones, reducing power consumption and false alarms while maintaining high accuracy in distinguishing user voice.

JP7757398B2Active Publication Date: 2025-10-21QUALCOMM INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023520368
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-05-05
Filing Date
2021-09-17
Publication Date
2025-10-21
Estimated Expiration
2041-09-17

AI Technical Summary

Technical Problem

Conventional self-voice activity detection techniques improve accuracy at the cost of increased power consumption and processing resources, leading to inefficiencies in portable devices.

Method used

A system utilizing a dynamic classifier that processes audio data from multiple microphones to distinguish user voice from other sounds, employing adaptive clustering and decision boundary adjustment to minimize power consumption while maintaining high accuracy.

Benefits of technology

The dynamic classifier effectively distinguishes user voice from other audio activity with low power consumption and high accuracy, reducing false alarms and improving device efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007757398000006
    Figure 0007757398000006
  • Figure 0007757398000007
    Figure 0007757398000007
  • Figure 0007757398000008
    Figure 0007757398000008
Patent Text Reader

Abstract

The device includes a memory configured to store instructions and one or more processors configured to execute the instructions. The one or more processors are configured to execute the instructions to receive audio data including first audio data corresponding to a first output of a first microphone and second audio data corresponding to a second output of a second microphone. The one or more processors are also configured to execute the instructions to provide the audio data to a dynamic classifier. The dynamic classifier is configured to generate a classification output corresponding to the audio data. The one or more processors are further configured to execute the instructions to determine whether the audio data corresponds to user voice activity based at least in part on the classification output.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority to commonly owned U.S. Provisional Patent Application No. 63 / 089,507, filed October 8, 2020, and U.S. Non-Provisional Patent Application No. 17 / 308,593, filed May 5, 2021, the entire contents of which are expressly incorporated herein by reference.

[0002] FIELD OF THE DISCLOSURE

[0002] This disclosure relates generally to self-voice activity detection. [Background technology]

[0003]

[0003] Advances in technology have resulted in smaller and more powerful computing devices. For example, there are now a variety of portable personal computing devices, including wireless telephones such as mobile phones and smartphones, tablets, and laptop computers, that are small, lightweight, and easily carried by users. These devices can communicate voice and data packets over wireless networks. Furthermore, many such devices incorporate additional functionality, such as digital still cameras, digital video cameras, digital recorders, and audio file players. Such devices can also process executable instructions, including software applications, such as web browser applications that can be used to access the Internet. Thus, these devices can contain significant computing power.

[0004] Such computing devices often incorporate functionality for receiving audio signals from one or more microphones. For example, the audio signals may represent user speech captured by the microphone, external sound captured by the microphone, or a combination thereof. By way of example, a headset device may include self-voice activity detection that attempts to distinguish between a user's speech (e.g., speech spoken by a person wearing the headset) and speech originating from other sources. For example, when a system including a headset device supports keyword activation, self-voice activity detection can reduce "false alarms" in which activation of one or more components or operations is initiated based on speech originating from nearby people (referred to as "non-user speech"). Reducing such false alarms improves the power consumption efficiency of the device. However, performing audio signal processing to distinguish between user and non-user speech also consumes power, and conventional techniques for improving a device's accuracy in distinguishing between user and non-user speech also tend to increase the device's power consumption and processing resource requirements. Summary of the Invention

[0005] According to one implementation of the present disclosure, a device includes a memory configured to store instructions and one or more processors configured to execute the instructions. The one or more processors are configured to execute the instructions to receive audio data including first audio data corresponding to a first output of a first microphone and second audio data corresponding to a second output of a second microphone. The one or more processors are also configured to execute the instructions to provide the audio data to a dynamic classifier. The dynamic classifier is configured to generate a classification output corresponding to the audio data. The one or more processors are further configured to execute the instructions to determine whether the audio data corresponds to user voice activity based at least in part on the classification output.

[0006] According to another implementation of the present disclosure, a method includes receiving, at one or more processors, audio data including first audio data corresponding to a first output of a first microphone and second audio data corresponding to a second output of a second microphone. The method further includes, at the one or more processors, providing the audio data to a dynamic classifier to generate a classification output corresponding to the audio data. The method also includes, at the one or more processors, determining whether the audio data corresponds to user voice activity based at least in part on the classification output.

[0007] According to another implementation of the present disclosure, a non-transitory computer-readable medium includes instructions that, when executed by one or more processors, cause the one or more processors to receive audio data including first audio data corresponding to a first output of a first microphone and second audio data corresponding to a second output of a second microphone. The instructions, when executed by the one or more processors, further cause the one or more processors to provide the audio data to a dynamic classifier to generate a classification output corresponding to the audio data. The instructions, when executed by the one or more processors, also cause the one or more processors to determine, based at least in part on the classification output, whether the audio data corresponds to user voice activity.

[0008] According to another implementation of the present disclosure, an apparatus includes means for receiving audio data including first audio data corresponding to a first output of a first microphone and second audio data corresponding to a second output of a second microphone. The apparatus further includes means for generating, in a dynamic classifier, a classification output corresponding to the audio data. The apparatus also includes means for determining, based at least in part on the classification output, whether the audio data corresponds to a user voice activity.

[0009]

[0009] Other aspects, advantages, and features of the present disclosure will become apparent upon review of the entire application, including the following sections. [Brief explanation of the drawings]

[0010] [Figure 1]

[0010] FIG. 1 is a block diagram of a particular illustrative aspect of a system operable to perform self-voice activity detection, in accordance with certain examples of the present disclosure. [Figure 2]

[0011] 1A-1C are diagrams of example aspects of operations associated with self-voice activity detection, in accordance with some examples of the present disclosure. [Figure 3]

[0012] 1 is a block diagram of an example aspect of a system operable to perform self-voice activity detection, in accordance with some examples of the present disclosure. [Figure 4]

[0013] 2 is a diagram of an exemplary aspect of the operation of components of the system of FIG. 1 in accordance with some examples of the present disclosure. [Figure 5]

[0014] FIG. 1 illustrates an example of an integrated circuit including a dynamic classifier for detecting user voice activity, in accordance with some examples of the present disclosure. [Figure 6]

[0015] 1 is a diagram of a mobile device including a dynamic classifier for detecting user voice activity, according to some examples of the present disclosure. [Figure 7]

[0016] 1 is a diagram of a headset including a dynamic classifier for detecting user voice activity, in accordance with some examples of the present disclosure. [Figure 8]

[0017] 1 is a diagram of a wearable electronic device including a dynamic classifier for detecting user voice activity, according to some examples of the present disclosure. [Figure 9]

[0018] 1 is a diagram of a voice-controlled speaker system including a dynamic classifier for detecting user voice activity, according to some examples of the present disclosure. [Figure 10]

[0019] 1 is a diagram of a camera including a dynamic classifier for detecting user voice activity, according to some examples of the present disclosure. [Figure 11]

[0020] 1 is a diagram of a headset, such as a virtual reality or augmented reality headset, including a dynamic classifier for detecting user voice activity, in accordance with some examples of the present disclosure. [Figure 12]

[0021] 1 is a diagram of a first example of a vehicle including a dynamic classifier for detecting user voice activity, in accordance with some examples of the present disclosure. [Figure 13]

[0022] FIG. 10 is a diagram of a second example vehicle including a dynamic classifier for detecting user voice activity, in accordance with some examples of the present disclosure. [Figure 14A]

[0023] 2 is a diagram of a specific implementation of a method of self-voice activity detection that may be performed by the device of FIG. 1 in accordance with some examples of the present disclosure. [Figure 14B]

[0024] 10 is a diagram of another particular implementation of a method of self-voice activity detection that may be performed by the device of FIG. 1, in accordance with some examples of the present disclosure. [Figure 15]

[0025] 1 is a block diagram of a particular illustrative example of a device operable to perform self-voice activity detection, in accordance with some examples of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0011]

[0026] Self-voice activity detection (“SVAD”), which reduces “false alarms” in which the activation of one or more components or operations results from non-user speech, can improve device power consumption efficiency by preventing the activation of such components or operations when a false alarm is detected. However, conventional audio signal processing techniques for improving SVAD accuracy also increase device power consumption and processing resources while performing the accuracy improvement techniques. Because SVAD processing typically operates continuously, even while the device is in low-power or sleep mode, the reduction in power consumption resulting from reducing false alarms using conventional SVAD techniques may be partially or completely offset by the increased power consumption associated with the SVAD processing itself.

[0012]

[0027] A system and method for self-voice activity detection using a dynamic classifier is disclosed. For example, in a headset implementation, audio signals may be received from a first microphone positioned to capture a user's voice and from a second microphone positioned to capture external sounds, such as to perform noise reduction and echo cancellation. The audio signals may be processed to extract a frequency domain feature set including interaural phase difference ("IPD") and interaural intensity difference ("IID").

[0013]

[0028] The dynamic classifier processes the extracted frequency-domain feature sets and generates an output indicating a classification of the feature sets. The dynamic classifier may perform adaptive clustering of the feature data and adjustment of a decision boundary between the two most discriminative categories in the feature data space to distinguish between feature sets corresponding to the user's voice activity and feature sets corresponding to other audio activity. In an illustrative example, the dynamic classifier is implemented using a self-organizing map.

[0014]

[0029] The dynamic classifier uses the extracted feature set to enable discrimination to actively respond and adapt to various conditions, such as environmental conditions in highly non-stationary situations, mismatched microphones, changes in user headset fitting, different user head-related transfer functions (“HRTFs”), direction-of-arrival (“DOA”) tracking of non-user signals, microphone noise floor, bias, and sensitivity across the frequency spectrum, or a combination thereof. In some implementations, the dynamic classifier responds to such variations and enables adaptive feature mapping that can reduce or minimize the number of thresholding parameters used and the amount of headset tuning by the customer. In some implementations, the dynamic classifier enables effective discrimination between user voice activity and other audio activity with high accuracy under varying conditions and relatively low power consumption compared to traditional SVAD systems that provide comparable accuracy.

[0015]

[0030] Certain aspects of the present disclosure are described below with reference to the drawings. In the description, common features are designated by common reference numerals. Various terms used herein are used to describe particular implementations only and are not intended to limit the implementations. For example, the singular forms "a," "an," and "the" are intended to include the plural unless the context clearly indicates otherwise. Furthermore, some features described herein are singular in some implementations and plural in other implementations. For illustrative purposes, FIG. 1 shows a device 102 including one or more processors ("processor" 190 in FIG. 1), indicating that in some implementations, the device 102 includes a single processor 190 and in other implementations, the device 102 includes multiple processors 190. For ease of reference herein, such features are generally introduced as "one or more" features and are subsequently referred to in the singular unless aspects relating to the multiple features are described.

[0016]

[0031] It will be further understood that the terms “comprise,” “comprises,” and “comprising” can be used interchangeably with “include,” “includes,” or “including.” Additionally, it will be understood that the term “wherein” can be used interchangeably with “where.” As used herein, “exemplary” may indicate an example, implementation, and / or aspect and should not be construed as limiting or as indicating a preference or preferred implementation. As used herein, ordinal terms (e.g., “first,” “second,” “third,” etc.) used to modify an element of a structure, component, operation, etc. do not by themselves indicate any priority or order of that element relative to another element, but rather merely distinguish that element from other elements having the same name (except for the use of the ordinal term). As used herein, the term “set” refers to one or more of a particular element, and the term “plurality” refers to a plurality (e.g., two or more) of a particular element.

[0017]

[0032] As used herein, "coupled" can include "communicatively coupled," "electrically coupled," or "physically coupled," and also (or alternatively) can include any combination thereof. Two devices (or components) may be directly or indirectly coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., wired networks, wireless networks, or combinations thereof), etc. Two electrically coupled devices (or components) may be included in the same device or different devices and, as illustrative, non-limiting examples, may be connected via electronic circuits, one or more connectors, or inductive coupling. In some implementations, two devices (or components) that are communicatively coupled, such as via electrical communication, may send and receive signals (e.g., digital or analog signals) directly or indirectly via one or more wires, buses, networks, etc. As used herein, "directly coupled" may include two devices that are coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) without any intervening components.

[0018]

[0033] In this disclosure, terms such as “determining,” “calculating,” “estimating,” “shifting,” and “adjusting” may be used to describe how one or more operations are performed. It should be noted that such terms should not be construed as limiting, and other techniques may be utilized to perform similar operations. Additionally, the terms “generating,” “calculating,” “estimating,” “using,” “selecting,” “accessing,” and “determining” referred to herein may be used interchangeably. For example, “generating,” “calculating,” “estimating,” or “determining” a parameter (or signal) may refer to actively generating, estimating, calculating, or determining the parameter (or signal), or may refer to using, selecting, or accessing a parameter (or signal) that has already been generated, such as by another component or device.

[0019]

[0034] 1 , a particular exemplary aspect of a system configured to perform self-voice activity detection using a dynamic classifier is disclosed, generally designated 100. System 100 includes a device 102 coupled to a first microphone 110, a second microphone 120, and a second device 160. Device 102 is configured to perform self-voice activity detection of sounds captured by microphones 110, 120 using a dynamic classifier 140. To illustrate, in an implementation in which device 102 corresponds to a headset, first microphone 110 (e.g., a “primary” microphone) may be configured to primarily capture the utterance of a user of device 102, such as a microphone positioned proximate the mouth of a wearer of device 102, and second microphone 120 (e.g., a “secondary” microphone) may be configured to primarily capture ambient sound, such as positioned proximate the ear of the wearer. In other implementations, such as when device 102 supports a standalone voice assistant that may be in the vicinity of multiple people (e.g., includes a loudspeaker with a microphone, as further described with reference to FIG. 11 ), device 102 may be configured to detect speech from the person closest to the primary microphone as self-voice activity, even if the person is relatively far from the primary microphone compared to a headset implementation. As used herein, the term “self-voice activity detection” is used interchangeably with “user voice activity detection” to refer to distinguishing the speech (e.g., voice or speech) of a user of device 102 (e.g., “user voice activity”) compared to sounds not originating from the user of the device (e.g., “other audio activity”).

[0020]

[0035] The device 102 includes a first input interface 114, a second input interface 124, one or more processors 190, and a modem 170. The first input interface 114 is coupled to the processor 190 and is configured to be coupled to the first microphone 110. The first input interface 114 is configured to receive a first microphone output 112 from the first microphone 110 and provide the first microphone output 112 to the processor 190 as first audio data 116.

[0021]

[0036] The second input interface 124 is coupled to the processor 190 and is configured to be coupled to the second microphone 120. The second input interface 124 is configured to receive the second microphone output 122 from the second microphone 120 and provide the second microphone output 122 to the processor 190 as second audio data 126.

[0022]

[0037] The processor 190 is coupled to the modem 170 and includes a feature extractor 130 and a dynamic classifier 140. The processor is configured to receive audio data 128 including first audio data 116 corresponding to the first output 112 of the first microphone 110 and second audio data 126 corresponding to the second output 122 of the second microphone 120. The processor 190 is configured to process the audio data 128 in the feature extractor 130 to generate feature data 132.

[0023]

[0038] In some implementations, the processor 190 is configured to process the first audio data 116 and the second audio data 126 before generating the feature data 132. In one example, the processor 190 is configured to perform echo cancellation, noise suppression, or both on the first audio data 116 and the second audio data 126. In some implementations, the processor 190 is configured to transform (e.g., Fourier transform) the first audio data 116 and the second audio data 126 into a transform domain before generating the feature data 132.

[0024]

[0039] The processor 190 is configured to generate feature data 132 based on the first audio data 116 and the second audio data 126. According to some aspects, the feature data 132 includes at least one interaural phase difference 134 between the first audio data 116 and the second audio data 126 and at least one interaural intensity difference 136 between the first audio data 116 and the second audio data 126. In a particular example, the feature data 132 includes an interaural phase difference (IPD) 134 for multiple frequencies and an interaural intensity difference (IID) 136 for multiple frequencies.

[0025]

[0040] The processor 190 is configured to process the feature data 132 in the dynamic classifier 140 to generate a classification output 142 of the feature data 132. In some implementations, the dynamic classifier 140 is configured to adaptively cluster a set (e.g., samples) of the feature data 132 based on whether a sound represented in the audio data 128 originates from a source closer to the first microphone 110 than to the second microphone 120. For example, the dynamic classifier 140 may be configured to receive a sequence of samples of the feature data 132 and adaptively cluster the samples in a feature space including IID frequency values ​​and IPD frequency values.

[0026]

[0041] The dynamic classifier 140 may also be configured to adjust a decision boundary between two most discriminative categories of the feature space to distinguish between a set of feature data corresponding to user voice activity (e.g., utterance 182 of user 180) and a set of feature data corresponding to other audio activity. By way of example, the dynamic classifier 140 may be configured to classify incoming feature data into one of two classes (e.g., class 0 or class 1), where one of the two classes corresponds to user voice activity and the other of the two classes corresponds to other audio activity. The classification output 142 may include a single bit or flag having one of two values: a first value (e.g., “0”) to indicate that the feature data 132 corresponds to one of the two classes, or a second value (e.g., “1”) to indicate that the feature data 132 corresponds to the other of the two classes.

[0027]

[0042] In some implementations, the dynamic classifier 140 performs clustering and vector quantization. For example, the clustering includes reducing (e.g., minimizing) the intra-cluster sum of squares, defined as:

[0028]

number

[0029] where C i represents cluster i, and p i represents the weight assigned to cluster i, and x j represents node j in the feature space, and μ i represents the center of gravity of cluster i. Cluster weight p i, can be determined by probabilistic factors such as prior cluster distributions, chance factors such as confidence measures assigned to each cluster likelihood, or any other factor that imposes some form of uneven bias towards different clusters. Vector quantization involves reducing (e.g., minimizing) error by quantizing an input vector to become a quantization weight vector defined by:

[0030]

number

[0031] where w i represents the quantization weight vector i.

[0032]

[0043] In some implementations, the dynamic classifier 140 is configured to perform competitive learning, in which quantization units compete to absorb new samples of feature data 132. The winning unit is then adjusted toward the new sample. For example, the weight vectors of each unit may be initialized for separation or randomly. For each new sample of received feature data, a determination is made as to which weight vector is closest to the new sample, based on, for example, Euclidean distance or dot-product similarity, as non-limiting examples. The weight vector closest to the new sample (the “winner” or best matching unit) may then be moved toward the new sample. For example, in Hebbian learning, the winner strengthens its correlation with the input, such as by adjusting the weights between two nodes in proportion to the product of the inputs to the two nodes.

[0033]

[0044] In some implementations, the dynamic classifier 140 includes local clusters in the presynaptic sheet connected to local clusters in the post-synaptic sheet, and the interconnections between adjacent neurons are strengthened through Hebbian learning to strengthen connections between correlated stimuli. The dynamic classifier 140 may include a Kohonen self-organizing map, where inputs are connected to every neuron in the post-synaptic sheet or map. Learning localizes the map in that different fields of absorption respond to different regions of the input space (e.g., feature data space).

[0034]

[0045] In particular implementations, the dynamic classifier 140 includes a self-organizing map 148. The self-organizing map 148 operates by initializing a weight vector and then, for each input t (e.g., each received set of feature data 132), determining the winning unit (or cell or neuron) according to the formula:

[0035]

number

[0036] We may find the winner v(t) as the unit with the smallest distance (e.g., Euclidean distance) to the input x(t). The weights of the winning unit and its neighbors are given by Δw i (t)=α(t)l(v,i,t)[x(t)-w v (t)], where Δw i where (t) represents the change in unit i, α(t) represents a learning parameter, and l(v,i,t) represents a neighborhood function around the winning unit, such as a Gaussian radial basis function. In some implementations, an inner product or another metric may be used as the similarity measure instead of Euclidean distance.

[0037]

[0046] In some implementations, the dynamic classifier 140 includes a variation of a Kohonen self-organizing map for adapting to a sequence of speech samples, as further described with reference to Figure 4. In one example, the dynamic classifier 140 may perform temporal sequence processing according to, for example, a temporal Kohonen map, in which an activation function with a time-constant modeling decay ("D") is defined for each unit and updated as follows:

[0038]

number

[0039] The winning unit is the unit with the greatest activity. As another example, the dynamic classifier 140 may use the difference vector y instead of the squared norm, i.e., y i (t,γ)=(1-γ)y i (t-1,γ)+γ(x(t)-w i The recurrent network may be implemented according to a recurrent self-organizing map using (t)), where γ represents a forgetting factor having a value between 0 and 1, and the winning unit is determined as the unit with the smallest difference vector as follows:

[0040]

number

[0041] The weight is Δw i (t)=α(t)l(v,i,t)[x(t)-y v (t,γ)] is updated as

[0042]

[0047] In some implementations, the processor 190 is configured to update the clustering operation 144 of the dynamic classifier 140 based on the feature data 132 and update the classification decision criterion 146 of the dynamic classifier 140. For example, as described above, the processor 190 is configured to adapt the clustering and the decision boundary between user voice activity and other audio activity based on incoming samples of the audio data 128, allowing the dynamic classifier 140 to adjust its operation based on changing conditions of the user 180, the environment, other conditions (e.g., microphone placement or adjustment), or any combination thereof.

[0043]

[0048] Although the dynamic classifier 140 is illustrated as including a self-organizing map 148, in other implementations, the dynamic classifier 140 may incorporate one or more other techniques for generating the classification output 142 instead of or in addition to the self-organizing map 148. By way of non-limiting example, the dynamic classifier 140 may include a restricted Boltzmann machine with an unsupervised configuration, an unsupervised autoencoder, an online variant of a Hopfield network, online clustering, or a combination thereof. As another non-limiting example, the dynamic classifier 140 may be configured to perform principal component analysis (e.g., sequentially fitting a set of orthogonal direction vectors to feature vector samples in a feature space, where each direction vector is selected to maximize the variance of the feature vector samples projected onto the direction vector in the feature space). As another non-limiting example, the dynamic classifier 140 may be configured to perform independent component analysis (e.g., determining a set of additive subcomponents of the feature vector samples in the feature space, assuming the subcomponents are statistically independent non-Gaussian signals).

[0044]

[0049] The processor 190 is configured to determine, based at least in part on the classification output 142, whether the audio data 128 corresponds to user voice activity and generate a user voice activity indicator 150 that indicates whether user voice activity has been detected. For example, the classification output 142 may indicate whether the feature data 132 is classified as one of two classes (e.g., class “0” or class “1”), but the classification output 142 may not indicate which class corresponds to user voice activity and which class corresponds to other audio activity. For example, based on how the dynamic classifier 140 is initialized and the feature data used to update the dynamic classifier 140, in some cases a classification output 142 with a value of “0” indicates user voice activity, while in other cases a classification output with a value of “0” indicates other audio activity. The processor 190 may determine which of the two classes indicates user voice activity and which of the two classes indicates other audio activity further based on at least one of the sign or magnitude of at least one value of the feature data 132, as further described with reference to FIG. 2.

[0045]

[0050] To illustrate, sound propagation of speech 182 from the mouth of user 180 to first microphone 110 and second microphone 120 may be detected in feature data 132, resulting in phase and signal strength differences (due to speech 182 arriving at first microphone 110 before second microphone 120) that may be distinguishable from phase and signal strength differences of sound from other audio sources. The phase and signal strength differences may be determined from IPD 134 and IID 136 in feature data 132 and used to map classification output 142 to user voice activity or other audio activity. Processor 190 may generate a user voice activity indicator 150 that indicates whether audio data 128 corresponds to user voice activity.

[0046]

[0051] In some implementations, the processor 190 is configured to initiate a voice command processing operation 152 in response to determining that the audio data 128 corresponds to user voice activity. In an illustrative example, the voice command processing operation 152 includes a voice activation operation such as keyword or key phrase detection, voiceprint authentication, natural language processing, one or more other operations, or any combination thereof. As another example, the processor 190 may process the audio data 128 to perform a first stage of keyword detection and may use the user voice activity indicator 150 to confirm that the detected keyword was spoken by the user 180 of the device 102 and not by a nearby person before initiating further processing of the audio data 128 via the voice command processing operation 152 (e.g., in a second stage of detection including more powerful voice activity recognition and speech recognition operations).

[0047]

[0052] Modem 170 is coupled to processor 190 and configured to enable communication with second device 160, such as via wireless transmission. In some examples, modem 170 is configured to transmit audio data 128 to second device 160 in response to determining that audio data 128 corresponds to user voice activity based on dynamic classifier 140. For example, in an implementation in which device 102 corresponds to a headset device wirelessly coupled to second device 160 (e.g., a Bluetooth connection to a mobile phone or computer), device 102 may send audio data 128 to second device 160 for performing voice command processing operations 152 in a voice activation system 162 of second device 160. In this example, device 102 offloads more computationally expensive processing (e.g., voice command processing operations 152) to be performed using the greater processing and power resources of second device 160.

[0048]

[0053] In some implementations, device 102 corresponds to or is included in one or various types of devices. In an illustrative example, processor 190 is integrated into a headset device that includes first microphone 110 and second microphone 120. When worn by user 180, the headset device is configured to position first microphone 110 closer to the user's mouth than second microphone 120 so as to capture speech 182 of user 180 with greater intensity and less delay at first microphone 110 compared to second microphone 120, as further described with reference to FIG. In other examples, processor 190 is integrated into at least one of a mobile phone or tablet computer device described with reference to Figure 6, a wearable electronic device described with reference to Figure 8, a voice-controlled speaker system described with reference to Figure 9, a camera device described with reference to Figure 10, or a virtual reality headset, mixed reality headset, or augmented reality headset described with reference to Figure 11. In another illustrative example, processor 190 is integrated into a vehicle that also includes first microphone 110 and second microphone 120, as further described with reference to Figures 12 and 13.

[0049]

[0054] During operation, the first microphone 110 is configured to capture speech 182 of a user 180, and the second microphone 120 is configured to capture ambient sound 186. In one example, speech 182 from a user 180 of the device 102 is captured by the first microphone 110 and the second microphone 120. Because the first microphone 110 is closer to the mouth of the user 180, the speech of the user 180 is captured by the first microphone 110 with higher signal strength and less delay compared to the second microphone 120. In another example, ambient sound 186 from one or more sound sources 184 (e.g., a conversation between two nearby people) may be captured by the first microphone 110 and the second microphone 120. Based on the position and distance of the sound source 184 relative to the first microphone 110 and the second microphone 120, the signal strength difference and relative delay between capturing the ambient sound 186 at the first microphone 110 and the second microphone 120 will be different from that for the speech 182 from the user 180.

[0050]

[0055] The first audio data 116 and the second audio data 126 are processed in a processor 190 by performing echo cancellation, noise suppression, frequency domain transformation, etc. The resulting audio data is processed in a feature extractor 130 to generate feature data 132 including an IPD 134 and an IID 136. The feature data 132 is input to a dynamic classifier 140 to generate a classification output 142, which is interpreted by the processor 190 as either user voice activity or other sound activity. The processor 190 generates a user voice activity indicator 150, such as a value of “0” to indicate that the audio data 128 corresponds to user voice activity or a value of “1” to indicate that the audio data 128 corresponds to other audio activity (or vice versa).

[0051]

[0056] The user voice activity indicator 150 may be used to determine whether to initiate a voice command processing operation 152 at the device 102. Alternatively, or additionally, the user voice activity indicator 150 may be used to determine whether to initiate generation of an output signal 135 (e.g., audio data 128) to a second device 160 for further processing in a voice activation system 162.

[0052]

[0057] Additionally, along with generating the classification output 142, the dynamic classifier 140 is updated based on the feature data 132, such as by adjusting the weights of the winning unit and its neighboring units to more closely resemble the feature data 132, updating the clustering operations 144, the classification criteria 146, or a combination thereof. In this way, the dynamic classifier 140 automatically adapts to changes in the user's speech, changes in the environment, changes in the characteristics of the device 102 or microphones 110, 120, or a combination thereof.

[0053]

[0058] Thus, the system 100 improves the performance of in-house voice activity detection by using a dynamic classifier 140 to distinguish between user voice activity and other audio activity with relatively low complexity, low power consumption, and high accuracy compared to conventional in-house voice activity detection techniques. Automatically adapting to user and environmental changes provides improved benefits by reducing or eliminating calibration to be performed by the user and improving the user's experience.

[0054]

[0059] In some implementations, the processor 190 provides the audio data 128 to the dynamic classifier 140 in the form of feature data 132 (e.g., frequency domain data) generated by the feature extractor 130, while in other implementations, the feature extractor 130 is omitted. In one example, the processor 190 provides the audio data 128 to the dynamic classifier 140 as a time series of audio samples, and the dynamic classifier 140 processes the audio data 128 to generate the classification output 142. In one exemplary implementation, the dynamic classifier 140 is configured to determine frequency domain data from the audio data 128 (e.g., generate the feature data 132) and use the extracted frequency domain data to generate the classification output 142.

[0055]

[0060] Although first microphone 110 and second microphone 120 are illustrated as being coupled to device 102, in other implementations, one or both of first microphone 110 or second microphone 120 may be integrated into device 102. While two microphones 110, 120 are illustrated, other implementations may include one or more additional microphones configured to capture user speech, one or more microphones configured to capture environmental sound, or both. Although system 100 is illustrated as including second device 160, in other implementations, second device 160 may be omitted and device 102 may perform the operations described as being performed in second device 160.

[0056]

[0061] 2 is a diagram of an example aspect of operations 200 associated with self-voice activity detection that may be performed by device 102 (e.g., processor 190) of FIG. 1. Feature extraction 204 is performed on input 202 to generate feature data 206. In one example, input 202 corresponds to audio data 128, feature extraction 204 is performed by feature extractor 130, and feature data 206 corresponds to feature data 132.

[0057]

[0062] A dynamic classifier 208 operates on the feature data 206 to generate a classification output 210. In one example, the dynamic classifier 208 corresponds to the dynamic classifier 140 and is configured to perform unsupervised real-time clustering based on the feature data 206 with highly dynamic decision boundaries for “self” vs. “other” labeling of the speech activation classes in the classification output 210. For example, the dynamic classifier 208 may divide the feature space into two classes: one class associated with the user's speech activity and the other class associated with other speech activity. The classification output 210 may include a binary indicator of which class is associated with the feature data 206. In one example, the classification output 210 corresponds to the classification output 142.

[0058]

[0063] The self / other association operation 212 generates a self / other indicator 218 based on the classification output 210 and a verification input 216. The verification input 216 may provide information associating each of the classes of the classification output 210 with a user voice activity (e.g., “self”) or other sound activity (e.g., “other”). For example, the verification input 216 may be generated based on at least one prior verification criterion 214, such as comparing the sign 230 of the phase difference (e.g., one or more values ​​of the IPD 134 over one or more specific frequency ranges indicating which microphone is closer to the source of the audio represented by the input 202), comparing the magnitude 232 of the intensity difference (e.g., one or more values ​​of the IID 136 over one or more specific frequency ranges indicating the relative distance of the source of the audio to the separate microphones), or a combination thereof. For example, a self / other association may determine that a classification output 210 value of "0" corresponds to feature data 206 that exhibits a negative sign 230 in one or more relevant frequency ranges, or a magnitude 232 less than a threshold amount in one or more relevant frequency ranges, or both, thereby filling the table so that "0" corresponds to "other" and "1" corresponds to "self."

[0059]

[0064] The self / other association operation 212 results in the generation of a self / other indicator 218 (e.g., a binary indicator having a first value (e.g., “0”) to indicate user voice activity or a second value (e.g., “1”) to indicate other sound activity, or vice versa). A wakeup / barge-in control operation 220 is responsive to the self / other indicator 218 to generate a signal 222 to a voice command process 224. For example, the signal 222 may have a first value (e.g., “0”) to indicate that the voice command process 224 should be executed on the input 202, the feature data 206, or both, to perform further voice command processing (e.g., to perform keyword detection, voice authentication, or both) when the input 202 corresponds to user voice activity, or may have a second value (e.g., “1”) to indicate that the voice command process 224 should not perform voice command processing when the input 202 corresponds to other sound activity.

[0060]

[0065] Dynamic classification, as described with reference to dynamic classifier 140 of FIG. 1 and dynamic classifier 208 of FIG. 2, responds only when the user speaks and suppresses the response whenever other interference (e.g., external speech) arrives, helping to improve SVAD accuracy with the goal of maximizing the self-keyword acceptance rate (“SKAR”) and other keyword rejection rate (“OKRR”). By using dynamic classification, various challenges associated with conventional SVAD processing are avoided or potentially reduced. For example, challenges of conventional SVAD processing that are avoided or reduced via dynamic classification implementations include noise and echo conditions (which can cause false activation and barge-in under harsh conditions), microphone mismatch and sensitivity, voice activation engine dependency, different user head-related transfer functions (HRTFs), different headset hardware effects, user behavior-based variations in occlusion and isolation levels, user characteristic similarity with other voice activities, and the eventual adverse effect on voice activation and response delay of user speech onset. To illustrate, conventional SVAD relies heavily on internal / external microphone calibration and sensitivity, direction of arrival of interfering speech, headset fitting and isolation variations, and non-stationary statistics of features, which can be adapted by dynamic classification operations.

[0061]

[0066] The use of dynamic classification allows for the use of extracted feature data 206 to discriminate and actively respond to and adapt to various conditions, such as environmental conditions in highly non-stationary situations, mismatched microphones, changes in user headset fitting, different user head-related transfer functions, direction of arrival ("DOA") tracking of non-user signals, and microphone noise floor, bias, and sensitivity across the frequency spectrum. Dynamic classification allows for adaptive feature mapping that can respond to such variations and reduce or minimize the number of thresholding parameters used and the amount of headset tuning by the customer.

[0062]

[0067] 3 is a block diagram of an example aspect of a system operable to perform self-voice activity detection according to some examples of the present disclosure, where a processor 190 includes an always-on power domain 303 and a second power domain 305, such as an on-demand power domain. In some implementations, a first stage 340 and a buffer 360 of a self-voice activity detector 320 are configured to operate in an always-on mode, and a second stage 350 of the self-voice activity detector 320 is configured to operate in an on-demand mode.

[0063]

[0068] The always-on power domain 303 includes a buffer 360, a feature extractor 130, and a dynamic classifier 140. The buffer 360 is configured to store the first audio data 116 and the second audio data 126 so that they are accessible for processing by the components of the self-voice activity detector 320.

[0064]

[0069] The second power domain 305 includes a voice command processing unit 370 in the second stage 250 of the self-voice activity detector 320 and also includes an activation circuit 330. In some implementations, the voice command processing unit 370 is configured to perform the voice command processing operation 152 of FIG. 1 or the voice command process 224 of FIG. 2.

[0065]

[0070] The first stage 240 of the self voice activity detector 320 is configured to generate at least one of a wakeup signal 322 or an interrupt 324 to initiate a voice command processing operation 152 (or a voice command process 224) in the voice command processing unit 370. In one example, the wakeup signal 322 is configured to transition the second power domain 305 from a low-power mode 332 to an active mode 334 to activate the voice command processing unit 370. In some implementations, the wakeup signal 322, the interrupt 324, or both correspond to the signal 222 of FIG. 2 .

[0066]

[0071] For example, activation circuit 330 may include or be coupled to power management circuitry, clock circuitry, headswitch or footswitch circuitry, buffer control circuitry, or any combination thereof. Activation circuit 330 may be configured to initiate power-up of second stage 350, such as by selectively applying or increasing the voltage of a power supply for second stage 350, second power domain 305, or both. As another example, activation circuit 330 may be configured to selectively gate or ungate a clock signal to second stage 350, such as to prevent or enable circuit operation without removing a power supply.

[0067]

[0072] The detector output 352 generated by the second stage 350 of the self-voice activity detector 320 is provided to an application 354. The application 354 may be configured to perform one or more actions based on the detected user speech. By way of illustration, the application 354 may correspond to a voice interface application, an integrated assistant application, a vehicle navigation and entertainment application, or a home automation system, as illustrative, non-limiting examples.

[0068]

[0073] By selectively activating the second stage 350 based on the results of processing audio data in the first stage 340 of the in-house voice activity detector 320, the overall power consumption associated with in-house voice activity detection, voice command processing, or both may be reduced.

[0069]

[0074] 4 is a diagram of an exemplary aspect of the operation of components of the system of FIG. 1, in accordance with some examples of the present disclosure. Feature extractor 130 is configured to receive a sequence of audio data samples 410, such as a sequence of consecutively captured frames of audio data 128, illustrated as a first frame (F1) 412, a second frame (F2) 414, and one or more additional frames including an Nth frame (FN) 416 (where N is an integer greater than 2). Feature extractor 130 is configured to output a sequence of sets of feature data 420, including a first set 422, a second set 424, and one or more additional sets including an Nth set 426.

[0070]

[0075] The dynamic classifier 140 is configured to receive a sequence 420 of sets of feature data and adaptively cluster a set (e.g., a second set 424) of the sequence 420 based at least in part on a previous set (e.g., a first set 422) of feature data in the sequence 420. As an illustrative, non-limiting example, the dynamic classifier 140 may be implemented as a temporal Kohonen map or a recurrent self-organizing map.

[0071]

[0076] During operation, feature extractor 130 processes first frame 412 to generate a first set of feature data 422, and dynamic classifier 140 processes first set of feature data 422 to generate a first classification output (C1) 432 of sequence of classification outputs 430. Feature extractor 130 processes second frame 414 to generate a second set of feature data 424, and dynamic classifier 140 processes second set of feature data 424 to generate a second classification output (C2) 434 based on second set of feature data 424 and based at least in part on first set of feature data 422. Such processing continues, including feature extractor 130 processing Nth frame 416 to generate an Nth set of feature data 426, and dynamic classifier 140 processing Nth set of feature data 426 to generate an Nth classification output (CN) 436. The Nth classification output 436 is based on the Nth set of feature data 426 and is based at least in part on one or more of the previous sets of feature data in the sequence 420 .

[0072]

[0077] By dynamically classifying based on one or more previous sets of feature data, the accuracy of classification by the dynamic classifier 140 may be improved for speech signals that may span multiple frames of audio data.

[0073]

[0078] FIG. 5 shows an implementation 500 of device 102 as an integrated circuit 502 that includes one or more processors 190. The integrated circuit 502 also includes an audio input 504, such as one or more bus interfaces, to allow audio data 128 to be received for processing. The integrated circuit 502 also includes a signal output 512, such as a bus interface, to allow transmission of an output signal, such as a user voice activity indicator 150. The integrated circuit 502 enables implementation of self-voice activity detection as a component within a system that includes a microphone, such as the mobile phone or tablet shown in FIG. 6, the headset shown in FIG. 7, the wearable electronic device shown in FIG. 8, the voice-controlled speaker system shown in FIG. 9, the camera shown in FIG. 10, the virtual reality, mixed reality, or augmented reality headset shown in FIG. 11, or the vehicle shown in FIG. 12 or 13.

[0074]

[0079] 6 shows, as an illustrative, non-limiting example, an implementation 600 in which the device 102 is a mobile device 602, such as a phone or tablet. The mobile device 602 includes a first microphone 110 positioned to primarily capture a user's speech, multiple second microphones 120 positioned to primarily capture environmental sounds, and a display screen 604. Components of the processor 190, including the feature extractor 130 and the dynamic classifier 140, are integrated into the mobile device 602 and are illustrated using dashed lines to indicate internal components that are generally invisible to a user of the mobile device 602. Although the processor 190 is illustrated as including the feature extractor 130, in other implementations, the feature extractor 130 is omitted, such as when the dynamic classifier 140 is configured to extract feature data during processing of the first audio data 116 and the second audio data 126, as described with reference to FIG. 1 . In particular examples, the dynamic classifier 140 operates to detect user voice activity, which is then processed to perform one or more actions on the mobile device 602, such as to launch a graphical user interface or, in some cases (e.g., via an integrated “smart assistant” application), display other information related to the user's speech on the display screen 604.

[0075]

[0080] 7 shows an implementation 700 in which the device 102 is a headset device 702. The headset device 702 includes a first microphone 110 positioned to primarily capture a user's speech and a second microphone 120 positioned to primarily capture environmental sounds. Components of the processor 190, including the feature extractor 130 and the dynamic classifier 140, are integrated into the headset device 702. In a particular example, the dynamic classifier 140 operates to detect user voice activity, which may cause the headset device 702 to perform one or more operations at the headset device 702, transmit audio data corresponding to the user voice activity to a second device (not shown), such as the second device 160 of FIG. 1, for further processing, or a combination thereof. Although the processor 190 is illustrated as including the feature extractor 130, in other implementations, the feature extractor 130 is omitted, such as when the dynamic classifier 140 is configured to extract feature data during processing of the first audio data 116 and the second audio data 126, as described with reference to FIG. 1.

[0076]

[0081] 8 shows an implementation 800 in which the device 102 is a wearable electronic device 802, illustrated as a “smart watch.” The feature extractor 130, the dynamic classifier 140, the first microphone 110, and the second microphone 120 are integrated into the wearable electronic device 802. In a particular example, the dynamic classifier 140 operates to detect user voice activity, which is then processed to perform one or more actions on the wearable electronic device 802, such as to launch a graphical user interface or possibly display other information related to the user's speech on a display screen 804 of the wearable electronic device 802. By way of illustration, the wearable electronic device 802 may include a display screen configured to display notifications based on user speech detected by the wearable electronic device 802. In a particular example, the wearable electronic device 802 includes a haptic device that provides a haptic notification (e.g., vibrates) in response to detection of the user voice activity. For example, a tactile notification may cause the user to turn on the wearable electronic device 802 to see a displayed notification indicating the detection of a keyword spoken by the user. Thus, the wearable electronic device 802 may alert a hearing-impaired user or a user wearing a headset that the user's voice activity has been detected. Although the wearable electronic device 802 is illustrated as including the feature extractor 130, in other implementations, the feature extractor 130 is omitted, such as when the dynamic classifier 140 is configured to extract feature data during processing of the first audio data 116 and the second audio data 126, as described with reference to FIG. 1 .

[0077]

[0082] 9 is an implementation 900 in which the device 102 is a wireless speaker and voice-activated device 902. The wireless speaker and voice-activated device 902 can have wireless network connectivity and is configured to perform assistant operations. A processor 190 including a feature extractor 130 and a dynamic classifier 140, the first microphone 110, the second microphone 120, or a combination thereof, is included in the wireless speaker and voice-activated device 902. Although the processor 190 is illustrated as including the feature extractor 130, in other implementations, the feature extractor 130 is omitted, such as when the dynamic classifier 140 is configured to extract feature data during processing of the first audio data 116 and the second audio data 126, as described with reference to FIG. 1 . The wireless speaker and voice-activated device 902 also includes a speaker 904. In operation, in response to receiving a verbal command identified as user speech via operation of dynamic classifier 140, wireless speaker and voice-activated device 902 can perform an assistant action, such as via execution of voice activation system 162 (e.g., an integrated assistant application). The assistant action may include adjusting the temperature, playing music, turning on lights, etc. For example, an assistant action is performed in response to receiving a keyword or key phrase (e.g., "hello assistant") followed by the command.

[0078]

[0083] In an illustrative example, when wireless speaker and voice-activated device 902 is positioned near a wall of a room (e.g., next to a window) and the first microphone 110 is positioned closer to the interior of the room compared to the second microphone 120 (e.g., the second microphone may be positioned closer to the wall or window than the first microphone 110), speech originating from inside the room may be identified as user voice activity, and sounds originating from outside the room (e.g., the speech of a person on the other side of a wall or window) may be identified as other audio activity. Because multiple people may be in the room, wireless speaker and voice-activated device 902 may be configured to identify speech from any of the multiple people as user voice activity (e.g., there may be multiple “users” of wireless speaker and voice-activated device 902). To illustrate, the dynamic classifier 140 may be configured to recognize feature data corresponding to speech originating from within a room as “own voice” even when the person speaking may be relatively far (e.g., several meters) from the wireless speaker and voice activated device 902 and closer to the first microphone 110 than the second microphone 120. In some implementations in which speech is detected from multiple people in a room, the wireless speaker and voice activated device 902 (e.g., the dynamic classifier 140) may be configured to identify speech from the person closest to the first microphone 110 as user voice activity (e.g., the own voice of the nearest user).

[0079]

[0084] 10 illustrates an implementation 1000 in which the device 102 is a portable electronic device corresponding to a camera device 1002. The feature extractor 130 and dynamic classifier 140, the first microphone 110, the second microphone 120, or a combination thereof, are included in the camera device 1002. During operation, in response to receiving a verbal command identified as user speech via operation of the dynamic classifier 140, the camera device 1002 can perform an action in response to the spoken user command, such as to adjust image or video capture settings, image or video playback settings, or image or video capture instructions, as illustrative examples. While the camera device 1002 is illustrated as including the feature extractor 130, in other implementations, the feature extractor 130 is omitted, such as when the dynamic classifier 140 is configured to extract feature data during processing of the first audio data 116 and the second audio data 126, as described with reference to FIG. 1.

[0080]

[0085] FIG. 11 shows an implementation 1100 in which the device 102 includes a portable electronic device corresponding to an extended reality (“XR”) headset 1102, such as a virtual reality (“VR”), augmented reality (“AR”), or mixed reality (“MR”) headset device. The feature extractor 130, the dynamic classifier 140, the first microphone 110, the second microphone 120, or a combination thereof, are integrated into the headset 1102. In particular aspects, the headset 1102 includes the first microphone 110 positioned to primarily capture the user's speech and the second microphone 120 positioned to primarily capture environmental sounds. User voice activity detection may be performed based on audio signals received from the first microphone 110 and the second microphone 120 of the headset 1102. While the headset 1102 is being worn, a visual interface device is placed in front of the user's eyes to enable the display of augmented reality or virtual reality images or scenes to the user. In particular examples, the visual interface device is configured to display a notification indicating user speech detected in the audio signal. Although the headset 1102 is illustrated as including the feature extractor 130, in other implementations, the feature extractor 130 is omitted, such as when the dynamic classifier 140 is configured to extract feature data during processing of the first audio data 116 and the second audio data 126, as described with reference to FIG.

[0081]

[0086] 12 shows an implementation 1200 in which the device 102 corresponds to or is integrated within a vehicle 1202, illustrated as a manned or unmanned aerial device (e.g., a package delivery drone). The feature extractor 130, the dynamic classifier 140, the first microphone 110, the second microphone 120, or a combination thereof, are integrated into the vehicle 1202. User voice activity detection may be performed based on audio signals received from the first microphone 110 and the second microphone 120 of the vehicle 1202, such as for delivery orders from an authorized user of the vehicle 1202. While the vehicle 1202 is illustrated as including the feature extractor 130, in other implementations, the feature extractor 130 is omitted, such as when the dynamic classifier 140 is configured to extract feature data during processing of the first audio data 116 and the second audio data 126, as described with reference to FIG. 1.

[0082]

[0087] 13 shows another implementation 1300 in which the device 102 corresponds to or is integrated within a vehicle 1302, illustrated as a car. The vehicle 1302 includes a processor 190 that includes a feature extractor 130 and a dynamic classifier 140. While the vehicle 1302 is illustrated as including the feature extractor 130, in other implementations, the feature extractor 130 is omitted, such as when the dynamic classifier 140 is configured to extract feature data during processing of the first audio data 116 and the second audio data 126, as described with reference to FIG. 1. The vehicle 1302 also includes a first microphone 110 and a second microphone 120. The first microphone 110 is positioned to capture the speech of an operator of the vehicle 1302. User voice activity detection may be performed based on audio signals received from the first microphone 110 and the second microphone 120 of the vehicle 1302. In some implementations, user voice activity detection may be performed based on audio signals received from internal microphones (e.g., the first microphone 110 and the second microphone 120), such as for voice commands from authorized passengers. For example, user voice activity detection may be used to detect voice commands from an operator of the vehicle 1302 (e.g., from a parent to set the volume to 5 or to set a destination for the autonomous vehicle) and ignore the voice of another passenger (e.g., a voice command from a child to set the volume to 10 or from another passenger discussing another location). In some implementations, user voice activity detection may be performed based on audio signals received from external microphones (e.g., the first microphone 110 and the second microphone 120), such as for authorized users of the vehicle.In particular implementations, in response to receiving a verbal command identified as user speech through operation of the dynamic classifier 140, the voice activation system 162 initiates one or more actions of the vehicle 1302 based on one or more keywords (e.g., "unlock," "start engine," "play music," "show weather forecast," or another voice command) detected in the output signal 135, such as by providing feedback or information via the display 1320 or one or more speakers (e.g., the speaker 1310).

[0083]

[0088] 14A, shown is a particular implementation of a method of user voice activity detection 1400. In a particular aspect, one or more operations of the method 1400 are performed by at least one of the feature extractor 130, the dynamic classifier 140, the processor 190, the device 102, the system 100, or a combination thereof of FIG.

[0084]

[0089] The method 1400 includes, at 1402, receiving, at one or more processors, audio data including first audio data corresponding to a first output of a first microphone and second audio data corresponding to a second output of a second microphone. For example, the feature extractor 130 of FIG. 1 receives audio data 128 including first audio data 116 corresponding to a first output of the first microphone 110 and second audio data 126 corresponding to a second output of the second microphone 126, as described with reference to FIG. 1.

[0085]

[0090] The method 1400 includes, at 1404, generating, at one or more processors, feature data based on the first audio data and the second audio data. For example, the feature extractor 130 of FIG. 1 generates the feature data 132 based on the first audio data 116 and the second audio data 126, as described with reference to FIG. 1. In another example, a dynamic classifier, such as the dynamic classifier 140 of FIG. 1, is configured to receive the first audio data 116 and the second audio data 126 and extract the feature data 132 during processing of the first audio data 116 and the second audio data 126.

[0086]

[0091] The method 1400 includes generating, at 1406, a classification output of the feature data in a dynamic classifier of the one or more processors. For example, the dynamic classifier 140 of FIG. 1 generates the classification output 142 of the feature data 132 as described with reference to FIG. 1.

[0087]

[0092] Method 1400 includes, at 1408, determining, at one or more processors, whether the audio data corresponds to a user voice activity based at least in part on the classification output. For example, processor 190 of FIG. 1 determines whether audio data 128 corresponds to a user voice activity based at least in part on the classification output 142, as described with reference to FIG. 1.

[0088]

[0093] The method 1400 improves the performance of in-house voice activity detection by using a dynamic classifier 140 to distinguish between user voice activity and other audio activity with relatively low complexity, low power consumption, and high accuracy compared to conventional in-house voice activity detection techniques. Automatically adapting to user and environmental changes provides improved benefits by reducing or eliminating calibration to be performed by the user and improving the user's experience.

[0089]

[0094] 14B, shown is a particular implementation of a method of user voice activity detection 1450. In a particular aspect, one or more operations of the method 1450 are performed by at least one of the dynamic classifier 140, the processor 190, the device 102, the system 100, or a combination thereof of FIG.

[0090]

[0095] The method 1450 includes receiving, at one or more processors, audio data including first audio data corresponding to a first output of a first microphone and second audio data corresponding to a second output of a second microphone, at 1452. In one example, as described with reference to FIG. 1, the feature extractor 130 of FIG. 1 receives audio data 128 including the first audio data 116 and the second audio data 126 corresponding to the second output of the second microphone 126.

[0091]

[0096] The method 1450 includes, at 1454, providing the audio data to a dynamic classifier in one or more processors to generate a classification output corresponding to the audio data. In one example, the feature extractor 130 of FIG. 1 generates feature data 132 based on the first audio data 116 and the second audio data 126, and the feature data 132 is processed by the dynamic classifier 140 to generate the classification output 142 according to the method 1400 of FIG. 14A as described in FIG. 1. In another example, the processor 190 provides the first audio data 116 and the second audio data 126 to the dynamic classifier 140, and the dynamic classifier 140 processes the first audio data 116 and the second audio data 126 to generate the classification output 142. In one exemplary implementation, the dynamic classifier 140 processes the first audio data 116 and the second audio data 126 to extract feature data 132 and determines a classification output 142 based on the feature data 132.

[0092]

[0097] The method 1450 includes, at 1456, determining, at the one or more processors, whether the audio data corresponds to a user voice activity based at least in part on the classification output. For example, the processor 190 of FIG. 1 determines whether the audio data 128 corresponds to a user voice activity based at least in part on the classification output 142, as described with reference to FIG. 1.

[0093]

[0098] The method 1450 improves the performance of in-house voice activity detection by using a dynamic classifier 140 to distinguish between user voice activity and other audio activity with relatively low complexity, low power consumption, and high accuracy compared to conventional in-house voice activity detection techniques. Automatically adapting to user and environmental changes provides improved benefits by reducing or eliminating calibration to be performed by the user and improving the user experience.

[0094]

[0099] 14A, 14B, or a combination thereof may be implemented by a field programmable gate array (FPGA) device, an application specific integrated circuit (ASIC), a processing unit such as a central processing unit (CPU), a DSP, a controller, another hardware device, a firmware device, or any combination thereof. As an example, the method 1400 of FIG. 14A, 14B, or a combination thereof may be performed by a processor executing instructions, such as those described with reference to FIG.

[0095]

[0100] 15, a block diagram of a particular example implementation of a device is shown, generally designated 1500. In various implementations, device 1500 may have more or fewer components than those illustrated in FIG. 15. In an example implementation, device 1500 may correspond to device 102. In an example implementation, device 1500 may perform one or more of the operations described with reference to FIGS. 1-14B.

[0096]

[0101] In particular implementations, device 1500 includes a processor 1506 (e.g., a central processing unit (CPU)). Device 1500 may include one or more additional processors 1510 (e.g., one or more DSPs). In particular aspects, processor 190 of FIG. 1 corresponds to processor 1506, processor 1510, or a combination thereof. Processor 1510 may include a speech and music coder-decoder (codec) 1508 including a voice coder (“vocoder”) encoder 1536, a vocoder decoder 1538, a feature extractor 130, a dynamic classifier 140, or a combination thereof.

[0097]

[0102] The device 1500 may include a memory 1586 and a codec 1534. The memory 1586 may include instructions 1556 executable by one or more additional processors 1510 (or processor 1506) to implement the functionality described with reference to the feature extractor 130, the dynamic classifier 140, or both. The device 1500 may include a modem 170 coupled to an antenna 1552 via a transceiver 1550.

[0098]

[0103] The device 1500 may include a display 1528 coupled to a display controller 1526. The speaker 1592, the first microphone 110, and the second microphone 120 may be coupled to a codec 1534. The codec 1534 may include a digital-to-analog converter (DAC) 1502, an analog-to-digital converter (ADC) 1504, or both. In particular implementations, the codec 1534 may receive analog signals from the first microphone 110 and the second microphone 120, convert the analog signals to digital signals using the analog-to-digital converter 1504, and provide the digital signals to a speech and music codec 1508. The speech and music codec 1508 may process the digital signals, which may be further processed by the feature extractor 130 and the dynamic classifier 140. In particular implementations, the speech and music codec 1508 may provide the digital signals to the codec 1534. The codec 1534 may convert the digital signal to an analog signal using the digital-to-analog converter 1502 and may provide the analog signal to the speaker 1592 .

[0099]

[0104] In certain implementations, the device 1500 may be included in a system-in-package or system-on-chip device 1522. In certain implementations, the memory 1586, the processor 1506, the processor 1510, the display controller 1526, the codec 1534, and the modem 170 are included in the system-in-package or system-on-chip device 1522. In certain implementations, the input device 1530 and the power supply 1544 are coupled to the system-on-chip device 1522. Additionally, in certain implementations, the display 1528, the input device 1530, the speaker 1592, the first microphone 110, the second microphone 120, the antenna 1552, and the power supply 1544 are external to the system-on-chip device 1522, as shown in FIG. In particular implementations, each of the display 1528, the input device 1530, the speaker 1592, the first microphone 110, the second microphone 120, the antenna 1552, and the power supply 1544 may be coupled to a component of the system-on-chip device 1522, such as an interface (e.g., the first input interface 114 or the second input interface 124) or a controller.

[0100]

[0105] The device 1500 may include a smart speaker, a speaker bar, a mobile communication device, a smartphone, a cellular phone, a laptop computer, a computer, a tablet, a personal digital assistant, a display device, a television, a game console, a music player, a radio, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a vehicle, a headset, an augmented reality headset, a virtual reality headset, an aviation vehicle, a home automation system, a voice-activated device, a wireless speaker and a voice-activated device, a portable electronic device, a car, a vehicle, a computing device, a communication device, an Internet of Things (IoT) device, a virtual reality (VR) device, a base station, a mobile device, or any combination thereof.

[0101]

[0106] In conjunction with the described implementations, the apparatus includes means for receiving audio data including first audio data corresponding to the first output of the first microphone and second audio data corresponding to the second output of the second microphone. For example, the means for receiving can correspond to first input interface 114, second input interface 124, feature extractor 130, dynamic classifier 140, processor 190, one or more processors 1510, one or more other circuits or components configured to receive audio data including the first audio data corresponding to the first output of the first microphone and the second audio data corresponding to the second output of the second microphone, or any combination thereof.

[0102]

[0107] The apparatus also includes means for generating feature data based on the first audio data and the second audio data. For example, the means for generating feature data may correspond to feature extractor 130, dynamic classifier 140, processor 190, one or more processors 1510, one or more other circuits or components configured to generate feature data, or any combination thereof.

[0103]

[0108] The apparatus further includes means for generating a classification output of the feature data at the dynamic classifier. For example, the means for generating the classification output may correspond to the dynamic classifier 140, the processor 190, the one or more processors 1510, one or more other circuits or components configured to generate the classification output at the dynamic classifier, or any combination thereof.

[0104]

[0109] The apparatus also includes means for determining, based at least in part on the classification output, whether the audio data corresponds to a user voice activity. For example, the means for determining may correspond to dynamic classifier 140, processor 190, one or more processors 1510, one or more other circuits or components configured to determine, based at least in part on the classification output, whether the audio data corresponds to a user voice activity, or any combination thereof.

[0105]

[0110] In conjunction with the described implementations, the apparatus includes means for receiving audio data including first audio data corresponding to the first output of the first microphone and second audio data corresponding to the second output of the second microphone. For example, the means for receiving can correspond to first input interface 114, second input interface 124, feature extractor 130, dynamic classifier 140, processor 190, one or more processors 1510, one or more other circuits or components configured to receive audio data including the first audio data corresponding to the first output of the first microphone and the second audio data corresponding to the second output of the second microphone, or any combination thereof.

[0106]

[0111] The apparatus further includes means for generating a classification output corresponding to the audio data at the dynamic classifier. For example, the means for generating the classification output may correspond to feature extractor 130, dynamic classifier 140, processor 190, one or more processors 1510, one or more other circuits or components configured to generate the classification output at the dynamic classifier, or any combination thereof.

[0107]

[0112] The apparatus also includes means for determining, based at least in part on the classification output, whether the audio data corresponds to a user voice activity. For example, the means for determining may correspond to dynamic classifier 140, processor 190, one or more processors 1510, one or more other circuits or components configured to determine, based at least in part on the classification output, whether the audio data corresponds to a user voice activity, or any combination thereof.

[0108]

[0113] In some implementations, a non-transitory computer-readable medium (e.g., a computer-readable storage device such as memory 1586) includes instructions (e.g., instructions 1556) that, when executed by one or more processors (e.g., one or more processors 1510 or processor 1506), cause the one or more processors to receive audio data (e.g., audio data 128) including first audio data (e.g., first audio data 116) corresponding to a first output of a first microphone (e.g., first microphone 110) and second audio data (e.g., second audio data 126) corresponding to a second output of a second microphone (e.g., second microphone 120). The instructions, when executed by one or more processors, also cause the one or more processors to provide the audio data to a dynamic classifier (e.g., dynamic classifier 140) to generate a classification output (e.g., classification output 142) corresponding to the audio data. In one example, the instructions, when executed by one or more processors, cause the one or more processors to generate feature data (e.g., feature data 132) based on the first audio data and the second audio data and process the feature data in a dynamic classifier. The instructions, when executed by the one or more processors, also cause the one or more processors to determine, based at least in part on the classification output, whether the audio data corresponds to a user voice activity.

[0109]

[0114] The present disclosure includes the following examples.

[0110]

[0115] Example 1. A device comprising one or more processors configured to receive audio data including first audio data corresponding to a first output of a first microphone and second audio data corresponding to a second output of a second microphone, generate feature data based on the first audio data and the second audio data, process the feature data in a dynamic classifier to generate a classification output of the feature data, and determine, based at least in part on the classification output, whether the audio data corresponds to user voice activity.

[0111]

[0116] Example 2. The device of Example 1, further comprising a first microphone and a second microphone, the first microphone coupled to the one or more processors and configured to capture user speech, and the second microphone coupled to the one or more processors and configured to capture ambient sound.

[0112]

[0117] Example 3. The device of Example 1, wherein the feature data includes at least one interaural phase difference between the first audio data and the second audio data and at least one interaural intensity difference between the first audio data and the second audio data.

[0113]

[0118] Example 4. The device of Example 3, wherein the one or more processors are further configured to transform the first audio data and the second audio data into a transform domain before generating the feature data, the feature data including interaural phase differences for a plurality of frequencies and interaural intensity differences for a plurality of frequencies.

[0114]

[0119] Example 5. The device of Example 1, wherein the dynamic classifier is configured to adaptively cluster the set of feature data based on whether a sound represented in the audio data originates from a source closer to the first microphone than to the second microphone.

[0115]

[0120] Example 6. The device of Example 1, wherein the one or more processors are further configured to update the clustering behavior of the dynamic classifier based on the feature data.

[0116]

[0121] Example 7. The device of Example 1, wherein the one or more processors are further configured to update classification decision criteria of the dynamic classifier.

[0117]

[0122] Example 8. The device of example 1, wherein the dynamic classifier includes a self-organizing map.

[0118]

[0123] Example 9. The device of Example 1, wherein the dynamic classifier is further configured to receive a sequence of sets of feature data and adaptively cluster the set of sequences based at least in part on previous sets of feature data in the sequence.

[0119]

[0124] Example 10. The device of Example 1, wherein the one or more processors are configured to determine whether the audio data corresponds to user voice activity further based on at least one of a sign or a magnitude of the at least one value of the feature data.

[0120]

[0125] Example 11. The device of Example 1, wherein the one or more processors are further configured to initiate a voice command processing operation in response to determining that the audio data corresponds to user voice activity.

[0121]

[0126] Example 12. The device of Example 11, wherein the one or more processors are configured to generate at least one of a wake-up signal or an interrupt to initiate a voice command processing operation.

[0122]

[0127] Example 13. The device of Example 12, wherein the one or more processors further include an always-on power domain including the dynamic classifier and a second power domain including the voice command processing unit, and wherein the wake-up signal is configured to transition the second power domain from the low power mode to activate the voice command processing unit.

[0123]

[0128] Example 14. The device of Example 1, further comprising a modem coupled to the one or more processors, the modem configured to transmit the audio data to the second device in response to a determination, based on the dynamic classifier, that the audio data corresponds to user voice activity.

[0124]

[0129] Example 15. The device of Example 1, wherein the one or more processors are integrated into a headset device including a first microphone and a second microphone, the headset device being configured, when worn by a user, to position the first microphone closer to the user's mouth than the second microphone so as to capture the user's speech with greater intensity and less delay at the first microphone compared to the second microphone.

[0125]

[0130] Example 16. The device of Example 1, wherein the one or more processors are integrated into at least one of a mobile phone, a tablet computing device, a wearable electronic device, a camera device, a virtual reality headset, or an augmented reality headset.

[0126]

[0131] Example 17. The device of Example 1, wherein the one or more processors are integrated into a vehicle, the vehicle further including a first microphone and a second microphone, the first microphone positioned to capture speech of an operator of the vehicle.

[0127]

[0132] Example 18. A method of voice activity detection comprising: receiving, at one or more processors, audio data including first audio data corresponding to a first output of a first microphone and second audio data corresponding to a second output of a second microphone; generating, at the one or more processors, feature data based on the first audio data and the second audio data; generating, at a dynamic classifier of the one or more processors, a classification output of the feature data; and determining, at the one or more processors, whether the audio data corresponds to user voice activity based at least in part on the classification output.

[0128]

[0133] Example 19. The method of Example 18, wherein the first microphone is configured to capture speech of the user and the second microphone is configured to capture ambient sound.

[0129]

[0134] Example 20. The method of Example 18, wherein the feature data includes at least one interaural phase difference between the first audio data and the second audio data and at least one interaural intensity difference between the first audio data and the second audio data.

[0130]

[0135] Example 21. The method of Example 20, further comprising transforming the first audio data and the second audio data into a transform domain before generating the feature data, the feature data including interaural phase differences for a plurality of frequencies and interaural intensity differences for a plurality of frequencies.

[0131]

[0136] Example 22. The method of Example 18, further comprising adaptively clustering, by a dynamic classifier, the set of feature data based on whether a sound represented in the audio data originates from a source closer to the first microphone than to the second microphone.

[0132]

[0137] Example 23. The method of Example 18, further comprising updating the clustering behavior of the dynamic classifier based on the feature data.

[0133]

[0138] Example 24. The method of Example 18, further comprising updating the classification decision criteria of the dynamic classifier.

[0134]

[0139] Example 25. The method of example 18, wherein the dynamic classifier includes a self-organizing map.

[0135]

[0140] Example 26. The method of Example 18, further comprising, in the dynamic classifier, receiving a sequence of sets of feature data and adaptively clustering the set of sequences based at least in part on previous sets of feature data in the sequence.

[0136]

[0141] Example 27. The method of Example 18, wherein determining whether the audio data corresponds to user voice activity is further based on at least one of a sign or a magnitude of the at least one value of the feature data.

[0137]

[0142] Example 28. The method of Example 18, further comprising initiating a voice command processing operation in response to determining that the audio data corresponds to user voice activity.

[0138]

[0143] Example 29. The method of Example 28, further comprising generating at least one of a wake-up signal or an interrupt to initiate a voice command processing operation.

[0139]

[0144] Example 30. The method of Example 29, wherein the wake-up signal is configured to transition a power domain out of a low power mode to initiate a voice command processing operation.

[0140]

[0145] Example 31. The method of Example 18, further comprising transmitting the audio data to the second device in response to determining, based on the dynamic classifier, that the audio data corresponds to user voice activity.

[0141]

[0146] Example 32. The method of Example 18, wherein the one or more processors are integrated into a headset device including a first microphone and a second microphone, the headset device, when worn by a user, positions the first microphone closer to the user's mouth than the second microphone to capture the user's speech with greater intensity and less delay at the first microphone compared to the second microphone.

[0142]

[0147] Example 33. The method of Example 18, wherein the one or more processors are integrated into at least one of a mobile phone, a tablet computing device, a wearable electronic device, a camera device, a virtual reality headset, an augmented reality headset, or a vehicle.

[0143]

[0148] Example 34. A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to receive audio data including first audio data corresponding to a first output of a first microphone and second audio data corresponding to a second output of a second microphone; generate feature data based on the first audio data and the second audio data; process the feature data in a dynamic classifier to generate a classification output of the feature data; and determine, based at least in part on the classification output, whether the audio data corresponds to a user voice activity.

[0144]

[0149] Example 35. The non-transitory computer-readable medium of Example 34, wherein the first microphone is configured to capture speech of a user and the second microphone is configured to capture ambient sound.

[0145]

[0150] Example 36. The non-transitory computer-readable medium of Example 34, wherein the feature data includes at least one interaural phase difference between the first audio data and the second audio data and at least one interaural intensity difference between the first audio data and the second audio data.

[0146]

[0151] Example 37. The non-transitory computer-readable medium of Example 34, wherein the instructions, when executed by the one or more processors, further cause the one or more processors to transform the first audio data and the second audio data into a transform domain before generating the feature data, the feature data including an interaural phase difference for a plurality of frequencies and an interaural intensity difference for a plurality of frequencies.

[0147]

[0152] Example 38. The non-transitory computer-readable medium of Example 34, wherein the instructions, when executed by the one or more processors, further cause the one or more processors to adaptively cluster, with a dynamic classifier, the set of feature data based on whether a sound represented in the audio data originates from a source closer to the first microphone than to the second microphone.

[0148]

[0153] Example 39. The non-transitory computer-readable medium of Example 34, wherein the instructions, when executed by the one or more processors, further cause the one or more processors to update a clustering operation of the dynamic classifier based on the feature data.

[0149]

[0154] Example 40. The non-transitory computer-readable medium of Example 34, wherein the instructions, when executed by the one or more processors, further cause the one or more processors to update classification decision criteria of the dynamic classifier.

[0150]

[0155] Example 41. The non-transitory computer-readable medium of Example 34, wherein the dynamic classifier includes a self-organizing map.

[0151]

[0156] Example 42. The non-transitory computer-readable medium of Example 34, wherein the instructions, when executed by one or more processors, further cause the one or more processors to receive, in a dynamic classifier, a sequence of sets of feature data and adaptively cluster the set of sequences based at least in part on previous sets of feature data in the sequences.

[0152]

[0157] Example 43. The non-transitory computer-readable medium of Example 34, wherein determining whether the audio data corresponds to user voice activity is further based on at least one of a sign or a magnitude of the at least one value of the feature data.

[0153]

[0158] Example 44. The non-transitory computer-readable medium of Example 34, wherein the instructions, when executed by the one or more processors, further cause the one or more processors to initiate a voice command processing operation in response to determining that the audio data corresponds to user voice activity.

[0154]

[0159] Example 45. The non-transitory computer-readable medium of Example 34, wherein the instructions, when executed by the one or more processors, further cause the one or more processors to generate at least one of a wake-up signal or an interrupt to initiate a voice command processing operation.

[0155]

[0160] Example 46. The non-transitory computer-readable medium of Example 45, wherein the wake-up signal is configured to transition the power domain out of the low-power mode to initiate a voice command processing operation.

[0156]

[0161] Example 47. The non-transitory computer-readable medium of Example 34, wherein the instructions, when executed by the one or more processors, further cause the one or more processors to transmit the audio data to a second device in response to a determination, based on the dynamic classifier, that the audio data corresponds to user voice activity.

[0157]

[0162] Example 48. The non-transitory computer-readable medium of Example 34, wherein the one or more processors are integrated into a headset device including a first microphone and a second microphone, the headset device, when worn by a user, positions the first microphone closer to the user's mouth than the second microphone to capture the user's speech with greater intensity and less delay at the first microphone compared to the second microphone.

[0158]

[0163] Example 49. The non-transitory computer-readable medium of Example 34, wherein the one or more processors are integrated into at least one of a mobile phone, a tablet computing device, a wearable electronic device, a camera device, a virtual reality headset, an augmented reality headset, or a vehicle.

[0159]

[0164] Example 50. An apparatus comprising: means for receiving audio data including first audio data corresponding to a first output of a first microphone and second audio data corresponding to a second output of a second microphone; means for generating feature data based on the first audio data and the second audio data; means for generating a classification output of the feature data in a dynamic classifier; and means for determining whether the audio data corresponds to user voice activity based at least in part on the classification output.

[0160]

[0165] Example 51. The apparatus of Example 50, wherein the feature data includes at least one interaural phase difference between the first audio data and the second audio data and at least one interaural intensity difference between the first audio data and the second audio data.

[0161]

[0166] Example 52. The apparatus of Example 50, further comprising means for transforming the first audio data and the second audio data into a transform domain prior to generating the feature data, the feature data including interaural phase differences for a plurality of frequencies and interaural intensity differences for a plurality of frequencies.

[0162]

[0167] Example 53. The apparatus of Example 50, further comprising means for adaptively clustering the set of feature data based on whether sounds represented in the audio data originate from a source closer to the first microphone than to the second microphone.

[0163]

[0168] Example 54. The apparatus of Example 50, further comprising means for updating the clustering operation of the dynamic classifier based on the feature data.

[0164]

[0169] Example 55. The apparatus of Example 50, further comprising means for updating the classification decision criteria of the dynamic classifier.

[0165]

[0170] Example 56. The apparatus of example 50, wherein the dynamic classifier includes a self-organizing map.

[0166]

[0171] Example 57. The apparatus of Example 50, further comprising means for initiating a voice command processing operation in response to determining that the audio data corresponds to user voice activity.

[0167]

[0172] Example 58. The apparatus of Example 50, further comprising means for generating at least one of a wake-up signal or an interrupt to initiate a voice command processing operation.

[0168]

[0173] Example 59. The apparatus of Example 50, further comprising means for transmitting the audio data to a second device in response to determining, based on the dynamic classifier, that the audio data corresponds to user voice activity.

[0169]

[0174] Example 60. The apparatus of Example 50, wherein the means for receiving audio data, the means for generating feature data, the means for generating a classification output, and the means for determining whether the audio data corresponds to user voice activity are integrated into a headset device including a first microphone and a second microphone, the headset device, when worn by a user, positions the first microphone closer to the user's mouth than the second microphone to capture the user's speech with greater intensity and less delay at the first microphone compared to the second microphone.

[0170]

[0175] Example 61. The apparatus of Example 50, wherein the means for receiving audio data, the means for generating feature data, the means for generating a classification output, and the means for determining whether the audio data corresponds to user voice activity are integrated into at least one of a mobile phone, a tablet computing device, a wearable electronic device, a camera device, a virtual reality headset, an augmented reality headset, or a vehicle.

[0171]

[0176] Example 62. A device including: a memory configured to store instructions; and one or more processors, the one or more processors configured to execute the instructions to: receive audio data including first audio data corresponding to a first output of a first microphone and second audio data corresponding to a second output of a second microphone; provide the audio data to a dynamic classifier configured to generate a classification output corresponding to the audio data; and determine, based at least in part on the classification output, whether the audio data corresponds to user voice activity.

[0172]

[0177] Example 63. The device of Example 62, further including a first microphone and a second microphone, the first microphone coupled to the one or more processors and configured to capture user speech, and the second microphone coupled to the one or more processors and configured to capture ambient sound.

[0173]

[0178] Example 64. The device of Examples 62 or 63, wherein the classification output is based on a gain difference between the first audio data and the second audio data, a phase difference between the first audio data and the second audio data, or a combination thereof.

[0174]

[0179] Example 65. The device of any one of Examples 62 to 64, wherein the one or more processors are further configured to generate feature data based on the first audio data and the second audio data and provide the feature data to the dynamic classifier, wherein the classification output is based on the feature data.

[0175]

[0180] Example 66. The device of Example 65, wherein the feature data includes at least one interaural phase difference between the first audio data and the second audio data and at least one interaural intensity difference between the first audio data and the second audio data.

[0176]

[0181] Example 67. The device of Example 65 or 66, wherein the one or more processors are configured to determine whether the audio data corresponds to user voice activity further based on at least one of a sign or a magnitude of the at least one value of the feature data.

[0177]

[0182] Example 68. The device of any one of Examples 65 to 67, wherein the one or more processors are further configured to transform the first audio data and the second audio data into a transform domain before generating the feature data, the feature data including interaural phase differences for the plurality of frequencies and interaural intensity differences for the plurality of frequencies.

[0178]

[0183] Example 69. The device of any one of Examples 65 to 68, wherein the dynamic classifier is configured to adaptively cluster the set of feature data based on whether sounds represented in the audio data originate from a source closer to the first microphone than to the second microphone.

[0179]

[0184] Example 70. The device of any one of Examples 62 to 69, wherein the one or more processors are further configured to update the clustering behavior of the dynamic classifier based on the audio data.

[0180]

[0185] Example 71. The device of any one of Examples 62 to 70, wherein the one or more processors are further configured to update the classification decision criteria of the dynamic classifier.

[0181]

[0186] Example 72. The device of any one of Examples 62 to 71, wherein the dynamic classifier includes a self-organizing map.

[0182]

[0187] Example 73. The device of any one of Examples 62 to 72, wherein the dynamic classifier is further configured to receive a sequence of sets of audio data and adaptively cluster the set of sequences based at least in part on previous sets of audio data in the sequence.

[0183]

[0188] Example 74. The device of any one of Examples 62 to 73, wherein the one or more processors are further configured to initiate a voice command processing operation in response to determining that the audio data corresponds to user voice activity.

[0184]

[0189] Example 75. The device of Example 74, wherein the one or more processors are configured to generate at least one of a wake-up signal or an interrupt to initiate a voice command processing operation.

[0185]

[0190] Example 76. The device of Example 75, wherein the one or more processors further include an always-on power domain including the dynamic classifier and a second power domain including the voice command processing unit, and wherein the wake-up signal is configured to transition the second power domain from the low power mode to activate the voice command processing unit.

[0186]

[0191] Example 77. The device of any one of Examples 62 to 76, further comprising a modem coupled to the one or more processors, the modem configured to transmit the audio data to the second device in response to a determination, based on the dynamic classifier, that the audio data corresponds to user voice activity.

[0187]

[0192] Example 78. The device of any one of Examples 62 to 77, wherein the one or more processors are integrated into a headset device including a first microphone and a second microphone, the headset device being configured, when worn by a user, to position the first microphone closer to the user's mouth than the second microphone to capture the user's speech with greater intensity and less delay at the first microphone compared to the second microphone.

[0188]

[0193] Example 79. The device of any one of Examples 62 to 77, wherein the one or more processors are integrated into at least one of a mobile phone, a tablet computing device, a wearable electronic device, a camera device, a virtual reality headset, or an augmented reality headset.

[0189]

[0194] Example 80. The device of any one of Examples 62 to 77, wherein the one or more processors are integrated into a vehicle, the vehicle further including a first microphone and a second microphone, the first microphone positioned to capture speech of an operator of the vehicle.

[0190]

[0195] Example 81. A method of voice activity detection, the method comprising: receiving, at one or more processors, audio data including first audio data corresponding to a first output of a first microphone and second audio data corresponding to a second output of a second microphone; providing, at the one or more processors, the audio data to a dynamic classifier to generate a classification output corresponding to the audio data; and determining, at the one or more processors, whether the audio data corresponds to user voice activity based at least in part on the classification output.

[0191]

[0196] Example 82. The method of Example 81, wherein the classification output is based on a gain difference between the first audio data and the second audio data, a phase difference between the first audio data and the second audio data, or a combination thereof.

[0192]

[0197] Example 83. The method of Example 81 or 82, wherein the dynamic classifier includes a self-organizing map.

[0193]

[0198] Example 84. The method of any one of Examples 81 to 83, wherein determining whether the audio data corresponds to user voice activity is further based on at least one of a sign or a magnitude of the at least one value of the feature data corresponding to the audio data.

[0194]

[0199] Example 85. The method of any one of Examples 81 to 84, further comprising initiating a voice command processing operation in response to determining that the audio data corresponds to user voice activity.

[0195]

[0200] Example 86. The method of Example 85, further comprising generating at least one of a wake-up signal or an interrupt to initiate a voice command processing operation.

[0196]

[0201] Example 87. The method of any one of Examples 81 to 86, further including transmitting the audio data to a second device in response to determining, based on the dynamic classifier, that the audio data corresponds to user voice activity.

[0197]

[0202] Example 88. A non-transitory computer-readable medium including instructions that, when executed by one or more processors, cause the one or more processors to receive audio data including first audio data corresponding to a first output of a first microphone and second audio data corresponding to a second output of a second microphone; provide the audio data to a dynamic classifier to generate a classification output corresponding to the audio data; and determine, based at least in part on the classification output, whether the audio data corresponds to a user voice activity.

[0198]

[0203] Example 89. The non-transitory computer-readable medium of Example 88, wherein the classification output is based on a gain difference between the first audio data and the second audio data, a phase difference between the first audio data and the second audio data, or a combination thereof.

[0199]

[0204] Example 90. An apparatus including: means for receiving audio data including first audio data corresponding to a first output of a first microphone and second audio data corresponding to a second output of a second microphone; means, in a dynamic classifier, for generating a classification output corresponding to the audio data; and means for determining, based at least in part on the classification output, whether the audio data corresponds to user voice activity.

[0200]

[0205] Example 91. The apparatus of example 90, wherein the classification output is based on a gain difference between the first audio data and the second audio data, a phase difference between the first audio data and the second audio data, or a combination thereof.

[0201]

[0206] Those skilled in the art will further appreciate that the various illustrative logical blocks, configurations, modules, circuits, and algorithm steps described in connection with the implementations disclosed herein may be implemented as electronic hardware, computer software executed by a processor, or a combination of both. Various illustrative components, blocks, configurations, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or processor-executable instructions depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, and such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.

[0202]

[0207] The steps of a method or algorithm described in connection with the implementations disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, a hard disk, a removable disk, a compact disk read-only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium may be integral to the processor. The processor and the storage medium may reside in an application-specific integrated circuit (ASIC). The ASIC may reside in a computing device or a user terminal. Alternatively, the processor and the storage medium may reside as discrete components in a computing device or a user terminal.

[0203]

[0208] The above description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the disclosed embodiments. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other embodiments without departing from the scope of the present disclosure. Thus, the present disclosure is not intended to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features defined by the following claims. The inventions described in the claims of the present application as originally filed are set forth below. [C1] a memory configured to store instructions; one or more processors, wherein the one or more processors: receiving audio data including first audio data corresponding to a first output of a first microphone and second audio data corresponding to a second output of a second microphone; providing the audio data to a dynamic classifier, the dynamic classifier configured to generate a classification output corresponding to the audio data; determining whether the audio data corresponds to a user voice activity based at least in part on the classification output; and a device configured to execute the instructions to perform the steps of: [C2] The device of C1, further comprising the first microphone and the second microphone, wherein the first microphone is coupled to the one or more processors and configured to capture a user's speech, and the second microphone is coupled to the one or more processors and configured to capture ambient sound. [C3] The device of C1, wherein the classification output is based on a gain difference between the first audio data and the second audio data, a phase difference between the first audio data and the second audio data, or a combination thereof. [C4] The device of C1, wherein the one or more processors are further configured to generate feature data based on the first audio data and the second audio data, wherein the audio data is provided to the dynamic classifier as the feature data, and the classification output is based on the feature data. [C5] The feature data is at least one interaural phase difference between the first audio data and the second audio data; at least one interaural intensity difference between the first audio data and the second audio data; 2. The device of claim 1, further comprising: [C6] The device of C4, wherein the one or more processors are configured to determine whether the audio data corresponds to the user voice activity further based on at least one of a sign or magnitude of at least one value of the feature data. [C7] 3. The device of claim 2, wherein the one or more processors are further configured to transform the first audio data and the second audio data to a transform domain before generating the feature data, the feature data including an interaural phase difference for a plurality of frequencies and an interaural intensity difference for a plurality of frequencies. [C8] The device of C4, wherein the dynamic classifier is configured to adaptively cluster a set of feature data based on whether a sound represented in the audio data originates from a source closer to the first microphone than to the second microphone. [C9] The device of C1, wherein the one or more processors are further configured to update a clustering operation of the dynamic classifier based on the audio data. [C10] The device of C1, wherein the one or more processors are further configured to update classification decision criteria of the dynamic classifier. [C11] The device of C1, wherein the dynamic classifier comprises a self-organizing map. [C12] The device of C1, wherein the dynamic classifier is further configured to receive a sequence of sets of audio data and adaptively cluster the set of sequences based at least in part on previous sets of audio data in the sequence. [C13] The device of C1, wherein the one or more processors are further configured to initiate a voice command processing operation in response to determining that the audio data corresponds to the user voice activity. [C14] The device of C13, wherein the one or more processors are configured to generate at least one of a wake-up signal or an interrupt to initiate the voice command processing operation. [C15] the one or more processors: an always-on power domain including the dynamic classifier; a second power domain including a voice command processing unit, wherein the wake-up signal is configured to transition the second power domain from a low power mode to activate the voice command processing unit. The device of C14, further comprising: [C16] The device of C1, further comprising a modem coupled to the one or more processors, the modem configured to transmit the audio data to a second device in response to a determination, based on the dynamic classifier, that the audio data corresponds to the user voice activity. [C17] The device of C1, wherein the one or more processors are integrated into a headset device including the first microphone and the second microphone, and the headset device is configured, when worn by a user, to position the first microphone closer to the user's mouth than the second microphone so as to capture the user's speech with greater intensity and less delay at the first microphone compared to the second microphone. [C18] The device of C1, wherein the one or more processors are integrated into at least one of a mobile phone, a tablet computing device, a wearable electronic device, a camera device, a virtual reality headset, or an augmented reality headset. [C19] The device described in C1, wherein the one or more processors are integrated into a vehicle, the vehicle further including the first microphone and the second microphone, the first microphone positioned to capture speech of an operator of the vehicle. [C20] 1. A method of voice activity detection, comprising: receiving, at one or more processors, audio data including first audio data corresponding to a first output of a first microphone and second audio data corresponding to a second output of a second microphone; providing, in the one or more processors, the audio data to a dynamic classifier to generate a classification output corresponding to the audio data; determining, at the one or more processors, whether the audio data corresponds to a user voice activity based at least in part on the classification output; A method comprising: [C21] The method of C20, wherein the classification output is based on a gain difference between the first audio data and the second audio data, a phase difference between the first audio data and the second audio data, or a combination thereof. [C22] The method of C20, wherein the dynamic classifier comprises a self-organizing map. [C23] The method of C20, wherein determining whether the audio data corresponds to the user voice activity is further based on at least one of a sign or a magnitude of at least one value of feature data corresponding to the audio data. [C24] The method of C20, further comprising initiating a voice command processing operation in response to determining that the audio data corresponds to the user voice activity. [C25] The method of C24, further comprising generating at least one of a wake-up signal or an interrupt to initiate the voice command processing operation. [C26] The method of C20, further comprising transmitting the audio data to a second device in response to a determination, based on the dynamic classifier, that the audio data corresponds to the user voice activity. [C27] A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to: receiving audio data including first audio data corresponding to a first output of a first microphone and second audio data corresponding to a second output of a second microphone; providing the audio data to a dynamic classifier to generate a classification output corresponding to the audio data; determining whether the audio data corresponds to a user voice activity based at least in part on the classification output; and A non-transitory computer-readable medium for causing [C28] The non-transitory computer-readable medium of C27, wherein the classification output is based on a gain difference between the first audio data and the second audio data, a phase difference between the first audio data and the second audio data, or a combination thereof. [C29] means for receiving audio data including first audio data corresponding to a first output of a first microphone and second audio data corresponding to a second output of a second microphone; means in a dynamic classifier for generating a classification output corresponding to the audio data; means for determining whether the audio data corresponds to a user voice activity based at least in part on the classification output; and An apparatus comprising: [C30] The apparatus of C29, wherein the classification output is based on a gain difference between the first audio data and the second audio data, a phase difference between the first audio data and the second audio data, or a combination thereof.

Claims

1. a memory configured to store instructions; one or more processors, wherein the one or more processors: receiving audio data including first audio data corresponding to a first output of a first microphone and second audio data corresponding to a second output of a second microphone; providing the audio data to a dynamic classifier, the dynamic classifier configured to generate a classification output corresponding to the audio data; determining whether the audio data corresponds to self-voice activity or other sound activity based at least in part on the classification output and a verification input, wherein the verification input is based on at least one previous verification criterion, the previous verification criterion being based on a magnitude of an intensity difference between the first audio data and the second audio data and a sign of a phase difference between the first audio data and the second audio data; a device configured to execute the instructions to perform the steps of:

2. 10. The device of claim 1, further comprising: the first microphone and the second microphone, the first microphone coupled to the one or more processors and configured to capture user speech; and the second microphone coupled to the one or more processors and configured to capture ambient sounds, the speech corresponding to the self-voice activity.

3. 10. The device of claim 1, wherein the classification output is based on a gain difference between the first audio data and the second audio data, a phase difference between the first audio data and the second audio data, or a combination thereof.

4. The one or more processors are further configured to generate feature data based on the first audio data and the second audio data, wherein the audio data is provided to the dynamic classifier as the feature data, and the classification output is based at least in part on the feature data, and the feature data comprises: at least one interaural phase difference between the first audio data and the second audio data; at least one interaural intensity difference between the first audio data and the second audio data; The device of claim 1 .

5. 5. The device of claim 4, wherein the one or more processors are further configured to transform the first audio data and the second audio data into a transform domain before generating the feature data, the feature data including interaural phase differences for a plurality of frequencies and interaural intensity differences for a plurality of frequencies.

6. 5. The device of claim 4, wherein the dynamic classifier is configured to adaptively cluster sets of feature data based on whether sounds represented in the audio data originate from a source closer to the first microphone than to the second microphone.

7. the one or more processors: a clustering operation of the dynamic classifier based on the audio data; or Classification decision criteria of the dynamic classifier The device of claim 1 , further configured to update:

8. The device of claim 1 , wherein the dynamic classifier comprises a self-organizing map.

9. 10. The device of claim 1, wherein the dynamic classifier is further configured to receive a sequence of sets of audio data and adaptively cluster the set of sequences based at least in part on previous sets of audio data within the sequence.

10. The device of claim 1 , wherein the one or more processors are further configured to initiate a voice command processing operation in response to determining that the audio data corresponds to the self-voice activity.

11. The one or more processors are configured to generate at least one of a wake-up signal or an interrupt to initiate the voice command processing operation, and the one or more processors preferably include: an always-on power domain including the dynamic classifier; a second power domain including a voice command processing unit, wherein the wake-up signal is configured to transition the second power domain from a low power mode to activate the voice command processing unit; The device of claim 10 further comprising:

12. 2. The device of claim 1, wherein the one or more processors are integrated into a headset device including the first microphone and the second microphone, and the headset device is configured, when worn by a user, to position the first microphone closer to the user's mouth than the second microphone so as to capture the user's speech with greater intensity and less delay at the first microphone compared to the second microphone.

13. the one or more processors are integrated into at least one of a mobile phone, a tablet computing device, a wearable electronic device, a camera device, a virtual reality headset, an augmented reality headset, or a vehicle that includes the first microphone and the second microphone; 2. The device of claim 1, wherein the first microphone is positioned to capture speech of an operator of the vehicle, the speech corresponding to the self-voice activity, and the audio data other than the speech corresponds to the other sound activity in the vehicle.

14. 1. A method of voice activity detection, comprising: receiving, at one or more processors, audio data including first audio data corresponding to a first output of a first microphone and second audio data corresponding to a second output of a second microphone; providing, in the one or more processors, the audio data to a dynamic classifier to generate a classification output corresponding to the audio data; determining, in the one or more processors, whether the audio data corresponds to self-voice activity or other sound activity based at least in part on the classification output and a verification input, wherein the verification input is based on at least one previous verification criterion, the previous verification criterion being based on a magnitude of an intensity difference between the first audio data and the second audio data and a sign of a phase difference between the first audio data and the second audio data; A method comprising:

15. 15. A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the method of claim 14.

Citation Information

Patent Citations

  • Hearing device with own-voice detection and related method

    JP2020102836A

  • Low-power voice command detector

    US20160267908A1

  • Peer to peer hearing system

    US20160360326A1

  • Hearing device comprising an own voice detector

    US20180146307A1

  • Voice activity detection for communication headset

    US20180350394A1