Digital signal processor-based continuous conversation
An always-on DSP in voice-enabled environments uses VAD and speaker verification to detect subsequent queries from the same user, reducing power consumption and maintaining a low-power state by avoiding the need for hot words.
Patent Information
- Application Number
- JP2024522103
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-10-13
- Filing Date
- 2021-12-15
- Publication Date
- 2026-02-19
- Estimated Expiration
- 2041-12-15
AI Technical Summary
Existing voice-enabled environments require users to speak a hot word for each subsequent query or command, which is cumbersome and inefficient in terms of computational resources and battery power consumption.
Implementing an always-on digital signal processor (DSP) to operate in a subsequent-query detection mode, using voice activity detection (VAD) and speaker verification to identify subsequent queries from the same user without requiring hot words, while maintaining a low-power state.
Reduces computational resource and battery power consumption by only activating a second processor for subsequent queries spoken by the same user, while allowing seamless detection of subsequent queries without hot words.
Smart Images

Figure 0007818079000001 
Figure 0007818079000002 
Figure 0007818079000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to digital signal processor-based continuous conversation. [Background technology]
[0002] A voice-enabled environment (e.g., a home, a workplace, a school, an automobile, etc.) allows a user to speak queries or commands aloud to a computer-based system, which addresses and responds to the query and / or performs a function based on the command. The voice-enabled environment may be implemented using a network of connected microphone devices distributed throughout various rooms or areas of the environment. The device may operate in a sleep state and, in response to detecting a hot word spoken by the user prior to the utterance, initiate a wake-up process to perform voice recognition on the utterance directed to the system. After receiving an answer addressed by the computer-based system, the user may speak a subsequent query or command. Requiring the user to speak a hot word for each subsequent query or command is not only unnatural but also cumbersome for the user. Summary of the Invention [Means for solving the problem]
[0003] One aspect of the disclosure provides a computer-implemented device that, when executed on data processing hardware of an assistant-enabled device, causes the data processing hardware to perform an operation, the operation including instructing an always-on first processor of the data processing hardware to operate in a subsequent-query detection mode in response to receiving a response to an initial query submitted to a digital assistant by a user of the assistant-enabled device. While the always-on first processor operates in the subsequent query detection mode, the operations also include receiving subsequent audio data captured by the assistant-enabled device in the environment of the assistant-enabled device at the always-on first processor; using a voice activity detection (VAD) model running on the always-on first processor to determine whether the VAD model detects voice activity in the subsequent audio data; and performing speaker verification on the subsequent audio data using a speaker identification (SID) model running on the always-on first processor to determine whether the subsequent audio data includes an utterance spoken by the same user who submitted the initial query to the digital assistant. When the VAD model detects voice activity in the subsequent audio data and the subsequent audio data includes an utterance spoken by the same user who submitted the initial query, Initiating a wake-up process in a second processor of the data processing hardware to determine whether the utterance includes a subsequent query directed to the digital assistant.
[0004] Implementations of the present disclosure may include one or more of the following optional features: In some implementations, instructing the always-on first processor of the data processing hardware to operate in the subsequent query detection mode causes the always-on first processor to begin executing a VAD and a SID model on the always-on first processor while operating in the subsequent query detection mode. In additional implementations, instructing the always-on first processor of the data processing hardware to operate in the subsequent query detection mode causes the always-on first processor to disable the hot word detection model while operating in the subsequent query detection mode. The SID model may include a text-independent SID model configured to extract a text-independent speaker identification vector from the subsequent audio data.
[0005] In some examples, initiating a wake-up process in the second processor causes the second processor to perform operations including: processing the subsequent audio data to generate a transcription of an utterance spoken by the same user who submitted the initial query; and performing query interpretation on the transcription to determine whether the utterance includes a subsequent query directed to the digital assistant. When the utterance includes a subsequent query directed to the digital assistant, the operations may further include instructing the digital assistant to perform an operation specified by the subsequent query, receiving a subsequent response from the digital assistant indicating the execution of the operation specified by the subsequent query, and presenting the subsequent response to output from the assistant-enabled device.
[0006] In other examples, initiating a wake-up process in the second processor causes the second processor to transmit subsequent audio data to a remote server over a network. In these examples, when the subsequent audio data is received by the remote server, the remote server performs operations including: processing the subsequent audio data to generate a transcription of an utterance spoken by the same user who submitted the initial query; performing query interpretation on the transcription to determine whether the utterance includes a subsequent query directed to the digital assistant; and, when the utterance includes a subsequent query directed to the digital assistant, instructing the digital assistant to perform an operation specified by the subsequent query. Further, in these examples, after instructing the digital assistant to perform the operation specified by the subsequent query, the operation may further include receiving a subsequent response from the digital assistant indicating the execution of the operation specified by the subsequent query, and presenting the subsequent response to output from the assistant-enabled device.
[0007] In some implementations, the operations also include, after receiving initial audio data corresponding to the initial query spoken by the user and submitted to the digital assistant, extracting from the initial audio data a first speaker identification vector that represents the characteristics of the initial query spoken by the user. In these implementations, performing speaker verification on the subsequent audio data includes extracting from the subsequent audio data a second speaker identification vector that represents the characteristics of the subsequent audio data using the SID model, and determining that the subsequent audio data contains speech spoken by the same user who submitted the initial query to the digital assistant when the first speaker identification vector matches the second speaker identification vector.
[0008] In some additional implementations, the operations also include, after receiving initial audio data corresponding to an initial query spoken by a user and submitted to the digital assistant, using the SID model to extract from the subsequent audio data a second speaker identification vector that represents a characteristic of the subsequent audio data; determining that the subsequent audio data contains speech spoken by the same user who submitted the initial query to the digital assistant when the first speaker identification vector matches the second speaker identification vector; and identifying the user who spoke the initial query as the respective enrolled user associated with one of the enrolled speaker vectors that matches the first speaker identification vector. In these additional implementations, performing speaker verification on the subsequent audio data includes using the SID model to extract from the subsequent audio data a second speaker identification vector that represents a characteristic of the subsequent audio data, and when the second speaker identification vector matches one of the enrolled speaker vectors associated with each enrolled user who spoke the initial query, determining that the subsequent audio data contains utterances spoken by the same user who submitted the initial query to the digital assistant.
[0009] In some examples, the operations further include instructing the always-on first processor to stop operating in the subsequent query detection mode and start operating in the hotword detection mode in response to determining at least one of the VAD model failing to detect voice activity in the subsequent audio data or the subsequent audio data failing to contain utterances spoken by the same user who submitted the initial query. Here, by instructing the always-on first processor to start operating in the hotword detection mode, the always-on first processor starts execution of the hotword detection model on the always-on first processor. Additionally or alternatively, by instructing the always-on first processor to stop operating in the subsequent query detection mode, the always-on first processor may disable or deactivate execution of the VAD and SID models on the always-on first processor.
[0010] The always-on first processor may include a digital signal processor, and the second processor may include an application processor. The Assistant-enabled device may include a battery-powered device in communication with one or more microphones configured to capture subsequent audio data and initial audio data corresponding to the initial query.
[0011] Another aspect of the present disclosure provides a system including data processing hardware and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform an operation, the operation including instructing an always-on first processor of the data processing hardware to operate in a subsequent-query detection mode in response to receiving a response to an initial query submitted to a digital assistant by a user of an assistant-enabled device. While the always-on first processor operates in the subsequent query detection mode, the operations also include receiving subsequent audio data captured by the assistant-enabled device in the environment of the assistant-enabled device at the always-on first processor; using a voice activity detection (VAD) model running on the always-on first processor to determine whether the VAD model detects voice activity in the subsequent audio data; and performing speaker verification on the subsequent audio data using a speaker identification (SID) model running on the always-on first processor to determine whether the subsequent audio data includes an utterance spoken by the same user who submitted the initial query to the digital assistant. When the VAD model detects voice activity in the subsequent audio data and the subsequent audio data includes an utterance spoken by the same user who submitted the initial query, Initiating a wake-up process in a second processor of the data processing hardware to determine whether the utterance includes a subsequent query directed to the digital assistant.
[0012] This aspect may include one or more of the following optional features: In some implementations, instructing the always-on first processor of the data processing hardware to operate in the subsequent query detection mode causes the always-on first processor to begin executing the VAD and SID models on the always-on first processor while operating in the subsequent query detection mode. In additional implementations, instructing the always-on first processor of the data processing hardware to operate in the subsequent query detection mode causes the always-on first processor to disable the hot word detection model while operating in the subsequent query detection mode. The SID model may include a text-independent SID model configured to extract a text-independent speaker identification vector from the subsequent audio data.
[0013] In some examples, initiating a wake-up process in the second processor causes the second processor to perform operations including: processing the subsequent audio data to generate a transcription of an utterance spoken by the same user who submitted the initial query; and performing query interpretation on the transcription to determine whether the utterance includes a subsequent query directed to the digital assistant. When the utterance includes a subsequent query directed to the digital assistant, the operations may further include instructing the digital assistant to perform an operation specified by the subsequent query, receiving a subsequent response from the digital assistant indicating the execution of the operation specified by the subsequent query, and presenting the subsequent response to output from the assistant-enabled device.
[0014] In other examples, initiating a wake-up process in the second processor causes the second processor to transmit subsequent audio data to a remote server over a network. In these examples, when the subsequent audio data is received by the remote server, the remote server performs operations including: processing the subsequent audio data to generate a transcription of an utterance spoken by the same user who submitted the initial query; performing query interpretation on the transcription to determine whether the utterance includes a subsequent query directed to the digital assistant; and, when the utterance includes a subsequent query directed to the digital assistant, instructing the digital assistant to perform an operation specified by the subsequent query. Further, in these examples, after instructing the digital assistant to perform the operation specified by the subsequent query, the operation may further include receiving a subsequent response from the digital assistant indicating the execution of the operation specified by the subsequent query, and presenting the subsequent response to output from the assistant-enabled device.
[0015] In some implementations, the operations also include, after receiving initial audio data corresponding to the initial query spoken by the user and submitted to the digital assistant, extracting from the initial audio data a first speaker identification vector that represents the characteristics of the initial query spoken by the user. In these implementations, performing speaker verification on the subsequent audio data includes extracting from the subsequent audio data a second speaker identification vector that represents the characteristics of the subsequent audio data using the SID model, and determining that the subsequent audio data contains speech spoken by the same user who submitted the initial query to the digital assistant when the first speaker identification vector matches the second speaker identification vector.
[0016] In some additional implementations, the operations also include, after receiving initial audio data corresponding to an initial query spoken by a user and submitted to the digital assistant, using the SID model to extract from the subsequent audio data a second speaker identification vector that represents a characteristic of the subsequent audio data; determining that the subsequent audio data contains speech spoken by the same user who submitted the initial query to the digital assistant when the first speaker identification vector matches the second speaker identification vector; and identifying the user who spoke the initial query as the respective enrolled user associated with one of the enrolled speaker vectors that matches the first speaker identification vector. In these additional implementations, performing speaker verification on the subsequent audio data includes using the SID model to extract from the subsequent audio data a second speaker identification vector that represents a characteristic of the subsequent audio data, and when the second speaker identification vector matches one of the enrolled speaker vectors associated with each enrolled user who spoke the initial query, determining that the subsequent audio data contains utterances spoken by the same user who submitted the initial query to the digital assistant.
[0017] In some examples, the operations further include instructing the always-on first processor to stop operating in the subsequent query detection mode and start operating in the hotword detection mode in response to determining at least one of the VAD model failing to detect voice activity in the subsequent audio data or the subsequent audio data failing to contain utterances spoken by the same user who submitted the initial query. Here, by instructing the always-on first processor to start operating in the hotword detection mode, the always-on first processor starts execution of the hotword detection model on the always-on first processor. Additionally or alternatively, by instructing the always-on first processor to stop operating in the subsequent query detection mode, the always-on first processor may disable or deactivate execution of the VAD and SID models on the always-on first processor.
[0018] The always-on first processor may include a digital signal processor, and the second processor may include an application processor. The Assistant-enabled device may include a battery-powered device in communication with one or more microphones configured to capture subsequent audio data and initial audio data corresponding to the initial query.
[0019] The details of one or more implementations of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]
[0020] [Figure 1A] FIG. 1 is a schematic diagram illustrating an example of an assistant-enabled device that detects ongoing conversations in a voice-enabled environment. [Figure 1B] FIG. 1 is a schematic diagram illustrating an example of an assistant-enabled device that detects ongoing conversations in a voice-enabled environment. [Figure 2]FIG. 1C is a schematic diagram of a digital signal processor running on the Assistant-enabled device of FIGS. 1A and 1B. [Figure 3A] 3 is a schematic diagram of the digital signal processor of FIG. 2 determining that there are no subsequent events in a subsequent query. [Figure 3B] 3 is a schematic diagram of the digital signal processor of FIG. 2 determining that there are no subsequent events in a subsequent query. [Figure 4] FIG. 1 is a schematic diagram of a speaker verification process. [Figure 5] FIG. 1 is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein. [Figure 6] 1 is a flowchart of an exemplary arrangement of operations for a method for detecting ongoing conversation in a voice-enabled environment. DETAILED DESCRIPTION OF THE INVENTION
[0021] Like reference symbols in the various drawings indicate like elements.
[0022] Voice-based interfaces, such as digital assistants, are becoming increasingly prevalent in a variety of devices, including, but not limited to, mobile phones and smart speakers / displays that include microphones for capturing audio. A common way to initiate a voice interaction with an assistant-enabled device (AED) is to speak a fixed phrase, e.g., a hotword, which, when detected by the AED in streaming audio, triggers the AED to begin a wake-up process and to begin recording and processing subsequent audio to confirm the user's spoken query. Hotwords are therefore a critical component in the overall digital assistant interface stack, as they allow the user to wake up the AED from a low-power state to a high-power state, allowing the AED to perform more expensive processing, such as full automatic speech recognition (ASR) or server-based ASR.
[0023] Often, a user will make a follow-up query after receiving a response to their initial query. For example, a user may speak the initial query, "Who is the 39th President of the United States?" to an AED, and the AED may respond, "Jimmy Carter," to the user. The user may then instruct the AED on a subsequent query, "How old is he?" Here, the continuous conversation, or subsequent query, mode keeps the AED in a high-power state for a predefined time after providing a response, keeping the microphone open and performing ASR to process any audio and determine whether any subsequent audio is intended for the AED. After the predefined time has elapsed, the AED returns to a low-power state. However, keeping the AED in a high-power state for a predefined time consumes significant computational resources and battery power. Furthermore, while the AED is in a high-power state, audio not intended for the AED may be captured (and then deleted) during the predefined time. Thus, the AED advantageously returns to a low power state after providing a response to the initial query, but is still able to detect any subsequent queries directed to the AED that do not include the hotword.
[0024] Therefore, implementations herein are directed to a method for processing subsequent queries while maintaining a low power state. In particular, a user submits an initial query including a hot word and a query to an AED having an always-on digital signal processor (DSP) operating in a hot word detection mode. When the DSP operating in the hot word detection mode detects the hot word in the initial query, the DSP initiates a wake-up process on a second processor to perform more expensive processing, such as a full ASR, to provide a response to the initial query. The second processor may include an application processor (AP) or other type of system-on-chip (SoC) processor.
[0025] In response to receiving a response to the initial query, the AED instructs the always-on DSP to operate in a subsequent query detection mode to detect any subsequent queries from the user who submitted the initial query. Here, the DSP receives subsequent audio data and determines whether voice activity is present and whether the subsequent audio data contains utterances spoken by the same user who submitted the initial query. When the first processor detects voice activity and detects the same user who submitted the initial query, the first processor (i.e., the always-on DSP) initiates a wake-up process on the second processor to perform more expensive processing, such as full ASR or server-based ASR, to provide a response to the subsequent query. Thus, while the DSP is operating in the subsequent query detection mode, the AED consumes significantly less computational resources and battery power than if the second processor were active and performing ASR, while still having the ability to detect subsequent queries spoken by the same user that do not contain the hotword. The DSP triggers the second processor to wake up and begin ASR only if the DSP detects subsequent queries spoken by the same user. Otherwise, when no subsequent queries are spoken by the same user, the AED remains in a low-power state and transitions the DSP to operate in hot word detection mode after a predetermined time, triggering the wake-up process only if the presence of a hot word is detected in the streaming audio data.
[0026] 1A and 1B , in some implementations, an exemplary system 100 includes an AED (i.e., user device) 102 associated with one or more users 10 and communicating with a remote system 110 via a network 104. The AED 102 may correspond to a computing device such as a mobile phone, computer (laptop or desktop), tablet, smart speaker / display, smart appliance, smart headphones, wearable, vehicle infotainment system, etc., and includes data processing hardware 103 and memory hardware 105. The AED 102 includes or is in communication with one or more microphones 106 for capturing streaming audio 118 in the AED's 102's environment, which may include speech 119, 129 spoken by the respective users 10. The remote system 110 may be a single computer, multiple computers, or a distributed system (e.g., a cloud environment) with scalable / elastic computing resources 112 (e.g., data processing hardware) and / or storage resources 114 (e.g., memory hardware).
[0027] The data processing hardware 103 of the AED 102 includes a first processor 200 and a second processor 300. As used herein, the first processor 200 includes an always-on DSP 200 (also referred to as DSP 200) configured to detect the presence of hot words 116 and / or subsequent queries in the streaming audio 118 without performing semantic analysis or speech recognition processing on the streaming audio 118. The DSP 200 may receive the streaming audio 118 including acoustic features extracted by an acoustic feature extractor from utterances 119 spoken by the users 10, 10a. As used herein, user 10a refers to the user 10a who spoke the initial utterance 119. The DSP (i.e., first processor) 200 may operate in a hot word detection mode (FIG. 1A) 210 and a subsequent query detection mode (i.e., subsequent mode) (FIG. 1B) 220. In some examples, due to storage / memory / processing constraints, the DSP 200 operates in only one of the hotword detection mode or the subsequent query detection mode at a time. That is, in these examples, either the hotword detection mode 210 is enabled and the subsequent detection mode 220 is disabled, or the hotword detection mode 210 is disabled and the subsequent detection mode 220 is enabled. However, in other examples, operation of the DSP 200 in the hotword detection mode 210 remains enabled while the DSP 200 simultaneously operates in the subsequent query detection mode. As will become apparent, the always-on processor 200 may operate in the hotword detection mode 210 and / or the subsequent query detection mode 220 while the AED 102 is in a low power state and the second processor 300 is in a sleep state. The second processor may include an application processor (AP) or another type of system-on-chip (SoC) processor that consumes more power than the always-on DSP when the second processor is awake.
[0028] 1A , in some implementations, the always-on DSP 200 operates in a hotword detection mode 210 while the subsequent query detection mode 220 is disabled and the second processor 300 is asleep. While operating in the hotword detection mode 210, the DSP 200 is configured to detect the presence of the hotword 116 “OK Google” in the streaming audio 118 and initiate a wake-up process on the second processor 300 to process the hotword 116 and / or the initial query 117 following the hotword 116 in the streaming audio 118. In the illustrated example, the utterance 119 includes the hotword 116 “OK Google” followed by the initial query 117 “What’s the weather like today?” The AED 102 can extract acoustic features from the streaming audio 118 and store the extracted acoustic features in a buffer in the memory hardware 105 for use in detecting whether the streaming audio 118 includes the presence of the hotword 116. The DSP 200 may generate a probability score indicating the presence of the hotword 116 in the acoustic features of the streaming audio 118 captured by the AED 102 and execute a hotword detection model configured to detect the hotword 116 in the streaming audio 118 when the probability score meets a hotword detection threshold. The DSP 200 may include multiple hotword detection models, each trained to detect different hotwords associated with specific terms / phrases. These hotwords may be predefined hotwords and / or custom hotwords assigned by the user 10. In some implementations, the hotword detection model includes a trained neural network-based model received from the remote system 110 via the network 104.
[0029] In the illustrated example, the DSP 200 may determine that the utterance 119, "OK Google, what's the weather today?" includes the hot word 116, "OK Google," when the DSP 200 detects acoustic features in the streaming audio 118 that are characteristic of the hot word 116. For example, the DSP 200 may generate MFCCs from the audio data and detect that the utterance 119, "OK Google, what's the weather today?" includes the hot word 116, "OK Google," based on classifying the MFCCs as including MFCCs similar to MFCCs characteristic of the hot word "OK Google" stored in the hot word detection model. As another example, the DSP 200 may generate mel-scale filter bank energies from the audio data and detect that the mel-scale filter bank energies include mel-scale filter bank energies similar to mel-scale filter bank energies characteristic of the hot word "OK Google" stored in the hot word detection model.
[0030] In response to detecting the presence of the hotword 116 in the streaming audio 118 corresponding to the utterance 119 spoken by the user 10a, the DSP 200 provides audio data 120 characterizing a hotword event to initiate a wake-up process on the second processor 300 to confirm the presence of the hotword 116. The audio data (interchangeably referred to as a hotword event) 120 includes a first portion 121 characterizing the hotword 116 and a second portion 122 characterizing the initial query 117. The second processor 300 can then execute a more robust hotword detection model to confirm whether the audio data 120 contains the hotword 116. Additionally or alternatively, the second processor 300 can perform speech recognition on the audio data 120 via an automatic speech recognition (ASR) model 310 to confirm whether the audio data 120 contains the hotword.
[0031] If the second processor 300 determines that the audio data 120 includes the hot word 116, it may execute the ASR model 310 to process the audio data 120 and generate a speech recognition result 315, and execute the natural language understanding (NLU) module 320 to perform semantic interpretation on the speech recognition result and determine that the audio data 120 includes an initial query 117 for the digital assistant 109 to perform an action. In this example, the ASR model 310 may process the audio data 120 to generate a speech recognition result 315 for "What's the weather like today?", and the NLU module 320 may identify the speech recognition result 315 as the initial query 117 for the digital assistant 109 to perform an action to fetch a response 192 (i.e., an answer to the initial query 117) indicating "It looks like it's going to be sunny today with a high of 76 degrees." The digital assistant 109 may provide the response 192 in response to the output from the AED 102. For example, the digital assistant 109 may audibly output the response 192 from the AED 102 as synthesized speech and / or display a text representation of the response 192 on the screen of the AED 102.
[0032] In some implementations, the ASR model 310 and the NLU module 320 are located on the remote system 110 in addition to or instead of the AED 102. When the DSP 200 triggers the wake-up of the second processor 300 in response to detecting the hot word 116 in the utterance 119, the second processor may transmit audio data 120 corresponding to the utterance 119 to the remote system 110 over the network 104. The AED 102 may transmit a first portion 121 of the audio data 120, including the hot word 116, for the remote system 110 to perform speech recognition via the ASR model 310 to confirm the presence of the hot word 116. Alternatively, the AED 102 may transmit only a second portion 122 of the audio data 120, corresponding to the first query 117 spoken in the utterance 119 after the hot word 116, to the remote system 110. The remote system 110 executes the ASR model 310 to generate speech recognition results 315 for the audio data 120. The remote system 110 may also execute an NLU module 320 to perform semantic interpretation on the speech recognition results 315 and identify an initial query 117 for the digital assistant 109 to perform an action. Alternatively, the remote system 110 may send the speech recognition results 315 to the AED 102, which executes the NLU module 320 to identify the initial query 117.
[0033] The digital assistant 109 may be located in the remote system 110 and / or the AED 102. The digital assistant 109 is configured to perform an operation specified by the initial query 117 from the second processor 300. In some examples, the digital assistant 109 accesses a search engine to fetch a response 192 associated with the initial query 117. In other examples, the digital assistant 109 accesses the memory hardware 105 of the AED 102 and / or the memory hardware 114 of the remote system to fetch the response 192 associated with the initial query 117. Alternatively, the digital assistant 109 performs an operation associated with the initial query 117 (i.e., "call mom").
[0034] In some implementations, the user 10a may have a subsequent query based on the response 192 to the initial query 117. That is, after receiving the response 192 "It looks like it's going to be sunny today with a high of 76 degrees," the user 10a may have a subsequent query asking about tomorrow's weather. Current techniques can provide subsequent queries without the user having to speak the hot word again by keeping the second processor 300 awake and processing all audio data following the initial query 117 for a predetermined period of time before returning to sleep. That is, the second processor 300 can continuously perform speech recognition and / or semantic interpretation on all subsequent audio data to determine whether the user has submitted a subsequent query for the digital assistant 109 to perform an action. While this technique is effective for recognizing subsequent queries spoken in utterances that do not include hot words, it is computationally expensive and consumes battery power because the second processor 300 must remain awake and continuously process all input audio data. In particular, the computational cost and battery power consumption is particularly wasteful in most cases where the user 10a does not submit subsequent queries.
[0035] In some implementations, the second processor 300 determines a first identified speaker vector 411. The first identified speaker vector 411 represents the voice characteristics of the user 10a who spoke the utterance 119. The AED 102 can store the first identified speaker vector 411 in the memory hardware 105, and / or the remote system 100 can store the first identified speaker vector 411 in the storage resource 114. Thereafter, as will become apparent, the first identified speaker vector 411 can be obtained to identify whether the user who spoke the subsequent query is the same user 10a who spoke the initial query 117.
[0036] 1B , in response to receiving a response 192 to the initial query 117 submitted by the user 10a, the digital assistant 109 instructs the DSP 200 to operate in a subsequent query detection mode 220. Here, by instructing the DSP 200 to operate in the subsequent detection mode 220, the hotword detection mode 210 is disabled and / or the second processor (e.g., AP) 300 may return to a sleep state. The second processor 300 may automatically return to a sleep state when the response 192 is output from the AED 102. While operating in the subsequent detection mode 220, the DSP 200 receives subsequent audio data 127 corresponding to a subsequent query 129 spoken by the user 10a and captured by the AED 102. In the illustrated example, the user 10a speaks the subsequent query 129, "How about tomorrow?" in response to the AED 102 providing the user 10a with a response 192, "It's going to be sunny today with a high of 76 degrees." Notably, the user 10a simply speaks the subsequent query 129 without speaking the hot word 116 twice, as the user 10a did when speaking the initial utterance 119 in FIG. 1A. However, in some other examples, the subsequent query 129 includes the hot word 116.
[0037] 1B and 2, during operation in the subsequent query detection mode 220, the DSP 200 executes a voice activity detection (VAD) model 222 and a speaker verification process 400. The VAD model 222 may be a model trained to determine whether voice activity is present in the subsequent audio data 127. That is, the VAD model 222 determines whether the subsequent audio data 127 contains voice activity, such as human speech, or non-voice activity audio (i.e., stereo, speaker, background noise, etc.). The VAD model 222 may be trained to output a voice activity score 224 indicating the likelihood that the subsequent audio data 127 contains voice activity. Here, the DSP 200 may determine that the subsequent audio data 127 contains voice activity when the voice activity score 224 meets a voice activity threshold. In some examples, the VAD model 222 outputs a binary voice activity indication, with a "1" indicating "Yes" and a "0" indicating "No," indicating that the subsequent audio data 127 contains ("Yes") or does not contain ("No") voice activity. The VAD model 222 can be trained to distinguish between human speech and synthetic / synthesized speech. The DSP 200 can be configured to operate as a state machine based on the activity of any model, such as the VAD model 222, the speaker verification process 400, or a hot word detection model. For example, while the DSP is in state 0, models A, B, and C are active, and when model A is triggered, the DSP may transition to state 1, where models B and D become active. The state machine can learn based on model outputs, user feedback, or pre-programmed.
[0038] The DSP 200 executes the speaker verification process 400 to determine a verification result 450 indicating whether the subsequent audio data 127 contains utterances spoken by the same user 10 a who submitted the initial query 117. In some examples, to conserve computation and battery power, execution of the speaker verification process 400 is conditional on the VAD model 222 first detecting voice activity. When the VAD model 222 detects voice activity in the subsequent audio data 127 (i.e., based on the voice activity score 224) and the verification result 450 output by the speaker verification process 400 determines that the subsequent audio data 127 contains utterances spoken by the same user 10 a who submitted the initial query 117, the DSP 200 provides a subsequent event 215 to the second processor 300 configured to wake the second processor 300 from a sleep state. Here, the subsequent event includes subsequent audio data 127, which causes the second processor 300 (e.g., application processor / CPU) to wake up and determine whether the subsequent audio data 127 includes a subsequent query 129 that specifies a subsequent action for the digital assistant 109 to perform.
[0039] Continuing with the example, while operating in subsequent query detection mode 220, DSP 200 determines that subsequent audio data 127 corresponding to subsequent query 129 contains voice activity and includes speech spoken by the same user 10a who submitted initial query 117. Accordingly, DSP 200 provides to second processor 300 a subsequent event 215 configured to wake second processor 300 from a sleep state and the subsequent audio data 127. Here, the determination by subsequent query detection mode 220 that voice activity is present in the subsequent audio data 127 and indicates that the subsequent audio data 127 was spoken by the same speaker who spoke initial query 117 provides that the subsequent audio data 127 likely includes the subsequent query 129. Thus, by utilizing the DSP 200, which consumes less power and computational resources than operating the second processor 300, the DSP 200 can act as a "gatekeeper" that triggers the output of a subsequent event 215 to wake the second processor 300 from a sleep state only if the subsequent audio data 127 contains voice activity and is likely spoken by the same user 10a who spoke the initial query 117. Otherwise, the second processor 300 is allowed to operate in a sleep state after the response 192 to the initial query 117 is output to the user 10a.
[0040] In response to the subsequent event 215 waking up the second processor 300, the second processor 300 can execute the ASR model 310 to generate a speech recognition result 325 for the subsequent audio data 127. The second processor 300 can also execute the NLU module 320 to perform semantic interpretation on the speech recognition result 325 and identify a subsequent query 129 to submit to the digital assistant 109. In the illustrated example, the second processor 300 determines that the subsequent audio data 127 corresponds to the subsequent query 129, "How's tomorrow?", and the assistant 109 obtains a subsequent response 193 to the subsequent query 129. Here, the second processor 300 provides the subsequent response 193 from the assistant 109 as output from the AED 102 in the form of synthesized speech and / or text indicating, "It's going to rain tomorrow and have a high of 68 degrees."
[0041] Alternatively, the second processor 300 may determine that a subsequent event 215 from the always-on DSP 200 indicates that the subsequent audio data 127 contains voice activity from the same user 10a who submitted the initial query 117, but that the subsequent audio data 127 is not directed to the AED 102 and therefore does not have a subsequent query 129 for the digital assistant 109 to execute. In this scenario, the second processor 300 may provide an indication 307 indicating that the subsequent audio data 127 does not have a subsequent query and return to a sleep state or remain awake for a predetermined period of time. In some examples, the second processor 300 may prompt the user 10a to repeat the subsequent query 129 in response to the indication 307 that the subsequent audio data 127 does not have a subsequent query 129.
[0042] As described above with reference to FIG. 1A , the ASR model 310 and the NLU module 320 may be located on the remote system 110 in addition to or instead of the AED 102. When the DSP 200 triggers the wake-up of the second processor 300 in response to detecting the subsequent event 215, the second processor 300 may transmit the subsequent audio data 127 to the remote system 110 over the network 104. The remote system 110 can then execute the ASR model 310 to generate speech recognition results 325 for the audio data 120. The remote system 110 can also execute the NLU module 320 to perform semantic interpretation on the speech recognition results 325 and identify a subsequent query 129 for the digital assistant 109 to perform an action. Alternatively, the remote system 110 may send the speech recognition results 325 to the AED 102, which then executes the NLU module 320 to identify the subsequent query 129.
[0043] 3A and 3B illustrate an audio environment 300 in which the DSP 200 determines that no subsequent event 215 exists in the audio data received after the initial query 117 spoken by the user 10a. Referring now to FIGS. 2 and 3A, in the audio environment 300, 300b after receiving a response to the initial query submitted by the user 10a, the digital assistant instructs the DSP to operate in a subsequent query detection mode 220. Here, by instructing the DSP 200 to operate in the subsequent detection mode 220, the hotword detection mode 210 may be disabled and / or the second processor 300 may return to a sleep state. In the illustrated example, while operating in the subsequent detection mode 220, the DSP 200 receives subsequent audio data 127 corresponding to a subsequent query 129, "How about next week?", spoken by another user 10, 10b, captured by the AED 102. In particular, the different user 10b who spoke the subsequent query 129 is different from the user 10a who spoke the initial query (FIG. 1A). Now, during operation of the subsequent query detection mode 220, the DSP 200 executes the VAD model 222 and determines that the subsequent audio data 127 contains voice activity. Even though the voice activity is from a different user 10b, the VAD model 222 indicates that voice activity is present in the subsequent audio data 127. Thus, the VAD 222 outputs a voice activity indication indicating "Yes." The VAD model 222 detecting voice activity in the subsequent data 127 can reset or increase the timeout period for operating in the subsequent detection mode 220 to perform the speaker verification process 400.
[0044] The DSP 200 also executes the speaker verification process 400 to determine a verification result 450 indicating whether the subsequent audio data 127 contains utterances spoken by the same user 10a who submitted the initial query 117 (FIG. 1A). Here, the speaker verification process 400 determines that the user 10b who spoke the subsequent query is not the same user as the user 10a who spoke the initial query. Therefore, the DSP 200 does not communicate the subsequent event 215 to the second processor 300, and therefore, the second processor 300 remains in a sleep state. Furthermore, if the DSP 200 does not detect the subsequent event 215 within a predetermined time, the DSP 200 returns to the hotword detection mode 210.
[0045] 2 and 3B, in the voice environment 300, 300a after receiving a response to the initial query submitted by the user 10a, the digital assistant instructs the DSP to operate in the subsequent query detection mode 220. Here, by instructing the DSP 200 to operate in the subsequent detection mode 220, the hotword detection mode 210 is disabled and / or the second processor 300 may return to a sleep state. In the illustrated example, while operating in the subsequent detection mode 220, the DSP 200 receives audio 132 from the audio source 130. The audio source 130 may be a television, a stereo, a speaker, or any other audio source. Here, during operation of the subsequent query detection mode 220, the DSP 200 executes the VAD model 222 and determines that the audio data 132 from the audio source 130 does not contain voice activity. Therefore, the VAD 222 outputs a voice activity indication indicating "No." The VAD 222 may output a "No" voice activity indication after a timeout period of not detecting any voice activity in the audio data 132. Thus, the DSP 200 does not communicate subsequent events 215 to the second processor 300, and thus the second processor 300 remains in a sleep state. The DSP 200 may now return to the hotword detection mode 210 without executing the speaker verification process 400 and stop operation of the subsequent query detection mode 220. That is, the DSP 200 may return to the hotword detection mode 210 and stop operation of the VAD 222 and the speaker verification process 400.
[0046] 4, the DSP 200 executes a speaker verification process 400 to determine a verification result 450 indicating whether the subsequent audio data 127 of the subsequent query 129 (FIG. 1B) contains utterances spoken by the same user 10 who submitted the initial query 117 (FIG. 1A). The DSP 200 or the AP 300 may extract a first identified speaker vector 411 from at least one of the first portion 121 of the audio data 120 characterizing the hot word 116 or the second portion 122 of the audio data 120 characterizing the initial query 117 spoken by the user 10a. Thereafter, while the DSP 200 is in the subsequent query detection mode 220, the speaker verification process 400 first identifies the user 10 who spoke the subsequent query by extracting a second identified speaker vector 412 representing the voice features of the subsequent query from the subsequent audio data 127. In some examples, the first and second identification speaker vectors 411, 412 each include a respective set of speaker identification vectors associated with the users 10 who spoke the initial and subsequent queries.
[0047] The speaker verification process 400 may execute a speaker identification model 410 configured to receive the subsequent audio data 127 as input and generate a second identified speaker vector 412 as output. The speaker identification model 410 may be a neural network model trained under machine or human supervision to output the second identified speaker vector 412. The second identified speaker vector 412 output by the speaker identification model 410 may include an N-dimensional vector having values corresponding to speech features of the subsequent query spoken by the user 10. In some examples, the second identified speaker vector 412 is a d-vector. The speaker identification model 410 may include a text-independent speaker identification model configured to extract a text-independent speaker identification vector from the subsequent audio data 127. That is, the speaker identification model 410 extracts speech features of the speaker identification vector regardless of the content of the subsequent query represented by the subsequent audio data 127. The speaker identification model 410 may similarly receive the initial audio data 120 as input and produce a first identified speaker vector 411 as output.
[0048] Once the second identified speaker vector 412 is output from the speaker identification model 410, the speaker verification process 400 determines whether the second identified speaker vector 412 matches the first speaker identification vector 411 or a reference speaker vector 435 associated with the user 10a who spoke the initial query 117. The reference speaker vector 435 may be stored in the AED 102 (e.g., in the memory hardware 105) and associated with an enrolled user account 430 among the AED 102's multiple enrolled user accounts 430a-n. That is, the AED 102 may have multiple different enrolled user accounts (i.e., enrolled users) 430, each with a respective reference speaker vector 435 corresponding to the voice features of the user associated with the enrolled user account 430. The user 10a who spoke the initial query 117 may be an enrolled user having the respective reference speaker vector 435. Thus, the reference speaker vector 435 associated with the user 10a may be identified based on the first identified speaker vector 411 extracted from the initial audio data 120. In additional examples, the reference speaker vector 435 associated with the user 10a is identified using other techniques.
[0049] The speaker identification model 410 can generate a reference speaker vector 435 for each enrolled user account 430 during the voice enrollment process. For example, a user may speak multiple phrases such that the speaker identification model 410 generates a reference speaker vector 435 that represents the user's voice characteristics. In some examples, the reference speaker vector 435 is a more accurate representation of the user's voice characteristics compared to the first identified speaker vector 411. Thus, utilizing the reference speaker vector 435 provides a more accurate estimation of whether a user who spoke a subsequent query matches the user who spoke the initial query.
[0050] Each reference speaker vector 435 may be used as a reference vector corresponding to a unique identifier voiceprint representing the voice characteristics of each user of the registered user account 430. Here, the comparator 420 may generate a score for the comparison indicating the likelihood that the subsequent query corresponds to the identification of the registered user account 430a associated with the user 10a who spoke the initial query 117. If the score meets a threshold, the identification is accepted. If the score does not meet the threshold, the comparator 420 may reject the identification. In some implementations, the comparator 420 calculates the respective cosine distances between the second identified speaker vector 412 and the reference speaker vector 435 associated with the first registered user account 430a, and determines that the second identified speaker vector 412 matches the reference speaker vector 435 when the respective cosine distances meet a cosine distance threshold. Alternatively, the comparator 420 may calculate the respective cosine distances between the first and second identified speaker vectors 411, 412 to determine whether a match exists.
[0051] If the speaker verification process 400 determines that the second identified speaker vector 412 matches the reference speaker vector 435 associated with the first enrolled user account 430a, the speaker verification process 400 identifies the user 10 who spoke the subsequent query as the first enrolled user account 430a associated with the user 10a who spoke the initial query 117. In the illustrated example, the comparator 420 determines a match based on the respective cosine distances between the reference speaker vector 435 associated with the first enrolled user account 430a and the second identified speaker vector 412. If the speaker verification process 400 determines that the user who spoke the subsequent query 129 is the same as the user who spoke the initial query 117, the DSP 200 triggers the output of a subsequent event 215 to cause the AP 300 to wake up from its sleep state and process the subsequent query 129.
[0052] Conversely, if the speaker verification process 400 determines that the second identified speaker vector 412 does not match either the reference speaker vector 435 or the first identified speaker vector 411 associated with the user 10a who spoke the initial query 117, then the process determines that the user 10 who spoke the subsequent query 129 is different from the user 10 who spoke the initial query 117. Thus, the DSP 200 forgoes detecting subsequent events, thereby allowing the AP 300 to remain in a sleep state.
[0053] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform tasks. In some examples, a software application may be referred to as an "application," "app," or "program." Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0054] Non-transitory memory may be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by a computing device. Non-transitory memory may be volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.
[0055] 5 is a schematic diagram of an exemplary computing device 500 that may be used to implement the systems and methods described herein. Computing device 500 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown, their connections and relationships, and their functionality are intended to be exemplary only and do not limit the implementation of the invention described and / or claimed herein.
[0056] Computing device 500 includes a processor 510, a memory 520, a storage device 530, a high-speed interface / controller 540 connecting to memory 520 and a high-speed expansion port 550, and a low-speed interface / controller 560 connecting to a low-speed bus 570 and storage device 530. Each of components 510, 520, 530, 540, 550, and 560 are interconnected using various buses and may be mounted on a common motherboard or in other manners as appropriate. Processor 510 can process instructions for execution within computing device 500, including instructions stored in memory 520 or on storage device 530, for displaying graphical information for a graphical user interface (GUI) on an external input / output device, such as a display 580 coupled to high-speed interface 540. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and multiple types of memory, as appropriate. Also, multiple computing devices 500 may be connected, each providing a portion of the required operations (eg, as a bank of servers, a group of blade servers, or a multi-processor system).
[0057] The memory 520 stores information non-transiently within the computing device 500. The memory 520 may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. The non-transient memory 520 may be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by the computing device 500. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.
[0058] The storage device 530 is capable of providing mass storage for the computing device 500. In some implementations, the storage device 530 is a computer-readable medium. In various different implementations, the storage device 530 may be a floppy disk device, a hard disk device, an optical disk device, or an array of devices including a tape device, a flash memory or other similar solid-state memory device, or devices in a storage area network or other configuration. In additional embodiments, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as the memory 520, the storage device 530, or memory on the processor 510.
[0059] The high-speed controller 540 manages bandwidth-intensive operations for the computing device 500, while the low-speed controller 560 manages less bandwidth-intensive operations. Such allocation of duties is merely exemplary. In some implementations, the high-speed controller 540 is coupled to the memory 520, the display 580 (e.g., through a graphics processor or accelerator), and a high-speed expansion port 550 that can accept various expansion cards (not shown). In some implementations, the low-speed controller 560 is coupled to the storage device 530 and the low-speed expansion port 590. The low-speed expansion port 590, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), may be coupled, for example, through a network adapter, to one or more input / output devices such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or router.
[0060] Computing device 500, as shown, may be implemented in a number of different forms. For example, computing device 500 may be implemented as a standard server 500a, or multiple times in a group of such servers 500a, as a laptop computer 500b, or as part of a rack server system 500c.
[0061] 6 is a flowchart of an exemplary configuration of operations for a method 600 of detecting an ongoing conversation in a voice-enabled environment. Method 600 may include a computer-implemented method executed on data processing hardware 103 of assistant-enabled device 102. At operation 602, the method includes instructing always-on first processor 200 of data processing hardware 103 to operate in subsequent query detection mode 220 in response to receiving a response 192 to an initial query 117 submitted to digital assistant 109 by user 10 of assistant-enabled device 102. While always-on first processor 200 operates in subsequent query detection mode 220, method 600 performs operations 604-610. At operation 604, method 600 includes receiving, in the always-on first processor 200, subsequent audio data 215 captured by assistant-enabled device 102 in the environment of assistant-enabled device 102.
[0062] At operation 606, the method 600 includes using a voice activity detection (VAD) model 222 running on the always-on first processor 200 to determine whether the VAD model 222 detects voice activity in the subsequent audio data 215. At operation 608, the method 600 includes performing speaker verification on the subsequent audio data 215 using a speaker identification (SID) model 400 running on the always-on first processor 200 to determine whether the subsequent audio data 215 contains utterances spoken by the same user 10 who submitted the initial query 117 to the digital assistant 109. In operation 610, when the VAD model 222 detects voice activity in the subsequent audio data 215 and the subsequent audio data 215 includes an utterance spoken by the same user 10 who submitted the initial query 117, the method 600 also includes initiating a wake-up process in the second processor 300 of the data processing hardware 103 to determine whether the utterance includes a subsequent query 129 directed to the digital assistant 109.
[0063] Various implementations of the systems and techniques described herein may be realized in digital electronic and / or optical circuitry, integrated circuitry, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be special purpose or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0064] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language and / or in an assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives the machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0065] The processes and logic flows described herein may be implemented by one or more programmable processors, also referred to as data processing hardware, that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be implemented by special-purpose logic circuitry, such as FPGAs (field-programmable gate arrays) and ASICs (application-specific integrated circuits). Processors suitable for executing computer programs include, by way of example, both general-purpose and special-purpose microprocessors, as well as any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer also includes one or more mass storage devices, e.g., magnetic disks, magneto-optical disks, or optical disks, for storing data, or is operably coupled to receive data from or transfer data to one or more mass storage devices, or both. A computer need not, however, have such devices. Computer-readable media suitable for storing computer program instructions and data include, by way of example, all forms of non-volatile memory, media, and memory devices, including semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices, magnetic disks, e.g., internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0066] To enable interaction with a user, one or more aspects of the present disclosure may be implemented on a computer having a display device, e.g., a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen, for displaying information to the user, and optionally a keyboard and pointing device, e.g., a mouse or trackball, by which the user can provide input to the computer. Other types of devices may also be used to enable interaction with the user; for example, feedback provided to the user may be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback, and input from the user may be received in any form, including acoustic, speech, or tactile input. Additionally, the computer may interact with the user by sending documents to and receiving documents from a device used by the user, e.g., by sending a web page to a web browser on the user's client device in response to a request received from the web browser.
[0067] Although several implementations have been described, it will be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims. [Explanation of symbols]
[0068] 10 users 100 systems 102 Assistant-enabled devices 103 Data Processing Hardware 104 Network 105 Memory Hardware 106 microphones 109 Digital Assistant 110 Remote Systems 112 Scalable / Elastic Computing Resources 114 Storage Resources 116 Hot Words 117 First Query 118 Streaming Audio 119 utterances 120 Audio Data 121 First Part 122 Second Part 127 Subsequent Audio Data 129 utterances 129 Subsequent Queries 130 Audio Sources 132 Audio 192 responses 193 Subsequent Responses 200 First Processor 200 DSP 210 Hotword Detection Mode 212 Hotword detection model 215 Subsequent Events 220 Subsequent Query Detection Mode 222 Voice Activity Detection (VAD) Model 224 Voice Activity Score 300 Second Processor 300 Audio Environment 310 ASR model 315 Voice Recognition Results 320 Natural Language Understanding (NLU) Module 325 Voice Recognition Results 400 Speaker Verification Process 410 Speaker Identification Model 411 First discriminant speaker vector 412 Second Discriminant Speaker Vector 420 Comparator 430 Registered User Accounts 435 Reference Speaker Vector 450 Verification Results 500 computing devices 510 processor 520 memory 530 Storage Devices 540 High-Speed Interface / Controller 550 High-Speed Expansion Port 560 Low-Speed Interface / Controller 570 Slow Bus 580 Display 590 Low-Speed Expansion Port 600 ways
Claims
1. A computer-implemented method that, when executed on data processing hardware (103) of an assistant-enabled device (AED) (102), causes the data processing hardware (103) to perform operations, the operations comprising: In response to receiving a response (192) to an initial query (117) submitted to a digital assistant (109) by a user (10) of the AED (102), instructing an always-on first processor (200) of the data processing hardware (103) to operate in a subsequent query detection mode (220); While the always-on first processor (200) operates in the subsequent query detection mode (220), receiving, in the environment of the AED (102), subsequent audio data (127) captured by the AED (102) at the always-on first processor (200); using a voice activity detection (VAD) model (222) executing on the always-on first processor (200) to determine whether the VAD model (222) detects voice activity in the subsequent audio data (127); Performing speaker verification on the subsequent audio data (127) using a speaker identification (SID) model (410) running on the always-on first processor (200) to determine whether the subsequent audio data (127) includes utterances spoken by the same user (10) who submitted the initial query (117) to the digital assistant (109); When the VAD model (222) detects voice activity in the subsequent audio data (127) and the subsequent audio data (127) includes the utterance spoken by the same user (10) who submitted the initial query (117), initiate a wake-up process in the second processor (300) of the data processing hardware (103) to determine whether the utterance includes a subsequent query (129) directed to the digital assistant (109); 11. A computer-implemented method comprising:
2. 2. The computer-implemented method of claim 1, wherein instructing the always-on first processor of the data processing hardware to operate in the subsequent query detection mode causes the always-on first processor to begin executing the VAD and SID model on the always-on first processor while operating in the subsequent query detection mode.
3. 2. The computer-implemented method of claim 1, wherein instructing the always-on first processor of the data processing hardware to operate in the subsequent query detection mode causes the always-on first processor to disable a hotword detection model while operating in the subsequent query detection mode.
4. Initiating the wake-up process in the second processor (300) causes the second processor (300) to: processing the subsequent audio data (127) to generate a transcription of the utterance spoken by the same user (10) who submitted the initial query (117); performing query interpretation on the transcription to determine whether the utterance includes the subsequent query (129) directed to the digital assistant (109); 4. The computer-implemented method (600) of claim 1, wherein the computer-implemented method performs operations including:
5. When the utterance includes the subsequent query (129) directed to the digital assistant (109), the operation: Instructing the digital assistant (109) to perform an action specified by the subsequent query (129); Receiving a subsequent response (193) from the digital assistant (109) indicating performance of the action specified by the subsequent query (129); presenting said subsequent response (193) as an output from said AED (102); 5. The computer-implemented method (600) of claim 4, further comprising:
6. Initiating the wake-up process in the second processor (300) causes the second processor (300) to transmit the subsequent audio data (127) to a remote server over a network, and upon receipt of the subsequent audio data (127) by the remote server, the remote server (110) processing the subsequent audio data (127) to generate a transcription of the utterance spoken by the same user (10) who submitted the initial query (117); performing query interpretation on the transcription to determine whether the utterance includes the subsequent query (129) directed to the digital assistant (109); When the utterance includes the subsequent query (129) directed to the digital assistant (109), instructing the digital assistant (109) to perform an action specified by the subsequent query (129); 4. The computer-implemented method (600) of claim 1, further comprising:
7. After the action instructs the digital assistant (109) to perform the action specified by the subsequent query (129), Receiving a subsequent response (193) from the digital assistant (109) indicating performance of the action specified by the subsequent query (129); presenting said subsequent response (193) as an output from said AED (102); 7. The computer-implemented method (600) of claim 6, further comprising:
8. After the operation receives initial audio data (120) corresponding to the initial query (117) spoken by the user (10) and submitted to the digital assistant (109), extracting from the initial audio data (120) a first speaker identification vector (411) characteristic of the initial query (117) spoken by the user (10); further comprising performing speaker verification on the subsequent audio data (127); extracting from the subsequent audio data (127) a second speaker identification vector (412) characteristic of the subsequent audio data (127) using the SID model; When the first speaker identification vector (411) matches the second speaker identification vector (412), determining that the subsequent audio data (127) includes the utterance spoken by the same user (10) who submitted the initial query (117) to the digital assistant (109); Including, A computer-implemented method (600) according to any one of claims 1 to 7.
9. After the operation receives initial audio data (120) corresponding to the initial query (117) spoken by the user (10) and submitted to the digital assistant (109), extracting from the initial audio data (120) a first speaker identification vector (411) characteristic of the initial query (117) spoken by the user (10); determining whether the first speaker identification vector (411) matches any enrolled speaker vectors stored in the AED (102), each enrolled speaker vector being associated with a different respective enrolled user (10) of the AED (102); When the first speaker identification vector (411) matches one of the enrolled speaker vectors, identifying the user (10) who spoke the initial query (117) as the respective enrolled user (10) associated with the one of the enrolled speaker vectors that matches the first speaker identification vector (411); further comprising performing speaker verification on the subsequent audio data (127); extracting from the subsequent audio data (127) a second speaker identification vector (412) characteristic of the subsequent audio data (127) using the SID model; When the second speaker identification vector (412) matches one of the enrolled speaker vectors associated with the respective enrolled user (10) who spoke the initial query (117), determining that the subsequent audio data (127) includes the utterance spoken by the same user (10) who submitted the initial query (117) to the digital assistant (109); Including, A computer-implemented method (600) according to any one of claims 1 to 7.
10. 10. The computer-implemented method (600) of any one of claims 1 to 9, wherein the SID model comprises a text-independent SID model configured to extract a text-independent speaker identification vector from the subsequent audio data (127).
11. 11. The computer-implemented method of claim 1, wherein the operations further include instructing the always-on first processor to cease operation in the subsequent query detection mode and commence operation in a hotword detection mode in response to determining that the VAD model has failed to detect voice activity in the subsequent audio data or that the subsequent audio data has failed to contain utterances spoken by the same user who submitted the initial query.
12. 12. The computer-implemented method (600) of claim 11, wherein the always-on first processor (200) disables or deactivates execution of the VAD and SID model (222, 410) on the always-on first processor (200) by instructing the always-on first processor (200) to stop operating in the subsequent query detection mode (220).
13. 13. The computer-implemented method (600) of claim 11 or 12, wherein instructing the always-on first processor (200) to begin operation in the hotword detection mode (210) causes the always-on first processor (200) to begin execution of a hotword detection model (212) on the always-on first processor (200).
14. the always-on first processor (200) includes a digital signal processor (DSP); the second processor (300) includes an application processor; A computer-implemented method (600) according to any one of claims 1 to 13.
15. 15. The computer-implemented method of claim 1, wherein the AED includes a battery-powered device in communication with one or more microphones configured to capture the subsequent audio data and initial audio data corresponding to the initial query.
16. data processing hardware (103); and memory hardware (105) in communication with the data processing hardware (103) and storing instructions that, when executed on the data processing hardware (103), cause the data processing hardware (103) to perform operations, the operations comprising: In response to receiving a response (192) to an initial query (117) submitted to a digital assistant (109) by a user (10) of the AED (102), instructing an always-on first processor (200) of the data processing hardware (103) to operate in a subsequent query detection mode (220); While the always-on first processor (200) operates in the subsequent query detection mode (220), receiving, in the environment of the AED (102), subsequent audio data (127) captured by the AED (102) at the always-on first processor (200); using a voice activity detection (VAD) model (222) executing on the always-on first processor (200) to determine whether the VAD model (222) detects voice activity in the subsequent audio data (127); Performing speaker verification on the subsequent audio data (127) using a speaker identification (SID) model (410) running on the always-on first processor (200) to determine whether the subsequent audio data (127) includes utterances spoken by the same user (10) who submitted the initial query (117) to the digital assistant (109); When the VAD model (222) detects voice activity in the subsequent audio data (127) and the subsequent audio data (127) includes the utterance spoken by the same user (10) who submitted the initial query (117), initiate a wake-up process in the second processor (300) of the data processing hardware (103) to determine whether the utterance includes a subsequent query (129) directed to the digital assistant (109); Including, Assistant Enabled Device (AED) (102).
17. 17. The AED (102) of claim 16, wherein instructing the always-on first processor (200) of the data processing hardware (103) to operate in the subsequent query detection mode (220) causes the always-on first processor (200) to begin executing the VAD and SID model (222, 410) on the always-on first processor (200) while operating in the subsequent query detection mode (220).
18. 18. The AED (102) of claim 16 or 17, wherein by instructing the always-on first processor (200) of the data processing hardware (103) to operate in the subsequent query detection mode (220), the always-on first processor (200) disables a hot word detection model (212) while operating in the subsequent query detection mode (220).
19. Initiating the wake-up process in the second processor (300) causes the second processor (300) to: processing the subsequent audio data (127) to generate a transcription of the utterance spoken by the same user (10) who submitted the initial query (117); performing query interpretation on the transcription to determine whether the utterance includes the subsequent query (129) directed to the digital assistant (109); 19. The AED (102) of any one of claims 16 to 18, which performs operations including:
20. When the utterance includes the subsequent query (129) directed to the digital assistant (109), the operation: Instructing the digital assistant (109) to perform an action specified by the subsequent query (129); Receiving a subsequent response (193) from the digital assistant (109) indicating performance of the action specified by the subsequent query (129); presenting said subsequent response (193) as an output from said AED (102); 20. The AED (102) of claim 19, further comprising:
21. Initiating the wake-up process in the second processor (300) causes the second processor (300) to transmit the subsequent audio data (127) to a remote server over a network, and upon receipt of the subsequent audio data (127) by the remote server, the remote server: processing the subsequent audio data (127) to generate a transcription of the utterance spoken by the same user (10) who submitted the initial query (117); performing query interpretation on the transcription to determine whether the utterance includes the subsequent query (129) directed to the digital assistant (109); When the utterance includes the subsequent query (129) directed to the digital assistant (109), instructing the digital assistant (109) to perform an action specified by the subsequent query (129); The AED (102) according to any one of claims 16 to 18, which causes the AED (102) to perform an operation including:
22. After the action instructs the digital assistant (109) to perform the action specified by the subsequent query (129), Receiving a subsequent response (193) from the digital assistant (109) indicating performance of the action specified by the subsequent query (129); presenting said subsequent response (193) as an output from said AED (102); 22. The AED (102) of claim 21, further comprising:
23. After the operation receives initial audio data (120) corresponding to the initial query (117) spoken by the user (10) and submitted to the digital assistant (109), extracting from the initial audio data (120) a first speaker identification vector (411) characteristic of the initial query (117) spoken by the user (10); further comprising performing speaker verification on the subsequent audio data (127); extracting from the subsequent audio data (127) a second speaker identification vector (412) characteristic of the subsequent audio data (127) using the SID model; When the first speaker identification vector (411) matches the second speaker identification vector (412), determining that the subsequent audio data (127) includes the utterance spoken by the same user (10) who submitted the initial query (117) to the digital assistant (109); Including, 23. An AED (102) according to any one of claims 16 to 22.
24. After the operation receives initial audio data (120) corresponding to the initial query (117) spoken by the user (10) and submitted to the digital assistant (109), extracting from the initial audio data (120) a first speaker identification vector (411) characteristic of the initial query (117) spoken by the user (10); determining whether the first speaker identification vector (411) matches any enrolled speaker vectors stored in the AED (102), each enrolled speaker vector being associated with a different respective enrolled user (10) of the AED (102); When the first speaker identification vector (411) matches one of the enrolled speaker vectors, identifying the user (10) who spoke the initial query (117) as the respective enrolled user (10) associated with the one of the enrolled speaker vectors that matches the first speaker identification vector (411); further comprising performing speaker verification on the subsequent audio data (127); extracting from the subsequent audio data (127) a second speaker identification vector (412) characteristic of the subsequent audio data (127) using the SID model; When the second speaker identification vector (412) matches one of the enrolled speaker vectors associated with the respective enrolled user (10) who spoke the initial query (117), determining that the subsequent audio data (127) includes the utterance spoken by the same user (10) who submitted the initial query (117) to the digital assistant (109); Including, 23. An AED (102) according to any one of claims 16 to 22.
25. 25. The AED (102) of any one of claims 16 to 24, wherein the SID model comprises a text-independent SID model configured to extract a text-independent speaker identification vector from the subsequent audio data (127).
26. 26. The AED (102) of any one of claims 16 to 25, wherein the operations further include instructing the always-on first processor (200) to cease operation in the subsequent query detection mode (220) and commence operation in a hotword detection mode (210) in response to the VAD model (222) determining at least one of: that the VAD model (222) failed to detect voice activity in the subsequent audio data (127) or that the subsequent audio data (127) failed to contain speech spoken by the same user (10) who submitted the initial query (117).
27. 27. The AED (102) of claim 26, wherein by instructing the always-on first processor (200) to stop operating in the subsequent query detection mode (220), the always-on first processor (200) disables or deactivates execution of the VAD and SID model (222, 410) on the always-on first processor (200).
28. 28. The AED (102) of claim 26 or 27, wherein instructing the always-on first processor (200) to begin operation in the hot word detection mode (210) causes the always-on first processor (200) to begin execution of a hot word detection model (212) on the always-on first processor (200).
29. the always-on first processor (200) includes a digital signal processor (DSP); the second processor (300) includes an application processor; 29. The AED (102) according to any one of claims 16 to 28.
30. 30. The AED (102) of any one of claims 16 to 29, wherein the AED (102) includes a battery-powered device in communication with one or more microphones configured to capture the subsequent audio data (127) and initial audio data (120) corresponding to the initial query (117).
Citation Information
Patent Citations
Voice Trigger for Digital Assistant
JP2016508007A
Speaker authentication method and speech recognition system
JP2019028464A
Speaker verification method and system
JP2019514045A
Speech interaction method, device and system
WO2020135811A1