frozen words
By introducing a frozen word mechanism and an acoustic feature detection model, users can manually terminate voice input, which solves the problems of delay and malfunction after the voice device is woken up, and achieves more efficient voice processing and resource management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GOOGLE LLC
- Filing Date
- 2021-11-17
- Publication Date
- 2026-05-08
AI Technical Summary
In existing technologies, voice devices may delay or incorrectly perform actions beyond the user's intent after detecting a hot word wake-up, and the microphone remains on to capture non-intended audio, resulting in wasted resources and action delays.
By introducing a freeze word mechanism into speech, users can say a freeze word to manually end their speech, triggering the hard microphone to shut down and prevent subsequent audio from being captured. The audio data is then processed by an acoustic feature detection model and a speech recognizer to identify the freeze word.
It effectively terminates speech processing, reduces resource waste, avoids unintended audio capture, and improves the timeliness and accuracy of action execution.
Smart Images

Figure CN116615779B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the term "freeze word". Background Technology
[0002] Voice-enabled environments (e.g., home, workplace, school, car, etc.) allow users to speak queries or commands aloud to a computer-based system, which responds and answers the queries and / or performs functions based on the commands. Voice-enabled environments can be implemented using a network of connected microphone devices distributed across various rooms or areas of the environment. Instead of utterances directed at another person in the environment, these devices can use hot words to help identify when a given utterance is directed at the system. Therefore, the devices can operate in sleep or hibernation mode and only wake up when a detected utterance includes a hot word. Once awakened, the device is able to proceed with more expensive processing, such as fully on-device automated speech recognition (ASR) or server-based ASR. Summary of the Invention
[0003] One aspect of this disclosure provides a method for detecting freeze words. The method includes receiving, at data processing hardware, audio data corresponding to a utterance spoken by a user and captured by a user device associated with the user. The method also includes processing the audio data by the data processing hardware using a speech recognizer to determine if the utterance includes a query for performing an action on a digital assistant. The speech recognizer is configured to trigger an endpoint of the utterance after a predetermined duration of non-speech in the audio data. Before the predetermined duration of non-speech in the audio data, the method includes detecting freeze words in the audio data by the data processing hardware. The freeze word follows a query in the utterance spoken by the user and captured by the user device. In response to detecting a freeze word in the audio data, the method includes triggering a hard microphone mute event at the user device by the data processing hardware to prevent the user device from capturing any audio following the freeze word.
[0004] Implementations of this disclosure may include one or more of the following optional features. In some implementations, frozen words include one of the following: predefined frozen words, including one or more fixed terms in a given language across all users; user-selected frozen words, including one or more terms specified by the user of the user device; or, action-specific frozen words associated with an operation to be performed by a digital assistant. In some examples, detecting frozen words in audio data includes: extracting audio features from the audio data; generating a frozen word confidence score by processing the extracted audio features using a frozen word detection model; executing the frozen word detection model on data processing hardware; and determining that the audio data corresponding to the utterance includes a frozen word when the frozen word confidence score meets a frozen word confidence threshold.
[0005] Detecting frozen words in audio data may include using a speech recognizer executed on data processing hardware to identify frozen words in the audio data. Optionally, the method may further include: in response to detecting a frozen word in the audio data: instructing the speech recognizer by the data processing hardware to stop any active processing of the audio data; and instructing a digital assistant by the data processing hardware to complete the execution of the operation.
[0006] In some embodiments, processing audio data to determine that the utterance includes a query to perform an operation on a digital assistant includes: processing the audio data using a speech recognizer to generate a speech recognition result of the audio data; and performing semantic interpretation on the speech recognition result of the audio data to determine that the audio data includes a query to perform an operation. In these embodiments, in response to detecting a frozen word in the audio data, the method further includes: modifying the speech recognition result of the audio data by data processing hardware by removing the frozen word from the speech recognition result; and instructing the digital assistant to perform the operation of the query request using the modified speech recognition result.
[0007] In some examples, before processing the audio data using a speech recognizer, the method further includes: detecting hot words in the audio data prior to the query using a hot word detection model by the data processing hardware; and, in response to the detection of hot words, triggering the speech recognizer by the data processing hardware to process the audio data by performing speech recognition on the hot words and / or one or more terms in the audio data following the hot words. In these examples, the method may also include the data processing hardware verifying the presence of hot words detected by the hot word detection model based on the detection of frozen words in the audio data. Optionally, detecting frozen words in the audio data may include executing a frozen word detection model on data processing hardware configured to detect frozen words in the audio data without performing speech recognition on the audio data. Here, the frozen word detection model and the hot word detection model may each comprise the same or different neural network-based models.
[0008] Another aspect of this disclosure provides a method for detecting frozen words. The method includes receiving at data processing hardware a first instance of audio data corresponding to a dictation-based query of audible content spoken by a user for a digital assistant to dictate. The dictation-based query is spoken by the user and captured by an assistant-enabled device associated with the user. The method also includes receiving at data processing hardware a second instance of audio data corresponding to utterances of audible content spoken by the user and captured by the assistant-enabled device. The method further includes processing the second instance of audio data by the data processing hardware using a speech recognizer to generate a transcription of the audible content. During the processing of the second instance of audio data, the method includes detecting frozen words in the second instance of audio data by the data processing hardware. Frozen words follow audible content in utterances spoken by the user and captured by the assistant-enabled device. In response to detecting frozen words in the second instance of audio data, the method includes providing the transcription of the audible content spoken by the user by the data processing hardware for output from the assistant-enabled device.
[0009] Implementations of this disclosure may include one or more of the following optional features. In some implementations, in response to detecting a frozen word in a second instance of audio data, the method further includes: initiating a hard microphone shutdown event at an assistant-enabled device by data processing hardware to prevent the assistant-enabled device from capturing any audio after the frozen word; stopping any active processing of the second instance of audio data by the data processing hardware; and stripping the frozen word from the end of the transcription by the data processing hardware before providing a transcription of audible content for output from the assistant-enabled device.
[0010] Optionally, the method may further include: processing a first instance of audio data using a speech recognizer by data processing hardware to generate a speech recognition result; and performing a semantic interpretation on the speech recognition result of the first instance of audio data by the data processing hardware to determine that the first instance of audio data includes a dictation-based query of audible content spoken by a user. In some examples, before initiating processing of a second instance of audio data to generate a transcription, the method further includes: determining, by the data processing hardware, a dictation-based query specifying a freeze word based on the semantic interpretation performed on the speech recognition result of the first instance of audio data; and increasing the termination timeout duration for terminating the audible content by a data processing hardware endpoint.
[0011] Another aspect of this disclosure provides a system for detecting frozen words. The system includes data processing hardware and memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include receiving audio data corresponding to a utterance spoken by a user and captured by a user device associated with the user. The operations also include processing the audio data using a speech recognizer to determine if the utterance includes a query for performing an operation on a digital assistant. The speech recognizer is configured to trigger the termination of the utterance after a predetermined duration of non-speech in the audio data. Before the predetermined duration of non-speech in the audio data, the operations include detecting frozen words in the audio data. The frozen word follows a query in the utterance spoken by the user and captured by the user device. In response to detecting a frozen word in the audio data, the operations include triggering a hard microphone mute event at the user device to prevent the user device from capturing any audio following the frozen word.
[0012] Implementations of this disclosure may include one or more of the following optional features. In some implementations, frozen words include one of the following: predefined frozen words, including one or more fixed terms across all users of a given language; user-selected frozen words, including one or more terms specified by a user of a user device; or, action-specific frozen words associated with an operation to be performed by a digital assistant. In some examples, detecting frozen words in audio data includes: extracting audio features from the audio data; generating a frozen word confidence score by processing the extracted audio features using a frozen word detection model; and determining that the audio data corresponding to the utterance includes a frozen word when the frozen word confidence score meets a frozen word confidence threshold. In these examples, the frozen word detection model is executed on data processing hardware.
[0013] Detecting frozen words in audio data may include using a speech recognizer executed on data processing hardware to identify frozen words in the audio data. Optionally, the operation may also include, in response to detecting a frozen word in the audio data: instructing the speech recognizer to stop any active processing of the audio data; and instructing the digital assistant to complete the execution of the operation.
[0014] In some embodiments, processing audio data to determine that the utterance includes a query for performing an operation on a digital assistant includes: processing the audio data using a speech recognizer to generate a speech recognition result of the audio data; and performing semantic interpretation on the speech recognition result of the audio data to determine that the audio data includes a query for performing an operation. In these embodiments, in response to detecting a frozen word in the audio data, the operation further includes: modifying the speech recognition result of the audio data by removing the frozen word from the speech recognition result; and instructing the digital assistant to perform the query request using the modified speech recognition result.
[0015] In some examples, before processing the audio data using a speech recognizer, the operation further includes: detecting hot words in the audio data prior to the query using a hot word detection model; and, in response to the detection of a hot word, triggering the speech recognizer to process the audio data by performing speech recognition on the hot word and / or one or more terms following the hot word in the audio data. In these examples, the operation also includes verifying the presence of hot words detected by the hot word detection model based on the detection of frozen words in the audio data. Optionally, detecting frozen words in the audio data may include executing a frozen word detection model on data processing hardware configured to detect frozen words in the audio data without performing speech recognition on the audio data. The frozen word detection model and the hot word detection model each include the same or different neural network-based models.
[0016] Another aspect of this disclosure provides a system for detecting frozen words. The system includes data processing hardware and memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include receiving a first instance of audio data corresponding to a dictation-based query of audible content spoken by a user for a digital assistant. The dictation-based query is spoken by the user and captured by an assistant-enabled device associated with the user. The operations also include receiving a second instance of audio data corresponding to utterances of audible content spoken by the user and captured by the assistant-enabled device. The operations further include processing the second instance of audio data using a speech recognition device to generate a transcription of the audible content. During the processing of the second instance of audio data, the operations include detecting frozen words in the second instance of audio data. Frozen words follow the audible content in the utterances spoken by the user and captured by the assistant-enabled device. In response to detecting a frozen word in the second instance of audio data, the operations include providing a transcription of the audible content spoken by the user for output from the assistant-enabled device.
[0017] Implementations of this disclosure may include one or more of the following optional features. In some implementations, in response to the detection of a frozen word in a second instance of audio data, the operation further includes: initiating a hard microphone shutdown event at an assistant-enabled device to prevent the assistant-enabled device from capturing any audio after the frozen word; stopping any active processing of the second instance of audio data; and stripping the frozen word from the end of the transcription before providing a transcription of audible content for output from the assistant-enabled device.
[0018] Optionally, the operation may further include: processing a first instance of audio data using a speech recognizer to generate a speech recognition result; and performing a semantic interpretation on the speech recognition result of the first instance of audio data to determine that the first instance of audio data includes a dictation-based query for dictating audible content spoken by a user. In some examples, before initiating processing of a second instance of audio data to generate a transcription, the operation further includes: determining, based on the semantic interpretation performed on the speech recognition result of the first instance of audio data, a dictation-based query specifying a freeze word; and an instruction terminator increasing the termination timeout duration for terminating the audible content.
[0019] Details of one or more embodiments of this disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, as well as from the claims. Attached Figure Description
[0020] Figure 1 This is an example system that includes devices configured to detect frozen words by an enabled assistant.
[0021] Figure 2A This is a schematic diagram of the first instance of a user utterance, which corresponds to a dictation-based query that specifies a freeze word to terminate the second instance of the user utterance.
[0022] Figure 2B This is a schematic diagram of an acoustic feature detector that terminates a utterance in response to the detection of a frozen word in the utterance.
[0023] Figure 3 This is a flowchart illustrating an example of the operational setup for a method used to detect frozen words in discourse.
[0024] Figure 4 This is a flowchart illustrating an example of the operational setup for a method used to detect frozen words in discourse.
[0025] Figure 5 This is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein.
[0026] Similar reference numerals in the various figures indicate similar elements. Detailed Implementation
[0027] Voice-based interfaces, such as digital assistants, are becoming increasingly prevalent across a wide range of devices, including but not limited to mobile phones and smart speakers / displays that include microphones for capturing speech. The general way to initiate voice interaction with an assistant-enabled device is to speak a fixed phrase, such as a hot word. When the voice-enabled device detects this fixed phrase in the streaming audio, it triggers a wake-up process to begin recording and processing subsequent speech to determine the user's spoken query. Therefore, hot words are a crucial component of the entire digital assistant interface stack because they allow users to wake their assistant-enabled devices from a low-power state, enabling the assistant-enabled devices to continue performing more expensive processing, such as fully automated speech recognition (ASR) or server-based ASR.
[0028] Queries spoken by users to assistant-enabled devices typically fall into two categories: conversational queries and non-conversational queries. Conversational queries are standard digital assistant queries that ask the digital assistant to perform actions such as "set a timer" or "remind me to buy the milk." Non-conversational queries, on the other hand, are dictation-based queries, which are longer forms of queries spoken by the user to transcribe emails, messages, documents, social media posts, or other content segments. For example, a user might say the query "send an email to Aleks saying," and then continue by saying the content of the email message that the digital assistant will transcribe / transcribe, which will then be sent from the user's email client to the recipient's (e.g., Aleks) email client.
[0029] ASR systems typically use terminators to determine when a user begins and ends a speech. Once labeled, they can process the audio portions representing the user's speech to generate speech recognition results, and in some cases, perform semantic interpretation on the speech recognition results to ascertain the query the user uttered. Terminators typically evaluate the duration of pauses between words to determine when a utterance begins or ends. For example, if a user says "what is..."<long pause> If the terminator specifies an incomplete phrase like "what is for dinner" (a long pause at dinner), the terminator can segment the speech input at the long pause, causing the ASR system to process only the incomplete phrase "what is" instead of the complete phrase "what is for dinner." If the terminator specifies an incorrect endpoint for the utterance, the processed utterance can be inaccurate and undesirable. Simultaneously, while allowing longer pauses between words to prevent premature termination in determining when a utterance begins or ends, the microphone of the user's assistant-enabled device remains on and may detect sounds not intended for the user's device. Furthermore, delaying the microphone's shutdown thus delays the execution of the action specified by the utterance. For example, if the user's utterance is "Call" directed to a digital assistant... The query for the "Mom (call Mom)" action will inevitably involve a delay in the digital assistant initiating the call while the terminator is waiting for the termination timeout duration to elapse to confirm that the user may have stopped speaking. In this scenario, the assistant-enabled device may also detect additional non-intended audio, which may lead to the execution of actions different from the user's intended action. This could result in wasted computing resources in interpreting and acting on additional audio detected because it is not possible to determine in time when the user may have finished speaking.
[0030] To mitigate the drawbacks associated with both excessively short termination timeout durations (e.g., potentially cutting off speech before the user has finished speaking) and excessively long termination timeout durations (e.g., increasing the chance of capturing unintended speech and adding delays for performing the utterance-specified action), embodiments of this paper involve freeze words that, when spoken at the end of a utterance, specify when the user has finished speaking to the assistant-enabled device. In a sense, a "freeze word" corresponds to the inverse of a hot word by allowing the user to manually terminate a utterance and initiate a hard microphone shutdown event to end a speech-based conversation or a long form of utterance. That is, while a hot word would trigger the assistant-enabled device to wake from sleep or hibernation to begin processing speech, a freeze word would perform this inverse by terminating all active processing of speech and disabling the microphone on the assistant-enabled device, thereby returning the assistant-enabled device to sleep or hibernation.
[0031] In addition to shutting down some or all of the ongoing voice processing, once a frozen word is detected, an assistant-enabled device can additionally disable or adjust future processing over a certain period of time to effectively reduce the device's responsiveness. For example, the hot word detection threshold can be temporarily increased to make it more difficult / less likely for a user to post a subsequent query within a certain time window after uttering the frozen word. In this scenario, the increased hot word detection threshold can be gradually reduced back to its default value over time. Alternatively, voice input can be disabled for a specific user after a frozen word is detected.
[0032] The assistant-enabled device executes an acoustic feature detection model configured to detect the presence of frozen words corresponding to utterances in the audio data without performing speech recognition or semantic interpretation on the audio data. Here, the acoustic feature detection model can be a neural network-based model trained to detect one or more frozen words. The assistant-enabled device can employ the same or different acoustic feature detection models to detect the presence of hot words in the audio data. In cases where the same acoustic feature detection model is used for both hot word detection and frozen word detection, the functionality of using only one of them at a time can be active. Notably, compared to ASR models, acoustic feature detection models can run on user devices due to their relatively compact size and lower processing requirements.
[0033] In some configurations, in addition to triggering a hard microphone mute event, frozen word detection in the audio data verifies the presence of recently detected hot words in the audio data while the assistant-enabled device is in sleep or hibernation mode. Here, detected hot words can be correlated with low hot word detection confidence scores, and subsequent detection of frozen words can be used as a proxy for verifying the presence of hot words in the audio data. In these configurations, the audio data can be buffered on the assistant-enabled device while performing hot word detection and frozen word detection, and once a frozen word is detected in the buffered audio data, the assistant-enabled device can initiate a wake-up process to perform speech recognition on the buffered audio data.
[0034] In some additional implementations, frozen word detection utilizes an automated speech recognizer currently executing on the device or server side to identify the presence of frozen words. The speech recognizer can be biased to identify one or more specific frozen words.
[0035] In some examples, language models can be used to determine whether frozen words are detected in audio. In these examples, language models can allow assistant-enabled devices to recognize that frozen words are actually part of a user's utterance / query rather than being spoken by the user to terminate the utterance / query. Furthermore, language models can allow for approximate matching of frozen words, where phrases are similar to the frozen word, and based on the language model score, the frozen word is unlikely to be part of a user's query / utterance.
[0036] Assistant-enabled devices can recognize one or more different types / categories of frozen words, such as, but not limited to, predefined frozen words, custom frozen words, user-selected frozen words, action-specific frozen words, and query-specific frozen words. Predefined frozen words can include phrases of one or more fixed terms used across all users in a given language. For example, for a session query “Call Mom right now” and “Tell me the temperature outside, thanks Google”, the phrases “right now” and “thanks Google” correspond to frozen words that allow the user to manually terminate the corresponding query.
[0037] User-selected frozen words can correspond to frozen words pre-specified by a specific user, for example, during the setup of a digital assistant. For instance, a user can choose a frozen word from a suggested list. Alternatively, a user can specify one or more custom terms to use as frozen words by typing or speaking them. In some cases, user-selected frozen words are active for specific types of queries. Here, a user can assign a different user-selected frozen word to a dictation-based query than to a session query. For example, in a dictation-based query, “Hey Google send a message to Aleks saying 'I'll be late for our meeting' The End.” In this example, “Hey Google” corresponds to the hot word, the phrase “send a message to Aleks saying” corresponds to the query for dictating to the digital assistant and sending a message to the recipient, the phrase “I'll be late for our meeting” corresponds to the content of the message, and the phrase “The End” includes the user-selected frozen word used to manually terminate the query. Therefore, upon detecting the freeze word "The End," the assistant-enabled device will immediately terminate the utterance and cause the speech recognizer to remove the freeze word from the dictated message before sending it to the recipient. Alternatively, the phrase "send a message to Aleks saying" could correspond to a query for a digital assistant to facilitate audio-based communication between the user and the recipient, where the message content "I'll be late for our meeting" is sent as a voice message to the recipient for audible playback on their device. It is noteworthy that when the freeze word "The End" is detected by the assistant-enabled device, the utterance will be immediately terminated, and the audio of the freeze word will be stripped from the voice message before being sent to the recipient.
[0038] Action-specific frozen words are associated with a specific action / feature specified by a query performed for a digital assistant. For example, a user saying the query “Hey Google broadcast I'm home end broadcast” includes the frozen word “end broadcast”, a broadcast action specific to the digital assistant. In this example, the term “broadcast I'm home” specifies the action of broadcasting an audible notification through one or more speakers to indicate to others that the user is at home. The audible notification may include a specific melody or ringtone that allows people who hear the audible notification to determine that the user is at home. In some implementations, action-specific frozen words are enabled in parallel with user-specified frozen words and / or predefined frozen words.
[0039] Query-specific frozen words can be specified as part of the query spoken by the user. For example, the following phrase: "Hey Google, dictate the following journal entry until I say I'm done".<contents of journal entry> The query "I'm done (Hi Google, dictate the following diary entry until I say I'm done, <content of diary entry> I'm done)" includes a dictation-based query for digital assistants to dictate what a user says about a diary entry. Additionally, the dictation-based query specifies a freeze word "I'm Done" before the user begins speaking the content of the diary entry. Here, the freeze word "I'm Done" instruction terminator, specified as part of the dictation-based query, waits for or at least extends the termination timeout duration to trigger termination until the freeze word "I'm Done" is detected. Extending the termination timeout duration allows for a pause that would otherwise trigger termination when the user speaks the content of the diary entry. In some examples, query-specific freeze words are enabled in parallel with user-specified freeze words and / or predefined freeze words.
[0040] refer to Figure 1In some implementations, example system 100 includes an assistant-enabled device (AED) 102 associated with one or more users 10 and communicating with a remote system 111 via network 104. AED 102 may correspond to a computing device such as a mobile phone, computer (laptop or desktop), tablet, smart speaker / monitor, smart appliance, smart headphones, wearable device, in-vehicle infotainment system, etc., and is equipped with data processing hardware 103 and memory hardware 105. AED 102 includes or communicates with one or more microphones 106 to capture speech from the corresponding user 10. Remote system 111 may be a single computer, multiple computers, or a distributed system (e.g., a cloud environment) with scalable / elastic computing resources 113 (e.g., data processing hardware) and / or storage resources 115 (e.g., memory hardware).
[0041] AED 102 includes an acoustic feature detector 110 configured to detect the presence of hot words 121 and / or frozen words 123 in streaming audio 118 without performing semantic analysis or speech recognition processing on the streaming audio 118. AED 102 also includes an acoustic feature extractor 112, which may be implemented as part of or separate from the acoustic feature detector 110. Acoustic feature extractor 112 is configured to extract acoustic features from utterance 119. For example, acoustic feature extractor 112 may receive streaming audio 118 corresponding to utterance 119 spoken by user 10, captured by one or more microphones 106 of AED 102, and extract acoustic features from audio data 120 corresponding to utterance 119. Acoustic features may include Mel-frequency cepstral coefficients (MFCCs) or filter bank energy calculated on a window of audio data 120 corresponding to utterance 119.
[0042] Acoustic feature detector 110 can receive audio data 120 including acoustic features extracted by acoustic feature extractor 112, and based on the extracted features, hot word classifier 150 is configured to classify whether the dialogue 119 includes a specific hot word 121 spoken by user 10. AED 102 can store the extracted acoustic features in a buffer of memory hardware 105, and hot word classifier 150 can use the acoustic features in the buffer to detect whether the audio data 120 includes hot word 121. Hot word classifier 150 can also be referred to as hot word detection model 150. AED 102 can include multiple hot word classifiers 150, each trained to detect different hot words associated with a specific term / phrase. These hot words can be predefined hot words and / or custom hot words assigned by user 10. In some embodiments, hot word classifier 150 includes a trained neural network-based model received from remote system 111 via network 104.
[0043] The acoustic feature detector 110 also includes a frozen word classifier 160, configured to classify whether the utterance 119 includes a frozen word 123 spoken by the user 10. The frozen word classifier 160 may also be referred to as a frozen word detection model 160. The AED 102 may include multiple frozen word classifiers 160, each trained to detect different frozen words associated with a specific term / phrase. As described above, frozen words may include predefined frozen words, user-selected frozen words, action-specific frozen words, and / or query-specific frozen words. Like the hot word classifier 150, the frozen word classifier 160 may include a trained neural network-based model received from the remote system 111. In some examples, the frozen word classifier 160 and the hot word classifier 150 are merged into the same neural network-based model. In these examples, the corresponding portions of the neural network models corresponding to the hot word classifier 150 and the frozen word classifier 160 are never active simultaneously. For example, when AED 102 is in sleep mode, hot word classifier 150 may be active to listen for hot word 121 in streaming audio 118, while frozen word classifier 160 may be inactive. Once hot word 121 is detected, triggering AED 102 to wake up and process subsequent audio, hot word classifier 150 may now be inactive, and frozen word classifier 160 may be active to listen for frozen word 123 in streaming audio 118. Classifiers 150 and 160 of acoustic feature detector 110 may run on a first processor (such as a digital signal processor (DSP)) and / or a second processor (such as an application processor (AP) or CPU) of AED 102, consuming more power than the first processor during operation.
[0044] In some implementations, the hot word classifier 150 is configured to identify hot words in the initial portion of utterance 119. In the example shown, if the hot word classifier 150 detects acoustic features in the audio data 120 that are characteristics of hot word 121, the hot word classifier 150 can determine that the utterance 119 “Ok Google, broadcast I'm home endbroadcast” includes the hot word 121 “Ok Google”. For example, the hot word classifier 150 can detect that the utterance 119 “Ok Google, broadcast I'm home end broadcast” includes the hot word 121 “Ok Google” based on generating an MFCC from the audio data and classifying the MFCC as including characteristics similar to those of the hot word “Ok Google” stored in the model of the hot word classifier 150. As another example, the hot word classifier 150 can detect that the utterance 119 “Ok Google, broadcast I'm homeend broadcast” includes the hot word 121 “Ok Google” based on the following operation: generating Mel-scale filter group energy from the audio data and classifying it as Mel-scale filter group energy including Mel-scale filter group energy that is similar to the characteristics of the hot word “Ok Google” stored in the model of the hot word classifier 150.
[0045] In phase A of the acoustic feature detector 110, when the hot word classifier 150 determines that the audio data 120 corresponding to utterance 119 includes hot word 121, the AED 102 can trigger a wake-up process to initiate speech recognition on the audio data 120 corresponding to utterance 119. For example, an automated speech recognition (ASR) engine 200 (interchangeably referred to as "speech recognizer" 200) running on the AED 102 can perform speech recognition or semantic interpretation on the audio data corresponding to utterance 119. The speech recognizer 200 may include an ASR model 210, a natural language understanding (NLU) module 220, and a terminator 230. The ASR model 210 can process the audio data 120 to generate a speech recognition result 215, and the NLU module 220 can perform semantic interpretation on the speech recognition result 215 to determine that the audio data 120 includes a query 122 for performing an operation on the digital assistant 109. In this example, ASR model 210 can process audio data 120 to generate a speech recognition result 215 of “broadcast I'm home end broadcast”, and NLU module 220 can recognize “broadcast I'm home” as a query 122 for digital assistant 109 to perform an audible notification, which is an audible output from one or more speakers informing others that the user is at home. Alternatively, query 122 can be a voice message broadcasting “I'm home” to digital assistant 109 as an audible output from one or more speakers. NLU module 220 can also be used to determine whether the presence of a frozen word detected in audio data 120 is actually part of query 122 and therefore not spoken by the user to end a utterance. Therefore, in scenarios where the frozen word is actually part of the utterance, NLU module 220 can ignore the detection of the frozen word. NLU module 220 can utilize language model scores in these scenarios.
[0046] In some implementations, a speech recognizer 200, located on a remote system 111, supplements or replaces the AED 102. When the hot word classifier 150 triggers the AED 102 to wake up in response to the detection of hot word 121 in utterance 119, the AED 102 can transmit audio data 120 corresponding to utterance 119 to the remote system 111 via network 104. The AED 102 can transmit a portion of the audio data including hot word 121 for the remote system 111 to confirm the presence of hot word 121 performed by the ASR model 210. Alternatively, the AED 102 can transmit only a portion of the audio data 120 corresponding to the portion of utterance 119 following hot word 121 to the remote system 111. The remote system 111 performs the ASR model 210 to generate a speech recognition result 215 of the audio data 120. The remote system 111 may also execute the NLU module 220 to perform semantic interpretation on the speech recognition result 215 to identify the query 122 for performing an operation on the digital assistant 109. Alternatively, the remote system 111 may transmit the speech recognition result 215 to the AED 102, and the AED 102 may execute the NLU module 220 to identify the query 122.
[0047] Continue to refer to Figure 1 Terminator 230 is configured to trigger the termination of a speech after a predetermined duration of non-speech in the audio data 120. Here, the predetermined duration of non-speech may correspond to a termination timeout duration, where terminator 230 will terminate the speech when at least the predetermined duration of non-speech is detected. That is, terminator 230 terminates the speech by making a hard microphone shutdown decision, instructing one or more microphones 106 at AED 102 to turn off and no longer capture streaming audio 118. The termination timeout duration is typically set to a default value that is long enough to prevent premature termination of the speech so that the content of the speech is not cut off before the user finishes speaking. However, while setting a longer termination timeout duration allows for longer pauses between words in the speech and prevents processing of incomplete phrases, the microphone of an assistant-enabled device remains on and may detect sounds not directed at an assistant-enabled device. Furthermore, delaying microphone shutdown thus delays the execution of the specified action / operation.
[0048] While the speech recognizer 200 is processing the audio data 120 and before the terminator 230 detects a predetermined duration of non-speech in the audio data, the frozen word classifier 160 simultaneously runs on the AED 102 and detects the frozen word 123 "end broadcast" in the audio data 120. Here, the frozen word 123 "end broadcast" follows the query 122 at the end of the utterance 119 spoken by the user 10 and corresponds to an action-specific frozen word 123. That is, the frozen word 123 "end broadcast" is specific to the action / operation of broadcasting a notification or message via a speaker. In some examples, the NLU module 220 provides instructions 222 to the acoustic feature detector 110 to activate / enable the frozen word 123 "end broadcast" in response to the query 122 that determines the speech recognition result 215 of the audio data 120 includes an operation to perform a broadcast for the digital assistant 109. In these examples, the acoustic feature detector 110 may activate / enable a frozen word detection model configured to detect the frozen word 123 "end broadcast".
[0049] In some implementations, the frozen word classifier 160 running on AED 102 is configured to identify the frozen word 123 at the end of utterance 119 without performing speech recognition or semantic interpretation. For example, in this example, if the frozen word classifier 160 detects acoustic features in the audio data 120 that are characteristics of hot word 121, the frozen word classifier 160 can determine that the utterance 119 “Ok Google, broadcast I'm home end broadcast” includes the frozen word 123 “endbroadcast”. For example, the frozen word classifier 160 can detect that the utterance 119 “Ok Google, broadcast I'm home end broadcast” includes the frozen word 123 “end broadcast” based on generating an MFCC from the audio data and classifying it as an MFCC that includes characteristics similar to those of the frozen word 123 stored in the model of the frozen word classifier 160. As another example, the frozen word classifier 160 can detect that the utterance 119 “OkGoogle, broadcast I'm home end broadcast” contains the frozen word 123 “end broadcast” based on the following operation: generating Mel-scale filter group energy from the audio data, and classifying the Mel-scale filter group energy as including Mel-scale filter group energy that is similar to the characteristics of the hot word “Ok Google” stored in the model of the hot word classifier 150. The frozen word classifier 160 can generate a frozen word confidence score by processing the extracted audio features in the audio data 120, and determine that the audio data 120 corresponding to the utterance 119 includes the frozen word 123 when the frozen word confidence score meets the frozen word confidence threshold.
[0050] In phase B of the acoustic feature detector 110, in response to the frozen word classifier 160 detecting a frozen word 123 in the audio data 120 before the terminator 230 detects a non-speech duration in the audio data, the AED 102 may trigger a hard microphone shutdown event 125 at the AED 102, which prevents the AED 102 from capturing any streaming audio 118 after the frozen word 123. For example, triggering the hard microphone shutdown event 125 may include the AED 102 deactivating one or more microphones 106. Thus, if the user 10 utters the frozen word 123 as a manual cue to indicate when the user 10 has completed the spoken query 122, the hard microphone shutdown event 125 is triggered without waiting for the termination timeout duration to elapse so that the terminator 230 can terminate the utterance. Conversely, triggering the hard microphone shutdown event 125 in response to the detection of the frozen word 123 would cause the AED 102 to instruct the terminator 230 and / or the ASR model 210 to terminate the utterance immediately. Triggering the hard microphone mute event 125 also causes the AED 102 to instruct the ASR system 200 to cease any active processing of audio data and instruct the digital assistant 109 to complete the execution of the operation. As a result, speech recognition accuracy is improved because the microphone 106 does not capture subsequent speech or background noise after the user utters the freeze word 123. Latency is also improved because the utterance 119 is manually terminated to allow the digital assistant 109 to begin executing the operation specified by query 122 without having to wait for the termination timeout duration to elapse. In the illustrated example, the ASR system 200 provides output 250 to the digital assistant 109, which instructs the digital assistant 109 to execute the operation specified by query 122. Output 250 may include instructions for performing the operation.
[0051] In some cases, output 250 may also include a speech recognition result 215 corresponding to the audio data 120 of utterance 119. These cases may occur when the query 122 recognized by the ASR system 200 corresponds to a search query, in which case the speech recognition result 215 of the search query 122 is provided as output 250 to a search engine (not shown) to retrieve search results. For example, the utterance 119 “Hey Google, tell me the weather for tomorrow now Google” may include the hot word “Hey Google,” the conversational search query 122 “tell me the weather for tomorrow,” and the final freeze word “now Google.” The ASR system 200 may process the audio data 120 to generate the speech recognition result 215 of utterance 119 and perform semantic interpretation on the speech recognition result 215 to recognize the search query 122. Continuing this example, in response to the frozen word classifier 160 detecting the frozen word "now Google", AED 102 can trigger a hard microphone mute event 125, and ASR system 200 can extract the phrase "now Google" from the end of speech recognition result 215 (e.g., transcription 225) and provide speech recognition result 215 as a search query to a search engine to retrieve search results for tomorrow's weather forecast. In this example, the frozen word "now Google" can include a predefined frozen word 123 common to all users of a given language, which manually triggers the hard microphone mute event 125 when spoken while speech recognition is active.
[0052] In some implementations, the digital assistant 109 is able to continue the conversation, wherein the microphone 106 can remain on to accept subsequent queries from the user after the digital assistant 109 outputs a response to a previous query. For example, using the example above, the digital assistant 109 can output search results for tomorrow's weather forecast as synthesized speech audibly, and then instruct the microphone 106 to remain on so that the user 10 can say the subsequent query without having to repeat hot word 121 as a prefix to the subsequent query. In this example, if the user 10 has no subsequent query, the user 10 saying the phrase "Thanks Google" (or another phrase of one or more fixed terms) can be used as a freeze word 123 to trigger a hard microphone off event. Keeping the microphone 106 on for a fixed duration to accept subsequent queries that the user 10 may or may not say inevitably requires increased processing, as speech processing is active while the microphone 106 is on, thus increasing power consumption and / or bandwidth usage. Therefore, when user 10 says a frozen word, it can trigger a hard microphone shutdown event to prevent the AED102 from capturing unconscious speech and provide power and bandwidth savings, as the AED102 can switch to a low-power sleep or hibernation state.
[0053] In some examples, if user 10 utters a freeze word to mute microphone 106 and end the ongoing conversation, AED 102 temporarily raises the hot word detection threshold and / or ignores subsequent speech uttered by the same user 10 for a certain period of time. AED 102 may store a reference speaker embedding for user 10, which indicates the user's speech characteristics that can be compared with a valid speaker embedding extracted from the utterance. For example, the valid speaker embedding may be text-dependent, where the embedding is extracted from the uttered hot word, while the reference speaker embedding may be extracted from the same hot word uttered by user 10 once or more during one or more previous interactions with digital assistant 109. When the valid speaker embedding extracted from subsequent utterances matches user 10's reference speaker embedding, the utterance may be ignored if subsequent utterances are provided shortly after user 10 utters the freeze word to trigger hard microphone mute.
[0054] Figure 2A and 2B This illustrates a first instance of ASR engine 200 receiving audio data 120a. Figure 2A ), which corresponds to a dictation-based query 122 for the digital assistant 109 to dictate audible content 124; and a second instance 120b of audio data 120 ( Figure 2B This corresponds to utterances 119 and 119b of audible content 124. See also Figure 2AAED 102 captures the first instance 119a of the utterance 119 spoken by user 10, which includes “Hey Google, dictate a message to Aleks until I say I'm done”. In this example, “Hey Google” corresponds to hot word 121, the phrase “dictate a message to Aleks” corresponds to a dictation-based query 122 for the digital assistant 109 to dictate a message to Aleks, and the phrase “until I say I'm done” specifies a freeze word to terminate the audible content of the message 124, where the phrase “I'm done” corresponds to freeze word 123.
[0055] Acoustic feature detector 110 receives streaming audio 118 captured by one or more microphones 106 of AED 102, corresponding to a first instance 119a of utterance 119. Hot word classifier 150 determines that streaming audio 118 includes hot word 121. For example, hot word classifier 150 determines that streaming audio 118 includes hot word 121 “Hey Google”. After hot word classifier 150 determines that streaming audio 118 includes hot word 121, AED 102 triggers a wake-up process to initiate speech recognition on a first instance 120a of audio data 120 corresponding to the first instance 119a of utterance 119.
[0056] The ASR 200 receives a first instance 120a of audio data 120 from the acoustic feature detector 110. The ASR model 210 can process the first instance 120a of audio data 120 to generate a speech recognition result 215. For example, the ASR model 210 receives the first instance 120a of audio data 120 corresponding to the utterance 119a “dictate a message to Aleks until I say I'm done” and generates the corresponding speech recognition result 215. The NLU module 220 can receive the speech recognition result 215 from the ASR model 210 and perform semantic interpretation on the speech recognition result 215 to determine whether the first instance 120a of audio data 120 includes a dictation-based query 122 for the digital assistant 109 to dictate audible content 124 spoken by the user 10. Specifically, the semantic interpretation performed by the NLU module 220 on the speech recognition result 215 identifies the phrase "dictate a message to Aleks" as a dictation-based query 122 for audible content 124 of a message (e.g., an electronic message or email) dictated by the digital assistant 109 to the recipient Aleks. In addition to messages, the dictation-based query 122 can be associated with dictating other types of content, such as audible content corresponding to diary entries or notes to be stored in a document.
[0057] In some implementations, ASR 200 further determines the freeze word 123 specified by the dictation-based query 122 based on a semantic interpretation performed on the speech recognition result 215 of a first instance of audio data 120. For example, in the illustrated example, NLU module 220 recognizes the phrase “until I say I'm done” as an instruction to set the phrase “I'm done” as the freeze word 123 for arguing the message to end the audible content 124. In some examples, NLU module 220 provides instruction 222 to acoustic feature detector 110 to activate / enable the freeze word 123 “I'm done” in response to determining that the speech recognition result 215 of the first instance of audio data 120a specifies the freeze word 123. In these examples, acoustic feature detector 110 may activate / enable a freeze word classifier (e.g., a freeze word detection model) 160 to detect the freeze word 123 “I'm done” in subsequent streaming audio 118 captured by AED 102.
[0058] In some examples, the frozen word classifier 160 and the hot word classifier 150 will never be active at the same time. Figure 2AThe dashed line surrounding the frozen word classifier 160 indicates that the frozen word classifier 160 is currently inactive, while the solid line surrounding the hot word classifier 150 indicates that the hot word classifier 150 is active. For example, before the NLU module 220 sends instruction 222 to the acoustic feature detector 110 to activate / enable the frozen word classifier 160 to detect the frozen word 123 “I'm done”, the hot word classifier 150 may be active (e.g., indicated by the solid line) to listen to the hot word 121 in the streaming audio 118 and the frozen word classifier 160 may be inactive (e.g., indicated by the dashed line). Once the NLU module 220 determines that the speech recognition result 215 includes a dictation-based query 122 and that the dictation-based query 122 specifies a frozen word 123, the NLU module 220 sends an instruction 222 to the acoustic feature detector 110 to activate the frozen word classifier 160 to detect the frozen word “I’m done” in subsequent streaming audio 118 and deactivate the hot word classifier 150.
[0059] In the example shown, the freeze word 123 “I'm done” corresponds to a query-specific freeze word specified as part of the query 122 spoken by user 10. It is noteworthy that the freeze word 123 “I'm done” is specified based on the dictation query 122 before the user begins speaking the audible content 124 of the message. Here, the freeze word “I'm done” specified as part of the dictation-based query instructs the terminator to wait or at least extend the termination timeout duration to trigger termination until the freeze word “I'm done” is detected. In some implementations, the NLU module 220 sends instruction 224 to the terminator 230 to increase the termination timeout duration. Once user 10 begins speaking the audible content 124 of the message, extending the termination timeout duration allows for a longer pause, which would otherwise trigger termination. In some examples, query-specific freeze words are enabled in parallel with action-specific freeze words (e.g., “EndMessage”) and / or user-specified freeze words (e.g., “The End”) and / or predefined freeze words (e.g., “Thanks Google”).
[0060] If the NLU module 220 determines that the speech recognition result 215 includes a dictation-based query 122 but the dictation-based query 122 does not specify a query-specific frozen word 123, the NLU module 220 will not send an instruction 222 to the acoustic feature detector 110 to activate / enable the frozen word classifier 160 to detect any query-specific frozen words, because the query 122 does not specify any frozen words. However, the NLU module 220 may still send an instruction 222 to the acoustic feature detector 110 to activate / enable the frozen word classifier 160 to detect at least one of action-specific frozen words, user-defined frozen words, or predefined frozen words. Optionally, the acoustic feature detector 110 may automatically activate / enable the frozen word classifier 160 to detect user-defined and / or predefined frozen words in subsequent streaming audio 118 when it detects a hot word 121 corresponding to the first instance 119a of utterance 119 in the streaming audio 118.
[0061] Now for reference Figure 2B After user 10 utters a first instance 119a of dictation-based query 122, conveying hot word 121 and specifying query-specific freeze word 123, user 10 then utters a second instance 119b of utterance 119 to convey audible content 124 of the message that user 10 wants digital assistant 109 to dictate, followed by query-specific freeze word 123 indicating that user 10 has completed uttering the audible content 124 of the message. Notably, user 10 does not need to prepend hot word 121 before the second instance 119b of utterance 119 because AED 102 is now awake and ASR 200 remains active in response to hot word classifier 150 detecting hot word 121 "Hey Google" in the first instance 119a of utterance 119. In the example shown, the second instance 119b of utterance 119 includes “Aleks, I'm running late, I'm done.” In this example, the phrase “Aleks, I'm running late” corresponds to the audible content 124 of the message, while the phrase “I'm done” corresponds to… Figure 2A The query-specific freeze word 123 is specified by the dictation-based query 122 in the first instance 119a of the utterance 119 spoken by user 10. Instead of the query-specific freeze word "I'm done" following the audible content 124, other types of freeze words 123 can follow the audible content 124 to similarly trigger the termination of the audible content 124.
[0062] The acoustic signature detector 110, executed on the AED 102, receives streaming audio 118 captured by one or more microphones 106 of the AED, corresponding to a second instance 119b of the utterance 119. The hot word classifier 150 is now inactive (e.g., as indicated by the dashed line) and the frozen word classifier 160 is now responding to the acoustic signature detector 110 from... Figure 2A The NLU module 220 in the system is active (e.g., as indicated by solid lines) upon receiving instruction 222 (for activating / enabling the frozen word classifier 160 to listen for the presence of a query-specific frozen word 123 in the streaming audio 118). The acoustic feature detector 110 utilizes the frozen word classifier 160 to determine whether the streaming audio 118 includes the frozen word 123. The acoustic feature detector 110 transmits a second instance 120b of audio data 120 to the ASR 200. The ASR 200 receives the second instance 120b of audio data 120, which corresponds to a second instance 119b of utterance 119 of audible content 124 spoken by user 10 and captured by AED 102. Furthermore, in response to the specified... Figure 2A The dictation-based query 122 of the frozen word 123 “I'm Done” in the first instance 119a of the utterance 119 receives instruction 224 from the NLU module 220, and the terminator 230 is applying an extended termination timeout duration.
[0063] ASR 200 processes a second instance 120b of audio data 120 to generate a transcription 225 of audible content 124. For example, ASR 200 generates a transcription 225 of audible content 124 “Aleks, I'm runninglate”. During the processing of the second instance 120b of audio data 120 at ASR 200, acoustic feature detector 110 detects a frozen word 123 in the second instance 120b of audio data 120. Specifically, a frozen word classifier (e.g., a frozen word detection model) 160 detects the presence of frozen word 123 in the second instance 120b of audio data 120. In the example shown, frozen word 123 includes the query-specific frozen word “I'm Done” to indicate the end of audible content 124. Frozen word 123 follows audible content 124 in a second instance 119a of utterance 119 spoken by user 10.
[0064] In response to the detection of a frozen word 123 in the second instance 120b of audio data 120, ASR 200 provides a transcription 225 of the audible content 124 spoken by user 10 for output from AED 102. AED 102 can output transcription 225 by transmitting transcription 225 to a receiver device (not shown) associated with receiver Aleks. In the case of transcription 225 dictating audible content 124 related to notes or diary entries, AED 102 can provide transcription 225 for output by storing transcription 225 in a document or sending transcription 225 to an associated application. Furthermore, AED 102 can output transcription 225 by displaying the transcription on the AED's graphical user interface (if available). Here, user 10 can view transcription 225 before sending it to the receiver device if user 10 wants to re-dictate transcription 225, correct any erroneous words in the transcription, and / or change any content of the message. Alternatively or concurrently, AED 102 may employ a text-to-speech (TTS) module to convert transcription 225 into synthesized speech for audible playback to user 10, enabling user 10 to confirm the transcription 225 that user 10 intends to send to the receiver device. In ASR 200 in remote system 111 ( Figure 1 In the configuration when the server is executed, ASR 200 can transmit transcription 225 to AED 102 and / or transmit transcription 225 to the receiver device. That is, ASR 200 provides transcription 225 "Aleks, I'm runninglate" of audible content 124 corresponding to the second instance 120b of audio data 120.
[0065] In some examples, in response to the detection of a frozen word 123 in a second instance 120b of audio data 120, acoustic feature detector 110 initiates / triggers a hard microphone shutdown event 125 at AED 102. The hard microphone shutdown event 125 prevents AED 102 from capturing any audio after the frozen word 123. That is, triggering the hard microphone shutdown event 125 at AED 102 may include AED 102 deactivating one or more microphones 106. Therefore, user 10 uttering the frozen word 123 as a manual cue indicating when user 10 has finished uttering the audible content 124 of the dictation-based query 122 triggers the hard microphone shutdown event 125 without waiting for the termination timeout duration to elapse, allowing terminator 230 to immediately terminate the second instance 119b of utterance 119. Alternatively, triggering the hard microphone shutdown event 125 in response to the detection of the frozen word 123 causes AED 102 to instruct terminator 230 and / or ASR model 210 to immediately terminate the utterance. Triggering the hard microphone shutdown event 125 also causes the AED 102 to instruct the ASR system 200 to stop any active processing of the second instance 120b of the audio data 120 and instruct the digital assistant 109 to complete the execution of the operation.
[0066] In some additional embodiments, supplementing or replacing the frozen word classifier 160 of the acoustic feature detector 110, the ASR system 200 detects the presence of frozen word 123 in a second instance 120b of audio data 120. That is, since the hot word classifier 150 has detected hot word 121 in a first instance 120a of audio data 120, the ASR 200 is already actively processing the second instance 120b of audio data 120b, so the ASR 200 is able to identify the presence of frozen word 123 in the second instance 120b of audio data 120. Therefore, the ASR 200 can be configured to initiate a hard microphone shutdown event 125 at the AED 200, stop active processing of the second instance 120b of audio data 120, and remove the identified frozen word 123 from the end of transcription 225. To further extend this capability of the ASR system 200, the frozen word classifier 160 can run on the AED 102 as a first-level frozen word detector, and the ASR system 200 can be used as a second-level frozen word detector to confirm the presence of frozen words detected by the frozen word classifier 160 in the audio data.
[0067] In some implementations, when processing the second instance 120b of audio data 120 to generate a transcription 225 of audible content 124, ASR 200 also transcribes the freeze word 123 to include in the transcription 225. For example, the transcription 225 of audible content 124 may include “Aleks, I'm running late I'm done”. Here, the transcription 225 of audible content 124 inadvertently includes the freeze word 123 “I'm done” as part of the audible content 124 of the message. That is, the user 10 does not intend for the digital assistant 109 to transcribe the freeze word 123 as part of the audible content 124 to be included in the transcription 225, but is instead told to specify the end of the audible content 124. Therefore, in response to the hard microphone shutdown event 125 initiated by the detection of the frozen word 123 at AED 102, ASR 200 may cause to remove the frozen word 123 from the end of the transcript 225 before providing the transcript 225 of audible content 124 for output from AED 102. Alternatively or additionally, ASR 200 may recognize the presence of the frozen word 123 at the end of the transcript 225 and remove the frozen word 123 from the end of the transcript 225 accordingly. In the example shown, ASR 200 removes the frozen word 123 from the transcript 225 “Aleks, I'm running late I'm done” before providing the transcript 225 to output 250. Therefore, after ASR 200 removes the frozen word 123 from the transcript 225, ASR 200 provides the transcript 225 “Aleks, I'm running late” for output 250 from AED 102.
[0068] Figure 3 This is a flowchart illustrating an exemplary arrangement of operations for method 300 for detecting frozen words. At operation 302, method 300 includes receiving audio data 120 at data processing hardware 113, the audio data 120 corresponding to utterance 119 spoken by user 10 and captured by user device 102 associated with user 10. At operation 304, method 300 includes processing the audio data 120 by the data processing hardware 113 using a speech recognizer 200 to determine that the utterance 119 includes a query 122 to perform an operation on digital assistant 109. The speech recognizer 200 is configured to trigger the termination of utterance 119 after a predetermined duration of non-speech in the audio data 120.
[0069] At operation 306, prior to a predetermined duration of non-speech in the audio data 120, method 300 includes detecting a freeze word 123 in the audio data 120 by data processing hardware 113. The freeze word 123 follows a query 122 in utterance 119 spoken by user 10 and captured by user device 102. At operation 308, in response to the detection of the freeze word 123 in the audio data 120, method 300 includes triggering a hard microphone mute event 125 at user device 102 by data processing hardware 113. The hard microphone mute event 125 prevents user device 102 from capturing any audio following the freeze word 123.
[0070] Figure 4 This is a flowchart illustrating an exemplary arrangement of operations for method 400 for detecting frozen words. At operation 402, method 400 includes a first instance 119a of receiving audio data 120 at data processing hardware 113, corresponding to a dictation-based query 122 for dictating audible content 124 spoken by user 10 to digital assistant 109. The dictation-based query 122 is spoken by user 10 and captured by assistant-enabled device (AED) 102 associated with user 10. At operation 404, method 400 includes a second instance 120b of receiving audio data 120 at data processing hardware 113, corresponding to utterance 119 of audible content 124 spoken by user 10 and captured by assistant-enabled device 102. At operation 406, method 400 includes processing the second instance 120b of audio data 120 by data processing hardware 113 using speech recognizer 200 to generate a transcription 225 of audible content 124.
[0071] At operation 408, during the processing of the second instance 120b of audio data 120, method 400 includes detecting a frozen word 123 in the second instance 120b of audio data 120 by data processing hardware 113. The frozen word 123 follows audible content 124 in utterance 119 spoken by user 10 and captured by assistant-enabled device 102. At operation 410, in response to the detection of frozen word 123 in the second instance 120b of audio data 120, method 400 includes providing a transcription 225 of the audible content 124 spoken by user 10 by data processing hardware 113 for use in output 250 from assistant-enabled device 102.
[0072] A software application (i.e., a software resource) can refer to computer software that instructs a computing device to perform a task. In some examples, a software application may be referred to as an "application," "app," or "program." Example applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0073] Non-transitory memory can be a physical device used to store programs (e.g., instruction sequences) or data (e.g., program state information) on a temporary or permanent basis for use by a computing device. Non-transitory memory can be volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., commonly used in firmware, such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase-change memory (PCM), and magnetic disks or magnetic tapes.
[0074] Figure 5 This is a schematic diagram of an example computing device 500 that can be used to implement the systems and methods described in this document. The computing device 500 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the ways in which the inventions described and / or claimed in this document can be implemented.
[0075] Computing device 500 includes a processor 510, memory 520, storage device 530, a high-speed interface 540 / high-speed controller connected to memory 520 and high-speed expansion port 550, and a low-speed interface 560 / low-speed controller connected to low-speed bus 570 and storage device 530. Each of components 510, 520, 530, 540, 550, and 560 is interconnected using various buses and may be mounted on a common motherboard or otherwise, as appropriate. Processor 510 may include data processing hardware 103, 113 of user device 102 or remote system 111. Processor 510 is capable of processing instructions for execution within computing device 500, including instructions stored in memory 520 or on storage device 530, to display graphical information for a graphical user interface (GUI) on an external input / output device such as display 580 coupled to high-speed interface 540. In other implementations, multiple processors and / or multiple buses may be used appropriately, along with multiple memories and memory types. In addition, multiple computing devices 500 can be connected, with each device providing a portion of the necessary operation (e.g., as a server group, blade server group, or multiprocessor system).
[0076] Memory 520 stores information non-transitory within computing device 500. Memory 520 may include memory hardware 105, 115 of user equipment 102 or remote system 111. Memory hardware may be computer-readable media, volatile memory cells, or non-volatile memory cells. Non-transitory memory may be a physical device used to store programs (e.g., instruction sequences) or data (e.g., program state information) on a temporary or permanent basis for use by computing device 500. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., commonly used for firmware, such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase-change memory (PCM), and magnetic disks or magnetic tapes.
[0077] Storage device 530 provides large-capacity storage for computing device 500. In some implementations, storage device 530 may be a computer-readable medium. In various implementations, storage device 530 may be a floppy disk device, hard disk device, optical disk device, magnetic tape device, flash memory or other similar solid-state storage device, or a device array, including devices in a storage area network or other configuration. In other implementations, a computer program product is tangibly embodied as an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer or machine-readable medium, such as memory 520, storage device 530, or memory on processor 510.
[0078] High-speed interface 540 manages bandwidth-intensive operations of computing device 500, while low-speed interface 560 manages less bandwidth-intensive operations. This allocation of responsibilities is merely exemplary. In some implementations, high-speed interface 540 is coupled to memory 520, display 580 (e.g., via a graphics processor or accelerator), and high-speed expansion port 550 which can accept various expansion cards (not shown). In some implementations, low-speed interface 560 is coupled to storage device 530 and low-speed expansion port 590. Low-speed expansion port 590, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, Wireless Ethernet), may be coupled to one or more input / output devices, such as keyboards, pointing devices, scanners, or networking devices, such as switches or routers, for example, via a network adapter.
[0079] As shown in the figure, the computing device 500 can be implemented in a variety of different forms. For example, it can be implemented as a standard server 500a or multiple times in a group of such servers 500a, as a laptop computer 500b, or as part of a rack server system 500c.
[0080] The various implementations of the systems and techniques described herein can be implemented in the form of digital electronic and / or optical circuits, integrated circuits, specially designed ASICs (Application-Specific Integrated Circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can be included in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be dedicated or general-purpose, coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to transfer data and instructions to the storage system, at least one input device, and at least one output device.
[0081] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages and / or in assembly / machine language. As used herein, the terms “machine-readable medium” and “computer-readable medium” mean any computer program product, non-transitory computer-readable medium, means and / or devices for providing machine instructions and / or data to a programmable processor (e.g., disks, optical disks, memory, programmable logic devices (PLDs)), including machine-readable media that receive machine instructions as machine-readable signals. The term “machine-readable signal” means any signal used to provide machine instructions and / or data to a programmable processor.
[0082] The processes and logical flows described in this specification can be executed by one or more programmable processors (also known as data processing hardware) that execute one or more computer programs to perform functions by manipulating input data and generating output. The processes and logical flows can also be executed by special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors of any kind of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic elements of a computer are a processor for executing instructions and one or more storage devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, or operatively coupled to receive data from or transfer data to said mass storage devices, or both, such as magnetic disks, magneto-optical disks, or optical disks. However, a computer does not necessarily have to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices like EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. Processors and memory can be supplemented by or incorporated into dedicated logic circuitry.
[0083] To provide interaction with the user, one or more aspects of this disclosure can be implemented on a computer having a display device or a touchscreen for displaying information to the user, and optionally a keyboard and pointing device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and a pointing device such as a mouse and trackball, through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual, auditory, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. Additionally, the computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending web pages to a web browser on the user's client device in response to a request received from a web browser.
[0084] Many implementations have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of this disclosure. Therefore, other implementations are also within the scope of the appended claims.
Claims
1. A method for detecting frozen words, comprising: Audio data is received at the data processing hardware, the audio data corresponding to words spoken by the user and captured by a user device associated with the user; The data processing hardware processes the audio data using a speech recognizer to determine that the utterance includes a query to perform an operation for a digital assistant, wherein the speech recognizer is configured to trigger the termination of the utterance after a predetermined duration of non-speech in the audio data; and Before the predetermined duration of non-speech components in the audio data: The data processing hardware detects frozen words in the audio data, the frozen words following the query in the utterance spoken by the user and captured by the user device; and In response to the detection of the frozen word in the audio data: The data processing hardware triggers a hard microphone shutdown event at the user equipment to prevent the user equipment from capturing any further audio data of the utterance following the frozen word; The data processing hardware modifies the speech recognition result by stripping the frozen word from the speech recognition result of the audio data; and The speech recognition results are provided by the data processing hardware and output from the user equipment.
2. The method according to claim 1, wherein, The frozen words include one of the following: Predefined frozen words, which include one or more fixed terms for a given language across all users; User-selected frozen words, the user-selected frozen words including one or more terms specified by the user of the user device; or Action-specific frozen words, which are associated with the action to be performed by the digital assistant.
3. The method according to claim 1, wherein, Detecting the frozen word in the audio data includes: Extract audio features from the audio data; A frozen word detection model is used to generate a frozen word confidence score by processing the extracted audio features; the frozen word detection model is executed on the data processing hardware; and When the confidence score of the frozen word meets the confidence threshold of the frozen word, it is determined that the audio data corresponding to the utterance includes the frozen word.
4. The method according to claim 1, wherein, Detecting the frozen word in the audio data includes: using the speech recognition device executed on the data processing hardware to identify the frozen word in the audio data.
5. The method of claim 1, further comprising responding to detecting the frozen word in the audio data: The data processing hardware instructs the speech recognizer to cease any active processing of the audio data; and The digital assistant executes the operation by being instructed by the data processing hardware.
6. The method according to claim 1, wherein, Processing the audio data to determine the utterance includes the query performed on the digital assistant, which includes: The speech recognizer is used to process the audio data to generate a speech recognition result of the audio data; and Perform semantic interpretation on the speech recognition results of the audio data to determine whether the audio data includes the query that performs the operation.
7. The method of claim 6, further comprising responding to detecting the frozen word in the audio data: The data processing hardware uses the voice recognition results to instruct the digital assistant to perform the operation of the query request.
8. The method of claim 1, further comprising, before processing the audio data using the speech recognizer: The data processing hardware uses a hot word detection model to detect hot words in the audio data preceding the query; and In response to the detection of the hot word, the data processing hardware triggers the speech recognizer to process the audio data by performing speech recognition on the hot word and / or one or more words in the audio data that follow the hot word.
9. The method according to claim 8, further comprising: The data processing hardware verifies the existence of the hot word detected by the hot word detection model based on the detection of the frozen word in the audio data.
10. The method according to claim 8, wherein: Detecting the frozen word in the audio data includes executing a frozen word detection model on the data processing hardware configured to detect the frozen word in the audio data without performing speech recognition on the audio data; and The frozen word detection model and the hot word detection model each include the same or different neural network-based models.
11. A method for detecting frozen words, the method comprising: A first instance of audio data is received at the data processing hardware, the first instance of audio data corresponding to a dictation-based query for audible content spoken by a digital assistant dictating a user, the dictation-based query being spoken by the user and captured by an assistant-enabled device associated with the user. A second instance of the audio data is received at the data processing hardware, the second instance of the audio data corresponding to the audible content spoken by the user and captured by the assistant-enabled device; The second instance of the audio data is processed by the data processing hardware using a speech recognition device to generate a transcription of the audible content; as well as During the processing of the second instance of the audio data: The data processing hardware detects frozen words in the second instance of the audio data, the frozen words being appended to the audible content of the utterance spoken by the user and captured by the assistant-enabled device; as well as In response to the detection of the frozen word in the second instance of the audio data: The data processing hardware removes the frozen words from the end of the transcription of the audible content spoken by the user; as well as The transcription of the audible content spoken by the user is provided by the data processing hardware for output from the assistant-enabled device.
12. The method of claim 11, further comprising responding to detecting the frozen word in the second instance of the audio data: The data processing hardware initiates a hard microphone shutdown event at the assistant-enabled device to prevent the assistant-enabled device from capturing any further audio of the utterance following the frozen word; and The data processing hardware stops any active processing of the second instance of the audio data.
13. The method of claim 11, further comprising: The first instance of the audio data being processed by the data processing hardware using the speech recognizer to generate a speech recognition result; as well as The data processing hardware performs semantic interpretation on the speech recognition result of the first instance of the audio data to determine that the first instance of the audio data includes the dictation-based query of dictating the audible content spoken by the user.
14. The method of claim 13, further comprising, before initiating processing of the second instance of the audio data to generate the transcription: The semantic interpretation performed by the data processing hardware based on the speech recognition result of the first instance of the audio data determines the frozen word specified in the dictation-based query; and The data processing hardware instruction terminator increases the termination timeout duration of the utterance used to terminate the audible content.
15. A system for detecting frozen words, the system comprising: Data processing hardware; as well as A memory hardware that communicates with the data processing hardware, the memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations, the operations including: Receive audio data, the audio data corresponding to utterances spoken by a user and captured by a user device associated with the user; The audio data is processed using a speech recognizer to determine that the utterance includes a query to perform an action for a digital assistant, wherein the speech recognizer is configured to trigger the termination of the utterance after a predetermined duration of non-speech in the audio data; and Before the predetermined duration of non-speech components in the audio data: Detecting frozen words in the audio data, the frozen words being appended to the query in the utterance spoken by the user and captured by the user device; and In response to the detection of the frozen word in the audio data: Trigger a hard microphone shutdown event at the user device to prevent the user device from capturing any further audio data of the utterance following the frozen word; The data processing hardware modifies the speech recognition result by stripping the frozen word from the speech recognition result of the audio data; and The speech recognition results are provided for output by the user equipment.
16. The system according to claim 15, wherein, The frozen words include one of the following: Predefined frozen words, which include one or more fixed terms for a given language across all users; User-selected frozen words, the user-selected frozen words including one or more terms specified by the user of the user device; or Action-specific frozen words, which are associated with the action to be performed by the digital assistant.
17. The system according to claim 15, wherein, Detecting the frozen word in the audio data includes: Extract audio features from the audio data; A frozen word detection model is used to generate a frozen word confidence score by processing the extracted audio features; the frozen word detection model is executed on the data processing hardware; and When the confidence score of the frozen word meets the confidence threshold of the frozen word, it is determined that the audio data corresponding to the utterance includes the frozen word.
18. The system according to claim 15, wherein, Detecting the frozen word in the audio data includes: using the speech recognition device executed on the data processing hardware to identify the frozen word in the audio data.
19. The system according to claim 15, wherein, The operation also includes responding to the detection of the frozen word in the audio data: The speech recognizer is instructed to cease any active processing of the audio data; and The digital assistant is instructed to perform the operation.
20. The system according to claim 15, wherein, Processing the audio data to determine the utterance includes the query performed on the digital assistant, which includes: The speech recognizer is used to process the audio data to generate a speech recognition result of the audio data; and Perform semantic interpretation on the speech recognition results of the audio data to determine whether the audio data includes the query that performs the operation.
21. The system according to claim 20, wherein, The operation also includes responding to the detection of the frozen word in the audio data: The digital assistant is instructed to perform the operation of the query request using the modified voice recognition results.
22. The system according to claim 15, wherein, The operation also includes, before processing the audio data using the speech recognizer: Use a hot word detection model to detect hot words in the audio data preceding the query; and In response to the detection of the hot word, the speech recognizer is triggered to process the audio data by performing speech recognition on the hot word and / or one or more words in the audio data that follow the hot word.
23. The system according to claim 22, wherein, The operation further includes: verifying the existence of the hot word detected by the hot word detection model based on the detection of the frozen word in the audio data.
24. The system according to claim 22, wherein: Detecting the frozen word in the audio data includes executing a frozen word detection model on the data processing hardware configured to detect the frozen word in the audio data without performing speech recognition on the audio data; and The frozen word detection model and the hot word detection model each include the same or different neural network-based models.
25. A system for detecting frozen words, the system comprising: Data processing hardware; as well as A memory hardware that communicates with the data processing hardware, the memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations, the operations including: A first instance of receiving audio data, the first instance of audio data corresponding to a dictation-based query for audible content spoken by a digital assistant dictating a user, the dictation-based query being spoken by the user and captured by an assistant-enabled device associated with the user; A second instance of the audio data is received, the second instance of the audio data corresponding to the audible content spoken by the user and captured by the assistant-enabled device; The second instance of the audio data is processed using a speech recognition device to generate a transcription of the audible content; and During the processing of the second instance of the audio data: In the second instance of the audio data, a frozen word is detected, the frozen word being appended to the audible content of the utterance spoken by the user and captured by the assistant-enabled device; and In response to the detection of the frozen word in the second instance of the audio data: The data processing hardware removes the frozen word from the end of the transcription of the audible content spoken by the user; and The transcription of the audible content spoken by the user is provided for output from the assistant-enabled device.
26. The system according to claim 25, wherein, The operation further includes, in response to detecting the frozen word in the second instance of the audio data: Initiate a hard microphone shutdown event at the device that enables the assistant to prevent the device from capturing any further audio of the utterance following the frozen word; as well as Stop any active processing of the second instance of the audio data.
27. The system according to claim 25, wherein, The operation also includes: The first instance of processing the audio data using the speech recognizer to generate a speech recognition result; and Perform semantic interpretation on the speech recognition result (215) of the first instance of the audio data to determine that the first instance of the audio data includes the dictation-based query of dictating the audible content spoken by the user.
28. The system according to claim 27, wherein, The operation further includes, before initiating processing of the second instance of the audio data to generate the transcription: The semantic interpretation performed on the speech recognition result of the first instance of the audio data determines the frozen word specified in the dictation-based query; as well as The instruction terminator increases the termination timeout duration of the utterance used to terminate the audible content.
Citation Information
Patent Citations
Speech endpointing based on word comparisons
CN105006235A
Speech recognition power management
US20140163978A1
Enhanced speech endpointing
US20170069309A1