Adapting hotword recognition based on personalized negation words
The cascaded hotword detection system with negative hotword classification addresses user-specific false positives by adapting the first-stage detector, enhancing detection accuracy and user experience.
Patent Information
- Application Number
- JP2023530712
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-11-20
- Filing Date
- 2021-11-11
- Publication Date
- 2025-08-14
- Estimated Expiration
- 2041-11-11
AI Technical Summary
Existing voice-enabled devices suffer from false positive and false negative hotword detections due to variations in user voices, accents, and acoustic environments, leading to undesirable device wake-ups and user experience issues.
A cascaded hotword detection system with a first-stage detector on the device and a second-stage detector on a remote server, combined with a negative hotword classification mechanism, adapts to user-specific false positives by updating the first-stage detector to prevent recurring false alarms.
Reduces false positive hotword detections by personalizing the hotword model, improving user experience through reduced unnecessary device wake-ups and power consumption.
Smart Images

Figure 0007723744000001 
Figure 0007723744000002 
Figure 0007723744000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to adapting hot word recognition based on personalized negation words. [Background technology]
[0002] A voice-enabled environment (e.g., a home, a workplace, a school, an automobile, etc.) allows a user to speak queries or commands aloud to a computer-based system, which then successfully answers or responds to the query and / or performs a function based on the command. A voice-enabled environment can be implemented using a network of connected microphone devices distributed throughout various rooms or areas of the environment. These devices may use hot words to help identify when a given utterance is directed to the system, as opposed to an utterance directed to another individual present in the environment. Thus, a device may operate in a sleep or hibernation state and wake up only if a detected utterance contains a hot word. Typically, a system used to detect hot words in streaming audio generates a probability score indicating the probability that the hot word is present in the streaming audio. If the probability score meets a predetermined threshold, the device initiates a wake-up process. Summary of the Invention [Means for solving the problem]
[0003] One aspect of the present disclosure provides a method for adapting hotword recognition based on personalized negative words. The method includes receiving, in data processing hardware, audio data characterizing a hotword event detected by a first-stage hotword detector in streaming audio captured by a user device. The method also includes processing, by the data processing hardware, the audio data using a second-stage hotword detector to determine whether a hotword is detected by the second-stage hotword detector in the first segment of the audio data. If no hotword is detected by the second-stage hotword detector in the first segment of the audio data, the method includes classifying, by the data processing hardware, the first segment of the audio data as containing a negative hotword that caused the first-stage hotword detector to falsely detect a hotword event in the streaming audio. Based on the first segment of audio data classified as containing a negative hotword, the method includes updating, by the data processing hardware, the first-stage hotword detector to prevent triggering of a hotword event in subsequent audio data containing the negative hotword.
[0004] Implementations of the present disclosure include one or more of the following optional features: In some implementations, the method further includes, if no hotword is detected by the second-stage hotword detector in the first segment of the audio data, inhibiting, by data cleaning hardware, a wake-up process at the user device to process the hotword and one or more other terms following the hotword in the streaming audio, and determining, by the data processing hardware, whether an immediate follow-up query was provided by the user of the user device after inhibiting the wake-up process at the user device. In these implementations, classifying the first segment of the audio data as including a negative hotword is further based on determining that no follow-up query was provided by the user of the user device after inhibiting the wake-up process.
[0005] In some examples, if a hotword is detected by the second-stage hotword detector in the first segment of the audio data, the method further includes processing, by the data processing hardware, a second segment of the audio data following the first segment of the audio data to determine whether the second segment of the audio data indicates a verbal query-type utterance. In these examples, if the second audio segment of the audio data does not indicate a verbal query-type utterance, the method also includes classifying, by the data processing hardware, the first segment of the audio data as including a negative word, and updating, by the data processing hardware, based on the first segment of the audio data classified as including a negative hotword, to prevent triggering of a hotword event in subsequent audio data that includes a negative hotword. In some implementations, the method further includes determining, by the data processing hardware, whether an immediate follow-up query has been provided by a user of the user device if the second audio segment of the audio data does not indicate a verbal query-type utterance. Wherein, classifying the first segment of audio data as including a negative hotword is further based on determining that no follow-up query was provided by a user of the user device.If the second audio segment of the audio data indicates a verbal query-type utterance, the method also includes receiving, in the data processing hardware, a negative interaction result indicating that a user of the user device has interacted negatively with the result of the verbal query-type utterance provided to the user device; classifying, by the data processing hardware, the first segment of audio data as including a negative hotword based on the received negative interaction result; and updating, by the data processing hardware, the first stage hotword detector to prevent detection of a hotword event in subsequent audio data including a negative hotword based on the first segment of audio data classified as including a negative hotword.
[0006] In some examples, after receiving audio data characterizing the hotword event detected by the first-stage hotword detector, the method further includes receiving, in the data processing hardware, a negative user interaction indicative of a user's inhibition of a wake-up process at the user device, wherein classifying the first segment of audio data as including a negative hotword is further based on the negative user interaction indicative of a user's inhibition of the wake-up process.
[0007] Optionally, updating the first-stage hotword detector to prevent triggering of a hotword event in subsequent audio data may include providing the user device with a first segment of audio data classified as containing a negative hotword. The user device is configured to retrain the first-stage hotword detector using the first segment of audio data classified as containing a negative hotword. In some implementations, the user device is configured to retrain the first-stage hotword detector by storing, in memory hardware of the user device, each instance of the first segment of audio data classified as containing a negative hotword and retraining the first-stage hotword detector based on a count of the number of instances of the first segment of audio data classified as containing a negative hotword stored in the memory hardware. In these implementations, the user device is further configured, before retraining the first-stage hotword detector, to determine that a corresponding confidence score associated with each instance of the first segment of audio data classified as containing a negative hotword does not satisfy a negative hotword threshold score and to determine that the number of instances exceeds a threshold number of instances.
[0008] Updating the first-stage hotword detector to prevent triggering of a hotword event in subsequent audio data may include providing a first segment of audio data classified as containing a negative hotword to a user device. The user device is configured to obtain an embedded representation of the first segment of audio data and store the embedded representation of the first segment of audio data in memory hardware of the user device. Additionally, the user device is configured to determine when subsequent audio data characterizing a hotword event detected by the first-stage hotword detector contains a negative hotword by calculating a reputation embedded representation for the audio data, determine a similarity score between the embedded representation of the first segment of audio data classified as a negative hotword and the reputation embedded representation for the subsequent audio data, and determine that the subsequent audio data contains a negative hotword if the similarity score satisfies a similarity score threshold.
[0009] In some implementations, the data processing hardware resides on a server in communication with the data processing hardware, and the first-stage hotword detector executes on a processor of the user device. Processing the audio data to determine if a hotword is detected by the second-stage hotword detector in the first segment of the audio data may include performing automatic speech recognition to determine if a hotword is recognized in the first segment of the audio data.
[0010] In some examples, the data processing resides on the user device. In these examples, the first-stage hotword detector may execute on a digital signal processor (DSP) of the data processing hardware, and the second-stage hotword detector executes on an application processor of the data processing hardware. The first-stage hotword detector may be configured to generate a probability score indicative of the presence of a hotword in audio characteristics of the streaming audio captured by the user device, and to detect a hotword event in the streaming audio if the probability score meets a hotword detection threshold of the first-stage hotword detector.
[0011] Another aspect of the present disclosure provides a system for adapting hotword recognition based on personalized negative words. The system includes data processing hardware and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include receiving audio data characterizing hotword events detected by a first-stage hotword detector in streaming audio captured by a user device. The operations also include processing the audio data using a second-stage hotword detector to determine whether hotwords are detected by the second-stage hotword detector in a first segment of the audio data. If no hotwords are detected by the second-stage hotword detector in the first segment of the audio data, the operations include classifying the first segment of the audio data as including negative hotwords that caused the first-stage hotword detector to falsely detect the hotword event in the streaming audio. Based on the first segment of audio data classified as including a negative hotword, the operations include updating the first stage hotword detector to prevent triggering a hotword event in subsequent audio data that includes the negative hotword.
[0012] Implementations of the present disclosure include one or more of the following optional features: In some implementations, the operations further include, if no hotword is detected by the second-stage hotword detector in the first segment of the audio data, inhibiting a wake-up process at the user device to process the hotword and one or more other terms following the hotword in the streaming audio, and determining whether an immediate follow-up query was provided by a user of the user device after inhibiting the wake-up process at the user device. In these implementations, the operation of classifying the first segment of the audio data as including a negative hotword is further based on determining that no follow-up query was provided by the user of the user device after inhibiting the wake-up process.
[0013] In some examples, if a hotword is detected by the second-stage hotword detector in the first segment of audio data, the operations further include processing a second segment of audio data following the first segment of audio data to determine whether the second segment of audio data indicates a verbal query-type utterance. In these examples, if the second audio segment of audio data does not indicate a verbal query-type utterance, the operations also include classifying the first segment of audio data as including a negative word and updating the first-stage hotword detector to prevent triggering of a hotword event in subsequent audio data including a negative hotword based on the first segment of audio data being classified as including a negative hotword. In some implementations, the operations further include determining whether an immediate follow-up query has been provided by the user of the user device if the second audio segment of audio data does not indicate a verbal query-type utterance. Here, classifying the first segment of audio data as including a negative hotword is further based on determining that a follow-up query has not been provided by the user of the user device. If the second audio segment of the audio data indicates a verbal query-type utterance, the operations also include receiving a negative interaction result indicating that a user of the user device engaged in a negative interaction with the result of the verbal query-type utterance provided to the user device; classifying the first segment of audio data as including a negative hotword based on the received negative interaction result; and updating the first stage hotword detector to prevent detecting a hotword event in subsequent audio data including a negative hotword based on the first segment of audio data classified as including a negative hotword.
[0014] In some examples, after receiving audio data characterizing the hotword event detected by the first-stage hotword detector, the operations further include receiving a negative user interaction indicative of user inhibition of a wake-up process at the user device, where classifying the first segment of audio data as including a negative hotword is further based on the negative user interaction indicative of user inhibition of the wake-up process.
[0015] Optionally, updating the first-stage hotword detector to prevent triggering of a hotword event in subsequent audio data may include providing the user device with a first segment of audio data classified as containing a negative hotword. The user device is configured to retrain the first-stage hotword detector using the first segment of audio data classified as containing a negative hotword. In some implementations, the user device is configured to retrain the first-stage hotword detector by storing, in memory hardware of the user device, each instance of the first segment of audio data classified as containing a negative hotword and retraining the first-stage hotword detector based on a count of the number of instances of the first segment of audio data classified as containing a negative hotword stored in the memory hardware. In these implementations, the user device is further configured, before retraining the first-stage hotword detector, to determine that a corresponding confidence score associated with each instance of the first segment of audio data classified as containing a negative hotword does not satisfy a negative hotword threshold score and to determine that the number of instances exceeds a threshold number of instances.
[0016] The operation of updating the first-stage hotword detector to prevent triggering of a hotword event in subsequent audio data may include an operation of providing a first segment of audio data classified as containing a negative hotword to a user device. The operation is configured to obtain an embedded representation of the first segment of audio data and store the embedded representation of the first segment of audio data in memory hardware of the user device. The user device is additionally configured to determine when subsequent audio data characterizing a hotword event detected by the first-stage hotword detector contains a negative hotword by calculating a reputation embedded representation for the audio data, determine a similarity score between the embedded representation of the first segment of audio data classified as a negative hotword and the reputation embedded representation for the subsequent audio data, and determine that the subsequent audio data contains a negative hotword if the similarity score satisfies a similarity score threshold.
[0017] In some implementations, the data processing hardware resides on a server in communication with the data processing hardware, and the first-stage hotword detector executes on a processor of the user device. Processing the audio data to determine if a hotword is detected by the second-stage hotword detector in the first segment of the audio data may include performing automatic speech recognition to determine if a hotword is recognized in the first segment of the audio data.
[0018] In some examples, the data processing resides on the user device. In these examples, the first-stage hotword detector may execute on a digital signal processor (DSP) of the data processing hardware, and the second-stage hotword detector executes on an application processor of the data processing hardware. The first-stage hotword detector may be configured to generate a probability score indicative of the presence of a hotword in audio characteristics of the streaming audio captured by the user device, and to detect a hotword event in the streaming audio if the probability score meets a hotword detection threshold of the first-stage hotword detector.
[0019] The details of one or more implementations of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will become apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]
[0020] [Figure 1] FIG. 1 is a schematic diagram of an example system for classifying negative hotwords and updating a first stage hotword detector to prevent detecting hotword events in audio containing negative hotwords. [Figure 2] FIG. 1 is a schematic diagram of a hot word detection architecture. [Figure 3] FIG. 2 is a schematic diagram of an example negative hotword classifier of the system of FIG. 1. [Figure 4] FIG. 1 is a schematic diagram of a user device that stores classification results for audio data classified as containing negative hot words. [Figure 5] FIG. 1 is a schematic diagram of an example user device that identifies the presence of personal negative hotwords in captured audio data to prevent triggering a hotword event in the audio data. [Figure 6]10 is a flowchart of an exemplary arrangement of operations for updating a first stage hotword detector to classify personalized negative hotwords and prevent triggering hotword events in audio data that includes personalized negative hotwords. [Figure 7] FIG. 1 is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION
[0021] Like reference symbols in the various drawings indicate like elements.
[0022] Voice-based interfaces, such as digital assistants, are becoming increasingly prevalent across a variety of devices, including, but not limited to, mobile phones and smart speakers / displays that include microphones for capturing voice. A common way to initiate a voice interaction with a voice-enabled device is to speak a fixed phrase, e.g., a hot word, which, when detected by the voice interaction device in streaming audio, triggers the voice-enabled device to begin a wake-up process to begin recording and processing subsequent voice to confirm the query spoken by the user. Therefore, as the primary entry point for a voice-based interface, it is important that hot word detection / recognition perform reliably in terms of both recall and precision so that the number of false wake-up events is minimized.
[0023] A false negative (also called a "false rejection") refers to the failure to detect a hot word in streaming audio spoken by a user when attempting to interact with a voice-based interface (e.g., a digital assistant). Here, the voice-enabled device fails to respond to the user and requests the user to attempt to invoke the interface again by re-speaking the hot word, often louder and with a different pronunciation, to ensure the hot word is detected. On the other hand, a false positive (also called a "false acceptance") generally refers to the detection of a hot word in streaming audio when the streaming audio does not actually contain the hot word, due to the streaming audio containing words / phrases that are phonetically similar to the hot word when spoken. A false positive causes the voice-enabled device to initiate a wake-up process even though the user did not intend to invoke the system, thereby surprising and / or confusing the user by reacting when the voice-enabled device should remain asleep.
[0024] A cascaded hotword detection architecture incorporates a first-stage hotword detector running on the device to detect the presence of hotwords in streaming audio and a second-stage hotword detector that confirms the presence of hotwords detected by the first-stage hotword detector. The second-stage hotword detector is associated with higher accuracy in detecting the presence of hotwords in streaming audio and therefore has higher power requirements than the first-stage hotword detector. Often, the second-stage hotword detector is implemented on a server that communicates with the first-stage hotword detector implemented on the voice-enabled device. Even during partial false negatives, if the first-stage hotword detector detects the presence of a hotword locally but the second-stage hotword detector on the server denies the presence of the hotword, even if the server ultimately suppresses the wake-up process, this has a negative impact on the user experience. That is, detection of a hotword by the first-stage hotword detector still causes a device wake-up and connection to the server that is noticeable to the user (e.g., a visible notification or flashing light), which is even more undesirable from a privacy and power conservation perspective. Therefore, it is desirable to eliminate the occurrence of partial false positive instances to improve the user experience.
[0025] Traditionally, voice-enabled devices use the same fixed hotword model for all users of a given language (or locale), which is periodically updated with new versions pushed to the device from the server. That is, the same hotword model is used to detect hotwords in streaming audio for all users, despite the presence of large variations across users' voices, accents, vocabulary, and / or the acoustic environments in which the voice-enabled devices operate. As a result, implementing stringent precision / recall requirements for detecting hotwords is nearly impossible when a single hotword model is shared across all users of a given language / locale.
[0026] For a given user and / or environment, false positive instances of a hotword are very likely to cluster. In a non-limiting example, a particular user's speaking the term "poodle" may cause a hotword detection model on a voice-enabled device to falsely detect the presence of the designated hotword "Hey Google," while a different user may cause the same hotword detection model implemented on another voice-enabled device to detect the designated hotword when the user speaks "doodle." The variability in these false positive instances across different users may be due to differences in users' pronunciations and / or the frequency of those terms in those users' respective vocabularies. Because the same false positive instances are likely to recur based on similar acoustic patterns for the same user in the same environment, a hotword detector should ideally learn to adapt to avoid repeating the same false positives multiple times when a user speaks a given thing with a pronunciation similar to the designated hotword.
[0027] Implementations herein are directed to personalizing a hotword detector on a user's voice-enabled device based on specific terms classified as negative hotwords that caused previous instances of false positive hotword detection. The specific terms classified as negative hotwords may be user-specific, such that any hotword detection is suppressed if an audio segment obtained from a specific user speaking the terms is detected by the hotword detector. Additionally or alternatively, the specific terms classified as negative hotwords may be device-specific, such that hotword detection is suppressed if an audio segment obtained from a user speaking the terms is detected by the hotword detector implemented on the specific device but not by hotword detectors implemented on other voice-enabled devices associated with the same user. This follows from the idea that devices located in some environments are more prone to false positive hotword detection than devices located in other environments due to acoustic variations, variations in vocabulary, and variations between users typically speaking in that environment.
[0028] 1 , in some implementations, an exemplary system 100 includes user devices 102 associated with one or more users 10 and communicating with a remote system via a network 104. The user devices 102 correspond to computing devices such as mobile phones, computers (laptop or desktop), tablets, smart speakers / displays, smart appliances, smart headphones, wearables, vehicle infotainment systems, etc., and include data processing hardware 103 and memory hardware 105. The user devices 102 include or communicate with one or more microphones 106 for capturing speech from the respective users 10. The remote system 110 can be a single computer, multiple computers, or a distributed system (e.g., a cloud environment) with scalable / elastic computing resources 112 (e.g., data processing hardware) and / or storage resources 114 (e.g., memory hardware).
[0029] The user device 102 includes a first-stage hotword detector 210 (also referred to as a hotword detection model) configured to detect the presence of hotwords in the streaming audio 118 without performing semantic analysis or speech recognition processing on the streaming audio 118. In some implementations, the user device 102 also includes an initial coarse hotword detector 205 ( FIG. 2 ) trained to initially detect the presence of hotwords before receiving the streaming audio 118 and invoking the first-stage (fine) hotword detector 210 to see if hotwords are detected in the streaming audio 118. In some implementations, the first-stage hotword detector 210 includes a trained neural network (e.g., a stored neural network) received from the remote system 110 over the network 104. The remote system 110 may push updates or new versions of the first-stage hotword detector 210 to a population of user devices 102 associated with users of a given local and / or specific language. That is, all user devices 102 associated with a US English-speaking user receive the same first-stage hot word detector 210, which may include a neural network trained to detect the presence of hot words in streaming audio 118 when spoken by an English-speaking user in the United States.
[0030] In some examples, a first-stage hotword detector 210 executing on the user device 102 is configured to detect the presence of the hotword "Hey Google" in the streaming audio 118 to initiate a wake-up process at the user device 102 to process the hotword and / or one or more other terms (e.g., queries or commands) following the hotword in the streaming audio 118. The first-stage hotword detector 210 may be configured to generate a probability score indicative of the presence of the hotword in audio features of the streaming audio 118 captured by the user device 102 and detect the hotword in the streaming audio 118 if the probability score meets a hotword detection threshold of the first-stage hotword detector 210. Thus, the first-stage hotword detector 210 may detect a hotword event in the streaming audio 118 captured by the user device 102 if the probability score generated for the audio features of the streaming audio 118 meets a hotword identification threshold.
[0031] In the illustrated example, user 10 speaks utterance 119 including a term / phrase (e.g., "Hey Poodle") captured as streaming audio 118 by user device 102 that has a similar pronunciation to a fixed term / phrase (e.g., "Hey Google") designated as a hot word that first-stage hot word detector 210 is trained to detect. In particular, user 10 may pronounce the term "Hey Poodle" in a way that makes it more difficult for the first-stage hot word detector to distinguish it from the designated hot word "Hey Google" than if it were spoken by another user using a slightly different pronunciation. As a result, first-stage hot word detector 210 executing on user device 102 may falsely detect the presence of the hot word by generating a probability score for audio features associated with "Hey Poodle" that meets a hot word detection threshold, thereby triggering the initiation of a wake-up process that was not intended by user 10.
[0032] If the probability score for the audio features associated with "Hey Poodle" meets the hotword detection threshold, the first-stage hotword detector 210 further transmits audio data 120 characterizing the hotword event to a second-stage hotword detector 220 executing on the remote system 110. In some examples, the audio data 120 is a direct representation of the streaming audio 118, while in other examples, the audio data 136 represents the streaming audio 118 after processing by the first-stage hotword detector 210 (e.g., to identify and / or isolate particular audio characteristics of the streaming audio 118 or to convert the streaming audio 118 into a format suitable for transmission and / or processing by the second-stage hotword detector 220). For example, the audio data 120 includes a first segment 121 chopped from the streaming audio 118 that includes relevant audio features associated with the presence of the hotword detected by the first-stage hotword detector 210. The audio data 120 also includes a second segment 122 that follows the first segment 121 and includes audio features captured by the user device 102 in the streaming audio 118. Typically, the first segment 121 is of a fixed duration sufficient to include audio features generally associated with a designated hotword. However, the second segment 122 may have a variable duration that includes audio captured by the user device 102 while the microphone 106 is open. The second segment 122 may capture query-type utterances that require further processing (e.g., automatic speech recognition and / or semantic interpretation) of one or more terms to identify a query or command within the audio data.In the exemplary scenario, the first segment 121 includes audio features captured by the user device that are not related to the user speaking the hotword, but are related to another term / phrase that the user 10 pronounces similarly to the hotword, so the user 10 does not intend to call the device 102 through voice, and therefore the second segment 122 is likely to not include query-type utterances, but instead to include background noise captured in the streaming audio 118 from the environment of the user device 102.
[0033] The first-stage hotword detector 210 and the second-stage hotword detector 220 cooperate to form a cascaded hotword detection architecture 220, whereby the second-stage hotword detector 220 is configured to determine whether a hotword detected by the first-stage hotword detector 210 is present in the audio data 120. Specifically, the second-stage hotword detector 220 executing on the remote system 110 processes the audio data 120 to determine whether a hotword is detected by the second-stage hotword detector in a first segment 121 of the audio data 120. In some examples, the second-stage hotword detector 220 is implemented as an automatic speech recognition (ASR) engine that performs speech recognition on the first segment 121 of the audio data 120 to determine whether a hotword is present. The second-stage hotword detector 220 may detect the presence of a hotword in the first segment 121 if the probability of recognizing the hotword meets a hotword detection threshold.
[0034] In other examples, the second-stage hotword detector 220 is similar to the first-stage hotword detector 210 in that the second-stage hotword detector 220 is a model implemented as a trained neural network configured to detect the presence of hotwords in the first audio segment 121 without performing semantic analysis or speech recognition processing. In these examples, the second-stage hotword detector 220 may be associated with a larger version of the hotword detection model used by the first-stage hotword detector 210 and may include a different neural network that is potentially more computationally intensive than the neural network of the first-stage hotword detector 210, thereby providing improved hotword detection accuracy over the first-stage hotword detector 210, which is limited by the resources of the user device 102. The second-stage hotword detector 220 may generate a probability score indicative of the presence of a hotword in the first segment of audio data 120 and may detect the presence of the hotword if the probability score meets a hotword detection threshold of the second-stage hotword detector 220. Here, the value of the hotword detection threshold in the second-stage hotword detector 220 may be the same as or different from the value of the hotword detection threshold in the first-stage hotword detector 210. In some examples, the value of the hotword detection threshold in the second-stage hotword detector 220 is set higher to require the second-stage hotword detector 220 to be more confident in determining whether a hotword is present in the first audio segment 121.
[0035] In some implementations, the second-stage hotword detector 220 executes on the user device 102 (e.g., data processing hardware 103) to implement the entire cascaded hotword detection architecture 200 on-device, without using the remote system 110. When executed on the user device 102, the second-stage hotword detector 220 can be implemented as an on-device ASR engine that detects the presence of hotwords by performing speech recognition on the first audio segment 121, or as a larger version of the hotword detection model implemented by the first-stage hotword detector 210 to detect the presence of hotwords in the first audio segment 121 without performing speech recognition.
[0036] 2 provides an example of the cascaded hotword detection architecture 200 of FIG. 1 , including a first-stage hotword detector 210, a second-stage hotword detector 220, and optionally an initial coarse hotword detector 205. In some examples, when the user device 102 is a battery-powered device, the data processing hardware 103 of the user device 102 collectively includes a first processor 60 (e.g., a digital signal processor (DSP)) and a second processor 70 (e.g., an application processor (AP)). The first processor 60 consumes less power during operation than the second processor 70 does during operation. As used herein, the first processor 60 may be interchangeably referred to as a DSP, and the second processor 70 may be interchangeably referred to as an “AP” or a “device SoC.” The initial coarse hotword detector 205 may run on the first processor 60, and the first-stage hotword detector 210 may run on the second processor 70. The second-tier hotword detector 220 may be executed on a server (e.g., the remote system 110) in communication with the user device 102 to provide server-side hotword verification that takes advantage of increased processing power at the server. Alternatively, the second-tier hotword detector 220 may be executed on the second processor 70 of the user device 102 to implement the entire cascaded hotword detection architecture on-device.
[0037] Generally, the coarse-grained hotword detector 205 resides on the dedicated DSP 60, includes a smaller model size than the model associated with the first-stage hotword detector 210, and is computationally efficient for coarsely screening the input streaming audio 118 for hotword detection. Therefore, the dedicated DSP 60 (e.g., the first processor) may be “always on” so that the coarse-grained hotword detector 205 is constantly running to coarsely screen hotword candidates within the multi-channel audio 118, while other components of the user device 102, including the main AP 70 (e.g., the second processor), are in a sleep state / mode to conserve battery life. Meanwhile, the first-stage hotword detector 210 resides on the main AP 70, includes a larger model size than the coarse-stage hotword detector 205, and provides more computational power than the coarse-stage hotword detector 205 to provide more accurate detection of hotwords initially detected by the coarse-grained hotword detector 205. Therefore, the first-stage hotword detector 210 may be more rigorous in determining whether hotwords are present within the audio 118. While the DSP 60 is "always on," the more power-hungry main AP 70 operates in sleep mode to preserve battery life until the coarse hotword detector 205 in the DSP 60 detects a candidate hotword in the streaming audio 118. Thus, once a candidate hotword is detected, the DSP 60 triggers the main AP 70 to transition from sleep mode to hotword detection mode to run the first stage hotword detector 210.
[0038] Upon receiving the streaming audio 118, the always-on DSP 60 executes / runs the coarse hotword detector 205 to determine whether hotwords are detected in the respective audio features of the streaming audio 118. In particular, the AP 70 may operate in a sleep mode when multi-channel audio is received at the DSP 60 and while the DSP 60 is processing the respective audio features of the streaming audio 118.
[0039] When the coarse hotword detector 205 detects a hotword in the streaming audio 118, the DSP 60 provides the truncated audio data 120 to the AP 70. In some examples, the DSP 60 providing the truncated audio data 120 to the AP 70 triggers / invokes the AP 70 to transition from sleep mode to hotword detection mode. In some implementations, the audio data 120 truncated from the streaming audio 118 includes a first segment 121 that characterizes the hotword detected by the coarse hotword detector 205 in the streaming audio 118. That is, the first segment includes a duration sufficient to safely contain the detected hotword. Additionally, the audio data 212 includes a second segment 122 following the first segment 121, which may include a duration of audio containing the spoken query. The coarse hotword detector 205 is optional, and the first stage hotword detector 210 may first detect hotword events in the streaming audio 118, crop the audio data 120 including the first and second segments 121, 122, and provide the cropped audio data 120 to the second stage hotword detector 220.
[0040] When a hotword is detected by the first-tier hotword detector 210 in the first segment 121 of the audio data 120, the AP 70 initiates a wake-up process at the user device 102 and provides the audio data 120 to the second-tier hotword detector 220 for processing to determine / confirm whether the second-tier hotword detector 220 detects a hotword in the first segment 121 of the audio data 120. In an example where the cascade hotword detection architecture 200 is implemented entirely on-device, the AP 70 simply passes the audio data 120 to the second-tier hotword detector 220, which also runs on the AP 70. In another example where the second-tier hotword detector 220 is implemented on the server 110, the AP 70 instructs the user device 102 to transmit the audio data 120 to the second-tier hotword detector 220 over the network 104.
[0041] 1 , in some examples, if no hotwords are detected by the second-tier hotword detector 220 in the first segment 121 of the audio data 120 (i.e., the probability score meets the hotword detection threshold), the negative hotword classifier 300 executing on the remote system 110 classifies the first segment 121 of the audio data 120 as containing a negative hotword (e.g., "Hey Poodle") that caused the first-tier hotword detector 210 to falsely detect a hotword event in the streaming audio 118. In configurations in which the second-tier hotword detector 220 executes on the user device 102, the negative hotword classifier 300 may also execute on the user device 102. The negative hotword classifier 300 may provide a classification result 170 to the negative hotword updater 400 indicating the classification of the first segment 121 of the audio data 120 as containing a negative hotword. The classification result 170 may provide a probability score 171 generated by the second-tier hotword detector 220 for the first audio segment 121 and / or other relevant information 172 related to the corresponding classification result 170 that may be useful for the negative hotword updater 400 to update the first-tier hotword detector 210. The other relevant information 172 may include, but is not limited to, a speaker identification score (e.g., a speaker embedding) that identifies the speaker characteristics of the user 10 who spoke the utterance 119, a timestamp of the hotword event indicating the time and / or day of the week, and a negative hotword confidence score 304 ( FIG. 3 ) that indicates the confidence for classifying the first segment 121 as a negative hotword. In the illustrated example, the negative hotword updater 400 executes at the user device 102 to personalize hotword detection at the user device 102 by updating the first-tier hotword detector 210 to prevent triggering hotword events in subsequent audio data that include negative hotwords.
[0042] Sending the classification result 170 to the negative hotword updater 400 may cause the negative hotword updater 400 to update the first-tier hotword detector 210 to prevent triggering a hotword event in subsequent audio data containing a negative hotword (e.g., "Hey Poodle"). In some implementations, if no hotword is detected by the second-tier hotword detector 220 in the first segment 121 of the audio data 120, the remote system 110 (or the user device 102) suppresses a wake-up process in the user device 102 to process the hotword and / or one or more other terms following the hotword in the streaming audio 118. In some implementations, the remote system 110 suppresses the wake-up process by sending an inhibit command 160 to the user device 102, causing the user device 102 to suppress the wake-up process. In other implementations, providing a classification result 170 indicating that the first segment 121 of the audio data 120 contains a negative hotword causes the user device 102 to suppress the wake-up process. In yet other implementations, the remote system 110 inhibits the wake-up process by not responding to the user device 102 (e.g., by closing the network connection) after receiving the audio data 120. The lack of response from the remote system 110 may cause the user device 102 to inhibit the wake-up process. That is, the user device 102 initiates the wake-up process only upon receiving confirmation from the second-stage hotword detector 220 that the hotword was present in the streaming audio 118, in some examples. The user device 102 may independently inhibit the wake-up process. For example, the user device 102 may automatically inhibit the wake-up process if the query or command following the hotword is empty (i.e., the streaming audio 118 following the hotword does not include a command or query directed to the user device 102).
[0043] In some scenarios, after the second-stage hotword detector 220 inhibits the wake-up process because it does not detect the presence of a hotword in the first segment 121 of the audio data 120, the negative hotword classifier 300 determines whether an immediate follow-up query was provided by the user 10 of the user device 102. This determination may be made when a subsequent hotword event detected by the first-stage hotword detector 210 is not received by the second-stage hotword detector 220. Here, the determination that the user 10 did not provide an immediate follow-up query serves as additional confirmation that the user 10 did not intend to speak a hotword in the previous utterance 119 and spoke a term ("Hey Poodle") having a pronunciation similar to the specific term / phrase ("Hey Google") designated as a hotword. Thus, classifying the first segment 121 of the audio data 120 as containing a negative hotword may be further based on a determination that a follow-up query was not received from the user device 102 after inhibiting the wake-up process.
[0044] In an additional example, if the second-stage hot word detector 220 detects the presence of a hot word (“Hey Google”) in the first segment 121 of the audio data 120 (i.e., the probability score meets the hot word detection threshold) even though the user 10 actually spoke another similar-sounding phrase (“Hey Poodle”), the second segment 122 of the audio data 120 (and optionally the first segment 121) is provided to the query processor 180. Here, the query processor 180 processes the second segment 122 of the audio data 120 to determine whether it indicates a verbal query-type utterance. In an example in which the second-stage hot word detector 220 is implemented as an ASR engine, the query processor 180 processes the resulting speech recognition results by performing semantic analysis to determine whether the second segment 122 indicates a query-type utterance. In another example, when the second-stage hotword detector 220 is implemented as a hotword detection model, the query processor 180 is implemented as an ASR engine that processes the second segment 122 of the audio data 120 by performing speech recognition and then performing semantic analysis on the speech recognition results. As used herein, a query-type utterance corresponds to an utterance directed to the user device 102, e.g., an utterance directed to a digital assistant interface to query the digital assistant to perform an operation or action. Thus, if the second segment 122 of the audio data 120 indicates a query-type utterance, there is a high likelihood that the second-stage hotword detector 220 was accurate in detecting the presence of a hotword in the first segment 121 of the audio data 120. Otherwise, if the query processor 180 determines that the second segment 122 does not indicate a query-type utterance, there is a high likelihood that the second-stage hotword detector 220 was inaccurate in detecting the presence of a hotword in the first segment 121.
[0045] The query processor 180 may provide a score 182 indicating whether the second segment 122 indicates query-type speech. In some examples, the score 182 is binary, with a score 182 of zero or one indicating query-type speech and a score 182 of the other of zero or one not indicating query-type speech. In other examples, the score 182 provides a likelihood (e.g., probability) that the second segment 122 indicates query-type speech. Here, if the score 182 does not meet a query-type speech threshold, the second segment 122 may not indicate query-type speech. In the illustrated example, the negative hotword classifier 300 may receive the score 182 as an input in addition to the determination made by the second-stage hotword detector 220 to determine whether the first segment 121 of the audio data 120 should be classified as containing negative hotwords.
[0046] Thus, when the negative hotword classifier 300 receives an indication from the query processor 180 that the second audio segment 122 of the audio data does not indicate a verbal query-type utterance, the negative hotword classifier 300 may classify the first segment 121 of the audio data 120 as containing a negative hotword, indicating that the second-stage hotword detector 220 provided a false acceptance. The negative hotword classifier 300 may additionally receive the probability score 171 generated by the second-stage hotword detector 220 for the first segment 121, whereby a probability score that only narrowly meets the hotword detection threshold may further bias the negative hotword classifier 300 to classify the first segment 121 as containing a negative hotword. Furthermore, after the query processor 180 determines that the second segment 122 does not indicate a query-type utterance, the negative hotword classifier 300 may also determine whether an immediate follow-up query was provided by the user 10 of the user device. As discussed above, the determination that the user 10 did not provide an immediate follow-up query serves as additional confirmation that the user did not intend to speak the hotword in the utterance 119, but rather spoke a term ("Hey Poodle") having a similar pronunciation to the particular term / phrase designated as the hotword ("Hey Google"). Thus, classifying the first segment 121 of the audio data 120 as including a negative hotword may be further based on the determination that no follow-up query was received from the user device 102.
[0047] In some examples, after receiving audio data 120 characterizing a hotword event detected by the first-stage hotword detector 210, the remote system 110 receives a negative user interaction 162 indicating user inhibition of the wake-up process at the user device 102. That is, a false acceptance instance by the first-stage hotword detector 210 in detecting a hotword event when the user speaks a negative hotword (“Hey Poodle”) may trigger the user device 102 to wake up first while waiting for the second-stage hotword detector 220 to confirm or deny the presence of the hotword. Here, the user device 102 may provide an audible and / or visual notification to notify the user that the user device 102 is waking up, and because the user 10 did not intend to trigger the wake-up process, the user 10 may provide the negative user interaction 162 to return the device 102 to sleep. For example, the user 10 may press a physical button on the user device, provide a gesture, or, if the user device 102 includes a display, select a graphic rendered within a graphical user interface displayed on the display that causes the user device 102 to go back to sleep. In some implementations, the negative hotword classifier 300 uses a negative user interaction 162 indicating user inhibition of a wake-up process at the user device as input for classifying the first segment 121 of the audio data 120 as including a negative hotword.
[0048] In some additional examples, if the query processor 180 determines that the second segment 122 of the audio data 120 indicates a verbal query-type utterance, the query processor 180 provides a query 185 including a transcription of the second segment 122 of the audio data 120 to the search engine 190 (or other downstream application). The search engine 190 then provides results 192 responsive to the query 185 back to the user device 102. Here, after the first-stage hot word detector 210 detected a false positive hot word event when the user said, "Hey Poodle," the query processor 180 may have identified the second segment 122 of the audio data 120 as indicating a query-type utterance, even though the second segment 122 corresponded to background voices or other background audio captured by the user device 102 in the streaming audio 118. This background audio may be captured in the streaming audio 118, and the query processor 180 may identify the query-type utterance and provide the corresponding query 185 to the search engine 190 to retrieve the results 192. The results 192 may be audibly and / or visually output by the user device 102 to the user 10, even if the user 10 never intended to invoke the user device 102. As a result, the user 10 may provide a negative user interaction 162 indicating that the user 10 negatively interacted with the results 192. For example, the user 10 may provide a verbal input indicating that the user 10 is confused by the results or a statement that the user 10 did not provide a query. Additionally or alternatively, the user 10 may provide an input indication indicating an instruction / command to reject the results 192.
[0049] In other scenarios, outcome 192 may be a prompt from the digital assistant stating for audible output from device 102 that user 10 needs to provide confirmation to perform the action, e.g., "You asked for the current weather, is that correct?", whereby a negative user interaction could be user 10 saying, "No, I did not ask about the weather." Similarly, outcome 192 may be a prompt requesting the user to repeat the query because query processor 180 was unsure of the query, e.g., "I did not understand your question, please repeat?", whereby a negative user interaction could be user 10 expressing confusion by uttering "Huh," user 10 affirmatively dismissing the prompt by saying "I did not ask anything," or simply the user failing to respond within a predetermined time period. Thus, the negative user interaction 162 may be provided to the negative hotword classifier 300 in addition to one or more of the other inputs discussed above, such as an indication that no hotwords were detected by the second stage hotword detector 220 in the first segment of the audio data 120, an indication that the second segment 122 of the audio data 120 is not associated with a query-type utterance, or an indication that the user 10 did not provide an immediate follow-up query after the wake-up process was suppressed.
[0050] 3 illustrates an example of the negative hotword classifier 300 of FIG. 1 receiving one or more input features 302 for determining whether the first segment 121 of audio data 120 should be classified as a negative hotword. If the negative hotword classifier 300 determines, based on the one or more input features 302, that the first segment 121 of audio should be classified as a negative hotword, the negative hotword classifier 300 generates a classification result 170 indicating the classification of the first segment 121 of audio data 120 as a negative hotword, as discussed above in FIG. 1. The one or more input features 302 received by the negative hotword classifier 300 may include, but are not limited to, whether the second hotword detector 220 detected the presence of a hotword in the first segment 121 of the audio data 120 and / or a probability score 171, whether an immediate follow-up query was received from the user device 102, whether the second segment 122 of the audio data 120 contains a query-type utterance (e.g., by providing a score 182 indicating whether the second segment 122 indicates a query-type utterance), and / or a probability score 171. The information includes whether an indication was provided, a transcription of the first segment 121 and / or the second segment 122 of the audio data 120, and a negative user interaction 162 indicating user inhibition of a wake-up process at the user device 102, and / or whether a negative user interaction 162 was received indicating that the user 10 negatively interacted with the results 192 in response to processing the second segment 122 (and / or optionally the first segment 121) of the audio data 120 as a query-type utterance (e.g., by providing a query 185 to a search engine 190 or other downstream application).
[0051] Some input features 302 may be weighted more heavily if the second-stage hotword detector 220 determines that the first segment 121 of the audio data 120 should be classified as a negative hotword. For example, the failure of the second-stage hotword detector 220 to detect the presence of a hotword in the first segment 121 is a strong indication that the first segment 121 contains a negative hotword that caused a false acceptance in the first-stage hotword detector 210. The magnitude of the probability score 171 may bias the classification result 170. For example, a probability score 171 that fails to meet the hotword detection threshold in the second-stage hotword detector 220 by a wide margin provides a higher likelihood of a negative hotword than a probability score 171 that fails to meet the hotword detection threshold by a small margin.
[0052] In some configurations, the negative hotword classifier 300 includes a trained classifier (which may include a neural network model trained via machine learning) configured to generate a negative hotword confidence score 304 indicating the likelihood that the first segment 121 of the audio data 120 contains a negative hotword. The classifier 300 may classify the first segment 121 as containing a negative hotword if the negative hotword confidence score 304 meets a confidence threshold. The negative hotword confidence score 304 may be included in the classification result 170 received by the negative hotword updater 400 of FIG. 1 for use in updating the first-stage hotword detector 210 to not detect hotword events in subsequent audio that contain negative hotwords. In some examples, the negative hotword confidence score 304 is a binary score indicating that the first segment 121 of the audio data 120 contains a negative hotword and therefore should be classified as a negative hotword, or that the first segment 121 does not contain a negative hotword.
[0053] 1 and 4 , in some examples, the negative hotword updater 400 updates the first-stage hotword detector 210 to prevent triggering a hotword event in subsequent audio data that includes negative hotwords by providing the user device 102 with the classification result 170 that includes the first segment 121 of the audio data 120. Here, the user device 102 may be configured to maintain the first-stage hotword detector 210 using the first segment 121 of the audio data 120 that is classified as including negative hotwords. For example, the first segment 121 of the audio data 120 may be labeled as a negative hotword and provided as a training input to the first-stage hotword detector 210 so that the first-stage hotword detector 210 learns not to detect the presence of the hotword (“Hey Google”) in subsequent audio data that includes the negative hotword (“Hey Poodle”). As used herein, retraining the first-stage hotword detector 210 may include retraining an existing hotword detector 210 running on the user device 102 that initially erroneously detected the hotword event, or may include a new first-stage hotword detector 210 that is later pushed to the user device 102.
[0054] Additionally, updating the first stage hotword detector 210 may also include updating the optional initial coarse hotword detector 205 running on the DSP 60 (FIG. 2) if the user device 102 includes a battery-powered device. Similar to the first stage hotword detector 210, updating the initial coarse hotword detector 205 may include retraining the initial coarse hotword detector 205 to prevent triggering hotword events for audio data that includes negative hotwords. In some examples, only the initial coarse hotword detector 205 is updated to prevent triggering hotword events in subsequent audio data that includes negative hotwords.
[0055] 4, the negative hotword updater 400 executing on the user device 102 stores in memory hardware 105 each instance of the first segment 121 of the audio data 120 that is classified by the negative hotword classifier 300 as containing a corresponding negative hotword. In the illustrated example, the user device 102 stores each instance in which the first segment 121 of the audio data 120 is classified as a negative hotword by storing a corresponding classification result 170 for each instance. Here, the classification result 170 includes the first segment 121 classified as containing a negative hotword, a probability score 171 indicating the likelihood that the first segment 121 contains the actual hotword, and other related information 172, such as a transcription of the utterance 119, a speaker identification score (e.g., speaker embedding) identifying speaker characteristics of the user 10 who spoke the utterance 119, a timestamp of the hotword event indicating the time and / or day of the week, and a negative hotword confidence score 304 ( FIG. 3 ) indicating the confidence for classifying the first segment 121 as a negative hotword. In the illustrated example, the user device 102 stores one or more classification results 170, each associated with a different corresponding negative hotword. For example, another classification result 170Aa-n, 170Ba-n, 170Ca-n may be stored for each of the negative hot words "Poodle," "Doodle," and "Noodle," which, when spoken by user 10, are pronounced similarly to the specified hot word "Hey Google."
[0056] In some implementations, the user device 102 (via the negative hotword updater 400) is configured to retrain the first-stage hotword detector 210 based on a count of the number of instances (e.g., the number of classification results 170) of the first segment 121 of the audio data 120 classified as containing a negative hotword stored in the memory hardware 105. Here, the number of instances of the audio data classified as containing the same hotword that meets a threshold number of instances may establish a pattern that the user 10 regularly speaks a negative hotword that is falsely detected as a designated hotword by the first-stage hotword detector 210. In some examples, the user device 102 requests a specified number of false acceptance instances resulting from the user 10 speaking the same term if the negative hotword confidence score 304 associated with the score is relatively low, e.g., if the negative hotword confidence score 304 only meets a narrow margin threshold.
[0057] Continuing with reference to FIG. 4 , in some examples, the negative hotword updater 400 may append an embedding 12 to each classification result 170 stored in the memory hardware 105. Here, the first-stage hotword detector 210 may calculate an embedding 12 for any audio data characterizing a hotword event detected by the first-stage hotword detector 210, and when the audio data 120 is subsequently classified as containing a negative hotword by the negative hotword classifier 300, the negative hotword updater 400 may append the embedding 12 to the corresponding instance of the classification result 170. In some implementations, the negative hotword updater 400 aggregates / averages the embeddings 12 stored in the memory hardware 105 for each corresponding negative hotword to generate a reference embedding 15 for each corresponding negative hotword. For example, a corresponding reference embedding 15 may be generated for each of the negative hotwords “Poodle,” “Noodle,” and “Doodle.”
[0058] 5 shows a schematic diagram 500 illustrating an example in which the user device 102 captures subsequent audio data 120 corresponding to another utterance 519 spoken by the user 10, which includes the term “My Poodle,” causing the first-stage hot word detector 210 to falsely detect another hot word event. The first-stage hot word detector 210 executing on the user device 102 calculates evaluation embedded representations 18 for the subsequent audio data 120 (e.g., a portion of the subsequent audio data 120 that characterizes the hot word event). The user device 102 may simultaneously access classification results 170 stored in the memory hardware 105, each of which may include embedded representations 12 for corresponding first segments 121 of the audio data 120 classified as one of the negative hot words (e.g., “Poodle,” “Noodle,” and “Doodle”). Additionally or alternatively, the user device 102 may access the corresponding reference embeddings 15 generated for each of the negative hotwords, as described above with reference to FIG.
[0059] In some implementations, the scorer 510 compares the evaluation embedding representations 18 calculated for the subsequent audio data 120 with all of the reference embeddings 15 generated and stored for each of the negative hot words. In these implementations, the reference embedding representations 12, 15 associated with each of the negative hot words "Poodle," "Noodle," and "Doodle" are all clustered together in an embedding representation space that is distinct from the clusters of reference embeddings associated with other negative hot words. Thus, the scorer 510 may determine a similarity score 515 between the reference embedding representations 12 calculated for each first segment 121 of the audio data 120 classified as one of the negative hot words and the evaluation embedding representations 18 for the subsequent audio data 120. Additionally or alternatively, the scorer 510 may determine a similarity score 515 between each corresponding reference embedding 15 representing an aggregate / average embedding for a corresponding one of the negative hot words (e.g., "Poodle," "Doodle," and "Noodle"). In some examples, each similarity score 515 is associated with a distance (e.g., cosine distance) between the evaluation embedding 18 and the reference embedding 12, 15 in the embedding space.
[0060] After the scorer 510 determines / generates the similarity scores 515, the classifier 520 may compare each similarity score 515 with a similarity score threshold, and if the similarity score 515 meets the similarity score threshold, the classifier 520 may determine / classify the subsequent audio data 120 as containing a negative hotword. In some examples, the similarity score threshold represents the maximum allowable cosine distance between embeddings associated with the same negative hotword. In some scenarios, when a similarity score 515 is calculated between the rating embedding 18 for the subsequent audio and the corresponding reference embedding 12 calculated for each instance of the first segment 121 of the audio data 120 that has been classified as a negative hotword, multiple similarity scores 515 may meet the similarity score threshold. To illustrate, in the illustrated example, the similarity score 515 between the rating embedding 18 and the corresponding reference embedding 12 classified as the negative hotword "Poodle" meets the similarity score threshold to indicate that the rating embedding 18 falls within the cluster of embeddings 12 classified as the negative hotword "Poodle" and outside the cluster of embeddings 12 classified as the other negative hotwords "Doodle" and "Noodle." Thus, if the classifier 520 determines that the similarity score 515 meets the similarity threshold, the classifier 520 determines that the subsequent audio data 120 contains a negative hotword. As a result, if the first-stage hotword detector 210 falsely detects a hotword event and triggers the initiation of a wake-up process in the user device, the classifier 520 may instruct the first-stage hotword detector 210 to refrain from detecting hotword events in the subsequent audio data or instruct the user device 102 to return to sleep.
[0061] 6 provides a flowchart of example operations of a method 600 for personalizing a hotword detector on a user device based on classifying as negative hotwords particular terms that caused previous instances of false positive hotword detection by the hotword detector. At operation 602, the method 600 includes receiving, at the data processing hardware 103, 112, audio data 120 characterizing hotword events detected by a first-stage hotword detector 210 in streaming audio 118 captured by the user device 102. The first-stage hotword detector 210 may execute on a digital signal processor (DSP) of the data processing hardware 103 of the user device 102 or may execute on an application processor of the data processing hardware 103 of the user device 102.
[0062] At operation 604, the method 600 includes processing, by the data processing hardware 103, 112, the audio data 120 using the second-stage hotword detector 220 to determine whether a hotword is detected by the second-stage hotword detector 220 in a first segment 121 of the audio data 120. The second-stage hotword detector 220 may be implemented as an ASR engine that performs automatic speech recognition to determine whether a hotword is recognized in the first segment 121. The second-stage hotword detector 220 may be implemented as a hotword detection model in other configurations, whereby the hotword detection model determines whether a hotword is detected in the first segment 121 without performing speech recognition.
[0063] At operation 606, if no hotword is detected by the second-tier hotword detector 220 in the first segment 121 of the audio data 120, the method 600 includes classifying, by the data processing hardware 103, 112, the first segment 121 of the audio data 120 as containing a negative hotword that caused the first-tier hotword detector 210 to falsely detect a hotword event in the streaming audio 118. At operation 608, the method 600 includes updating, by the data processing hardware 103, 112, the first-tier hotword detector 210 to prevent triggering a hotword event in subsequent audio data 120 that contains the negative hotword.
[0064] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform tasks. In some examples, a software application may be referred to as an "application," "app," or "program." Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0065] 7 is a schematic diagram of an exemplary computing device 700 that may be used to implement the systems and methods described in this disclosure. Computing device 700 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown, their connections and relationships, and their functionality are intended to be exemplary only and are not intended to limit the implementation of the invention(s) described and / or claimed in this document.
[0066] Computing device 700 includes a processor 710, a memory 720, a storage device 730, a high-speed interface / controller 740 that connects to memory 720 and a high-speed expansion port 750, and a low-speed interface / controller 760 that connects to a low-speed bus 770 and storage device 730. Each of components 710, 720, 730, 740, 750, and 760 are interconnected using various buses and may be mounted on a common motherboard or otherwise mounted as desired. Processor (e.g., data processing hardware) 710 can process instructions for execution within computing device 700, including instructions stored in memory 720 (e.g., memory hardware) or on storage device 730, for displaying graphical information for a graphical user interface (GUI) on an external input / output device, such as a display 780 coupled to high-speed interface 740. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and types of memory, as desired. Also, multiple computing devices 700 may be connected, each providing a portion of the required operations (e.g., a bank of servers, a group of blade servers, or a multiprocessor system). Processor 710 may include data processing hardware 103 present on user device 102 of FIG. 1 or data processing hardware 112 present on remote system 110 of FIG. 1.
[0067] Memory 720 stores information non-transitoryly within computing device 700. Memory 720 may include memory hardware 105 present on user device 102 of FIG. 1 or memory hardware 114 present on remote system 110 of FIG. 1. Memory 720 may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. Non-transitory memory 720 may be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (program state information) for use by computing device 700. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.
[0068] The storage device 730 can provide mass storage for the computing device 700. In some implementations, the storage device 730 is a computer-readable medium. In various different implementations, the storage device 730 can be a floppy disk device, a hard disk device, an optical disk device, or an array of devices including a tape device, a flash memory or other similar solid-state memory device, or a device in a storage area network or other configuration. In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as the memory 720, the storage device 730, or memory on the processor 710.
[0069] The high-speed controller 740 manages bandwidth-intensive operations for the computing device 700, and the low-speed controller 760 manages less bandwidth-intensive operations. Such role assignments are merely exemplary. In some implementations, the high-speed controller 740 is coupled to memory 720, a display 780 (e.g., via a graphics processor or accelerator), and a high-speed expansion port 750, which may accept various expansion cards (not shown). In some implementations, the low-speed controller 760 is coupled to a storage device 730 and a low-speed expansion port 790. The low-speed expansion port 790, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), may be coupled to one or more input / output devices, such as a keyboard, pointing device, scanner, or a networking device, such as a switch or router, for example, via a network adapter.
[0070] The computing device 700, as shown, can be implemented in several different forms. For example, it can be implemented as a standard server 700a, or multiple times within a group of servers 700a, as a laptop computer 700b, or as part of a rack server system 700c.
[0071] Various implementations of the systems and techniques described herein may be realized in digital electrical and / or optical circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be special-purpose or general-purpose, coupled to receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0072] These computer programs (also known as programs, software, software applications, or code) include machine language for programmable processors and can be implemented in high-level procedural and / or object-oriented languages, and / or assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0073] The processes and logic flows described herein can be implemented by one or more programmable processors, also referred to as data processing hardware, that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be implemented by special-purpose logic circuitry, such as an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit). Processors suitable for executing computer programs include, by way of example, both general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Generally, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or is operably coupled to receive data from or transmit data to them, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include, by way of example, all forms of non-volatile memory, media, and memory devices, including semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices, magnetic disks, e.g., internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0074] To provide for user interaction, one or more aspects of the present disclosure can be implemented on a computer having a display device, such as a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touchscreen, for displaying information to the user, and optionally a keyboard and pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices can be used to provide user interaction as well, and the feedback provided to the user can be any form of sensory feedback, such as visual, auditory, or tactile feedback, and input from the user can be received in any form, including acoustic, speech, or tactile input. Additionally, the computer can interact with the user by sending documents to and receiving documents from a device used by the user, for example, by sending a web page to the web browser in response to a request received from the web browser on the user's client device.
[0075] Although several implementations have been described, it will be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims. [Explanation of symbols]
[0076] 10 users 12 Embedded Representations, Reference Embedded Representations 15 Reference Embedding, Reference Embedding Representation 18 Evaluation Embedding Representation, Evaluation Embedding 60 First Processor, Dedicated DSP, DSP 70 Second processor, main AP 100 systems 102 User Devices 103 Data Processing Hardware 104 Network 105 Memory Hardware 106 microphones 110 Remote Systems 112 Computing resources, data processing hardware 114 Storage Resources, Memory Hardware 118 Streaming Audio, Input Streaming Audio, Multi-Channel Audio, Audio 119 utterances 120 Audio Data 121 First segment, first audio segment 122 Second Segment 136 Audio Data 160 Restraining Order 162 Negative User Interactions 170 Classification results 172 Related Information 180 Query Processor 182 score 185 queries 190 search engines 192 Results 200 Cascaded Hotword Detection Architecture 205 Initial Coarse Hotword Detector, Coarse Hotword Detector, Coarse-Stage Hotword Detector 210 first stage hot word detector, first stage (fine) hot word detector, hot word detector 220 Second Stage Hot Word Detector, Second Hot Word Detector 300 Negative Hot Word Classifier, Classifier 302 Input Features 304 Negative Hotword Confidence Score 400 Negative Hotword Uploader 500 Schematic 510 Scorer 515 Similarity Score 519 utterances 520 Classifier 700 computing devices 700a Server 700b laptop computer 700c Rack Server System 710 Processor, Components 720 Memory, Components, Non-Temporary Memory 730 Storage devices, components 740 High Speed Interface / Controller, Components, High Speed Interface 750 High-Speed Expansion Port, Components 760 Low-Speed Interface / Controller, Components 770 Slow Bus 780 Display
Claims
1. When executed on the data processing hardware (710), the data processing hardware (710) receiving audio data (120) characterizing hot word events detected by a first stage hot word detector (210) in streaming audio (118) captured by a user device (102); processing the audio data (120) using a second-stage hot word detector (220) to determine whether a hot word is detected by the second-stage hot word detector (220) within a first segment (121) of the audio data (120); If the hot word is not detected by the second stage hot word detector (220) in the first segment (121) of the audio data (120), classifying the first segment (121) of the audio data (120) as containing a negative hotword that caused the first stage hotword detector (210) to falsely detect the hotword event in the streaming audio (118); and updating the first stage hotword detector (210) to prevent triggering of the hotword event in subsequent audio data (120) that includes the negative hotword based on the first segment (121) of the audio data (120) classified as including the negative hotword, the method comprising: the operation is performed when the hot word is not detected by the second stage hot word detector (220) in the first segment (121) of the audio data (120); inhibiting a wake-up process in the user device (102) for processing the hot word and / or one or more other terms following the hot word in the streaming audio (118); determining whether an immediate follow-up query has been provided by a user of the user device (102) after inhibiting the wake-up process at the user device (102); further comprising classifying the first segment of the audio data as including the negative hotword is further based on determining that a follow-up query was not provided by the user of the user device after suppressing the wake-up process. A computer-implemented method (600).
2. The operation comprises: if the hot word is detected by the second stage hot word detector (220) in the first segment (121) of the audio data (120), processing a second segment (122) of the audio data (120) that follows the first segment (121) of the audio data (120) to determine whether the second segment (122) of the audio data (120) indicates a verbal query-type utterance; if the second segment (122) of the audio data (120) does not indicate the verbal query-type utterance; classifying the first segment (121) of the audio data (120) as containing the negative hotword; updating the first stage hotword detector (210) to prevent triggering of the hotword event in subsequent audio data (120) containing the negative hotword based on the first segment (121) of the audio data (120) classified as containing the negative hotword; further comprising:
10. The method (600) of claim 1.
3. The operation includes: determining if the second segment (122) of the audio data (120) does not indicate the verbal query-type utterance; further comprising an act of determining whether an immediate follow-up query has been provided by a user of the user device (102); the operation of classifying the first segment of the audio data as including the negative hotword is further based on determining that a follow-up query was not provided by the user of the user device.
3. The method (600) of claim 2.
4. The operation includes: if the second segment (122) of the audio data (120) indicates the verbal query-type utterance; receiving a negative interaction result (162) indicating that a user of the user device (102) has negatively interacted with a result (192) of the verbal query-type utterance provided to the user device (102); classifying the first segment (121) of the audio data (120) as including the negative hotword based on the received negative interaction result (162); updating the first stage hotword detector (210) to prevent detecting the hotword event in subsequent audio data (120) containing the negative hotword based on the first segment (121) of the audio data (120) classified as containing the negative hotword; further comprising:
4. The method (600) of claim 2 or 3.
5. After the operation receives the audio data (120) characterizing the hot word event detected by the first stage hot word detector (210), further comprising the act of receiving a negative user interaction (162) indicating user inhibition of a wake-up process at the user device (102); classifying the first segment of the audio data as including the negative hotword is further based on the negative user interaction indicating a user inhibition of the wake-up process.
5. The method (600) according to any one of claims 1 to 4.
6. 6. The method of claim 1, wherein updating the first-stage hotword detector to prevent triggering of the hotword event in subsequent audio data includes providing the first segment of the audio data classified as containing the negative hotword to the user device, and wherein the user device is configured to retrain the first-stage hotword detector using the first segment of the audio data classified as containing the negative hotword.
7. The user device (102) storing each instance of the first segment (121) of the audio data (120) classified as containing the negative hotword in memory hardware (720) of the user device (102); retraining the first stage hot word detector (210) based on a count of the number of instances of the first segment (121) of the audio data (120) classified as containing the negative hot word stored in the memory hardware (720); 7. The method of claim 6, further comprising: retraining the first stage hot word detector by:
8. Before the user device (102) retrains the first stage hot word detector (210), determining that a corresponding confidence score associated with each instance of the first segment (121) of the audio data (120) classified as including the negative hotword does not meet a negative hotword threshold score; Determining when the number of instances exceeds the threshold number of instances 8. The method (600) of claim 7, further comprising:
9. updating the first stage hotword detector (210) to prevent triggering of the hotword event in subsequent audio data (120) includes providing the first segment (121) of the audio data (120) classified as containing the negative hotword to the user device (102), wherein the user device (102) obtaining an embedded representation (12) of the first segment (121) of the audio data (120); configured to store the embedded representation (12) of the first segment (121) of the audio data (120) in memory hardware (720) of the user device (102); The user device (102) Compute an evaluation embedding representation (18) for the subsequent audio data (120); determining a similarity score (515) between the embedded representation (12) of the first segment (121) of the audio data (120) classified as the negative hotword and the rating embedded representation (18) for the subsequent audio data (120); If the similarity score (515) meets a similarity score threshold, determine that the subsequent audio data (120) contains the negative hotword. By doing so, configured to determine when the subsequent audio data (120) characterizing the hot word event detected by the first stage hot word detector (210) includes the negative hot word.
9. The method (600) according to any one of claims 1 to 8.
10. the data processing hardware (710) resides on a server (110) that communicates with the user device (102); the first stage hot word detector (210) executes on a processor of the user device (102); 10. The method (600) according to any one of claims 1 to 9.
11. 11. The method of claim 10, wherein processing the audio data to determine whether the hotword is detected by the second stage hotword detector in the first segment of the audio data comprises performing automatic speech recognition to determine whether the hotword is recognized in the first segment of the audio data.
12. 10. The method (600) of any one of claims 1 to 9, wherein the data processing hardware (710) resides on the user device (102).
13. the first stage hot word detector (210) is implemented on a digital signal processor (DSP) (60) of the data processing hardware (710); the second stage hot word detector (220) runs on an application processor (70) of the data processing hardware (710); 13. The method (600) of claim 12.
14. The first stage hot word detector (210) generating a probability score (171) indicative of the presence of the hotword in the audio characteristics of the streaming audio (118) captured by the user device (102); If the probability score (171) meets the hot word detection threshold of the first stage hot word detector (210), then the hot word event is detected in the streaming audio (118). It is configured as follows:
14. The method (600) according to any one of claims 1 to 13.
15. data processing hardware (710); and memory hardware (720) in communication with the data processing hardware (710), the memory hardware (720), when executed on the data processing hardware (710), causing the data processing hardware (710) to: processing the audio data (120) using a second-stage hot word detector (220) to determine whether a hot word is detected by the second-stage hot word detector (220) within a first segment (121) of the audio data (120); If the hot word is not detected by the second stage hot word detector (220) in the first segment (121) of the audio data (120), classifying, by the data processing hardware (710), the first segment (121) of the audio data (120) as containing a negative hot word that caused a first-stage hot word detector (210) to falsely detect a hot word event in the streaming audio (118); updating, by the data processing hardware (710), the first stage hotword detector (210) to prevent triggering of the hotword event in subsequent audio data (120) containing the negative hotword, based on the first segment (121) of the audio data (120) classified as containing the negative hotword; storing an instruction to execute the the operation is performed when the hot word is not detected by the second stage hot word detector (220) in the first segment (121) of the audio data (120); inhibiting a wake-up process in the user device (102) for processing the hot word and / or one or more other terms following the hot word in the streaming audio (118); determining whether an immediate follow-up query has been provided by a user of the user device (102) after inhibiting the wake-up process at the user device (102); further comprising classifying the first segment of the audio data as including the negative hotword is further based on determining that a follow-up query was not provided by the user of the user device after suppressing the wake-up process. System (100).
16. The action of claim 15, wherein if the hot word is detected by the second stage hot word detector in the first segment of the audio data, the action of processing a second segment (122) of the audio data (120) that follows the first segment (121) of the audio data (120) to determine whether the second segment (122) of the audio data (120) indicates a verbal query-type utterance; if the second segment (122) of the audio data (120) does not indicate the verbal query-type utterance; classifying the first segment (121) of the audio data (120) as containing the negative hotword; updating the first stage hotword detector (210) to prevent triggering of the hotword event in subsequent audio data (120) containing the negative hotword based on the first segment (121) of the audio data (120) classified as containing the negative hotword; further comprising:
16. The system (100) of claim 15.
17. the operation being: if the second segment (122) of the audio data (120) does not indicate the verbal query-type utterance; further comprising an act of determining whether an immediate follow-up query has been provided by a user of the user device (102); the operation of classifying the first segment of the audio data as including the negative hotword is further based on determining that a follow-up query was not provided by the user of the user device.
17. The system (100) of claim 16.
18. the operation being, if the second segment (122) of the audio data (120) indicates the verbal query-type utterance: receiving a negative interaction result (162) indicating that a user of the user device (102) has negatively interacted with a result (192) of the verbal query-type utterance provided to the user device (102); classifying the first segment (121) of the audio data (120) as including the negative hotword based on the received negative interaction result (162); updating the first stage hotword detector (210) to prevent detecting the hotword event in subsequent audio data (120) containing the negative hotword based on the first segment (121) of the audio data (120) classified as containing the negative hotword; further comprising:
18. A system (100) according to claim 16 or 17.
19. After the operation receives the audio data (120) characterizing the hot word event detected by the first stage hot word detector (210), further comprising the act of receiving a negative user interaction (162) indicating user inhibition of a wake-up process at the user device (102); classifying the first segment of the audio data as including the negative hotword is further based on the negative user interaction indicating a user inhibition of the wake-up process. A system (100) according to any one of claims 15 to 18.
20. 20. The system of claim 15, wherein updating the first-stage hotword detector to prevent triggering of the hotword event in subsequent audio data includes providing the first segment of the audio data classified as containing the negative hotword to the user device, and wherein the user device is configured to retrain the first-stage hotword detector using the first segment of the audio data classified as containing the negative hotword.
21. The user device (102) storing each instance of the first segment (121) of the audio data (120) classified as containing the negative hotword in memory hardware (720) of the user device (102); retraining the first stage hot word detector (210) based on a count of the number of instances of the first segment (121) of the audio data (120) classified as containing the negative hot word stored in the memory hardware (720); 21. The system (100) of claim 20, configured to retrain the first stage hot word detector (210) by:
22. Before the user device (102) retrains the first stage hot word detector (210), determining that a corresponding confidence score associated with each instance of the first segment (121) of the audio data (120) classified as including the negative hotword does not meet a negative hotword threshold score; Determining when the number of instances exceeds the threshold number of instances 22. The system (100) of claim 21, further configured to:
23. updating the first stage hotword detector (210) to prevent triggering of the hotword event in subsequent audio data (120) includes providing the first segment (121) of the audio data (120) classified as containing the negative hotword to the user device (102), wherein the user device (102) obtaining an embedded representation (12) of the first segment (121) of the audio data (120); configured to store the embedded representation (12) of the first segment (121) of the audio data (120) in memory hardware (720) of the user device (102); The user device (102) Compute an evaluation embedding representation (18) for the subsequent audio data (120); determining a similarity score (515) between the embedded representation (12) of the first segment (121) of the audio data (120) classified as the negative hotword and the rating embedded representation (18) for the subsequent audio data (120); If the similarity score (515) meets a similarity score threshold, determine that the subsequent audio data (120) contains the negative hotword. By doing so, configured to determine when the subsequent audio data (120) characterizing the hot word event detected by the first stage hot word detector (210) includes the negative hot word.
23. A system (100) according to any one of claims 15 to 22.
24. the data processing hardware (710) resides on a server (110) that communicates with the user device (102); the first stage hot word detector (210) executes on a processor of the user device (102); A system (100) according to any one of claims 15 to 23.
25. 25. The system (100) of claim 24, wherein the operation of processing the audio data (120) to determine whether the hot word is detected by the second stage hot word detector (220) in the first segment (121) of the audio data (120) comprises the operation of performing automatic speech recognition to determine whether the hot word is recognized in the first segment (121) of the audio data (120).
26. 24. The system (100) of any one of claims 15 to 23, wherein the data processing hardware (710) resides on the user device (102).
27. the first stage hot word detector (210) is implemented on a digital signal processor (DSP) (60) of the data processing hardware (710); the second stage hot word detector (220) runs on an application processor (70) of the data processing hardware (710); 27. The system (100) of claim 26.
28. The first stage hot word detector (210) generating a probability score (171) indicative of the presence of the hotword in the audio characteristics of the streaming audio (118) captured by the user device (102); Detecting the hotword event in the streaming audio if the probability score satisfies the hotword detection threshold of the first stage hotword detector. It is configured as follows:
28. A system (100) according to any one of claims 15 to 27.
Citation Information
Patent Citations
Method and device for outputting information
CN111640426A
Acoustic model generating device, method for the same, and program
JP2014092750A
An embedded system for building space-saving speech recognition with user-definable constraints.
JP2015520409A
Speech recognition power management
JP2016505888A
Recognition device, recognition method, and recognition program
JP2020016784A