Hotword threshold auto-tuning
The cascaded hotword detection system dynamically adjusts thresholds based on false acceptance and rejection rates, improving accuracy and user experience by adapting to individual device and user-specific conditions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- GOOGLE LLC
- Filing Date
- 2024-07-01
- Publication Date
- 2026-05-12
AI Technical Summary
Existing voice-responsive systems struggle with setting optimal hotword detection thresholds due to varying acoustic environments and user tolerances, leading to false acceptances and rejections, which are not tailored to individual devices or users.
A cascaded hotword detection system using a first-stage detector on a user device and a second-stage remote detector for verification, dynamically adjusting the hotword detection threshold based on false acceptance and rejection rates to adapt to specific environments and user preferences.
Improves hotword detection accuracy by individually determining and adjusting thresholds for each device, reducing false activations and rejections, enhancing user experience by minimizing unnecessary device wake-ups.
Smart Images

Figure 0007857347000001 
Figure 0007857347000002 
Figure 0007857347000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to automatic tuning of hotword thresholds.
Background Art
[0002] A voice-responsive environment (such as a home, workplace, school, vehicle, etc.) enables a user to speak a query or command to a computer-based system, and the computer-based system to respond and answer to the query and / or perform functions based on the command. The voice-responsive environment can be implemented using a network of connected microphone devices distributed in various rooms or areas of the environment. These devices help identify when a given utterance is directed to the system using a hotword, rather than an utterance directed to another individual present in the environment. Thus, the devices may operate in a sleep or hibernation state and wake up only when the detected utterance contains a hotword. Typically, systems used to detect hotwords in streaming audio generate a probability score indicating the likelihood that a hotword exists in the streaming audio. When the probability score meets a predetermined threshold, the device initiates a wake-up process.
Summary of the Invention
Means for Solving the Problems
[0003] One aspect of the present disclosure provides a method for automatic tuning of a hotword threshold. The method includes, in data processing hardware, receiving audio data from a user device running a first-stage hotword detector that characterizes hotwords detected by the first-stage hotword detector in streaming audio captured by the user device. The first-stage hotword detector is configured to generate a confidence score indicating whether a hotword is present in the audio features of the streaming audio captured by the user device, and to detect a hotword in the streaming audio when the confidence score satisfies the hotword detection threshold of the first-stage hotword detector.
[0004] The method also includes processing audio data using a second-stage hotword detector with data processing hardware to determine whether a hotword is detected in the audio data by the second-stage hotword detector. If the hotword is not detected in the audio data by the second-stage hotword detector, the method includes identifying a misacceptance instance in the first-stage hotword detector indicating that the first-stage hotword detector misaccepted the hotword in the streaming audio.
[0005] The method also includes determining, by data processing hardware, whether the alien acceptance rate associated with the first-stage hotword detector of the user device satisfies an alien acceptance rate threshold. The alien acceptance rate is based on several alien acceptance instances identified by the first-stage hotword detector within the alien acceptance period. If the alien acceptance rate associated with the first-stage hotword detector satisfies the alien acceptance rate threshold, the method includes adjusting the hotword detection threshold of the first-stage hotword detector by the data processing hardware.
[0006] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, the method further includes, when a hotword is not detected in the audio data by the second-stage hotword detector, the data processing hardware suppressing a wake-up process on the user device for processing the hotword and / or one or more other words following the hotword in the streaming audio. In some examples, when a hotword is detected in the audio data by the second-stage hotword detector, the method further includes, when the data processing hardware determines whether subsequent audio data characterizing a spoken query following the hotword in the streaming audio is received from the user device. When subsequent audio data characterizing a spoken query is not received from the user device, the method may include, by the data processing hardware, identifying a misacceptance instance in the first-stage hotword detector indicating that the first-stage hotword detector misdetected the hotword in the streaming audio.
[0007] Optionally, the method further includes processing the speech query by data processing hardware when subsequent audio characterizing the speech query is received from the user device. The user device may be configured to initiate a wake-up process to process the hotword and / or one or more other words following the hotword in the streaming audio when the first-stage hotword detector detects the hotword in the streaming audio. Adjusting the hotword detection threshold of the first-stage hotword detector includes, in some examples, increasing the value of the hotword detection threshold.
[0008] The method may further include, when receiving audio data characterizing a hotword detected in the streaming audio by the first-stage hotword detector, receiving a near-miss indicator from the user device indicating that the first-stage hotword detector detected a hotword in the streaming audio within a threshold period after the first-stage hotword detector generated a previous accuracy score that failed to satisfy the hotword detection threshold by a threshold margin. The previous accuracy score indicates whether the hotword is present in previous audio features of the streaming audio captured by the user device.
[0009] When a hotword is detected in audio data by a second-stage hotword detector, the method may include, by data processing hardware, identifying, based on near miss indicators, instances in the first-stage hotword detector that indicate the first-stage hotword detector failed to initially detect the hotword within previous audio features of the streaming audio, and by data processing hardware determining whether the rejection rate associated with the first-stage hotword detector of the user device satisfies a rejection rate threshold. The rejection rate is based on several rejection instances identified by the first-stage hotword detector within the rejection period. When the rejection rate associated with the first-stage hotword detector satisfies the rejection rate threshold, the method may include, by data processing hardware, adjusting the hotword detection threshold of the first-stage hotword detector. In some examples, adjusting the hotword detection threshold includes decreasing the hotword detection threshold of the first-stage hotword detector.
[0010] Another aspect of the present disclosure provides another method for performing automatic tuning of a hotword threshold. The method includes receiving streaming audio captured by one or more microphones communicating with the data processing hardware of a user device. The method also includes the data processing hardware generating a confidence score indicating whether a hotword is present in the audio features of the streaming audio using a first-stage hotword detector. The method also includes the data processing hardware determining whether the confidence score satisfies a hotword detection threshold.
[0011] When the accuracy score satisfies the hotword detection threshold, the method includes data processing hardware detecting the hotword in the streaming audio, and data processing hardware transmitting audio data characterizing the hotword detected in the streaming audio using the first-stage hotword detector to a remote computing device running a second-stage hotword detector. The remote computing device is configured to determine whether the hotword is detected in the audio data by the second-stage hotword detector, and, if the hotword is not detected in the audio data by the second-stage hotword detector, to identify instances of misacceptance in the first-stage hotword detector indicating that the first-stage hotword detector misaccepted the hotword in the streaming audio. When the misacceptance rate, based on several misacceptance instances identified in the first-stage hotword detector during the misacceptance period, satisfies the misacceptance rate threshold, the method includes data processing hardware adjusting the hotword detection threshold of the first-stage hotword detector.
[0012] This embodiment may include one or more of the following optional features: Adjusting the hotword detection threshold of the first-stage hotword detector may include increasing the value of the hotword detection threshold. In some examples, when the accuracy score satisfies the hotword detection threshold, the method includes the data processing hardware initiating a wake-up process on the user device for processing the hotword and / or one or more other words that follow the hotword in the streaming audio. When the hotword is not detected in the audio data by the second-stage hotword detector, the method may include the data processing hardware suppressing the wake-up process on the user device.
[0013] In some examples, the method further includes, when the accuracy score satisfies the hotword detection threshold, determining a near miss marker by the data processing hardware indicating that a previous accuracy score that failed to satisfy the hotword detection threshold by a threshold margin was generated within a threshold period prior to the first-stage hotword detector detecting a hotword in the streaming audio. The method may also include, by the data processing hardware, transmitting the near miss marker to a remote computing device.
[0014] The remote computing device may be configured to identify a rejection case in the first-stage hotword detector based on a near-miss indicator when a hotword is detected in the audio data by the second-stage hotword detector. A rejection case indicates that the first-stage hotword detector failed to initially detect the hotword within previous audio features of the streaming audio. When the rejection rate based on several rejection cases identified in the first-stage hotword detector within the rejection period satisfies the rejection threshold, the method, in some implementations, includes adjusting the hotword detection threshold of the first-stage hotword detector by the data processing hardware. Optionally, adjusting the hotword detection threshold includes decreasing the value of the hotword detection threshold.
[0015] Another aspect of this disclosure provides a system for performing automatic tuning of a hotword threshold. The system includes data processing hardware and memory hardware that communicates with the data processing hardware. The memory hardware stores instructions, which, when executed on the data processing hardware, cause the data processing hardware to perform an action. The action includes receiving audio data from a user device running a first-stage hotword detector, which characterizes hotwords detected by the first-stage hotword detector in streaming audio captured by the user device. The first-stage hotword detector is configured to generate a confidence score indicating whether a hotword exists in the audio features of the streaming audio captured by the user device, and to detect a hotword in the streaming audio when the confidence score satisfies the hotword detection threshold of the first-stage hotword detector.
[0016] The operation also includes processing the audio data using the second-stage hotword detector to determine whether the hotword is detected in the audio data by the second-stage hotword detector. If the hotword is not detected in the audio data by the second-stage hotword detector, the operation includes identifying a misacceptance instance in the first-stage hotword detector, indicating that the first-stage hotword detector misaccepted the hotword in the streaming audio.
[0017] The operation also includes determining whether the alien acceptance rate associated with the user device's first-stage hotword detector satisfies the alien acceptance rate threshold. The alien acceptance rate is based on several alien acceptance instances identified by the first-stage hotword detector within the alien acceptance period. If the alien acceptance rate associated with the first-stage hotword detector satisfies the alien acceptance rate threshold, the operation includes adjusting the hotword detection threshold of the first-stage hotword detector.
[0018] Implementations of this disclosure may include one or more of the following optional features. In some implementations, the operation further includes suppressing a wake-up process on the user device for processing the hotword and / or one or more other words following the hotword in the streaming audio when the hotword is not detected in the audio data by the second-stage hotword detector. In some examples, the operation further includes determining whether subsequent audio data characterizing a voice query following the hotword in the streaming audio is received from the user device when the hotword is detected in the audio data by the second-stage hotword detector. If subsequent audio data characterizing a voice query is not received from the user device, the operation may include identifying a misacceptance instance in the first-stage hotword detector indicating that the first-stage hotword detector misdetected the hotword in the streaming audio.
[0019] Optionally, the operation further includes processing the speech query when subsequent audio characterizing the speech query is received from the user device. The user device may be configured to initiate a wake-up process to process the hotword and / or one or more other words following the hotword in the streaming audio when the first-stage hotword detector detects the hotword in the streaming audio. Adjusting the hotword detection threshold of the first-stage hotword detector includes, in some examples, increasing the value of the hotword detection threshold.
[0020] The operation may further include receiving a near miss marker from the user device indicating that the first-stage hotword detector detected a hotword in the streaming audio within a threshold period after the first-stage hotword detector generated a previous accuracy score that failed to satisfy the hotword detection threshold by a threshold margin. The previous accuracy score indicates whether the hotword is present in previous audio features of the streaming audio captured by the user device.
[0021] When a hotword is detected in the audio data by the second-stage hotword detector, the operation may include identifying, based on near miss indicators, instances of rejection in the first-stage hotword detector indicating that the first-stage hotword detector initially failed to detect the hotword within previous audio features of the streaming audio, and determining whether the rejection rate associated with the first-stage hotword detector on the user device satisfies a rejection rate threshold. The rejection rate is based on several rejection instances identified by the first-stage hotword detector within the rejection period. If the rejection rate associated with the first-stage hotword detector satisfies the rejection rate threshold, the operation may include adjusting the hotword detection threshold of the first-stage hotword detector. In some examples, adjusting the hotword detection threshold includes decreasing the hotword detection threshold of the first-stage hotword detector.
[0022] Another aspect of this disclosure provides another system for performing automatic tuning of a hotword threshold. The system includes data processing hardware for a user device and memory hardware that communicates with the data processing hardware. The memory hardware stores instructions, which, when executed on the data processing hardware, cause the data processing hardware to perform an action. The action includes receiving streaming audio captured by one or more microphones that communicate with the data processing hardware. The action also includes generating a confidence score indicating whether a hotword is present in the audio features of the streaming audio using a first-stage hotword detector. The action includes determining whether the confidence score satisfies a hotword detection threshold.
[0023] When the accuracy score satisfies the hotword detection threshold, the operation includes detecting the hotword in the streaming audio and sending the audio data characterizing the hotword detected in the streaming audio using the first-stage hotword detector to a remote computing device running a second-stage hotword detector. The remote computing device is configured to determine whether the hotword is detected in the audio data by the second-stage hotword detector, and, if the hotword is not detected in the audio data by the second-stage hotword detector, to identify a case of misacceptance in the first-stage hotword detector indicating that the first-stage hotword detector misdetected the hotword in the streaming audio.
[0024] When the alien acceptance rate, based on several alien acceptance instances identified in the first-stage hotword detector within the alien acceptance period, satisfies the alien acceptance rate threshold, the operation includes adjusting the hotword detection threshold of the first-stage hotword detector.
[0025] This embodiment may include one or more of the following optional features: Adjusting the hotword detection threshold of the first-stage hotword detector may include increasing the value of the hotword detection threshold. In some examples, when the accuracy score satisfies the hotword detection threshold, the operation includes initiating a wake-up process on the user device to process the hotword and / or one or more other words that follow the hotword in the streaming audio. When the hotword is not detected in the audio data by the second-stage hotword detector, the operation may include suppressing the wake-up process on the user device.
[0026] In some examples, the operation further includes determining a near miss flag indicating that a previous confidence score that cannot satisfy the hot word detection threshold by the threshold margin when the confidence score satisfies the hot word detection threshold was generated within a threshold period before detecting the hot word in the streaming audio by the first stage hot word detector. The operation may include sending the near miss flag to a remote computing device.
[0027] The remote computing device may be configured to identify a false rejection case in the first stage hot word detector based on the near miss flag when the hot word is detected in the audio data by the second stage hot word detector. The false rejection case indicates that the first stage hot word detector could not initially detect the hot word in previous audio features of the streaming audio. When a false rejection rate based on some false rejection cases identified within the false rejection period in the first stage hot word detector satisfies a false rejection threshold, the operation, in some implementations, includes adjusting the hot word detection threshold of the first stage hot word detector. Optionally, adjusting the hot word detection threshold includes decreasing the value of the hot word detection threshold.
[0028] Details of one or more implementations of the present disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims.
Brief Description of the Drawings
[0029] [Figure 1] It is a schematic diagram of an exemplary system for performing hot word threshold automatic tuning. [Figure 2] It is a schematic diagram of an exemplary component of a hot word detection threshold adjuster. [Figure 3] It is a schematic diagram of a hot word detection threshold adjuster that increments an acceptance count for others. [Figure 4] This is a schematic diagram illustrating an example of accepting non-humans. [Figure 5A] This is a schematic diagram illustrating an example of a case where the person refuses to cooperate. [Figure 5B] This is a schematic diagram illustrating an example of a case where the person refuses to cooperate. [Figure 6] This is a flowchart illustrating an example configuration of how to perform automatic tuning of hotword thresholds. [Figure 7] This is a flowchart of another exemplary configuration for how to perform automatic threshold tuning. [Figure 8] This is a schematic diagram of an exemplary computing device that may be used to implement the systems and methods described herein. [Modes for carrying out the invention]
[0030] Similar reference numerals in various drawings indicate similar elements.
[0031] Voice-enabled devices (e.g., user devices running voice assistants) allow users to speak queries or commands aloud, respond to those queries, and / or perform functions based on those commands. Through the use of "hotwords" (also called "keywords," "attention words," "wake-up phases / words," "trigger phases," or "voice action initiation commands"), which are predetermined words / phrases spoken to attract the attention of a voice-enabled device, the voice-enabled device can distinguish between utterances directed at the system (i.e., to initiate a wake-up process to process one or more words following the hotword in the utterance) and utterances directed at individuals in the environment. Typically, voice-enabled devices operate in a sleep state to conserve power and do not process input audio data unless input audio data follows a spoken hotword. As an example, while in sleep mode, a voice-enabled device captures input audio via one or more microphones and uses a trained hotword detector to detect whether a hotword is present in the input audio. When a hotword is detected in the input audio, the voice-enabled device initiates a wake-up process to process the hotword and / or any other words in the input audio that follow the hotword.
[0032] Hotword detection is like finding a needle in a haystack, because the hotword detector must constantly listen to the streaming audio and trigger precisely and instantaneously when the presence of a hotword is detected in the streaming audio, while ignoring most of it. To cope with the complexity of detecting whether a hotword is present in a continuous audio stream, neural networks are commonly used by hotword detectors. Typically, the neural network generates a confidence score based on the received streaming audio, indicating whether a hotword is present in the streaming audio. The hotword detector determines whether the confidence score satisfies a detection threshold. If the confidence score satisfies the detection threshold, the hotword detector determines that a hotword is present in the streaming audio. The hotword detector may then initiate the device wake-up process.
[0033] The hotword detection threshold is typically set to a predetermined value that balances the false acceptance rate with the false rejection rate. False acceptance occurs when the hotword detector detects a hotword (i.e., the accuracy score satisfies the hotword detection threshold), but the streaming audio does not actually contain the hotword. Despite false acceptance, the hotword detector will initiate the wake-up process on the voice-enabled device, even if the user did not intend to call the device. False rejection, on the other hand, occurs when the streaming audio contains a hotword, but the hotword detector determines that the hotword is not present in the streaming audio (i.e., the accuracy score cannot satisfy the hotword detection threshold). False rejection by the hotword detector is frustrating for the user because they must then attempt to call the voice-enabled device by speaking the hotword again, usually louder, and / or asking the user to walk closer to the device, in order to ensure that the spoken hotword is not falsely rejected again. Therefore, selecting a hotword detection threshold is extremely difficult due to the wide variety of devices, environments, and users. Traditionally, detection thresholds have not been tailored to each individual device. However, each device may encounter significantly different acoustic environments. For example, a device near a television, which is often left on, may encounter considerably more false acceptances than the same device with the same hotword detection threshold would encounter in a quiet office. Furthermore, each user may have considerably different tolerances for false rejection and false acceptance. That is, one user may tolerate a moderate number of false acceptances, while another user may not tolerate the same number.
[0034] The implementations described herein relate to a hotword detection threshold tuner system that dynamically adjusts the hotword detection threshold of a user device running a first-stage hotword detector to adapt the hotword detector to the environment on a case-by-case basis. Herein, the term “hotword detection threshold” refers to a value or confidence score that streaming audio must satisfy in order for the hotword detector to determine / detect the presence of a given hotword within the audio features of the streaming audio, and thus trigger a wake-up process on the user device. The user device’s first-stage hotword detector detects a hotword in the streaming audio based on a first confidence score indicating whether the hotword is present within the audio features of the streaming audio. In this case, the first confidence score satisfies the hotword detection threshold associated with the first-stage hotword detector, thereby causing the user device to send audio data characterizing the hotword detected by the first-stage hotword detector to a remote second-stage hotword detector for verification. For example, the user device sends the audio data to a server running the second-stage hotword detector via the internet. The second-stage hotword detector may use a more accurate hotword detection model than the one used by the first-stage hotword detector running on the user device to detect whether a hotword is present in the audio. The second-stage hotword detector processes the audio data to determine whether a hotword is detected by the second-stage hotword detector. If a hotword is not detected by the second-stage hotword detector, the system identifies a case of misacceptance in the first-stage hotword detector, indicating that the first-stage hotword detector misaccepted the hotword. The system determines whether the misacceptance rate associated with the first-stage hotword detector satisfies the misacceptance rate threshold and adjusts the hotword detection threshold of the first-stage hotword detector accordingly.For example, the system may increase the hotword detection threshold value to lower the sensitivity of the first-stage hotword detector so that future instances of other-acceptance cases decrease or disappear.
[0035] Accordingly, the systems described herein include a cascaded hotword detection technique that uses multiple models to improve accuracy and verify and / or confirm hotword detection. The system individually determines the rates of false acceptance and false rejection for each device and, accordingly, adapts a hotword detection threshold based on the false acceptance and false rejection rates for each device.
[0036] Referring to Figure 1, in some implementations, the exemplary system 100 includes one or more user devices 102, each associated with a corresponding user 10, and communicating with a remote system 110 via a network 104. Each user device 102 may correspond to a computing device such as a mobile phone, computer, smart speaker, smart appliance, smart headphones, or wearable, and each user device 102 is equipped with data processing hardware 103 and memory hardware 105. The user device 102 includes or communicates with one or more microphones 106 for capturing utterances from the corresponding user 10. The remote system 110 may be a single computer, multiple computers, or a distributed system (e.g., a cloud environment) having scalable / elastic computing resources 112 (e.g., data processing hardware) and / or storage resources 114 (e.g., memory hardware). In some implementations, the user device 102 receives a trained neural network 130 (e.g., a memorized neural network) from a remote system 110 via a network 104 and runs the trained neural network 130 to detect hotwords in the streaming audio 118. The trained neural network 130 resides in the first-stage hotword detector 120 (also called a hotworder) of the user device 102, which is configured to detect whether hotwords are present in the streaming audio 118 without performing semantic analysis or speech recognition processing on the streaming audio 118.
[0037] In the illustrated example, when user 10 speaks an utterance 119 containing a hotword (e.g., "Hey Google") which is captured as streaming audio 118 by user device 102, the first-stage hotword detector 120 running on user device 102 is configured to detect the presence of the hotword in the utterance 119 and to initiate a wake-up process on user device 102 to process the hotword and / or one or more other words (e.g., a query or command) that follow the hotword in the utterance 119. That is, user device 102 may be configured to initiate a wake-up process to process the hotword and / or one or more words that follow the hotword in the streaming audio 118 when the first-stage hotword detector 120 detects the hotword in the streaming audio 118.
[0038] The first-stage hotword detector 120 generates a confidence score 132 (e.g., from a neural network 130) indicating whether a hotword exists within the audio features of the streaming audio 118 captured by the user device 102. The first-stage hotword detector 120 detects a hotword in the streaming audio 118 when the confidence score 132 satisfies the first-stage hotword detector 120's hotword detection threshold 134. When the confidence score 132 satisfies the hotword detection threshold 134, the first-stage hotword detector 120 transmits audio data 136 representing the streaming audio 118 to the second-stage hotword detector 140 running on the remote system 110. In some examples, audio data 136 represents the streaming audio 118 itself, while in other examples, audio data 136 represents the streaming audio 118 after it has been processed by the first-stage hotword detector 120 (for example, to identify and / or isolate specific audio characteristics of the streaming audio 118, or to convert the streaming audio 118 into a format suitable for transmission and / or processing by the second-stage hotword detector 140). For example, audio data 136 may be cut from the streaming audio 118 to include the relevant segment containing audio features associated with a hotword detected by the first-stage hotword detector 120.
[0039] The second-stage hotword detector 140 is configured to detect whether a hotword is present in the audio data 136, similar to the first-stage hotword detector 120. The second-stage hotword detector 140 differs from the first-stage hotword detector 120. For example, the second-stage hotword detector 140 may include a different neural network that is more computationally intensive than the neural network 130 of the first-stage hotword detector 120. The second-stage hotword detector 140 may provide improved accuracy compared to the first-stage hotword detector 120, which is limited by the resources of the user device 102.
[0040] The second-stage hotword detector 140 processes the audio data 136 to determine whether a hotword exists within the audio data 136. The second-stage hotword detector 140 may generate an accuracy score to compare with a hotword detection threshold, similar to the first-stage hotword detector 120, or it may determine whether a hotword exists using a completely different method. When the second-stage hotword detector 140 does not detect a hotword in the audio data 136, the hotword detection threshold adjusters 200, 200a~b (Figure 2) identify a misacceptance case 210 in the first-stage hotword detector 120, indicating that the first-stage hotword detector 120 misaccepted a hotword in the streaming audio 118.
[0041] Referring to Figure 2, the hotword detection threshold tuner 200 maintains the other-acceptance count 220. The hotword detection threshold tuner 200 increments the other-acceptance count 220 in response to identifying other-acceptance cases 210. Based on the other-acceptance count 220, the hotword detection threshold tuner determines the current other-acceptance rate 230. The other-acceptance rate 230 represents several other-acceptance cases 210 identified by the hotword detection threshold tuner 200 within an other-acceptance period. For example, the other-acceptance period may be 1 hour, 4 hours, or 24 hours, etc. The other-acceptance count 220 may include only the number of other-acceptance cases 210 within the most recent other-acceptance period. Thus, the other-acceptance rate 230 indicates how often the first-stage hotword detector 120 incorrectly determines that a hotword is present in the streaming audio 118.
[0042] The hotword detection threshold adjuster 200 may determine whether the other-acceptance rate 230 satisfies the other-acceptance rate threshold 240. For example, if the other-acceptance period is 1 hour and the other-acceptance rate threshold 240 is 3 per hour, the other-acceptance rate 230 satisfies the other-acceptance rate threshold 240 when the hotword detection threshold adjuster 200 has identified 3 or more other-acceptance cases 210 in the most recent 1 hour.
[0043] Referring again to Figure 1, when the other-acceptance rate 230 associated with the first-stage hotword detector 120 satisfies the other-acceptance rate threshold 240, the hotword detection threshold tuner 200 adjusts the hotword detection threshold 134 of the first-stage hotword detector 120. In some implementations, the hotword detection threshold tuner 200 runs on the remote system 110 (i.e., is the hotword detection threshold tuner 200a) and sends a hotword detection threshold tuning command 150 to the first-stage hotword detector 120. When the tuning command 150 is received by the user device 102, it causes the user device 102 to adjust the hotword detection threshold 134 of the first-stage hotword detector 120. In other implementations, the hotword detection threshold tuner 200 runs on the user device 102 (i.e., is the hotword detection threshold tuner 200b) and receives an marker 142 of an other-acceptance case 210 from the second-stage hotword detector 140 running on the remote system 110. In this case, the user device 102 maintains the other-acceptance count 220 and determines the current other-acceptance rate 230. The hotword detection threshold tuner 200b provides the first-stage hotword detector 120 with a hotword detection threshold tuning command 150 to adjust the hotword detection threshold 134 based on the other-acceptance rate threshold 240 and the current other-acceptance rate 230.
[0044] In some implementations, when the alien acceptance rate 230 exceeds the alien acceptance rate threshold 240, the hotword detection threshold tuner 200 increases the value of the hotword detection threshold 134. That is, the accuracy score 132 required to detect the presence of a hotword in the streaming audio 118 increases, thereby reducing the frequency of alien acceptance instances 210. In some examples, the hotword detection threshold tuner 200 adjusts or modifies the alien acceptance rate threshold 240 based on the adjusted hotword detection threshold 134. In some configurations, the user 10 of the user device 102 may set and / or adjust the alien acceptance rate threshold 240.
[0045] In some implementations, when a hotword is not detected in the audio data 136 by the second-stage hotword detector 140, the remote server 110 suppresses the wake-up process on the user device 102. The wake-up process allows the user device 102 to process the hotword and / or one or more other words (e.g., queries or commands) that follow the hotword in the streaming audio 118. In some implementations, the remote system 110 suppresses the wake-up process by sending a suppression command 162 to the user device 102, causing the user device 102 to suppress the wake-up process. In other implementations, the remote system 110 suppresses the wake-up process by sending an indicator 164 to the user device 102 indicating that the second-stage hotword detector 140 could not confirm the presence of a hotword in the audio data 136, thereby causing the user device 102 to suppress the wake-up process (i.e., remain in sleep mode or return to sleep mode). In other implementations, the remote system 110 suppresses the wake-up process by not responding to the user device 102 after receiving the audio data 136 (for example, by closing the network connection). The lack of response from the remote system 110 may cause the user device 102 to suppress the wake-up process. That is, in some examples, the user device 102 starts the wake-up process only after receiving confirmation from the second-stage hotword detector 140 that a hotword was present in the streaming audio 118. The user device 102 may independently suppress the wake-up process. For example, the user device 102 may automatically suppress the wake-up process when the query or command following the hotword is empty (i.e., the streaming audio 118 following the hotword does not contain a command or query directed to the user device 102).In this case, the user device 102 may detect the other party acceptance case 210 and notify the hotword detection threshold adjuster to increment the other party acceptance count 220.
[0046] Next, referring to Figure 3, in some examples, when a hotword is detected in audio data 136 by the second-stage hotword detector 140, the remote system 110 determines whether subsequent audio data 136 characterizing a voice query following the hotword in streaming audio 118 is received from the user device 102. If subsequent audio data 136 characterizing a voice query is not received from the user device 102, the hotword detection threshold adjuster 200 identifies a misacceptance case 210 in the first-stage hotword detector 120, indicating that the first-stage hotword detector 120 misaccepted the hotword in streaming audio 118. That is, in some implementations, the hotword detection threshold adjuster 200 identifies a misacceptance case 210 based on the absence of a query or command in the subsequent audio data 136 following the detected hotword. For example, when audio that is not intended to trigger the wake-up process (such as ambient noise in the environment, e.g., from a television) unintentionally or unintentionally triggers hotword detection, the hotword detection threshold adjuster 200 can identify the other acceptance case 210 if there are no subsequent queries or commands (which should occur during an intentional wake-up command).
[0047] In some examples, the remote system processes the voice query when subsequent audio data 136 characterizing the voice query is received from the user device 102 (i.e., after both the first-stage hotword detector 120 and the second-stage hotword detector 140 have detected the presence of the hotword in the streaming audio 118). In these examples, processing the query may include passing the audio data 136 to a speech recognition system to transcribe the voice query. The remote system 110 may use the transcription to perform natural language understanding and / or provide the transcription to a search engine and / or other applications in order to process the query.
[0048] In some implementations, the remote system 110 does not include a second-stage hotword detector 140, but instead executes a query / command processor 430 (Figure 4) configured to perform speech recognition on the audio data 136 to verify whether the first-stage hotword detector 120 accurately detected the presence of a hotword in the streaming audio 118. That is, in some implementations, after the first-stage hotword detector 120 detects the presence of a hotword in the streaming audio 118, it sends the audio data 136 to the remote server to process subsequent queries from the user 10. In this case, the hotword detection threshold adjuster 200 may identify other-acceptance cases 210 in scenarios where the processor 430 cannot recognize a hotword in the received audio data 136, and in scenarios where the processor 430 determines that the subsequent audio data 136 received from the user device 102 is empty (i.e., the subsequent audio data 136 does not contain a query or command). In this case, both results in the hotword detection threshold adjuster 200 incrementing the other-acceptance count 220.
[0049] Referring to Figure 4, schematic diagram 400 depicts television 410 emitting playback audio 420 containing the utterance "Hey you all!". Due to the phonetic similarity between the utterance "Hey you all!" and the hotword "Hey Google", the first-stage hotword detector 120 determines an accuracy score 132 that satisfies the hotword detection threshold, thereby detecting the presence of the hotword in the streaming audio 118 representing playback audio 420a from television 410. The second-stage hotword detector 140 may, as discussed above, verify whether the hotword was accurately detected by the first-stage hotword detector 120. The second-stage hotword detector 140 may also notify the hotword detection threshold adjuster 200 when it is unable to detect the hotword, resulting in the adjuster 200 identifying a misidentification case 210 and incrementing the misidentification rate 230. On the other hand, the second-stage hotword detector 140 may also incorrectly detect a hotword in the utterance "Hey, you all" and pass the corresponding audio to the query processor 430. In this case, the query processor 430 may perform speech recognition on the audio data and determine that the hotword was incorrectly detected by the hotword detectors 120 and 140, respectively. In addition to or instead of this, the processor 430 may determine that no subsequent audio data 136 containing a query or command is received after the hotword was incorrectly detected. In any of these scenarios, the query processor 430 may notify the hotword detection threshold adjuster 200 to identify the other acceptance case 210.
[0050] Referring again to Figures 1 and 2, in some implementations, the hotword detection threshold tuner 200 identifies a rejection case 250, which indicates a case where the first-stage hotword detector 120 failed to detect the presence of a hotword in the streaming audio 118 when the hotword was present. In response, the hotword detection threshold tuner 200 increments the rejection count 260 and determines the current rejection rate 270. When the rejection rate 270 satisfies the rejection threshold 280, the hotword detection threshold tuner 200 adjusts the hotword detection threshold 134. In this case, the tuner 200 provides the first-stage hotword detector 120 with a hotword detection threshold tuning command 150 to reduce the hotword detection threshold and thereby increase the sensitivity of the first-stage hotword detector 120 to detect hotwords in the streaming audio 118.
[0051] The hotword detection threshold adjuster 200 may identify a user denial case 250 in response to the hotword detection threshold adjuster 200 receiving a near miss indicator 510 indicating that the first-stage hotword detector 120 detected a hotword in the streaming audio within a threshold period after the first-stage hotword detector 120 has generated a previous accuracy score that failed to satisfy the hotword detection threshold by a threshold margin. For example, the first-stage hotword detector 120 running on user device 102 may fail to detect a hotword in the first utterance spoken by the user. In this case, the first-stage hotword detector 120 may determine an accuracy score equal to 0.7, which fails to satisfy the hotword detection threshold set to 0.75. The near miss threshold may be set to a value less than the hotword detection threshold (0.65), so that the range of values between the near miss threshold (0.65) and the hotword detection threshold (0.75) corresponds to the "threshold margin". For example, the near miss threshold may be set to 0.65, so that any streaming audio 118 associated with an accuracy score greater than or equal to the near miss threshold of 0.65 but less than the hotword detection threshold of 0.75 cannot satisfy the hotword detection threshold by a threshold margin. Continuing this example, in a subsequent attempt by user 10 to call user device 102, the first-stage hotword detector 120 accurately detects the hotword in a second utterance spoken by user 10 within a threshold period (e.g., 5 seconds). The hotword detection threshold adjuster 200 may identify a user denial case 250 after receiving a near miss marker 510 and confirming that the second-stage hotword detector 140 also detected the presence of the hotword in the second utterance, provided that the first-stage hotword detector 120 determines the accuracy score associated with the first utterance that failed to satisfy the hotword detection threshold by a threshold margin, and subsequently detects the hotword in the second utterance within the threshold period.In particular, the more accurate second-stage hotword detector 140 may have detected the presence of a hotword in the first utterance, but the first-stage hotword detector 120 never sent the corresponding audio data 136 to the second-stage hotword detector 140 because the accuracy score generated by the first-stage hotword detector 120 related to the first utterance did not satisfy the hotword detection threshold. The hotword detection threshold adjuster 200 may determine whether the rejection rate 270 (based on the rejection count 260) satisfies the rejection threshold 280, and if so, adjust the hotword detection threshold 134 of the first-stage hotword detector 120. In some cases, the first-stage hotword detector 120 provides a near-miss indicator 510 to the hotword detection threshold adjuster 200, which then identifies the person denies the request only after receiving confirmation that the second-stage hotword detector 140 has detected a hotword in the audio data 136.
[0052] Next, referring to Figure 5A, schematic Figure 500a depicts a near miss example that functions as a proxy for user denial case 250, where user 10 is speaking a first utterance 119a ("Hey, Google") which is received by a first-stage hotword detector 120 on user device 102 (not shown). In this case, the first-stage hotword detector 120 generates a probability score 132 that does not satisfy the hotword detection threshold 134 but satisfies the near miss threshold 520. For example, when the hotword detection threshold 134 is 0.75 and the near miss threshold 520 is 0.65 (i.e., less than but generally close to the hotword detection threshold 134), a probability score of 0.70 (or any other value between 0.65 and 0.75) may not satisfy the hotword detection threshold 134 by a threshold margin because this probability score satisfies the near miss threshold 520.
[0053] Schematic diagram 500b in Figure 5B depicts a scenario where, within a threshold period (e.g., 5, 10, or 30 seconds) after receiving the first utterance 119a, user 10 makes a second utterance 119b ("Hey, Google!") in another attempt to call and wake up the user device. This utterance 119b may be spoken more forcefully and / or with greater annunciation (because the user device 102 was unable to initiate the wake-up process in response to the previous utterance 119a). In this case, both the first-stage hotword detector 120 and the second-stage hotword detector 140 identify the presence of a hotword in the streaming audio 118 related to the second utterance 119b. Even though the confidence score 132 for the first utterance 119a calculated by the first-stage hotword detector 120 failed to satisfy the hotword detection threshold 134, the hotword detection threshold adjuster 200 receives a near-miss indicator 510, which surrogately indicates that the first-stage hotword detector 120 failed to detect the hotword in the streaming audio 118 because the confidence score 132 for the first utterance satisfied the near-miss threshold 520 and the second utterance 119b was within the threshold period. After the second-stage hotword detector 140 confirms the presence of the hotword in the second utterance 119b, the hotword detection threshold adjuster 200 may identify a self-rejection case 250 and increment the self-rejection count 260. The hotword detection threshold adjuster 200 may lower the hotword detection threshold 134 of the first-stage hotword detector 120 in response to the rejection rate 270 satisfying the rejection threshold 280.
[0054] In some examples, the hotword detection threshold adjuster 200 may adjust the hotword detection threshold 134 of the first-stage hotword detector 120 based on a combined value representing the hotword usage frequency, the other-acceptance count 220, and the self-rejection count 260, by applying a predefined threshold to this combined value. For example, the combined value is the ratio of other-acceptances to self-rejections (because the other-acceptance count 220 and the self-rejection count 260 are generally inversely correlated). In other examples, the hotword detection threshold adjuster 200 adjusts the hotword detection threshold 134 of the first-stage hotword detector 120 based on information collected from other user devices 102 running the first-stage hotword detector 120. In these examples, the remote system 110 estimates a multivariate distribution of hotword usage frequency, false acceptance count 220 (or false acceptance rate 230), and true rejection count 260 (or true rejection rate 270) from a large population of user devices 102, and identifies outliers within the distribution to trigger threshold tuning of the outliers by the hotword detection threshold tuner 200. That is, devices with false acceptance count 220 or true rejection count 260 that are sufficiently outside the general population may be candidates for threshold tuning.
[0055] Figure 6 is a flowchart of an exemplary configuration of the operation of Method 600 for automatic tuning of the hotword threshold. In operation 602, Method 600 includes receiving audio data 136 from a user device 102 running a first-stage hotword detector 120 in data processing hardware 112, characterizing the hotwords detected by the first-stage hotword detector 120 in streaming audio 118 captured by the user device 102. The first-stage hotword detector 120 is configured to generate a confidence score 132 indicating whether a hotword is present in the audio features of the streaming audio 118 captured by the user device 102, and to detect a hotword in the streaming audio 118 when the confidence score 132 satisfies the hotword detection threshold 134 of the first-stage hotword detector 120.
[0056] In operation 604, method 600 includes processing audio data 136 using a second-stage hotword detector 140 with data processing hardware 112 to determine whether a hotword is detected in the audio data 136 by the second-stage hotword detector 140. If a hotword is not detected in the audio data 136 by the second-stage hotword detector 140, method 600 includes, in operation 606, having data processing hardware 112 identify a misacceptance case 210 in the first-stage hotword detector 120 indicating that the first-stage hotword detector 120 misaccepted the hotword in the streaming audio 118.
[0057] Method 600 includes, in operation 608, determining by the data processing hardware 112 whether the alien acceptance rate 230 associated with the first-stage hotword detector 120 of the user device 102 satisfies the alien acceptance rate threshold 240. The alien acceptance rate 230 is based on several alien acceptance instances 210 identified by the first-stage hotword detector 120 within the alien acceptance period. If the alien acceptance rate 230 associated with the first-stage hotword detector 120 satisfies the alien acceptance rate threshold 240, Method 600, in operation 610, adjusts the hotword detection threshold 134 of the first-stage hotword detector 120 by the data processing hardware 112.
[0058] Figure 7 is a flowchart of another exemplary configuration of the operation of Method 700 for automatic tuning of hotword thresholds. In operation 702, Method 700 includes receiving streaming audio 118 captured by one or more microphones 106 communicating with the data processing hardware 103 of the user device 102. In operation 704, Method 700 includes the data processing hardware 103 using a first-stage hotword detector 120 to generate a confidence score 132 indicating whether a hotword is present in the audio features of the streaming audio 118.
[0059] In operation 706, method 700 includes determining whether the accuracy score 132 satisfies the hotword detection threshold 134 using the data processing hardware 103. If the accuracy score 132 satisfies the hotword detection threshold 134, method 700 includes in operation 708 the data processing hardware 103 detecting a hotword in the streaming audio 118, and in operation 710 the data processing hardware 103 transmitting audio data 136 characterizing the hotword detected in the streaming audio 118 using the first-stage hotword detector 120 to a remote computing device 110 running the second-stage hotword detector 140.
[0060] In operation 712, the remote computing device 110 is configured to determine whether a hotword is detected in the audio data 136 by the second-stage hotword detector 140. In operation 714, if the hotword is not detected in the audio data 136 by the second-stage hotword detector 140, the remote computing device is configured to identify a misacceptance instance 210 in the first-stage hotword detector 120 indicating that the first-stage hotword detector 120 misaccepted the hotword in the streaming audio 118. If the misacceptance rate 230 based on several misacceptance instances 210 identified in the first-stage hotword detector 120 during the misacceptance period satisfies the misacceptance rate threshold 240, the method 700 includes, in operation 716, adjusting the hotword detection threshold 134 of the first-stage hotword detector 120 by the data processing hardware 103.
[0061] Figure 8 is a schematic diagram of an exemplary computing device 800 that may be used to implement the systems and methods described herein. The computing device 800 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown herein, their connections and relationships, and their functions are intended to be illustrative only and are not intended to limit the implementations of the invention described and / or claimed herein.
[0062] The computing device 800 includes a processor 810, memory 820, storage device 830, a high-speed interface / controller 840 connected to memory 820 and high-speed expansion port 850, and a low-speed bus 870 and a low-speed interface / controller 860 connected to storage device 830. Components 810, 820, 830, 840, 850, and 860 are interconnected using different buses and may be mounted on a common motherboard or in other configurations as needed. The processor 810 can process instructions to be executed within the computing device 800, including instructions stored in memory 820 or on storage device 830 for displaying graphical information for a graphical user interface (GUI) on an external input / output device such as a display 880 coupled to the high-speed interface 840. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and memory types as needed. Alternatively, multiple computing devices 800 may be connected, with each device providing the necessary operational components (for example, as a server bank, a group of blade servers, or a multiprocessor system).
[0063] Memory 820 stores information non-temporarily within the computing device 800. Memory 820 may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. Non-temporarily, memory 820 may be a physical device used to store programs (e.g., sequences of instructions) or data (e.g., program state information) temporarily or permanently for use by the computing device 800. Examples of non-volatile memory include, but are not limited to, flash memory (typically used for firmware such as boot programs) and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase-change memory (PCM), and disks or tapes.
[0064] The storage device 830 can provide large-capacity storage to the computing device 800. In some implementations, the storage device 830 is a computer-readable medium. In various different implementations, the storage device 830 may be a floppy disk device, a hard disk device, an optical disk device, or a tape device, or it may be flash memory or other similar solid-state memory device, or it may be an array of devices including devices in a storage area network or other configuration. In further implementations, a computer program product is tangibly embodied within the information carrier. The computer program product includes instructions that, when executed, perform one or more of the methods described above. The information carrier is a computer-readable medium or a machine-readable medium, such as memory 820, the storage device 830, or memory on the processor 810.
[0065] The high-speed controller 840 manages bandwidth-intensive operations of the computing device 800, while the low-speed controller 860 manages lower bandwidth-intensive operations. Such a division of roles is merely illustrative. In some implementations, the high-speed controller 840 is coupled to memory 820, to display 880 (e.g., through a graphics processor or accelerator), and to high-speed expansion port 850, which may accept various expansion cards (not shown). In some implementations, the low-speed controller 860 is coupled to storage device 830 and low-speed expansion port 890. The low-speed expansion port 890, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, Wireless Ethernet), may be coupled to one or more input / output devices, such as a keyboard, pointing device, scanner, or networking devices such as switches and routers (e.g., through a network adapter).
[0066] The computing device 800 may be implemented in several different forms as shown in the figure. For example, the computing device 800 may be implemented as a standard server 800a, may be implemented multiple times within a group of such servers 800a, may be implemented as a laptop computer 800b, or may be implemented as part of a rack server system 800c.
[0067] Various implementations of the systems and techniques described herein can be realized as digital electronic circuits and / or optical circuits, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations as one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be dedicated or general-purpose, that receives data and instructions from a memory system, at least one input device, and at least one output device, and transmits data and instructions to them.
[0068] A software application (i.e., a software resource) can refer to computer software that causes a computing device to perform a task. In some cases, a software application may be called an “application,” “app,” or “program.” Examples of applications, though not limited to these, include system diagnostic applications, system administration applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0069] These computer programs (also known as programs, software, software applications, or code) contain machine instructions for a programmable processor and can be implemented in a procedural and / or object-oriented high-level programming language and / or assembly language / machine code. Herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, non-temporary computer-readable medium, apparatus and / or device (e.g., magnetic disks, optical disks, memory, programmable logic devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including machine-readable medium that receives machine instructions as machine-readable signals. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0070] The processes and logical flows described herein can be implemented by one or more programmable processors, also called data processing hardware, which execute one or more computer programs to perform functions by acting on input data and producing outputs. Processes and logical flows can also be implemented by dedicated logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). Examples of processors suitable for executing computer programs include one or more processors from both general-purpose and dedicated microprocessors, and any type of digital computer. Generally, a processor receives instructions and data from read-only memory or random-access memory, or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or is operablely coupled to receive data from or transfer data thereto, or both. However, a computer is not required to have such devices. Suitable computer-readable media for storing computer program instructions and data include, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and all forms of non-volatile memory, media, and memory devices, including CD-ROM and DVD-ROM disks. Processors and memory can be complemented by or integrated into dedicated logic circuits.
[0071] To enable interaction with a user, one or more aspects of the present disclosure may be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touchscreen, and optionally a keyboard and pointing device, such as a mouse or trackball, by which the user can input to the computer. Interaction with the user may also be enabled using other types of devices, for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or haptic feedback, and input from the user may be received in any form, including acoustic input, voice input, or haptic input. In addition, the computer may interact with the user by sending documents to and receiving documents from devices used by the user, for example, by sending web pages to a web browser on the user's client device in response to requests received from that web browser.
[0072] Several implementation forms have been described above. However, it should be understood that various modifications may be made without departing from the spirit and scope of this disclosure. Therefore, other implementation forms are included within the scope of the attached claims. [Explanation of Symbols]
[0073] 10 users 100 Systems 102 User Devices 103 Data Processing Hardware 104 Network 105 Memory Hardware 106 Microphone 110 Remote systems, remote servers, remote computing devices 112 Computing resources, data processing hardware 114 Memory Resources 118 Streaming Audio 119 utterances 119a First utterance 119b Second utterance 120 First-stage hotword detector 130 pre-trained neural networks 132 Accuracy Score 134 Hot word detection threshold 136 Audio Data 140 Second-stage hotword detector 142 signs 150 Hotword Detection Threshold Tuning Instructions 162 Restraining Order 164 signs 200 Hotword Detection Threshold Adjuster 200a Hotword Detection Threshold Adjuster 200b Hotword Detection Threshold Adjuster 210 Case Studies of Accepting Others 220 Acceptance Count 230 Others Acceptance Rate 240 Other Acceptance Rate Threshold 250 Cases of Refusal by the Individual 260 Count of user refusal 270 Rejection Rate 280 Self-rejection threshold, self-rejection rate threshold 400 Schematic Diagram 410 TV 420 Playback Audio 420a Audio Playback 430 query / command processors, query processors 500a Schematic Diagram 500b Schematic Diagram 510 Near Miss Sign, Near Miss Indicator 520 Near Miss Threshold 600 ways 700 methods 800 computing devices, systems 800a Server 800b Laptop Computer 800c Rack Server System 810 Processors, Components, and Data Processing Hardware 820 Non-temporary memory, components, memory hardware 830 Storage devices, components 840 High-Speed Interfaces / Controllers, Components 850 High-Speed Expansion Ports, Components 860 Low-Speed Interface / Controller, Components 870 Slow Bus 880 displays 890 Low-speed expansion port
Claims
1. A method performed by a computer on data processing hardware, wherein the data processing hardware includes: Receiving a near miss marker and audio data characterizing a hotword detected by a hotword detector in streaming audio captured by a user device, wherein the near miss marker indicates that the hotword detector detected the hotword in the streaming audio within a threshold period after generating a previous accuracy score that failed to satisfy the hotword detection threshold of the hotword detector by a threshold margin. Processing the audio data to confirm that the hotword was correctly detected by the hotword detector within the streaming audio, Based on the confirmation that the near miss marker and the hotword were correctly detected by the hotword detector within the streaming audio, the system determines whether the rejection rate, which is based on the number of rejection cases in which the hotword detector failed to detect the hotword within the audio features of the streaming audio and was identified by the hotword detector during the rejection period, satisfies the rejection threshold, and if it does, adjusts the hotword detection threshold of the hotword detector. A computer implementation method that causes an operation including the following to be performed.
2. The computer implementation method according to claim 1, wherein adjusting the hotword detection threshold of the hotword detector includes reducing the hotword detection threshold of the hotword detector.
3. The computer implementation method according to claim 1, wherein processing the audio data includes performing speech recognition on the audio data to confirm that the hotword was correctly detected by the hotword detector in the streaming audio when the hotword is recognized in the audio data.
4. The computer implementation method according to claim 1, wherein processing the audio data includes processing the audio data without performing semantic analysis or speech recognition processing on the audio data in order to confirm that the hotword has been correctly detected by the hotword detector in the streaming audio.
5. The aforementioned hotword detector Generate an accuracy score indicating the presence of the hotword in the audio features of the streaming audio captured by the user device. The system is configured to detect the hotword in the streaming audio when the accuracy score satisfies the hotword detection threshold of the first-stage hotword detector. The computer implementation method according to claim 1.
6. The computer implementation method according to claim 5, wherein the aforementioned prior accuracy score indicates that the hotword exists within the prior audio features of the streaming audio captured by the user device.
7. The aforementioned operation is, The hotword detector identifies instances of user denial in the hotword detector that indicate it failed to detect the hotword within the previous audio features of the streaming audio, The method further includes determining whether the self-rejection rate associated with the hotword detector satisfies the self-rejection rate threshold, Adjusting the hotword detection threshold of the hotword detector is based on determining whether the self-rejection rate associated with the hotword detector satisfies the self-rejection rate threshold. The computer implementation method according to claim 6.
8. The computer implementation method according to claim 1, wherein the hotword detector operates on the user device.
9. The computer implementation method according to claim 1, wherein the hotword detector includes a neural network trained to detect the presence of the hotword in the streaming audio without performing semantic analysis or speech recognition processing on the streaming audio.
10. It is a system, Data processing hardware and The system comprises memory hardware that communicates with the data processing hardware, the memory hardware stores instructions, and when an instruction is executed on the data processing hardware, it causes the data processing hardware to perform an operation, the operation being: Receiving a near miss marker and audio data characterizing a hotword detected by a hotword detector in streaming audio captured by a user device, wherein the near miss marker indicates that the hotword detector detected the hotword in the streaming audio within a threshold period after generating a previous accuracy score that failed to satisfy the hotword detection threshold of the hotword detector by a threshold margin. Processing the audio data to confirm that the hotword was correctly detected by the hotword detector within the streaming audio, Based on the confirmation that the near miss marker and the hotword were correctly detected by the hotword detector within the streaming audio, the system determines whether the rejection rate, which is based on the number of rejection cases in which the hotword detector failed to detect the hotword within the audio features of the streaming audio and was identified by the hotword detector during the rejection period, satisfies the rejection threshold, and if it does, adjusts the hotword detection threshold of the hotword detector. including, system.
11. The system according to claim 10, wherein adjusting the hotword detection threshold of the hotword detector includes reducing the hotword detection threshold of the hotword detector.
12. The system according to claim 10, wherein processing the audio data includes performing speech recognition on the audio data to confirm that the hotword was correctly detected by the hotword detector in the streaming audio when the hotword is recognized in the audio data.
13. The system according to claim 10, wherein processing the audio data includes processing the audio data without performing semantic analysis or speech recognition processing on the audio data in order to confirm that the hotword has been correctly detected by the hotword detector in the streaming audio.
14. The aforementioned hotword detector Generate an accuracy score indicating the presence of the hotword in the audio features of the streaming audio captured by the user device. The system is configured to detect the hotword in the streaming audio when the accuracy score satisfies the hotword detection threshold of the first-stage hotword detector. The system according to claim 10.
15. The system according to claim 14, wherein the aforementioned prior accuracy score indicates that the hotword is present in the prior audio features of the streaming audio captured by the user device.
16. The aforementioned operation is, The hotword detector identifies instances of user denial in the hotword detector that indicate it failed to detect the hotword within the previous audio features of the streaming audio, The method further includes determining whether the self-rejection rate associated with the hotword detector satisfies the self-rejection rate threshold, Adjusting the hotword detection threshold of the hotword detector is based on determining whether the self-rejection rate associated with the hotword detector satisfies the self-rejection rate threshold. The system according to claim 15.
17. The hotword detector operates on the user device, according to claim 10.
18. The system according to claim 10, wherein the hotword detector includes a neural network trained to detect the presence of the hotword in the streaming audio without performing semantic analysis or speech recognition processing on the streaming audio.