Hot phrase triggering based on sequence of detections
A flexible hotphrase detection system with a two-stage verification process enhances user interaction by allowing for variations in spoken commands, addressing the inflexibility of conventional models.
Patent Information
- Application Number
- JP2024215360
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-12-10
- Filing Date
- 2024-12-10
- Publication Date
- 2025-12-24
- Estimated Expiration
- 2041-11-21
AI Technical Summary
Conventional low-power hot phrase detection models require exact command sequences and lack flexibility, making it difficult for users to interact naturally with always-on digital assistants due to the inability to accommodate variations in spoken commands.
Implement a first-stage hotphrase detector that recognizes multiple hotwords in an utterance and integrates them to detect a complete hotphrase, followed by a second-stage verification using a more powerful model to confirm the detected hotphrase, allowing for flexible command recognition.
Enables users to interact more naturally with digital assistants by accommodating variations in spoken commands, reducing the need for exact sequences and improving the flexibility of hotphrase detection.
Smart Images

Figure 0007791972000001 
Figure 0007791972000002 
Figure 0007791972000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to hot phrase triggers based on sequences of detection. [Background technology]
[0002] A voice-enabled environment allows a user to simply speak a query or command a load, causing the digital assistant to receive the query and respond and / or implement the command. A voice-enabled environment (e.g., a home, a workplace, a school, etc.) can be implemented using a network of connected microphone devices distributed throughout various rooms and / or areas of the environment. Through such a network of microphones, a user has the power to verbally query a digital assistant from essentially anywhere in the environment, without having to have a computer or other device in front of or even near them. These devices can use hot words to help distinguish when a given utterance is directed to the system, as opposed to utterances directed to another individual entity in the environment. Thus, a device can operate in a sleep or hibernation state and wake up only when a detected utterance contains the hot word. Once woken up, the device can proceed to perform more expensive processing, such as fully on-device automatic speech recognition (ASR) or server-based ASR. For example, while cooking in the kitchen, a user can say the designated hotword "Hey computer" to trigger and wake up the voice-enabled device and then ask a digital assistant running on the voice-enabled device to "set the timer for 20 minutes"; in response, the digital assistant will confirm (e.g., in the form of a synthesized voice output) that the timer has been set and will alert the user (e.g., in the form of an alarm from an audio speaker or other audible alert) when the timer has expired after 20 minutes. Summary of the Invention [Means for solving the problem]
[0003] One aspect of the present disclosure provides a method for detecting hot phrases. The method includes receiving, at data processing hardware of a user device associated with the user, audio data corresponding to an utterance spoken by the user and captured by the user device. The utterance includes a command for the digital assistant to perform an action. During each of a plurality of fixed long-term windows of the audio data, the method includes: determining, by the data processing hardware, whether any of the trigger words in the set of trigger words associated with the hot phrase are detected in the audio data during the corresponding fixed long-term window using a hot phrase detector configured to detect each trigger word in the set of trigger words associated with the hot phrase; when one of the trigger words in the set of trigger words associated with the hot phrase is detected in the audio data during the corresponding fixed long-term window, determining, by the data processing hardware, whether each other trigger word in the set of trigger words associated with the hot phrase is also detected in the audio data; and when each other trigger word in the set of trigger words is also detected in the audio data, identifying, by the data processing hardware, the hot phrase in the audio data corresponding to the utterance. The method also includes the step of triggering an automatic speech recognizer (ASR) to perform speech recognition on the audio data when the hot phrase is identified in the audio data corresponding to the utterance by the data processing hardware.
[0004] Another aspect of the present disclosure provides a system for detecting hot phrases in audio data. The system includes data processing hardware and memory hardware in communication with the data processing hardware, the memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform an operation. The operation includes receiving audio data corresponding to an utterance spoken by a user and captured by a user device associated with the user. The utterance includes a command for the digital assistant to perform an operation. During each of a plurality of fixed long-term windows of the audio data, the operation also includes: using a hot phrase detector configured to detect each trigger word in a set of trigger words associated with the hot phrase to determine whether any of the trigger words in the set of trigger words are detected in the audio data during the corresponding fixed long-term window; when one of the trigger words in the set of trigger words associated with the hot phrase is detected in the audio data during the corresponding fixed long-term window, determining whether each other trigger word in the set of trigger words associated with the hot phrase is also detected in the audio data; and when each other trigger word in the set of trigger words is also detected in the audio data, identifying the hot phrase in the audio data corresponding to the utterance. The operations also include triggering an automatic speech recognizer (ASR) to perform speech recognition on the audio data when a hot phrase is identified in the audio data corresponding to the utterance.
[0005] The details of one or more implementations of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]
[0006] [Figure 1] FIG. 1 is a diagram of an exemplary system including a hot phrase detector for detecting hot phrases in audio. [Figure 2]FIG. 2 is a diagram of an example hot phrase detector of FIG. 1. [Figure 3] 1 is a flowchart of an exemplary arrangement of operations for a method for detecting hot phrases in audio. [Figure 4] FIG. 1 is a schematic diagram of an exemplary computing device that can be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION
[0007] Like reference symbols in the various drawings indicate like elements.
[0008] Through such a network of one or more assistant-enabled devices, a user has the power to speak queries or instructions aloud and have the digital assistant field, respond to queries, and / or carry out commands. Ideally, a user should be able to interact with a digital assistant as if they were talking to another person by speaking queries / commands into the assistant-enabled device. However, due to the fact that continuously performing full speech recognition on an assistant-enabled device with limited resources, such as a smartphone or smartwatch, it is difficult for a digital assistant to be constantly responsive to the user.
[0009] Thus, these assistant-enabled devices typically operate in a sleep or hibernation state, where a low-power hot word model can detect predefined hot words in the sound without performing speech recognition. Upon detecting a predefined hot word in a spoken utterance, the assistant-enabled device can wake up and proceed to perform more expensive processing, such as full on-device automatic speech recognition (ASR) or server-based ASR. To alleviate the need for users to speak predefined hot words and thus create an experience that supports always-on voice, some current efforts focus on directly activating digital assistants for a narrow set of common phrases (e.g., "set timer," "turn down volume," etc.). While in a low-power state, the assistant-enabled device can run a low-power model, such as a compact hot phrase (or warm word) model, or a low-power speech recognizer capable of detecting / recognizing fixed hot phrases in the sound. When a fixed hot phrase is detected / recognized by the low-power model, the voice-enabled device triggers and wakes up a higher-power, more accurate model to verify the presence of the fixed phrase in the sound.
[0010] One challenge with hot phrase detection models is that they are inflexible because they require the user to speak the exact command that the hot phrase model is trained to recognize. That is, the user must speak the exact hot phrase that the hot phrase model expects, without the ability to accommodate variations / flexibility in different phrases. In many scenarios, the sequence of words for a given command is not always spoken consecutively in an utterance, making it difficult to express the given command in a hot phrase. For example, when executing a command to send a text message, a user may say, "Send a message to John that I'm going to be late." Here, the command includes a fixed portion and several variable portions that are difficult to detect / recognize using conventional low-power hot phrase detection models. Therefore, conventional low-power hot phrase detection models lack flexibility and support only a limited number of different hot phrases.
[0011] Implementations herein relate to enabling more flexible hotword detection models capable of operating at low power while enabling users to interact more naturally with an always-on, flexible assistant-enabled device (AED). More specifically, the AED can execute a first-stage hotphrase detector that either executes a single hotword detection model configured to detect multiple different hotwords in an utterance, or executes a set of parallel hotword detection models, each configured to detect a corresponding hotword in the utterance. When the set of hotword detection models detects multiple hotwords in a given utterance, the first-stage hotphrase detector can integrate the multiple hotwords to detect a complete hotphrase. That is, when multiple hotwords are detected in an expected order within a predefined time window, a complete hotphrase can be detected, thereby enabling the AED to wake up from a low-power state to execute a second-stage hotphrase detector and verify the detected hotphrase. The second-stage hot phrase detector can be used to verify the hot words detected by the first stage and / or to enable recognition of parameters within a predefined time window that were not detected / recognized by the first-stage hot phrase detector. These parameters can include, for example, intermediate words / terms that the hot word model was not trained to detect, but that are otherwise distributed in the utterance spoken as part of the issued query / command.
[0012] The hot phrase detector can be activated / initialized to detect multiple hot words / trigger words based on the context related to the currently used application and / or the content displayed on the screen of the AED. For example, if a user sees "Send a message" and "Answer a call" displayed on the screen, the hot phrase detector can activate the words send / answer / call / message.
[0013] FIG. 1 illustrates an exemplary system 100 including an assistant-enabled device (AED) 104 running a digital assistant 109 with which a user 102 can interact through voice. In the illustrated example, the AED 104 corresponds to a smart speaker. However, the AED 104 may include other computing devices, such as, but not limited to, a smartphone, tablet, smart display, desktop / laptop, smartwatch, smart appliance, headphones, or in-vehicle infotainment device. The AED 104 includes data processing hardware 10 and memory hardware 12 that stores instructions that, when executed on the data processing hardware 10, cause the data processing hardware 10 to perform operations. The AED 104 includes an array of one or more microphones 16 configured to capture acoustic sounds, such as speech, directed toward the AED 104. The AED 104 may also include or communicate with an audio output device (e.g., a speaker) capable of outputting sound for playback to the user 102.
[0014] In the illustrated example, the user 102, near the AED 104, speaks an utterance 110, "Send a message to John: 'I'm going to be late.'" The microphone 16 of the AED 104 receives the utterance 110 and processes audio data 202 corresponding to the utterance 110. Initial processing of the audio data may include filtering the audio data and converting the audio data from an analog signal to a digital signal. Once the AED 104 processes the audio data, the AED may store the audio data in a buffer in the memory hardware 12 for further processing. Using the audio data in the buffer, the AED 104 may use a hot phrase detector 200 to detect whether the audio data 202 contains a hot phrase. More specifically, the hot phrase detector is configured to detect, in the audio data, each trigger word in a set of trigger words associated with a hot phrase during a fixed long-term window 220 of the audio data 202. Thus, the hot phrase detector 200 is configured to identify trigger words contained in the audio data without performing speech recognition on the audio data. In the illustrated example, the hot phrase detector 200 can determine that the utterance 110, "Send a message to John saying, 'I'm going to be late,'" contains the hot phrase 210, "Send a <..> message that <..>" if, during a fixed long-term window 220 of the audio data 202, the hot phrase detector 200 detects acoustic features in the audio data that are characteristic of each of the trigger words "send," "message," and "say." The acoustic features may be Mel-Frequency Cepstral Coefficients (MFCCs), which represent the short-term power spectrum of the utterance 110, or may be Mel-scale filter bank energies for the utterance 110. While the example represents each trigger word as a complete word, trigger words may also include subwords or parts of words.
[0015] As used herein, hot phrases 210 refer to a narrow set of trigger words (e.g., warm words) that the AED 104 is configured to recognize / detect in sounds without performing voice recognition to directly trigger the respective action. That is, hot phrases 210 serve a dual purpose: they wake up the AED 104 from a low-power state (e.g., sleep or hibernation) and issue a command that specifies an action for the digital assistant 109 to perform. In an example, the hot phrase 210 "send a <..> message that <..>" allows the user to call the AED 104 and trigger the performance of the respective action (e.g., send message content to the recipient) without the user having to prefix the utterance 110 with a predefined wake-up phrase (e.g., hot word, wake-up word) that first wakes the AED 104 to process the subsequent sounds corresponding to the command / query.
[0016] In particular, the hot phrase detector 200 is configured to detect the hot phrase 210 so long as each trigger word in the set of trigger words associated with the hot phrase 210 is detected in the audio data in a sequence that matches a predefined order associated with the hot phrase 210 within a fixed long-term window 220. That is, in addition to the fixed portion corresponding to the set of trigger words, the hot phrase detector 200 is configured to detect that the utterance 110 may also include some variable portion not associated with the hot phrase, such as words / terms spoken by the user 102 between the first trigger word (e.g., "send") and the last trigger word (e.g., "say"). As such, the hot phrase detector 200 does not require the user 102 to speak the exact command that the hot phrase detector 200 is trained to detect. That is, the hot phrase detector 200 has the ability to accommodate variation / flexibility in different phrases associated with the hot phrase, thereby eliminating the need for the user to consecutively speak the set of trigger words and allowing the user to embed open-ended parameters inside the hot phrase. While some hot phrases (e.g., “volume up,” “volume down,” “next track,” “set timer,” “stop alarm,” etc.) are typically spoken consecutively in an utterance, the hot phrase detector 200 disclosed herein is still capable of detecting sequences of trigger words for hot phrases that are not necessarily spoken consecutively, thereby enabling the AED 104 to detect a wider range of hot phrases 210. For example, in the example of FIG. 1 , the user 102 can communicate the same command to cause the digital assistant 109 to perform an action by speaking a slightly different utterance: “Send a nice message to my colleague John saying, ‘I’m going to be late.’” Here, this utterance still includes the set of trigger words associated with the hot phrase 110, “send a message that…,” but has different types of words / terms spoken by the user 102 between the first trigger word (e.g., “send”) and the last trigger word (e.g., “that”).Thus, the hot phrase detector 200 can still detect the hot phrases 210 and call and wake the AED 104 to trigger the performance of the respective action (e.g., sending the message content to the recipient).
[0017] The hot phrase detector 200 can operate / run continuously on the AED 104 while the AED 104 is in a low-power state listening for each trigger word in a set of trigger words in the streaming audio. When the AED 104 comprises a battery-powered device such as a smartphone, the hot phrase detector 200 can run on low-power hardware such as a digital signal processor (DSP) chip. The hot phrase detector 200 can operate / run on the application process (AP) / CPU of other types of AEDs, but can consume less power and require less processing than performing voice recognition.
[0018] When the hot phrase detector 200 identifies a hot phrase 210 in the audio data 202 by detecting each trigger word in the set of trigger words during a fixed long-term window 220 of the audio data 202, the AED 104 can trigger a wake-up process to begin voice recognition on the audio data 202 corresponding to the utterance 110. For example, an automatic speech recognizer (ASR) 116 running on the AED 104 can perform voice recognition on the audio data 202 as a verification step to confirm the presence of the hot phrase 210 in the audio data 202. The hot phrase detector 200 can rewind the audio data buffered in the memory hardware 12 when or before the first trigger word is detected and provide the audio data 202 beginning when or before the first trigger word is detected to the ASR 116 for processing thereon. In this manner, the buffered audio data 202 provided to the ASR 116 can include any introductory sounds beginning before the first trigger word. The duration of the introductory sounds may depend on the specific hot phrase 210 based on when the first trigger word is expected to occur in relation to other terms in a given utterance. The audio data 202 provided to the ASR 116 includes a portion corresponding to the introductory sounds, as well as a fixed long-term window 220 that characterizes the detected set of trigger words and a subsequent portion 222 that includes the message content "I'm going to be late."
[0019] Here, ASR 116 generates a transcription 120 of utterance 110 by processing audio data 202 and determines whether each trigger word in a set of trigger words associated with hot phrase 210 is recognized in transcription 120. ASR 116 can also process a portion 222 of audio data 202 corresponding to the content of a message, such as "I'm going to be late," after the last trigger word (e.g., "says") for inclusion in transcription 120. When ASR 116 determines that each trigger word in the set of trigger words is recognized in transcription 120, ASR 116 can provide transcription 120 to query processing 180 and perform query interpretation on transcription 120 to identify commands for digital assistant 109 and perform actions. Query processing 180 can receive transcription 120 of utterance 110 and execute a dedicated model configured to classify the likelihood that utterance 110 corresponds to a query / command-like utterance directed to digital assistant 109. Query processing 180 may additionally or alternatively perform query interpretation through a natural language processing (NLP) layer to perform intent classification. In an example, query interpretation performed on transcript 120 by query processing 180 may identify a command to send a message to a receiving device associated with John and provide the portion of transcript 120 containing the message content "I'm going to be late" to a messaging application for transmission to a receiving device associated with John.
[0020] On the other hand, if the ASR 116 determines that one or more of the trigger words in the set of trigger words are not recognized in the transcript 120, the ASR 116 determines that a missed trigger event occurred in the hot phrase detector 200 and, therefore, that the hot phrase 210 was not spoken in the user's utterance 110. In the example shown, the ASR 116 commands the AED 104 to inhibit the wake-up process and return to a low power state upon determining the missed trigger event. In some examples, when one or more of the trigger words detected by the hot phrase detector 200 are misrecognized by the ASR, the AED 104 performs a refinement process to fine-tune the hot phrase detector based on each trigger word misrecognized by the ASR.
[0021] Optionally, the ASR 116 may run on a remote server (not shown) that communicates over a network with the AED 104. In some examples, a computationally more powerful second-stage hot phrase detector verifies the presence of hot phrases 210 in the audio data 202 in addition to or instead of the verification performed by the ASR 116.
[0022] Referring to FIG. 2 , in some implementations, the hot phrase detector 200 includes a trigger word detection model 205 that is trained to detect each trigger word in a set of trigger words associated with a hot phrase 210. Audio data 202 converted from streaming audio captured by the microphone 16 of the AED 104 is buffered and provided to the trigger word detection model 205. The buffer may reside on the memory hardware 12. The model 205 is configured to output a confidence score 207 for a range of supported trigger words that includes the set of trigger words associated with the hot phrase 210. The range of supported trigger words may include other trigger words for different sets of trigger words associated with one or more additional hot phrases. Some trigger words may belong to multiple sets of trigger words. For example, the trigger word “message” may also belong to another set of trigger words associated with another hot phrase, “dictate a message: <..>.” In some examples, the model 205 includes a fixed-window audio model having several neural network layer blocks configured to process audio frames to generate a sound classification (e.g., a confidence score 207) every Nms. Here, the neural network layer blocks may include convolutional blocks. At each of a number of time steps, the output layer of the model may output a confidence score 207 for each supported trigger word. Thus, each trigger word supported by the model may be referred to as a target class. Once the model 205 outputs respective confidence scores 207 for trigger words that satisfy the trigger word confidence threshold, the hot phrase detector 200 detects respective trigger events 260 indicating the presence of the trigger word in the audio data 202 and buffers the respective trigger events 260 in a buffer.
[0023] In the illustrated example, each respective trigger event 260 in the buffer indicates a respective confidence score 207 for the corresponding trigger word, and a respective timestamp 209 indicates when the corresponding trigger word was detected in the audio data 202. For example, when the trigger word confidence threshold is equal to 0.7, the model 205 may detect each trigger event 260 for the trigger word “send” at zero (0) milliseconds (ms), indicating when the current fixed long time window 220 begins, if the model 205 outputs a respective confidence score equal to 0.95; for the trigger word “message” at 300 ms, if the model 205 outputs a respective confidence score equal to 0.8; and for the trigger word “say” at 1000 ms, if the model 205 outputs a respective confidence score equal to 0.85. Notably, the hot phrase detector 200 does not initiate a wake-up process in response to detecting a trigger event 260 for each individual trigger word.
[0024] The hot phrase detector 200 is further configured to execute a trigger word integration routine 280 each time the trigger word detection model 205 detects a respective trigger event 260. Here, the routine 280 is configured to determine whether a respective trigger event 260 is present in the buffer for each other corresponding trigger word in the set of trigger words, and, when a respective trigger event 260 is also present in the buffer for each other corresponding trigger word in the set of trigger words, to determine a hot phrase confidence score 282 indicating the likelihood that the utterance spoken by the user contains the hot phrase 210. In some examples, the hot phrase detector 200 identifies a hot phrase in the audio data 202 when the hot phrase confidence score 282 satisfies a hot phrase confidence threshold.
[0025] The routine 280 can be configured to determine a hot phrase confidence score 282 based on the respective trigger word confidence score 207 and the respective timestamp 209 indicated by the respective trigger event 260 in the buffer for each corresponding trigger word in the set of trigger words. In practice, each trigger event 260 can include multiple respective timestamps 209 indicating when the trigger word confidence score 207 exceeds the trigger word confidence threshold, allowing for combined consecutive detections using multiple techniques. For example, the timestamp 209 associated with the highest trigger word confidence score 207 can be indicated by the trigger event 260 stored in the buffer. Executing the trigger word integration routine 280 can include executing a neural network-based model. The neural network-based model can include a sequence-based machine learning model, such as a model having a recurrent neural network (RNN) architecture. In other examples, executing the trigger word integration routine 280 includes executing a grammar- or heuristic-based model. The routine 280 also considers sequences in which trigger words are detected during the fixed long-term window 220. That is, the sequence of the set of trigger words detected in the audio data 202 must match a predefined order associated with the hot phrase 210 in order to identify the hot phrase 210. For example, in the illustrated example, upon receiving a trigger event 260 indicating the detection of the trigger word "saying," the routine 280 can use the respective timestamps 209 in the buffer to determine that the trigger word "message" was detected after the trigger word "send" and before the trigger word "saying."
[0026] In some examples, the hot phrase confidence score 282 generated by the routine 280 is further based on the respective time duration between each pair of adjacent trigger words in the set of trigger words detected in the audio data. For example, the routine 280 can compare each respective time duration to a corresponding baseline expected time duration between pairs of adjacent trigger words for a particular phrase. That is, in the hot phrase 210 "send a <..> message that," the baseline expected time duration between the trigger words "send" and "message" is shorter than the baseline expected time duration between the trigger words "message" and "that." The routine 280 may also constrain the hot phrase based on a maximum time duration between particular pairs of trigger words.
[0027] The grammar (e.g., target classes / trigger words) for the trigger word detection model 205 can be constructed manually or learned / trained. When learning, AED queries for a particular vertical or intent can be used. For example, to represent commands to dictate a message in a hands-free manner and send the message to a recipient, a query transcription in which a user speaks a command to dictate and send a message can be utilized by the trigger word detection model 205 to learn a minimum set of trigger words that cover the largest portion of the query transcription for the send message command. That is, the minimum set of trigger words that cover the largest portion of the transcription is associated with the trigger words that occur most frequently in the transcription. In another example, when building the trigger word detection model 205 to support a low-power command to play music, a query transcription for the play music command can be obtained and the minimum set of trigger words that cover the largest portion of the transcription in the resulting query transcription can be identified. Notably, the trigger word detection model 205 can be built on-device and / or on a per-user basis. As a result, the trigger word detection model 205 is built to detect personalized hot phrases spoken by one user and / or multiple users of a particular AED. Trigger word detection models for common / generic hot phrases can also be discovered on the server side and pushed to a population of AEDs.
[0028] 3 is a flowchart of an example arrangement of operations for a method 300 for detecting hot phrases 210 in audio data 202 during a fixed long-term window 220 of the audio data 202. At operation 302, the method 300 includes receiving, at data processing hardware 10 of a user device 104 associated with the user 102, audio data 202 corresponding to an utterance 110 spoken by the user 102 and captured by the user device 104. The utterance 110 includes a command for the digital assistant 109 to perform an action. The user device 104 may include an assistant-enabled device (AED) that executes the digital assistant 109.
[0029] Operations 304, 306, and 308 of method 300 are performed during each of a plurality of fixed long-term windows 220 of audio data 202. In operation 304, method 300 includes determining, by data processing hardware 10, using hot phrase detector 200 configured to detect each trigger word in a set of trigger words associated with hot phrases 210, whether any of the trigger words in the set of trigger words are detected in audio data 202 during the corresponding fixed long-term window 220. In operation 306, when one of the trigger words in the set of trigger words associated with hot phrases 210 is detected in audio data 202 during the corresponding fixed long-term window 220, method 300 also includes determining, by data processing hardware 10, whether each other trigger word in the set of trigger words associated with hot phrases 210 is also detected in the audio data. At operation 308, when each other trigger word in the set of trigger words is also detected in the audio data, the method 300 also includes identifying, by the data processing hardware, hot phrases in the audio data corresponding to the utterance.
[0030] At operation 310, method 300 also includes triggering an automatic speech recognizer (ASR) to perform speech recognition on the audio data when a hot phrase is identified in the audio data corresponding to the utterance by data processing hardware 10. Here, the ASR can process sounds beginning when or before a first trigger word is detected to generate a transcript 120 for the utterance and determine whether each trigger word in the set of trigger words is found in the transcript 120.
[0031] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform a task. In some examples, a software application may be referred to as an "application," an "app," or a "program." Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0032] Non-transitory memory may be a physical device used to store programs (e.g., sequences of instructions) or data (e.g., program state information) on a temporary or permanent basis for use by a computing device. Non-transitory memory may be volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.
[0033] 4 is a schematic diagram of an exemplary computing device 400 that can be used to implement the systems and methods described herein. Computing device 400 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown herein, their connections and relationships, and their functionality are intended to be exemplary only and are not intended to limit the implementation of the invention(s) described and / or claimed herein.
[0034] Computing device 400 includes a processor 410, memory 420, a storage device 430, a high-speed interface / controller 440 connected to memory 420 and a high-speed expansion port 450, and a low-speed interface / controller 460 connected to a low-speed bus 470 and storage device 430. Each of components 410, 420, 430, 440, 450, and 460 are interconnected using various buses and may be mounted on a common motherboard or in other suitable manners. Processor 410 can process instructions for execution within computing device 400, including instructions stored in memory 420 or storage device 430, to display graphical information for a graphical user interface (GUI) on an external input / output device, such as a display 480 coupled to high-speed interface 440. In other implementations, multiple processors and / or multiple buses, along with multiple memories and multiple types of memory, may be used as appropriate. Multiple computing devices 400 may also be connected (e.g., as a bank of servers, a group of blade servers, or a multiprocessor system), with each device performing a portion of the required operations.
[0035] The memory 420 stores information non-transiently within the computing device 400. The memory 420 may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. The non-transient memory 420 may be a physical device used to store programs (e.g., sequences of instructions) or data (e.g., program state information) on a temporary or permanent basis for use by the computing device 400. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), as well as disk or tape.
[0036] The storage device 430 is capable of providing mass storage for the computing device 400. In some implementations, the storage device 430 is a computer-readable medium. In various different implementations, the storage device 430 may be a floppy disk device, a hard disk device, an optical disk device, or an array of devices including a tape device, a flash memory or other similar solid-state memory device, or a device in a storage area network or other configuration. In further implementations, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer- or machine-readable medium, such as the memory 420, the storage device 430, or memory on the processor 410.
[0037] The high-speed controller 440 manages bandwidth-intensive operations for the computing device 400, while the low-speed controller 460 manages lower-bandwidth-intensive operations. Such load arrangements are merely exemplary. In some implementations, the high-speed controller 440 is coupled to the memory 420, the display 480 (e.g., through a graphics processor or accelerator), and the high-speed expansion port 450, which can accept various expansion cards (not shown). In some implementations, the low-speed controller 460 is coupled to the storage device 430 and the low-speed expansion port 490. The low-speed expansion port 490, which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices such as a keyboard, pointing device, scanner, etc., or a network device such as a switch or router, for example, through a network adapter.
[0038] The computing device 400 can be implemented in several different forms, as shown in the figure. For example, the computing device 400 can be implemented as a standard server 400a, or multiple times in a group of such servers 400a, as a laptop computer 400b, or as part of a rack server system 400c.
[0039] Various implementations of the systems and techniques described herein can be realized in digital electrical and / or optical circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor that can be coupled to receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device, whether special purpose or general purpose.
[0040] These computer programs (also known as programs, software, software applications, or code) contain machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, and include a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0041] The processes and logic flows described herein may be implemented by one or more programmable processors, also referred to as data processing hardware, that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be implemented by special purpose logic circuitry, such as, for example, an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). Processors suitable for executing computer programs include, by way of example, both general-purpose and special purpose microprocessors, and any one or more processors of any type of digital computer. Generally, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as, for example, magnetic, magneto-optical, or optical disks, or be operatively coupled to receive data from or transmit data to the mass storage devices, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal or removable hard disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0042] To achieve user interaction, one or more aspects of the present disclosure can be implemented on a computer having a display device, such as a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen, for displaying information to the user, and optionally a keyboard and pointing device, such as a mouse or trackball, for allowing the user to provide input to the computer. Other types of devices can be used to achieve user interaction as well. For example, feedback to the user can be any form of sensory feedback, such as visual feedback, audio feedback, or tactile feedback, and input from the user can be received in any form, including acoustic, voice, or tactile input. Additionally, the computer can interact with the user by sending documents to and receiving documents from a device used by the user, for example, by sending web pages to a web browser on the user's client device in response to a request received from the web browser.
[0043] Although several implementations have been described, it will be understood that various modifications can be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims. [Explanation of symbols]
[0044] 10 Data Processing Hardware 12 Memory Hardware 16 microphones 100 systems 102 users 104 User Devices, Assistant-Enabled Devices, and AEDs 109 Digital Assistant 110 utterances 116 Automatic speech recognizer, ASR 120 Transcription 180 Query Processing 200 Hot Phrase Detector 202 Audio Data 205 Trigger Word Detection Model 207 Trigger Word Confidence Score 209 Timestamp 210 Hot Phrases 220 Fixed Long Window 222 parts 250 Hot Phrase Events 260 Trigger Event 280 Trigger Word Integration Routine 282 Hot Phrase Confidence Score 300 ways 400 computing devices 400a Standard Server 400b laptop computer 400c Rack Server System 410 processor 420 memory 430 Storage Devices 440 High-Speed Interface / Controller 450 High-Speed Expansion Port 460 Low-Speed Interface / Controller 470 Slow Bus 480 display 490 Low-Speed Expansion Port
Claims
1. 1. A computer-implemented method that runs on data processing hardware, comprising: receiving audio data corresponding to an utterance spoken by a user; The utterance is: A command for the digital assistant to perform an action; and a set of trigger words; one or more other words spoken between a first trigger word in said set of trigger words and a last trigger word in said set of trigger words; determining that the first trigger word in the set of trigger words has been detected in the audio data; after determining that the first trigger word in the set of trigger words has been detected in the audio data, determining that each other trigger word in the set of trigger words has also been detected in the audio data; triggering an automatic speech recognizer (ASR) to perform speech recognition on the audio data based on determining that the first trigger word and each other trigger word in the set of trigger words has been detected in the audio data; A method of causing the data processing hardware to perform operations including:
2. triggering ASR to perform speech recognition on the audio data, generating a transcription of the utterance by processing the audio data; performing query interpretation on the transcription to identify that the transcription includes the command for the digital assistant to perform an action; 2. The method of claim 1, comprising:
3. generating a transcription of the utterance, rewinding the audio data buffered in memory hardware in communication with the data processing hardware when or before the first trigger word in the set of trigger words is detected in the audio data; processing the audio data beginning when or before the first trigger word in a sequence of trigger words is detected to generate the transcription of the utterance; 3. The method of claim 2, comprising:
4. The method of claim 2 , wherein the transcription includes the one or more other words between the first trigger word in the set of trigger words and the last trigger word in the set of trigger words.
5. The operation is determining that each other trigger word in the set of trigger words has been detected in the audio data during a fixed long-term window beginning when the first trigger word in the set of trigger words is detected in the audio data; 2. The method of claim 1, wherein the step of triggering the ASR to perform speech recognition is based on determining that each other trigger word in the set of trigger words has been detected in the audio data during the fixed long-time window.
6. determining that the first trigger word in the set of trigger words has been detected in the audio data, using a hot phrase detector to generate a trigger word confidence score indicative of the likelihood of the first trigger word being present in the audio data; detecting the first trigger word in the audio data when the trigger word confidence score satisfies a trigger word confidence threshold; buffering in memory hardware in communication with the data processing hardware the audio data and a trigger event for the first trigger word detected in the audio data, the trigger event indicating the trigger word confidence score and a timestamp indicating when the first trigger word was detected in the audio data; 2. The method of claim 1, comprising:
7. the operations further comprising: performing a trigger word integration routine based on determining that the first trigger word in the set of trigger words is detected in the audio data; the trigger word integration routine: determining that for each other corresponding trigger word in the set of trigger words, a respective trigger event is also buffered in the memory hardware; configured to determine, for each other corresponding trigger word in the set of trigger words, a hot phrase confidence score indicative of the likelihood that the utterance spoken by the user contains the set of trigger words when a respective trigger event is also buffered in the memory hardware; 7. The method of claim 6, wherein triggering an ASR to perform speech recognition on the audio data comprises triggering an ASR to perform speech recognition on the audio data when the trigger word confidence score satisfies a trigger word confidence threshold.
8. The method of claim 7 , wherein the trigger word integration routine is performed via a neural network-based model.
9. The method of claim 7 , wherein the trigger word integration routine is performed via a discovery-based model.
10. The method of claim 1 , wherein the data processing hardware is located on a user device.
11. 1. A system comprising: data processing hardware; memory hardware in communication with the data processing hardware; The memory hardware stores instructions that, when executed on the data processing hardware, receiving audio data corresponding to an utterance spoken by a user; The utterance is: A command for the digital assistant to perform an action; and a set of trigger words; one or more other words spoken between a first trigger word in said set of trigger words and a last trigger word in said set of trigger words; determining that the first trigger word in the set of trigger words has been detected in the audio data; after determining that the first trigger word in the set of trigger words has been detected in the audio data, determining that each other trigger word in the set of trigger words has also been detected in the audio data; triggering an automatic speech recognizer (ASR) to perform speech recognition on the audio data based on determining that the first trigger word and each other trigger word in the set of trigger words has been detected in the audio data; A system for causing the data processing hardware to perform operations including:
12. triggering ASR to perform speech recognition on the audio data, generating a transcription of the utterance by processing the audio data; performing query interpretation on the transcription to identify that the transcription includes the command for the digital assistant to perform an action; The system of claim 11 , comprising:
13. generating a transcription of the utterance, rewinding the audio data buffered in memory hardware in communication with the data processing hardware when or before the first trigger word in the set of trigger words is detected in the audio data; processing the audio data beginning when or before the first trigger word in a sequence of trigger words is detected to generate the transcription of the utterance; The system of claim 12, comprising:
14. 13. The system of claim 12, wherein the transcription includes the one or more other words between the first trigger word in the set of trigger words and the last trigger word in the set of trigger words.
15. The operation is determining that each other trigger word in the set of trigger words has been detected in the audio data during a fixed long-term window beginning when the first trigger word in the set of trigger words is detected in the audio data; 12. The system of claim 11, wherein the step of triggering the ASR to perform speech recognition is based on determining that each other trigger word in the set of trigger words has been detected in the audio data during the fixed long-time window.
16. determining that the first trigger word in the set of trigger words has been detected in the audio data, using a hot phrase detector to generate a trigger word confidence score indicative of the likelihood of the first trigger word being present in the audio data; detecting the first trigger word in the audio data when the trigger word confidence score satisfies a trigger word confidence threshold; buffering in memory hardware in communication with the data processing hardware the audio data and a trigger event for the first trigger word detected in the audio data, the trigger event indicating the trigger word confidence score and a timestamp indicating when the first trigger word was detected in the audio data; The system of claim 11 , comprising:
17. the operations further comprising: performing a trigger word integration routine based on determining that the first trigger word in the set of trigger words is detected in the audio data; the trigger word integration routine: determining that for each other corresponding trigger word in the set of trigger words, a respective trigger event is also buffered in the memory hardware; configured to determine, for each other corresponding trigger word in the set of trigger words, a hot phrase confidence score indicative of the likelihood that the utterance spoken by the user contains the set of trigger words when a respective trigger event is also buffered in the memory hardware; 17. The system of claim 16, wherein triggering an ASR to perform speech recognition on the audio data comprises triggering an ASR to perform speech recognition on the audio data when the trigger word confidence score satisfies a trigger word confidence threshold.
18. 20. The system of claim 17, wherein the trigger word integration routine includes running the trigger word integration routine via a neural network-based model.
19. 20. The system of claim 17, wherein the trigger word integration routine includes running the trigger word integration routine via a discovery-based model.
20. The system of claim 11 , wherein the data processing hardware is on a user device.
Citation Information
Patent Citations
Speech processing system
JP2006215499A
Multi-channel speech recognition for vehicle environment
JP2019133156A
Processing method for waking up application program, apparatus, and storage medium
JP2019185011A
Information processing apparatus, information processing method and program
JP2020012954A
Electronic apparatus, control method and program
JP2020160387A