Short-lived repeating voice commands

The system addresses awkward interactions and efficiency issues in voice command systems by using warm words and repeating warm words for long-term operations, enhancing user experience and reducing false positives through low-power speech recognition.

JP2025536079APending Publication Date: 2025-10-30GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025527735
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-14
Filing Date
2023-11-02
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Existing voice command systems in voice-enabled environments require repeated hotwords to initiate commands, leading to awkward interactions and increased false positives due to multiple activated words, degrading efficiency and user experience.

Method used

Implementing a system where a set of warm words, including repeating warm words, are activated for controlling long-term operations, allowing users to issue commands without initial hotwords, with detection based on acoustic features and low-power speech recognition, reducing false positives and improving user interaction.

Benefits of technology

Enhances user experience by enabling natural, conversational interactions with reduced false positives and improved efficiency by limiting activated words, thus optimizing processing costs and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536079000001_ABST
    Figure 2025536079000001_ABST
Patent Text Reader

Abstract

A method (300) for detecting short-lived repetitive voice commands includes activating a set of one or more warm words (112) each associated with a respective action for controlling a long-lived operation performed by a digital assistant (105). While the digital assistant is performing the long-lived operation, the method includes receiving audio data (202) and detecting a warm word from the set of one or more activated warm words within the audio data. In response to detecting the warm word, the method includes performing a respective action associated with the detected warm word and activating a set of one or more repeating warm words associated with the detected warm word. The method further includes receiving additional audio data, detecting a repeating warm word from the set of one or more activated repeating warm words within the additional audio data, and performing a respective action associated with the detected repeating warm word.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to short-lived repeating voice commands. [Background technology]

[0002] In a voice-enabled environment, a user simply speaks a query or command out loud, and the digital assistant will address and respond to the query and / or cause the command to be carried out. A voice-enabled environment (e.g., home, work, school, etc.) may be implemented using a network of connected microphone devices located in various rooms or environmental areas. Through such a network of microphones, a user has the ability to verbally query a digital assistant from essentially anywhere in the environment, without having to have a computer or other device in front of or nearby. For example, while cooking in the kitchen, a user may ask the digital assistant, "Please set the timer for 20 minutes," and in response, the digital assistant will confirm that the timer is set (in the form of a synthesized voice output) and alert the user (e.g., in the form of an alarm or other audible alert from an audio speaker) when the timer reaches 20 minutes. Often, there is more than one way to ask a digital assistant to perform an action. For example, a user may say "more time" or "longer," both of which may indicate that the user wants to add additional time to the timer beyond the initial 20-minute timer, and in response, the digital assistant can detect the specific action of increasing the timer's time and then add time to the timer. Summary of the Invention

[0003] One aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform an operation, the operation including activating a set of one or more warm words, each associated with a respective action for controlling a long-term operation performed by the digital assistant. While the digital assistant is performing the long-term operation and the set of one or more warm words is activated, the operation also includes receiving audio data corresponding to a first utterance captured by the assistant-enabled device and detecting a warm word from the set of one or more activated warm words in the audio data. In response to detecting a warm word from the set of one or more activated warm words, the operation also includes performing a respective action associated with the detected warm word to control the long-term operation, and activating a set of one or more repeating warm words associated with the detected warm word. Here, each of the set of one or more activated repeating warm words is associated with a respective action for controlling the long-term operation. The operations further include receiving additional audio data corresponding to the second utterance captured by the assistant-enabled device, and detecting a repeating warm word from the set of one or more activated repeating warm words within the additional audio data. In response to detecting the repeating warm word from the set of one or more repeating warm words, the operations also include performing a respective action associated with the detected repeating warm word to control long-term behavior.

[0004] Embodiments of the present disclosure may include one or more of the following optional features. In some embodiments, activating a set of one or more repeated warm words includes activating a respective repeated warm word model for execution on the assistant-enabled device for each corresponding repeated warm word in the set of one or more activated repeated warm words. Here, detecting repeated warm words from the set of one or more activated repeated warm words in the additional audio data includes detecting the repeated warm words in the additional audio data without performing speech recognition on the additional audio data using the respective repeated warm word models activated for the corresponding detected repeated warm words. In these embodiments, detecting repeated warm words in the additional audio data may include extracting audio features of the additional audio data, generating repeated warm word confidence scores by processing the extracted audio features using the respective repeated warm word models activated for the corresponding detected repeated warm words, and determining that the additional audio data corresponding to the second utterance includes the corresponding repeated warm word if the repeated warm word confidence score satisfies a repeated warm word confidence threshold. Additionally or alternatively, detecting warm words from the set of one or more activated warm words in the audio data includes detecting the warm words in the audio data without performing speech recognition on the audio data using respective warm models activated for the corresponding warm words, wherein each repeated warm word model activated for a corresponding detected repeated warm word in the additional audio data is derived from a respective warm word model activated for detecting the corresponding warm word in the audio data.

[0005] In some examples, activating the set of one or more repeating warm words includes executing a speech recognizer on the assistant-enabled device, and detecting repeating warm words from the set of one or more activated repeating warm words in the additional audio data includes recognizing the repeating warm words in the additional audio data using a speech recognizer executing on the assistant-enabled device, where the speech recognizer is biased to recognize the repeating warm words in the set of one or more activated repeating warm words. In some implementations, the operation includes receiving a user input indication indicating a selection of one or more available repeating warm words to add to the set of one or more repeating warm words in a user interface associated with the assistant-enabled device, the user interface presenting a list of available repeating warm words to add to and / or remove from the set of one or more repeating warm words associated with the warm word, and to be activated when the warm word is detected in the audio data. In these embodiments, each corresponding available repeating warm word in the list of available repeating warm words presented by the user interface associated with the assistant-enabled device can be displayed by the user interface as a respective graphical element that can be selected to add or remove the corresponding available repeating warm word from a set of one or more warm words associated with the warm word, and to be activated when the warm word is detected in the audio data.

[0006] In some examples, the operations further include obtaining a voice command log including previous instances of a user of the assistant-enabled device speaking a repeat voice command for controlling a long-term operation immediately after detecting a warm word in the previous audio data. Here, at least one repeat warm word in the set of one or more activated warm words is learned based on the repeat voice command. In these examples, the operations may further include prompting the user of the assistant-enabled device to add the repeat voice command included in the voice command log to a set of one or more repeat warm words associated with the detected warm word. In response to receiving a user indication indicating approval to add the repeat voice command to the set of one or more repeat warm words, the operations also include associating the repeat voice command with the detected warm word and activating the repeat voice command as a repeat warm word associated with the detected warm word.

[0007] In some implementations, each action for controlling a long-term operation associated with at least one repeating warm word in the set of one or more activated repeating warm words is the same as each action performed in response to detecting the associated warm word from the set of one or more activated warm words in the audio data. In other implementations, each action for controlling a long-term operation associated with at least one repeating warm word in the set of one or more activated repeating warm words is different from each action performed in response to detecting the associated warm word from the set of one or more activated warm words in the audio data. In some examples, activating the set of one or more repeating warm words associated with the detected warm word includes activating the set of one or more repeating warm words for a predetermined duration.

[0008] In some implementations, after performing the respective action associated with the detected repeating warm word to control the long-term operation, the operations further include maintaining activation of at least one of the repeating warm words in the set of one or more repeating warm words associated with the detected warm word. In some examples, after performing the respective action associated with the detected repeating warm word to control the long-term operation, the operations further include deactivating the set of one or more repeating warm words based on receiving a command that conflicts with the long-term operation or based on failing to detect additional repeating warm words in the activated set of one or more repeating warm words after a predetermined duration has elapsed. In some implementations, in response to detecting a warm word from the activated set of one or more warm words, the operations further include activating a set of one or more warm words each associated with a respective action for controlling a long-term operation different from the long-term operation.

[0009] Another aspect of the present disclosure provides a system including data processing hardware and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed by the data processing hardware, cause the data processing hardware to perform operations, including activating a set of one or more warm words, each associated with a respective action for controlling a long-term operation performed by the digital assistant. While the digital assistant is performing the long-term operation and the set of one or more warm words is activated, the operation also includes receiving audio data corresponding to a first utterance captured by the assistant-enabled device and detecting a warm word from the set of one or more activated warm words in the audio data. In response to detecting a warm word from the set of one or more activated warm words, the operation also includes performing a respective action associated with the detected warm word to control the long-term operation and activating a set of one or more repeating warm words associated with the detected warm word. Here, each of the set of one or more activated repeating warm words is associated with a respective action for controlling the long-term operation. The operations further include receiving additional audio data corresponding to the second utterance captured by the assistant-enabled device, and detecting a repeating warm word from the set of one or more activated repeating warm words within the additional audio data. In response to detecting the repeating warm word from the set of one or more repeating warm words, the operations also include performing a respective action associated with the detected repeating warm word to control long-term behavior.

[0010] This aspect may include one or more of the following optional features. In some implementations, activating one or more sets of repeated warm words includes activating a respective repeated warm word model for execution on the assistant-enabled device for each corresponding repeated warm word in the set of one or more activated repeated warm words. Here, detecting repeated warm words from the set of one or more activated repeated warm words in the additional audio data includes detecting the repeated warm words in the additional audio data without performing speech recognition on the additional audio data using the respective repeated warm word models activated for the corresponding detected repeated warm words. In these implementations, detecting repeated warm words in the additional audio data may include extracting audio features of the additional audio data, generating repeated warm word confidence scores by processing the extracted audio features using the respective repeated warm word models activated for the corresponding detected repeated warm words, and determining that the additional audio data corresponding to the second utterance contains the corresponding repeated warm word if the repeated warm word confidence score satisfies a repeated warm word confidence threshold. Additionally or alternatively, detecting warm words from the set of one or more activated warm words in the audio data includes detecting the warm words in the audio data without performing speech recognition on the audio data using respective warm models activated for the corresponding warm words, wherein each repeated warm word model activated for a corresponding detected repeated warm word in the additional audio data is derived from a respective warm word model activated for detecting the corresponding warm word in the audio data.

[0011] In some examples, activating the set of one or more repeating warm words includes executing a speech recognizer on the assistant-enabled device, and detecting repeating warm words from the set of one or more activated repeating warm words in the additional audio data includes recognizing the repeating warm words in the additional audio data using a speech recognizer executing on the assistant-enabled device, where the speech recognizer is biased to recognize the repeating warm words in the set of one or more activated repeating warm words. In some implementations, the operation includes receiving a user input indication indicating a selection of one or more available repeating warm words to add to the set of one or more repeating warm words in a user interface associated with the assistant-enabled device, the user interface presenting a list of available repeating warm words to add to and / or remove from the set of one or more repeating warm words associated with the warm word, and to be activated when the warm word is detected in the audio data. In these embodiments, each corresponding available repeating warm word in the list of available repeating warm words presented by the user interface associated with the assistant-enabled device can be displayed by the user interface as a respective graphical element that can be selected to add or remove the corresponding available repeating warm word from a set of one or more warm words associated with the warm word, and to be activated when the warm word is detected in the audio data.

[0012] In some examples, the operations further include obtaining a voice command log including previous instances of a user of the assistant-enabled device speaking a repeat voice command for controlling a long-term operation immediately after detecting a warm word in the previous audio data. Here, at least one repeat warm word in the set of one or more activated warm words is learned based on the repeat voice command. In these examples, the operations may further include prompting the user of the assistant-enabled device to add the repeat voice command included in the voice command log to a set of one or more repeat warm words associated with the detected warm word. In response to receiving a user indication indicating approval to add the repeat voice command to the set of one or more repeat warm words, the operations also include associating the repeat voice command with the detected warm word and activating the repeat voice command as a repeat warm word associated with the detected warm word.

[0013] In some implementations, each action for controlling a long-term operation associated with at least one repeating warm word in the set of one or more activated repeating warm words is the same as each action performed in response to detecting the associated warm word from the set of one or more activated warm words in the audio data. In other implementations, each action for controlling a long-term operation associated with at least one repeating warm word in the set of one or more activated repeating warm words is different from each action performed in response to detecting the associated warm word from the set of one or more activated warm words in the audio data. In some examples, activating the set of one or more repeating warm words associated with the detected warm word includes activating the set of one or more repeating warm words for a predetermined duration.

[0014] In some implementations, after performing the respective action associated with the detected repeating warm word to control the long-term operation, the operations further include maintaining activation of at least one of the repeating warm words in the set of one or more repeating warm words associated with the detected warm word. In some examples, after performing the respective action associated with the detected repeating warm word to control the long-term operation, the operations further include deactivating the set of one or more repeating warm words based on receiving a command that conflicts with the long-term operation or based on failing to detect additional repeating warm words in the activated set of one or more repeating warm words after a predetermined duration has elapsed. In some implementations, in response to detecting a warm word from the activated set of one or more warm words, the operations further include activating a set of one or more warm words each associated with a respective action for controlling a long-term operation different from the long-term operation.

[0015] The details of one or more embodiments of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will become apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]

[0016] [Figure 1A] FIG. 1 is a schematic diagram of an example system including a user controlling long-term behavior using recurring warm words that each specify a respective action for a digital assistant to perform. [Figure 1B] FIG. 1 is a schematic diagram of an example system including a user controlling long-term behavior using recurring warm words that each specify a respective action for a digital assistant to perform. [Figure 1C]FIG. 1 is a schematic diagram of an example system including a user controlling long-term behavior using recurring warm words that each specify a respective action for a digital assistant to perform. [Figure 2] FIG. 1 is a schematic diagram of an iterative warm word detection process. [Figure 3] 1 is an exemplary GUI rendered on the screen of a user device. [Figure 4] 1 is a flowchart of an exemplary arrangement of method operations for detecting repetitive warm words. [Figure 5] FIG. 1 is a schematic diagram of an exemplary computing device that can be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION

[0017] Like reference symbols in the various drawings indicate like elements.

[0018] A user's interaction with an Assistant-enabled device is designed to be primarily, but not exclusively, through voice input. As a result, the Assistant-enabled device needs some way to identify when any given utterance in the surrounding environment is directed at the device, rather than directed at an individual in the environment or originating from a non-human source (e.g., a television or music player). One way to achieve this is through the use of hot words, reserved by agreement among users in the environment as a predetermined word or words to be spoken to attract the device's attention. In the exemplary environment, the hot word used to attract the Assistant's attention is the words "OK, computer." As a result, whenever the words "OK, computer" are spoken, they are picked up by the microphone and transmitted to a hot word detector, which performs speech understanding techniques to determine whether the hot word was spoken and, if so, waits for the next command or query. Thus, utterances directed to an Assistant-enabled device take the general form [hot word][query], where the "hot word" in this example is "OK computer" and the "query" can be any question, command, declaration, or other request that can be voice-recognized, analyzed, and processed by the system, either alone or in conjunction with a server over a network.

[0019] When a user provides several hotword-based commands to an Assistant-enabled device, such as a mobile phone or smart speaker, the user's interaction with the phone or speaker can be awkward. The user may say, "OK, computer, play my study playlist." The phone or speaker may start playing the first song in the playlist. The user may want to turn up the volume of the music and say, "OK, computer, turn up the volume." To turn up the volume again, the user may say, "OK, computer, turn up the volume." To alleviate the need to continually repeat hotwords before speaking commands, the Assistant-enabled device may be configured to recognize / detect a limited set of hot phrases or warm words to directly trigger the respective actions. In this example, the warm word "turn up the volume" serves two purposes: as a hotword and as a command. Therefore, instead of saying "OK, computer, turn up the volume," the user can simply say, "Turn up the volume" to invoke the Assistant-enabled device to trigger the respective action.

[0020] While the spoken command "turn up the volume" itself may be a natural way to communicate a user's intent to increase the volume of music, a user may find it unnatural to be limited to interacting with a phone or speaker as the only way to adjust the volume. For example, if a user wants to continue adjusting the volume, rather than repeatedly speaking "turn up the volume," the user may speak alternative voice commands such as "more," "one more," or "louder," which are associated with "turn up the volume" (i.e., triggering the same action) but which are more natural to repeat within a limited time after the user has spoken "turn up the volume."

[0021] A set of warm words can be activated to control long-term operations. As used herein, a long-term operation refers to an application or event executed by the digital assistant that may last for an extended duration or a discrete period and that can be controlled by the user while the application or event is ongoing. For example, if the digital assistant sets a timer for 30 minutes, the timer is a long-term operation from the time the timer is set until the timer expires, or until the resulting alert is acknowledged after the timer expires. In this case, a warm word such as "stop the timer" can be activated to allow the user to stop the timer by simply saying "stop the timer" without first speaking a hot word. Similarly, a command to instruct the digital assistant to play music from a streaming music service is a long-term operation while the digital assistant is streaming music from the streaming music service through a playback device. In this case, the set of active warm words can be "pause," "pause music," "volume up," "volume down," "next," "previous," etc. to control the playback of the music the digital assistant is streaming through the playback device. Long-running actions can include multi-step dialog queries such as "make a restaurant reservation," in which different sets of warm words are activated depending on the given stage of the multi-step dialog. For example, a digital assistant can prompt a user to select from a list of restaurants, whereby a set of warm words each including a respective identifier (e.g., the name or number of a restaurant in the list) can be activated to complete the action of selecting a restaurant from the list and making a reservation for that restaurant.

[0022] Additionally, each warm word in the set of warm words for controlling long-term actions can include one or more repeating voice commands associated with the warm word and active for a limited time after the user invokes the warm word. For example, the warm word "tell me a joke" for a long-term action of querying a digital assistant can have the repeating warm word "other" associated with it. Similarly, the warm word "turn on the lights" for a long-term action of operating a smart light bulb can have the repeating warm word "brighter" associated with it.

[0023] One challenge with repeated warm words is limiting the number of words / phrases that are simultaneously activated so as not to degrade quality and efficiency. For example, the number of false positives, which indicates when an assistant-enabled device incorrectly detects / recognizes one of the active words, increases significantly as the number of simultaneously activated repeated warm words increases. Embodiments herein are directed to activating a set of one or more repeated warm words associated with an ongoing long-term action, which are activated for a period of time after a user invokes an active warm word to control the long-term action. That is, the active repeated warm word is associated with a high likelihood of being spoken by the user after the first warm word to control the long-term action. The warm word detector runs on the assistant device and can consume low power.

[0024] By associating repeated warm words with specific active warm words so that the repeated warm words are warm-word dependent, the accuracy of triggering the respective actions upon detecting the repeated warm words is improved because the assistant device is biased to detect the repeated warm words when the specific active warm words are spoken. Additionally, the assistant device not only wakes up less often and potentially connects to a server less often, but also reduces the number of false positives, improving processing costs. Furthermore, the user's experience with the digital assistant is improved because the user's commands for controlling the execution of long-running actions reflect a more natural, conversational interaction.

[0025] 1A-1C illustrate exemplary systems 100, 100a-c for activating a warm word 112 for an action to control a long-lasting operation and for activating a repeating warm word 112R associated with an initial warm word 112 spoken by a user 102 in an initial command for controlling the long-lasting operation. Briefly, as described in more detail below, an assistant-enabled device (AED) 104 starts playing music 122 in response to an utterance 106 spoken by a user 102, such as "OK computer, play music." While the AED 104 is executing the long-lasting operation of the music 122 as playback audio from the speaker 18, the AED 104 can detect / recognize an active warm word 112, such as "Turn up the volume" (FIG. 1B), spoken by the user 102 as an action to control the long-lasting operation, e.g., as a command to increase the volume of the playback audio of the music 122. Because the AED 104 detects the active warm word 112 (i.e., turn up the volume), the AED 104 activates a set of repeating warm words 112R associated with the detected active warm word 112, "turn up the volume," to address follow-up queries that may include one of the repeating warm words 112R for a limited time after detecting the active warm word 112.

[0026] Systems 100a-100c include an AED 104 running a digital assistant 105 that a user 102 may interact with by speaking. In the illustrated example, the AED 104 corresponds to a smart speaker with which the user 102 may interact. However, the AED 104 may include other computing devices, such as, but not limited to, a smartphone, tablet, smart display, desktop / laptop, smartwatch, smart glasses / headset, smart appliance, headphones, or vehicle infotainment device. The AED 104 includes data processing hardware 10 and memory hardware 12 that stores instructions that, when executed on the data processing hardware 10, cause the data processing hardware 10 to perform operations. The AED 104 includes an array of one or more microphones 16 configured to capture sounds, such as speech, directed at the AED 104. The AED 104 may also include or communicate with an audio output device (e.g., a speaker) 18 that may output audio, such as music 122 and / or synthesized voice, from the digital assistant 105. Additionally, the AED 104 may include or be in communication with one or more cameras 19 configured to capture images within the environment and output image data.

[0027] In some configurations, the AED 104 communicates with a user device 50 associated with the user 102. In the illustrated example, the user device 50 includes a smartphone with which the user 102 may interact. However, the user device 50 may include other computing devices, such as, but not limited to, a smartwatch, a smart display, smart glasses, a smartphone, smart glasses / headset, a tablet, a smart appliance, headphones, a computing device, a smart speaker, or other assistant-enabled devices. The user device 50 may include at least one microphone 52 resident on the user device 50 that communicates with the AED 104. In these configurations, the user device 50 may also communicate with one or more microphones 16 present on the AED 104. Additionally, the user 102 may control and / or configure the AED 104 and may interact with the digital assistant 105 using an interface 300, such as a graphical user interface (GUI) 300 ( FIG. 3 ) rendered for display on the screen of the user device 50.

[0028] 1A shows a user 102 speaking an utterance 106, "Ok computer, play music," near an AED 104. The microphone 16 of the AED 104 receives the utterance 106 and processes audio data 202 corresponding to the utterance 106. Initial processing of the audio data 202 may include filtering the audio data 202 and converting the audio data 202 from an analog signal to a digital signal. Once the AED 104 processes the audio data 202, the AED may store the audio data 202 in a buffer in the memory hardware 12 for further processing. With the audio data 202 in the buffer, the AED 104 may use a hot word detector 108 to detect whether the audio data 202 contains a hot word. The hot word detector 108 is configured to identify hot words contained in the audio data 202 without performing voice recognition on the audio data 202.

[0029] In some implementations, the hot word detector 108 is configured to identify hot words in an early portion of the utterance 106. In this example, the hot word detector 108 may determine that the utterance 106, "Ok computer, play music," includes the hot word 110, "ok computer," if the hot word detector 108 detects acoustic features in the audio data 202 that are characteristic of the hot word 110. The acoustic features may be Mel-Frequency Cepstral Coefficients (MFCCs), which are a representation of the short-term power spectrum of the utterance 106, or may be Mel-scale filter bank energy of the utterance 106. For example, the hot word detector 108 may detect that the utterance 106, "Ok computer, play music," includes the hot word 110, "ok computer," based on generating MFCCs from the audio data 202 and classifying the MFCCs as including MFCCs similar to MFCCs that are characteristic of the hot word "ok computer" stored in a hot word model of the hot word detector 108. As another example, the hot word detector 108 may detect that the utterance 106 "Ok computer, play music" contains the hot word 110 "ok computer" based on generating mel-scale filter bank energy from the audio data 202 and classifying the mel-scale filter bank energy as containing mel-scale filter bank energy similar to mel-scale filter bank energy that is characteristic of the hot word "ok computer" stored in the hot word model of the hot word detector 108.

[0030] When the hot word detector 108 determines that the audio data 202 corresponding to the utterance 106 includes the hot word 110, the AED 104 may trigger a wake-up process to begin speech recognition on the audio data 202 corresponding to the utterance 106. For example, a speech recognizer 116 executing on the AED 104 using an automatic speech recognition model 117 (FIG. 2) may perform speech recognition or semantic interpretation on the audio data 202 corresponding to the utterance 106. For example, the speech recognizer 116 executing on the AED 104 may perform speech recognition or semantic interpretation on the audio data 202 corresponding to the utterance 106. The speech recognizer 116 may perform speech recognition on the portion of the audio data 202 that follows the hot word 110. In this example, the speech recognizer 116 may identify the words "play music" in the command 118.

[0031] In some examples, the AED 104 is configured to communicate with a remote system 130 over the network 120. The remote system 130 may include remote resources, such as remote data processing hardware 132 (e.g., a remote server or CPU) and / or remote memory hardware 134 (e.g., a remote database or other storage hardware). The AED 104 may utilize the remote resources to perform various functions related to speech processing and / or synthesis playback communication. In some implementations, the speech recognizer 116 is located on the remote system 130 in addition to or instead of the AED 104. When the hot word detector 108 triggers the AED 104 to wake up in response to detecting a hot word 110 in an utterance 106, the AED 104 may transmit initial audio data 202 corresponding to the utterance 106 to the remote system 130 over the network 120. Here, the AED 104 may transmit a portion of the initial audio data 202 including the hot word 110 to the remote system 130 to confirm the presence of the hot word 110. Alternatively, the AED 104 may send only the portion of the initial audio data 202 corresponding to the portion of the utterance 106 after the hot word 110 to the remote system 130, which executes the speech recognizer 116 to perform speech recognition and returns a transcription of the initial audio data 202 to the AED 104.

[0032] 1A-2 , the AED 104 may further include an NLU module 124 that performs semantic interpretation on the utterance 106 to identify queries / commands intended for the AED 104. Specifically, the NLU module 124 identifies words in the utterance 106 identified by the speech recognizer 116 and performs semantic interpretation to identify any voice commands in the utterance 106. The NLU module 124 of the AED 104 (and / or the remote system 130) may identify the words "play music" as a command specifying a long-running action of the digital assistant 105 (i.e., play music 122). In the illustrated example, the digital assistant 105 begins executing the long-running action of playing music 122 as playback audio (e.g., track 1) from the speaker 18 of the AED 104. In some implementations, the AED 104 verifies the identity of the user 102 before beginning execution of the long-running action of playing music 122. The digital assistant 105 may stream music 122 from a streaming service (not shown), or the digital assistant 105 may instruct the AED 104 to play music stored on the AED 104. An exemplary long-term operation includes music playback, although long-term operations may also include other types of media playback, such as videos, podcasts, and / or audiobooks. Long-term operations may also include home automation (e.g., adjusting light levels, controlling the thermostat, etc.).

[0033] 1A , the AED 104 notifies the user 102 (e.g., Barb) who spoke the utterance 106 that a long-lasting action will be performed, that a warm word 112 will be activated, and / or that the user 102 can speak any of the warm words 112 to command the AED 104 to perform the respective action to control the long-lasting action. For example, the digital assistant 105 may generate an audible output synthesized voice 123 from the speaker 18 of the AED 104 stating, "Hey Barb, you can speak music playback controls without saying 'Ok computer.'" In an additional example, the digital assistant 105 may provide a notification to a user device 50 (e.g., a smartphone) linked to the identified user's user account to inform the identified user 102 (e.g., Barb) which warm word 112 is currently active to control the long-lasting action.

[0034] The AED 104 (and / or server 120) may include an action identifier 126, a warm word selector 220, and a warm word detector 210. The action identifier 126 may be configured to identify one or more long-term actions currently being performed by the digital assistant 105. For each long-term action currently being performed by the digital assistant 105, the warm word selector 220 may activate a corresponding warm word 112 associated with each respective action for controlling the long-term action. The warm word detector 210 may receive the activated warm word 112 and, as described in more detail below, detect the warm word 112 when streaming audio captured by the AED 104 without performing speech recognition on the captured audio.

[0035] In some examples, the warm word selector 220 accesses a registry or table (e.g., stored in memory hardware 12) that associates the identified long-running action with corresponding warm words 112 that are highly correlated with the long-running action. For example, if the long-running action corresponds to a set timer function, the associated warm words 112 available for the warm word selector 220 to activate include the warm word 112 "stop timer" to instruct the digital assistant 105 to stop the timer. Similarly, for the long-running action "call [contact name]," the associated warm words 112 include "hang up" and / or "end call," which are warm word(s) 112 to end an ongoing call. In the illustrated example, for a long-duration action of playing music 122, the associated warm words 112 available for activation by the warm word selector 220 include the warm words 112 "next," "pause," "previous," "volume up," and "volume down," each associated with a respective action for controlling the playback of music 122 from the speaker 18 of the AED 104. Thus, the warm word selector 220 may activate these warm words 112 (i.e., active warm words 112A) while the digital assistant 105 is performing the long-duration action and deactivate these warm words 112 (e.g., inactive warm words 112N) when the long-duration action ends. Similarly, different warm words 112 may be activated / deactivated depending on the state of the ongoing long-duration action. For example, if the user speaks "pause" to pause the playback of the music 122, the warm word selector 220 may activate the warm word 112 of "play" to resume playing the music 122.In some configurations, instead of accessing the registry, the warm word selector 220 examines code associated with long-running applications (e.g., a music application running in the foreground or background of the AED 104) to identify any warm words 112 that the application's developer wants the user 102 to be able to speak to interact with the application, and the respective actions for each warm word 112. The warm words 112 in the registry may also relate to follow-up queries that the user 102 (or a typical user) tends to issue following a given query, e.g., "Ok computer, next track." In some implementations, the warm words 112 that the warm word selector 220 activates are added to existing / active warm words 112. Alternatively, the warm word selector 220 may selectively activate / deactivate warm words 112 based on the processing capabilities of the AED 104 to keep the total number of warm words 112 low.

[0036] Each warm word 112 in the registry may include a set of associated repeating warm words 112R, and one or more repeating warm words 112R in the set of repeating warm words 112R may be associated with the same respective action to control the long-term action as the warm word 112. For example, if the long-term action corresponds to an adjustable lighting function and the associated warm word 112 is "turn on the lights," the associated repeating warm words 112R available for the warm word selector 220 to activate include the repeating warm words 112R "again," "further," "brighter," "higher," and "more lights" to instruct the digital assistant 105 to turn on the lights. However, at least one of the repeating warm words 112R may be associated with an action related to controlling a long-term behavior as a warm word 112, such as "darker" and / or "lower" in examples, so that a user can return to a lower lighting level after speaking the warm word 112 "turn on the lights" or one of the repeating warm words 112R related to also causing the lighting level to be increased. Similarly, in the case of a long-term behavior of issuing a query and the warm word 112 "tell me a joke," the associated repeating warm words 112R include "another" and / or "one more" to issue an additional query. In other examples, the respective actions for controlling a long-term behavior associated with at least one repeating warm word 112R are different from the respective actions for controlling a long-term behavior by the associated warm word 112. For example, if the long-term action is scanning the content, the warm word 112 "fast forward" includes not only the associated repeated warm word 112R "faster" to continue the fast-forward action, but also the associated repeated warm word 112R "stop" to end the fast-forward action.

[0037] In some implementations, for each warm word 112 (and repeated warm word 112R), the AED 104 further receives a respective warm word model 114 configured to detect the corresponding warm word 112 in the streaming audio without performing speech recognition. For example, the AED 104 (and / or the server 130) further includes one or more warm word models 114, where the warm word models 114 may be stored in the memory hardware 12 of the AED 104 or in remote memory hardware 134 on the server 130. If stored on the server 130, the AED 104 can activate the warm word model 114 (via the warm word selector 220) so that the AED 104 can request the server 130 to retrieve the warm word model 114 for the corresponding warm word 112 and provide the retrieved warm word model 114. A warm word detector 210 running on the AED 210 may receive the activated warm word models 114 to detect utterances of corresponding warm words 112 when streaming audio captured by the AED 104 without performing speech recognition on the captured audio. Furthermore, a single warm word model 114 may be able to detect all of the warm words 112 (and repeated warm words 112R) associated with a long-term activity when streaming audio. In some implementations, the warm word model 114 for each corresponding repeated warm word 112R is derived from the warm word model 114 for the corresponding warm word 112.

[0038] In some configurations, the AED 104 receives code associated with an application loaded on the AED 104 (e.g., a music application running in the foreground or background of the AED 104) to identify any warm words 112 and associated warm word models 114 that the application's developer desires the user 102 to speak to interact with the application, as well as respective actions for each warm word 112. In another example, the AED 104 receives, for at least one warm word 112 in the active warm words 112A of at least one user 102, a respective warm word model 114 via a warm word application programming interface (API) running on the AED 104, configured to detect the corresponding warm word 112 when streaming audio without performing speech recognition. The warm words 112 in the registry may also relate to follow-up queries that the user 102 (or a typical user) tends to issue following a given query, e.g., "Ok computer, play my music playlist."

[0039] In additional embodiments, activating the warm word 112A causes the AED 104 to run the speech recognizer 116 in a low-power and low-fidelity state. Here, the speech recognizer 116 is constrained or biased to recognize only the active warm word 112A when spoken in speech captured by the AED 104. Because the speech recognizer 116 recognizes only a limited number of terms / phrases, the number of parameters of the speech recognizer 116 can be significantly reduced, thereby reducing the memory requirements and number of calculations required to recognize the active warm word in speech. Therefore, the low-power and low-fidelity characteristics of the speech recognizer 116 may be suitable for execution on a digital signal processor (DSP). In these embodiments, the speech recognizer 116 running on the AED 104 may recognize the speech of the warm word 112 when streaming audio captured by the AED 104, instead of using the warm word model 114. Here, the low-power speech recognizer 116 may be activated upon detecting a warm word 112, or may always run on-device in the background while long-term operation is performed by the AED 104. In these implementations, each user 102 explicitly grants the digital assistant 105 permission to perform speech recognition, and each user 102 has the option to revoke the granted permission at any time. In some examples, the detection of a warm word 112 by the corresponding warm word model 114 is confirmed by the speech recognizer 116, which performs speech recognition on the audio data.

[0040] 1B , while the digital assistant 105 is performing a long duration operation playing music 122, the user 102 speaks an utterance 146 that includes one of the warm words 112 from a set of activated warm words 112A. In the illustrated example, the user 102 speaks the active warm word 112A, "Turn up the volume." Without performing speech recognition on the captured audio, the AED 104 can apply the activated warm word model 114 to the set of warm words 112A to identify whether the utterance 146 includes any of the active warm words 112A. The active warm words 112 can be "next," "pause," "previous," "turn up the volume," and "turn down the volume." The warm word detector 210 compares the audio data corresponding to the utterance 146 with the activated warm word models 114 corresponding to the active warm words 112, "next," "pause," "previous," "turn up the volume," and "turn down the volume," and determines that the activated warm word model 114 for the warm word 112, "turn up the volume," detects the warm word 112, "turn up the volume," in the utterance 146, without performing speech recognition on the audio data. After the warm word 112, "turn up the volume," is detected in the additional audio 202 corresponding to the additional utterance 146, the digital assistant 105 running on the AED 104 begins performing the respective action associated with the detected warm word 112, "turn up the volume," which increases the volume of the music 122. In some implementations, in response to detecting the hot word 110 and the command 118, the AED 104 activates the repeating warm word 112R associated with the command 118 "play music" in addition to the warm word 112 associated with the command 118 "play music," thereby allowing the warm word detector 210 to detect the repeating warm word 112R without first detecting the associated warm word 112 in the audio.

[0041] In some implementations, the AED 104 identifies a warm word 112 that is not present in the set of one or more activated warm words 112 but whose model is nevertheless stored in the warm word model 114. In this case, the AED 104 may provide an indication to the user device 50 to display in the GUI 300 that the warm word 112 is not present in the set of one or more activated warm words 112 (e.g., deactivate the set). For example, the user 102 may say "play" when music 122 is playing. The AED 104 may identify the warm word 112N as "play." Because the warm word 112 "play" is not present in the set of one or more activated warm words 112A, the AED 104 does not perform any action. However, the user device 50 may display an indication in the GUI 300 that the warm word "play" is the inactive warm word 112N, and the active warm words 112A are "next," "pause," "previous," "volume up," and "volume down."

[0042] The warm word detector 210 may detect that the associated utterance 146 contains one of the warm words 112 from a set of one or more activated warm words 112 by extracting audio features of the audio data associated with the utterance 146. Each activated warm word model 114 may generate a corresponding warm word confidence score by processing the extracted audio features and comparing the corresponding warm word confidence score to a warm word confidence threshold. For example, the warm word models 114 may collectively generate a corresponding warm word confidence score for each of the active warm words 112: “play,” “next,” “pause,” “previous,” “volume up,” and “volume down.” In some implementations, the speech recognizer 116 generates a warm word confidence score for each portion of the processed audio data associated with the utterance 146. If the warm word confidence score meets the threshold, the warm word model 114 determines that the audio data corresponding to the utterance 146 includes the warm word 112 in the set of one or more activated warm words 112. For example, if the warm word confidence score generated by the warm word model 114 (or the speech recognizer 116) is 0.9 and the warm word confidence threshold is 0.8, the AED 104 determines that the audio data corresponding to the utterance 146 includes the warm word 112.

[0043] In some implementations, if the warm word confidence score is within a range below the threshold, the digital assistant 105 may generate an audible output synthesized voice 123 from the speaker 18 of the AED 104 requesting that the user 102 confirm or repeat the warm word 112. In these implementations, if the user 102 confirms that they spoke the warm word 112, the AED may use the audio data to update the corresponding warm word model 114.

[0044] In addition to performing the respective actions associated with the warm word 112 “turn up the volume” detected by the warm word detector 210, the warm word selector 220 may activate corresponding repeating warm words 112R each associated with the detected warm word 112 “turn up the volume” for a limited time after detecting the active warm word 112 “turn up the volume.” Here, each of the activated set of one or more repeating warm words 112R is associated with a respective action of turning up the volume of the music 112. In some implementations, in response to the warm word detector 210 detecting the warm word 112 “turn up the volume,” the warm word selector 220 activates a set of warm words 112A each associated with a respective action for controlling a long-term action different from the long-term action of playing music 122. For example, the different long-term action may control lighting in the user's 102's environment; if the volume of the music 122 is turned up (e.g., at a party), the user 102 may be more likely to adjust the lights up or down.

[0045] In some implementations, the warm word selector activates the repeating warm words 112R for a predetermined duration (e.g., 10 seconds). Here, once the predetermined duration has elapsed, the warm word selector 220 deactivates the repeating warm words 112R. For example, if the warm word detector 220 fails to detect the repeating warm words 112R during an utterance within the predetermined duration, the set of repeating warm words 112R is deactivated and will no longer be detected in the audio. Conversely, if the warm word detector 220 detects one of the repeating warm words 112R during an utterance within the predetermined duration, the warm word selector 210 may reset the predetermined duration and wait to see if the user 102 speaks one of the repeating warm words 112R again.

[0046] 1C, for the warm word 112 "Turn up the volume," the associated repeating warm words 112R available for activation by the warm word selector 220 include the repeating warm words 112R "Turn up," "Louder," "Further," and "Higher," each associated with a respective action for controlling the playback of music 122 from the speaker 18 of the AED 104. Thus, the warm word selector 220 activates these repeating warm words 112R (i.e., active warm words 112A) in addition to the other active warm words 112A "Next," "Pause," "Previous," "Turn up the volume," and "Turn down the volume" while the digital assistant 105 is performing a long-duration operation, and deactivates these repeating warm words 112R (e.g., inactive warm words 112N) when the long-duration operation ends.

[0047] As shown, while the digital assistant 105 continues to perform a long-duration operation playing music 122, the user 102 speaks an utterance 147 that includes one of the repeating warm words 112R from the set of activated repeating warm words 112R. In the illustrated example, the user 102 speaks the active repeating warm word 112R, "Turn it up." Without performing speech recognition on the captured audio, the AED 104 can apply the activated warm word model 114 to the set of repeating warm words 112R to identify whether the utterance 147 includes any of the active warm words 112A and / or any of the active repeating warm words 112R. The active repeating warm words 112R can be "turn it up," "louder," "further," and "higher." The warm word detector 210 compares the audio data corresponding to the utterance 147 with the activated warm word models 114 corresponding to the active repeating warm words 112R, "turn it up," "louder," "further," and "higher," and determines that the warm word model 114 activated for the repeating warm word 112R, "turn it up," detects the repeating warm word 112, "turn it up," in the utterance 146 without performing speech recognition on the audio data. As described above, each warm word model 114 activated for the repeating warm word 112R, "turn it up," can be derived from the warm word model 114 activated to detect the corresponding warm word 112, "turn it up" in the audio data. After detecting the repeating warm word 112R, "turn it up," in the additional audio 202 corresponding to the utterance 147, the digital assistant 105 executing on the AED 104 performs the respective action associated with the detected warm word 112, "turn it up," to increase the volume of the music 122.

[0048] As described above with respect to warm words 112, the warm word detector 210 may detect that the associated utterance 147 contains one of the repeated warm words 112R from a set of one or more activated repeated warm words 112R by extracting audio features of the audio data associated with the utterance 147. Each activated warm word model 114 may generate a corresponding repeated warm word confidence score by processing the extracted audio features and comparing the corresponding repeated warm word confidence score to a repeated warm word confidence threshold. For example, the warm word models 114 may collectively generate a corresponding repeated warm word confidence score for each of the active repeated warm words 112R, including "turn it up," "louder," "further," and "higher." In some implementations, the speech recognizer 116 generates a repeated warm word confidence score for each portion of the processed audio data associated with the utterance 147. If the repeated warm word confidence score meets the threshold, the warm word detector 210 determines that the audio data corresponding to the utterance 147 includes a repeated warm word 112R among the set of one or more activated repeated warm words 112R. For example, if the repeated warm word confidence score generated by the warm word detector 210 (or the speech recognizer 116) is 0.9 and the repeated warm word confidence threshold is 0.8, the AED 104 determines that the audio data corresponding to the utterance 147 includes a repeated warm word 112R.

[0049] In some implementations, if the repeated warm word confidence score is within a range below the threshold, the digital assistant 105 may generate an audible output synthesized voice 123 from the speaker 18 of the AED 104 requesting that the user 102 confirm or repeat the repeated warm word 112R. In these implementations, if the user 102 confirms that they have spoken the repeated warm word 112R, the AED 104 may use the audio data to update the corresponding warm word model 114.

[0050] In some embodiments, after performing each action associated with the detected repeated warm word 112R "turn it up," the warm word selector 220 maintains activation of at least one of the repeating warm words 112R in the set of repeating warm words 112R associated with the detected warm word 112 "turn up." For example, the warm word selector 220 maintains activation of the repeating warm word 112R "turn it up" so that the user 102 can continue to speak "turn it up...turn it up...turn it up." In some embodiments, the AED 104 may adjust parameters to increase the volume in response to detecting the repeating warm word 112R. For example, the AED 104 may lower a detection threshold to detect subsequent warm words 112 (and / or repeating warm words 112R). Additionally or alternatively, the AED 104 may adjust parameters to control long-term operation. For example, the AED 104 may detect that the user 102 always issues two sequences of the repeated warm word 112R "turn it up" and may adjust the volume up so that in future interactions with the user 102, the AED 104 adjusts the volume up more for each instance of the repeated warm word 112R "turn it up."

[0051] In some implementations, after the AED 104 performs each action of increasing the volume of the music 122 associated with the detected repeating warm word 112R, "turn it up," the warm word selector 220 deactivates the set of one or more repeating warm words 112R. For example, the warm word selector 220 deactivates the set of repeating warm words 112R based on receiving a command that conflicts with long-term operation. Here, the command can be a manual entry into the GUI 300 of the user device 50, a voice command, or other action that stops long-term operation. In another example, the warm word selector 220 deactivates the set of repeating warm words 112R based on failing to detect a repeating warm word 112R in the activated set of one or more repeating warm words 112R after a predetermined duration has elapsed since the user last spoke the initial warm word 112 or one of the repeating warm words 112R. In some examples, the predetermined duration is variable and depends on the detected repeated warm word 112R, the detected warm word 112, and / or is personalized to the particular user 102 who spoke the utterance.

[0052] Referring now to FIG. 3 , a GUI 300 executing on a user device 50 may render to display an identifier of the long action (e.g., “Play Track 1”), an identifier of the AED 104 (e.g., smart speaker) currently performing the long action, and / or the identity of the active user 102 (e.g., Verb) that initiated the long action. In some implementations, the identity of the active user 102 includes an image 304 of the active user 102. Here, the user 102 may speak any of the active warm words 112 displayed on the GUI 300 to perform a respective action to control the long action. The GUI 300 of FIG. 3 is shown executing on a user device 50 in communication with the AED 104; in other implementations, the GUI 300 executes on the screen of the AED 104.

[0053] The user device 50 may also render graphical elements 302 for display in the GUI 300 to perform the respective action associated with each active warm word 112A, such as to play music 122 from the speaker 18 of the AED 104. In other words, the GUI 300 allows the user 102 to select which warm words 112 should be added to the set of active warm words 112A (FIGS. 1A-1C). In the illustrated example, the GUI 300 presents a list of available warm words 112 displayed as respective graphical elements that, when spoken by the user 102, may be selected to add or remove the corresponding available warm word from the set of each active warm word 112A that the user desires the AED 104 to detect and ultimately perform the respective action specified by the corresponding warm word. As shown, the warm words 112 include the active warm words 112 "next," "pause," "previous," "turn up the volume," and "turn down the volume," as well as the inactive warm word 112N "play." Additionally, in response to the warm word detector 220 detecting the warm word 112 "turn up the volume," the GUI 300 presents a list of available repeating warm words 112R for highlighting / selecting a graphical element corresponding to "turn up the volume" to add to and / or remove from the set of one or more repeating warm words 112R associated with the warm word 112 "turn up the volume." Specifically, the list includes graphical elements 302 for the repeating warm words 112R indicated as "louder," "more," "up," "higher," "pause," "previous," "next," and "quieter."As described above, the activated repeating warm words 112R may correspond to the detected warm word 112, "turn up the volume," the same action of turning up the volume of the music 122, and / or a different action of turning down the volume of the music (e.g., the repeating warm word 112R, "quieter"). In the illustrated example, the repeating warm words 112R, "louder," "further," "turn up," "higher," and "quieter," are selected for detection in the audio and activated (via the warm word detector 220) for the AED 104. In some implementations, the list of repeating warm words 112R for each corresponding warm word 112 is predetermined by the application manufacturer. In other implementations, the list of repeating warm words 112R is added manually by the user 102 or learned by the warm word detector 210 based on previous interactions with the user 102.

[0054] In some implementations, the GUI 300 receives a user input indication indicating a selection of one or more of the available repeating warm words 112R in the displayed list of repeating warm words 112R to add / remove a set of one or more repeating warm words 112R. The GUI 300 may receive the user input indication via any one of touch, voice, gesture, gaze, and / or an input device (e.g., a mouse or stylist). For example, the user 102 may provide a user input indication indicating a selection of a “pause” control (e.g., by touching a graphical button in the GUI 300 for the repeating warm word 112R “pause”) to cause the AED 104 (via the warm word selector 210) to add the repeating warm word 112R “pause” to the active warm words 112A for detection by the AED 104.

[0055] As shown, the GUI 300 displays a graphical element 306 for the user 102 to customize the list of warm words 112. Here, the warm word detector 210 (or the speech recognizer 116) may maintain a voice command log 212 (FIG. 2) containing previous instances in which the user 102 spoke a repeat voice command to control long-term operation immediately after the warm word detector 210 detected a warm word 112A in the previous audio data 202. In these examples, the repeat voice command may indicate that the user 102 associates the repeat voice command with the detected warm word 112A. Here, at least one of the repeating warm words 112R in the set of activated warm words 112A is learned based on the repeat voice command. In these examples, the repeat voice command may be used to train the warm word model 114 for each of the learned repeating warm words 112R.

[0056] As shown, the graphical element 306 includes selectable options for the user 102 to add the word(s) "skip" and / or "more and more" to the list of repeating warm words 112R and / or to the active warm words 112A. Here, the AED 104 may have learned the repeating voice command "more and more," and the graphical element 306 operates to present a prompt to prompt the user 102 to add the repeating voice command "more and more" to the list of repeating warm words 112R. Additionally, the graphical element 306 includes a text box for the user 102 to manually enter, by typing or speaking, the word to add to the list of repeating warm words 112R.

[0057] 4 is a flowchart of an exemplary arrangement of operations of a method 400 for detecting repeated warm words. At operation 402, the method 400 includes activating a set of one or more warm words 112, each associated with a respective action for controlling a long-term operation performed by the digital assistant 105. While the digital assistant 105 is performing the long-term operation and the set of one or more warm words 112 is activated, the method 400 also includes, at operation 404, receiving audio data 202 corresponding to a first utterance 146 captured by the assistant-enabled device 104. At operation 406, the method 400 also includes detecting a warm word 112 from the set of one or more activated warm words 112 in the audio data 202.

[0058] In response to detecting the warm word 112 from the set of one or more activated warm words 112, the method 400 also includes, at operation 408, performing a respective action associated with the detected warm word 112 to control a long-term operation, and activating a set of one or more repeating warm words 113 associated with the detected warm word 112. Here, each of the activated sets of one or more repeating warm words 113 is associated with a respective action for controlling a long-term operation. At operation 510, the method 400 includes receiving additional audio data 202 corresponding to a second utterance 147 captured by the assistant-enabled device 104. At operation 412, the method 400 also includes detecting a repeating warm word 113 from the set of one or more activated repeating warm words 113 in the additional audio data 202. In response to detecting a recurring warm word 113 from the set of one or more recurring warm words 113, the method 400 also includes, at operation 412, performing a respective action associated with the detected recurring warm word 113 to control long-term behavior.

[0059] 5 is a schematic diagram of an exemplary computing device 500 that can be used to implement the systems and methods described herein. Computing device 500 is intended to represent various types of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown here, their connections and relationships, and their functionality are for illustrative purposes only and are not intended to limit the scope of the present invention as described and / or claimed herein.

[0060] Computing device 500 includes a processor 510, memory 520, a storage device 530, a high-speed interface / controller 540 connecting to memory 520 and a high-speed expansion port 550, and a low-speed interface / controller 560 connecting to a low-speed bus 570 and storage device 530. Each of the components 510, 520, 530, 540, 550, and 560 are interconnected using various buses and may be mounted on a common motherboard or otherwise mounted as desired. Processor 510 processes instructions for execution within computing device 500, including instructions stored in memory 520 or storage device 530, and can display graphical information for a graphical user interface (GUI) on an external input / output device, such as a display 580 connected to high-speed interface 540. In other implementations, multiple processors and / or multiple buses may be used as needed, along with multiple memories and memory types. Also, multiple computing devices 500 may be connected, each providing a portion of the required operations (eg, as a bank of servers, a group of blade servers, or a multi-processor system).

[0061] The memory 520 stores information non-transiently within the computing device 500. The memory 520 may be a computer-readable medium, a volatile memory unit(s), or a non-volatile memory unit(s). The non-transient memory 520 may be a physical device used to temporarily or permanently store programs (e.g., instruction sequences) or data (e.g., program state information) for use by the computing device 500. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.

[0062] Storage device 530 can provide mass storage for computing device 500. In some embodiments, storage device 530 is a computer-readable medium. In various different implementations, storage device 530 can be a device array including a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or a device in a storage area network or other configuration. In further embodiments, a computer program product is tangibly embodied on an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as memory 520, storage device 530, or memory on processor 510.

[0063] High-speed controller 540 manages bandwidth-intensive operations for computing device 500, and low-speed controller 560 manages low-bandwidth-intensive operations. This allocation of roles is merely exemplary. In some implementations, high-speed controller 540 is coupled to memory 520 (e.g., via a graphics processor or accelerator), to display 580, and to high-speed expansion port 550, which can accept various expansion cards (not shown). In some implementations, low-speed controller 560 is coupled to storage device 530 and low-speed expansion port 590. Low-speed expansion port 590, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), may be connected, for example, via a network adapter, to one or more input / output devices, such as a keyboard, pointing device, scanner, or network devices, such as a switch or router.

[0064] The computing device 500, as shown, can be implemented in many different forms. For example, it may be implemented as a standard server 500a, or multiple times within a group of such servers 500a, as a laptop computer 500b, or as part of a rack server system 500c.

[0065] Various implementations of the systems and techniques described herein may be realized in digital electronic and / or optical circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementation in one or more computer programs executable and / or interpretable by a programmable system including at least one programmable processor, which may be special purpose or general purpose, coupled to receive data and instructions from and transmit data and instructions to a storage system, at least one input device, and at least one output device.

[0066] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform a task. In some embodiments, a software application may be referred to as an "application," "app," or "program." Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.

[0067] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in a high-level procedural programming language and / or an object-oriented programming language and / or an assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0068] The processes and logic flows described herein may be implemented by one or more programmable processors, also referred to as data processing hardware, executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be implemented by special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). Processors suitable for executing computer programs include, for example, both general-purpose and special-purpose microprocessors and any one or more processors of any kind of digital computer. Generally, a processor receives instructions and data from a read-only memory, a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or is operably connected to receive data from or transmit data to them, or both. A computer need not, however, have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0069] To interact with a user, one or more aspects of the present disclosure can be implemented in a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, or a touch screen, for displaying information to the user, and optionally a keyboard and pointing device, such as a mouse or trackball, through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user. For example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user can be received in any form, such as acoustic input, voice input, or tactile input. Furthermore, a computer can interact with a user by sending documents to and receiving documents from a device used by the user, for example, by sending a web page to a web browser on the user's client device in response to a request received from the web browser.

[0070] Although several embodiments have been described, it will be understood that various modifications can be made without departing from the spirit and scope of the disclosure. Accordingly, other embodiments are within the scope of the following claims.

Claims

1. A computer-implemented method (400) that, when executed by data processing hardware (10), causes the data processing hardware (10) to perform operations, the operations comprising: activating a set of one or more warm words (112) each associated with a respective action for controlling a long-term operation performed by the digital assistant (105); While the digital assistant (105) is performing the long-term operation and the set of one or more warm words (112) is activated, Receiving audio data (202) corresponding to a first utterance (146) captured by an assistant-enabled device (104); detecting a warm word (112) from the set of one or more activated warm words (112) within the audio data (202); In response to detecting the warm word (112) from the set of one or more activated warm words (112), performing the respective actions associated with the detected warm words (112) to control the long-term operation; and activating a set of one or more repeating warm words (112) associated with the detected warm words (112), each of the activated sets of one or more repeating warm words (112) being associated with a respective action for controlling the long-term operation, the operation further comprising: While the digital assistant (105) is performing the long-term operation and the set of one or more warm words (112) is activated, Receiving additional audio data (202) corresponding to a second utterance (147) captured by the assistant-enabled device (104); and detecting a repeating warm word (112) from the set of one or more activated repeating warm words (112) within the additional audio data (202); and in response to detecting the repetitive warm word (112) from the set of one or more repetitive warm words (112), performing the respective action associated with the detected repetitive warm word (112) to control the long-term operation.

2. activating the set of one or more repeating warm words (112) includes, for each corresponding repeating warm word (112) in the set of one or more activated repeating warm words (112), activating a respective repeating warm word model (114) for execution on the assistant-enabled device (104); 2. The method of claim 1, wherein detecting the repeated warm words (112) from the set of one or more activated repeated warm words (112) in the additional audio data (202) comprises detecting the repeated warm words (112) in the additional audio data (202) using the respective activated repeated warm word models (114) for the corresponding detected repeated warm words (112) without performing speech recognition on the additional audio data (202).

3. Detecting the repeating warm word (112) in the additional audio data (202) comprises: extracting audio features from the additional audio data (202); generating a repeated warm word confidence score by processing the extracted audio features using the respective repeated warm word models (114) activated for the corresponding detected repeated warm words (112); and determining that the additional audio data (202) corresponding to the second utterance (147) includes the corresponding repeated warm word (112) if the repeated warm word confidence score meets a repeated warm word confidence threshold.

4. detecting the warm word (112) from the set of one or more activated warm words (112) in the audio data (202) includes detecting the warm word (112) in the audio data (202) using each activated warm model for the corresponding warm word (112) without performing speech recognition on the audio data (202); 3. The method of claim 2, wherein the respective repeated warm word models activated for the corresponding detected repeated warm words in the additional audio data are derived from the respective warm word models activated to detect the corresponding warm words in the audio data.

5. Activating the set of one or more repeating warm words (112) includes executing a speech recognizer (116) on the assistant-enabled device (104), wherein the speech recognizer (116) is biased to recognize the repeating warm words (112) within the activated set of one or more repeating warm words (112); 5. The method of claim 1, wherein detecting the repeating warm word (112) from the set of one or more activated repeating warm words (112) in the additional audio data (202) comprises recognizing the repeating warm word (112) in the additional audio data (202) using the speech recognizer (116) running on the assistant-enabled device (104).

6. The operations further include receiving a user input indication of one or more of the available repeating warm words (112) to be added to the set of one or more repeating warm words (112) associated with the warm word (112) and / or to be removed from the set of one or more repeating warm words (112) associated with the warm word (112), and to be activated when the warm word (112) is detected in the audio data (202). The method (400) of any one of claims 1 to 5 includes receiving a user input indication of one or more of the available repeating warm words (112) to be added to the set of one or more repeating warm words (112).

7. 7. The method of claim 6, wherein each corresponding available repeating warm word in the list of available repeating warm words presented by the user interface associated with the assistant-enabled device is displayed by the user interface as a respective graphical element that can be selected to add or remove the corresponding available repeating warm word from the set of one or more warm words associated with the warm word, and that is activated when the warm word is detected in the audio data.

8. The operation further comprises: Immediately after detecting a warm word (112) in the previous audio data (202), obtaining a voice command log (212) including previous instances of a user of the assistant-enabled device (104) repeatedly speaking voice commands to control the long-term operation; 8. The method (400) of claim 1, wherein at least one recurring warm word (112) in the set of one or more activated warm words (112) is learned based on the recurring voice command.

9. The operation further comprises: Prompting a user of the assistant-enabled device to add a repeating voice command included in the voice command log to a set of one or more repeating warm words associated with the detected warm word; associating the repeating voice command (118) with the detected warm word (112) in response to receiving a user indication of approval to add the repeating voice command (118) to the set of one or more repeating warm words (112); 10. The method of claim 8, further comprising: activating the repeat voice command as a repeat warm word associated with the detected warm word.

10. 10. The method (400) of claim 1, wherein the respective action for controlling the long-term operation associated with at least one recurring warm word (112) in the set of one or more activated recurring warm words (112) is the same as the respective action performed in response to detecting the associated warm word (112) from the set of one or more activated warm words (112) in the audio data (202).

11. 11. The method (400) of claim 1, wherein the respective action for controlling the long-term behavior associated with at least one recurring warm word (112) in the set of one or more activated recurring warm words (112) is different from the respective action performed in response to detecting the associated warm word (112) from the set of one or more activated warm words (112) in the audio data (202).

12. 12. The method (400) of claim 1, wherein activating the set of one or more repeating warm words (112) associated with the detected warm word (112) comprises activating the set of one or more repeating warm words (112) for a predetermined duration.

13. The operation further comprises:

13. The method (400) of claim 1, comprising maintaining activation of at least one of the recurring warm words (112) in the set of one or more recurring warm words (112) associated with the detected recurring warm word (112) after performing the respective action associated with the detected recurring warm word (112) to control the long-term behavior.

14. The operation further comprises, after performing the respective actions associated with the detected recurring warm words (112) to control the long-term operation: receiving a command (118) that conflicts with the long-term operation; or failing to detect additional repeating warm words (112) in the activated set of one or more repeating warm words (112) after a predetermined duration has elapsed; The method (400) of any one of claims 1 to 13, comprising deactivating the set of one or more recurring warm words (112) based on:

15. 15. The method of claim 1, wherein in response to detecting the warm word from the set of one or more activated warm words, the action further comprises activating a set of one or more warm words each associated with a respective action for controlling a long-term operation different from the long-term operation.

16. A system (100), comprising: data processing hardware (10); and memory hardware (12) in communication with the data processing hardware (10), The memory hardware (12) stores instructions that, when executed on the data processing hardware (10), cause the data processing hardware (10) to perform operations, the operations being: activating a set of one or more warm words (112) each associated with a respective action for controlling a long-term operation performed by the digital assistant (105); While the digital assistant (105) is performing the long-term operation and the set of one or more warm words (112) is activated, Receiving audio data (202) corresponding to a first utterance (146) captured by an assistant-enabled device (104); detecting a warm word (112) from the set of one or more activated warm words (112) within the audio data (202); In response to detecting the warm word (112) from the set of one or more activated warm words (112), performing the respective actions associated with the detected warm words (112) to control the long-term operation; and activating a set of one or more repeating warm words (112) associated with the detected warm words (112), each of the activated sets of one or more repeating warm words (112) being associated with a respective action for controlling the long-term operation, the operation further comprising: While the digital assistant (105) is performing the long-term operation and the set of one or more warm words (112) is activated, Receiving additional audio data (202) corresponding to a second utterance (147) captured by the assistant-enabled device (104); and detecting a repeating warm word (112) from the set of one or more activated repeating warm words (112) within the additional audio data (202); and in response to detecting the repetitive warm word (112) from the set of one or more repetitive warm words (112), executing the respective action associated with the detected repetitive warm word (112) to control the long-term operation.

17. activating the set of one or more repeating warm words (112) includes, for each corresponding repeating warm word (112) in the set of one or more activated repeating warm words (112), activating a respective repeating warm word model (114) for execution on the assistant-enabled device (104); 17. The system of claim 16, wherein detecting the repeated warm word (112) from the set of one or more activated repeated warm words (112) in the additional audio data (202) comprises detecting the repeated warm word (112) in the additional audio data (202) using the respective activated repeated warm word model (114) for the corresponding detected repeated warm word (112) without performing speech recognition on the additional audio data (202).

18. Detecting the repeating warm word (112) in the additional audio data (202) comprises: extracting audio features from the additional audio data (202); generating a repeated warm word confidence score by processing the extracted audio features using the respective repeated warm word models (114) activated for the corresponding detected repeated warm words (112); and determining that the additional audio data (202) corresponding to the second utterance (147) includes the corresponding repeated warm word (112) if the repeated warm word confidence score meets a repeated warm word confidence threshold.

19. detecting the warm word (112) from the set of one or more activated warm words (112) in the audio data (202) includes detecting the warm word (112) in the audio data (202) using each activated warm model for the corresponding warm word (112) without performing speech recognition on the audio data (202); 20. The system of claim 17, wherein the respective repeated warm word models activated for the corresponding detected repeated warm words in the additional audio data are derived from the respective warm word models activated to detect the corresponding warm words in the audio data.

20. Activating the set of one or more repeating warm words (112) includes executing a speech recognizer (116) on the assistant-enabled device (104), wherein the speech recognizer (116) is biased to recognize the repeating warm words (112) within the activated set of one or more repeating warm words (112); 20. The system (100) of any one of claims 16 to 19, wherein detecting the repeated warm word (112) from the set of one or more activated repeated warm words (112) in the additional audio data (202) includes recognizing the repeated warm word (112) in the additional audio data (202) using the speech recognizer (116) running on the assistant-enabled device (104).

21. The operations further include receiving a user input indication of one or more of the available repeating warm words (112) to be added to the set of one or more repeating warm words (112) associated with the warm word (112) and / or to be removed from the set of one or more repeating warm words (112) associated with the warm word (112), and to be activated when the warm word (112) is detected in the audio data (202). The system (100) of any one of claims 16 to 20 includes receiving a user input indication of one or more of the available repeating warm words (112) to be added to the set of one or more repeating warm words (112).

22. 22. The system (100) of claim 21, wherein each corresponding available repeating warm word (112) in the list of available repeating warm words (112) presented by the user interface (300) associated with the assistant-enabled device (104) is displayed by the user interface (300) as a respective graphical element (306) that can be selected to add or remove the corresponding available repeating warm word (112) from the set of one or more warm words (112) associated with the warm word (112), and to be activated when the warm word (112) is detected in the audio data (202).

23. The operation further comprises: Immediately after detecting a warm word (112) in the previous audio data (202), obtaining a voice command log (212) including previous instances of a user of the assistant-enabled device (104) repeatedly speaking voice commands to control the long-term operation; 23. The system (100) of claim 16, wherein at least one recurring warm word (112) in the set of one or more activated warm words (112) is learned based on the recurring voice command.

24. The operation further comprises: Prompting a user of the assistant-enabled device to add a repeating voice command included in the voice command log to a set of one or more repeating warm words associated with the detected warm word; associating the repeating voice command (118) with the detected warm word (112) in response to receiving a user indication of approval to add the repeating voice command (118) to the set of one or more repeating warm words (112); activating the repeat voice command as an additional repeat warm word associated with the detected warm word.

25. 25. The system (100) of claim 16, wherein the respective action for controlling the long-term operation associated with at least one recurring warm word (112) in the set of one or more activated recurring warm words (112) is the same as the respective action performed in response to detecting the associated warm word (112) from the set of one or more activated warm words (112) in the audio data (202).

26. 26. The system (100) of claim 16, wherein the respective action for controlling the long-term operation associated with at least one recurring warm word (112) in the set of one or more activated recurring warm words (112) is different from the respective action performed in response to detecting the associated warm word (112) from the set of one or more activated warm words (112) in the audio data (202).

27. 27. The system (100) of claim 16, wherein activating the set of one or more repeating warm words (112) associated with the detected warm word (112) comprises activating the set of one or more repeating warm words (112) for a predetermined duration.

28. The operation further comprises:

28. The system (100) of claim 16, further comprising: maintaining activation of at least one of the recurring warm words (112) in the set of one or more recurring warm words (112) associated with the detected warm word (112) after performing the respective action associated with the detected recurring warm word (112) to control the long-term operation.

29. The operation further comprises, after performing the respective actions associated with the detected recurring warm words (112) to control the long-term operation: receiving a command (118) that conflicts with the long-term operation; or failing to detect additional repeating warm words (112) of said activated set of one or more repeating warm words (112) after a predetermined duration has elapsed; The system (100) of any one of claims 16 to 28, comprising deactivating the set of one or more recurring warm words (112) based on:

30. 30. The system (100) of any one of claims 16 to 29, wherein in response to detecting the warm word (112) from the set of one or more activated warm words (112), the operation further comprises activating a set of one or more warm words (112) each associated with a respective action for controlling a long-term operation different from the long-term operation.

Citation Information

Patent Citations

  • Contextual Hot Words

    JP2020503568A

  • Speaker dependent follow up actions and warm words

    WO2022125279A1