Multi-Assistant Warm Words

The MAD's warm word arbitration routine optimizes the selection of warm words for multiple digital assistants, addressing resource inefficiencies and false positives, ensuring efficient control of long-term operations.

JP2025537020APending Publication Date: 2025-11-12GOOGLE LLC
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2025527757
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-14
Filing Date
2023-10-31
Publication Date
2025-11-12

AI Technical Summary

Technical Problem

In voice-enabled environments with multiple digital assistants, the challenge lies in managing overlapping warm words that lead to increased false positives and resource redundancy, as well as inefficient use of computational resources due to simultaneous activation of multiple assistants.

Method used

A multi-assistant device (MAD) executes a warm word arbitration routine to select a final set of warm words for detection, considering active sets associated with each digital assistant, resource availability, and ambient context, enabling efficient detection and control of long-term operations.

Benefits of technology

This approach reduces false positives and conserves resources by enabling only a single instance of duplicate warm words, allowing seamless control of long-term operations with reduced power consumption and improved efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025537020000001_ABST
    Figure 2025537020000001_ABST
Patent Text Reader

Abstract

A method (300) for using multi-assistant warm words includes receiving, for each digital assistant (105) in a group of digital assistants enabled on a multi-assistant device (104), a respective active set (111) of warm words (110), each of which specifies a respective action to be performed. The method also includes executing a warm word arbitration routine to enable a final set (113) of warm words for detection based on the respective active sets of warm words, wherein each warm word in the final set of warm words is selected from the respective active sets of warm words for at least one digital assistant. While the final set of warm words is enabled for detection, the method includes receiving audio data (402) corresponding to the utterance, detecting warm words from the final set of warm words, and instructing the digital assistant associated with the detected warm word to perform the respective action specified by the detected warm word.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS The present disclosure relates to multi-assistant warm words. [Background technology]

[0002] In a voice-enabled environment, a user simply speaks a query or command out loud, and the digital assistant will address and respond to the query and / or cause the command to be carried out. Voice-enabled environments (e.g., home, work, school, etc.) can be implemented using a network of connected microphone devices distributed throughout various rooms and / or environmental areas. Through such a network of microphones, a user has the ability to verbally query a digital assistant from essentially anywhere in the environment, without having to have a computer or other device in front of or even nearby. For example, while cooking in the kitchen, a user may ask the digital assistant to "set the timer for 20 minutes," and in response, the digital assistant confirms that the timer has been set (e.g., in the form of a synthesized voice output) and then alerts the user (e.g., in the form of an alarm or other audible alert from an audio speaker) when the timer reaches 20 minutes. Often, there are multiple digital assistants activated simultaneously on a given device that a user in a given environment can query / command to perform various actions. Each of these digital assistants may have its own set of phrases (i.e., warm words) that can be detected in speech without full speech recognition and may overlap with one or more phrases of other digital assistants activated on the device. For example, a user of multiple digital assistants may speak the command "play music," which may correspond to both a music digital assistant and a browser digital assistant. In response, the appropriate digital assistant may stream a music playlist to the user through an audio speaker. Summary of the Invention

[0003] One aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform an operation including receiving, for each digital assistant in a group of digital assistants enabled for simultaneous execution on a multi-assistant device (MAD), a respective active set of warm words, each specifying a respective action to be performed by each digital assistant. Based on the respective active sets of warm words associated with each digital assistant in the group of digital assistants, the method also includes executing, by a multi-assistant interface running on the MAD, a warm word arbitration routine to enable a final set of warm words for detection by the MAD. Each corresponding warm word in the final set of warm words enabled for detection by the MAD is selected from the respective active sets of warm words for at least one digital assistant in the group of digital assistants. While the final set of warm words is enabled for detection by the MAD, the method further includes receiving audio data corresponding to the utterance captured by the MAD, detecting warm words from the final set of warm words in the audio data, and instructing a digital assistant from the group of digital assistants associated with the detected warm words to perform a respective action specified by the detected warm word.

[0004] Embodiments of the present disclosure may include one or more of the following optional features. In some embodiments, receiving a respective active set of warm words associated with the digital assistant includes, for at least one warm word in the respective active set of warm words, receiving a respective warm word model configured to detect the corresponding warm word in streaming audio without performing voice recognition via a warm word application programming interface (API) running on the MAD. In some examples, for a corresponding digital assistant of a digital assistant in the group of digital assistants, the operation further includes receiving a user command specifying a long-term action to be performed by the corresponding digital assistant, and performing the long-term action specified by the voice command via the corresponding digital assistant. Here, receiving a respective active set of warm words associated with the digital assistant includes receiving a respective active set of warm words associated with the corresponding digital assistant in response to the corresponding digital assistant performing the long-term action. In these examples, each warm word in the respective active set of warm words may be associated with a respective action for controlling the long-term action performed by the corresponding digital assistant.

[0005] In some embodiments, the operations further include discovering a new digital assistant within the group of digital assistants enabled for simultaneous execution on the MAD, and the multi-assistant interface performs a warm word arbitration routine in response to discovering the new digital assistant within the group of digital assistants. In some examples, the operations further include determining that the digital assistant has been removed from the group of digital assistants enabled for simultaneous execution on the MAD. Here, the multi-assistant interface performs a warm word arbitration routine in response to determining that the digital assistant has been removed from the group of digital assistants. In some embodiments, for a corresponding digital assistant of a digital assistant within the group of digital assistants, the operations further include determining whether to add or remove a warm word from the respective active sets of warm words associated with the corresponding digital assistant. In these embodiments, the multi-assistant interface performs a warm word arbitration routine in response to determining whether to add or remove a warm word from the respective sets of warm words.

[0006] In some examples, the operations further include determining a change in the ambient context, and the multi-assistant interface executes a warm word arbitration routine in response to determining the change in the ambient context. In some embodiments, the operations further include obtaining effective warm word constraints. In these embodiments, the effective warm word constraints include at least one of the following: the availability of memory and computing resources on the MAD for detecting warm words, the computational requirements for enabling each warm word in each active set of warm words associated with each digital assistant in the group of digital assistants, an acceptable false acceptance rate tolerance, or an acceptable false rejection rate tolerance. Here, the number of warm words in the final set of warm words enabled for detection by the MAD is based on the obtained active warm word constraints. In some examples, executing the warm word arbitration routine includes identifying any shared warm words corresponding to warm words present in at least two of the active sets of warm words, and determining the final set of warm words is based on assigning a higher priority to warm words identified as shared warm words.

[0007] In some embodiments, executing the warm word arbitration routine includes determining the frequency of detection of the warm word by MAD for each warm word in each active set of warm words for each digital assistant in the group of digital assistants. In these embodiments, determining the final set of warm words is based on the determined frequency of detection of each warm word in each active set of warm words for each digital assistant in the group of digital assistants. In some examples, executing the warm word arbitration routine includes determining the time when the warm word was most recently detected by MAD for each warm word in each active set of warm words for each digital assistant in the group of digital assistants. Here, determining the final set of warm words is based on the determined time when each warm word in each active set of warm words for each digital assistant in the group of digital assistants was most recently detected.

[0008] In some embodiments, the operations further include receiving a voice command instructing the MAD to enable the first digital assistant and the second digital assistant to run simultaneously on the MAD. Here, the voice command is spoken by a user of the MAD and captured by the MAD in the streaming audio. In these embodiments, after receiving the voice command, the operations further include enabling the first digital assistant and the second digital assistant to run simultaneously on the MAD, wherein the group of digital assistants includes the first digital assistant and the second digital assistant. In some examples, the operations further include receiving a multi-assistant configuration request from a software application running on the MAD or on another device that communicates with the MAD to enable the first digital assistant and the second digital assistant to run simultaneously on the MAD. In these examples, after receiving the multi-assistant configuration request, the operations also include enabling the first digital assistant and the second digital assistant to run simultaneously on the MAD, wherein the group of digital assistants includes the first digital assistant and the second digital assistant.

[0009] Another aspect of the present disclosure provides a system including data processing hardware and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform an operation, including receiving, for each digital assistant in a group of digital assistants enabled for simultaneous execution on a multi-assistant device (MAD), a respective active set of warm words, each specifying a respective action to be performed by each digital assistant. Based on the respective active sets of warm words associated with each digital assistant in the group of digital assistants, the method also includes executing, by a multi-assistant interface running on the MAD, a warm word arbitration routine to enable a final set of warm words for detection by the MAD. Each corresponding warm word in the final set of warm words enabled for detection by the MAD is selected from the respective active sets of warm words for at least one digital assistant in the group of digital assistants. While the final set of warm words is enabled for detection by the MAD, the method further includes receiving audio data corresponding to the utterance captured by the MAD, detecting warm words from the final set of warm words in the audio data, and instructing a digital assistant from the group of digital assistants associated with the detected warm words to perform a respective action specified by the detected warm word.

[0010] This aspect may include one or more of the following optional features. In some implementations, receiving a respective active set of warm words associated with the digital assistant includes, for at least one warm word in the respective active set of warm words, receiving a respective warm word model configured to detect the corresponding warm word in streaming audio without performing voice recognition via a warm word application programming interface (API) running on the MAD. In some examples, for a corresponding digital assistant of a digital assistant in the group of digital assistants, the operation further includes receiving a user command specifying a long-term action to be performed by the corresponding digital assistant, and performing the long-term action specified by the voice command via the corresponding digital assistant. Here, receiving a respective active set of warm words associated with the digital assistant includes receiving a respective active set of warm words associated with the corresponding digital assistant in response to the corresponding digital assistant performing the long-term action. In these examples, each warm word in the respective active set of warm words may be associated with a respective action for controlling the long-term action performed by the corresponding digital assistant.

[0011] In some embodiments, the operations further include discovering a new digital assistant within the group of digital assistants enabled for simultaneous execution on the MAD, and the multi-assistant interface performs a warm word arbitration routine in response to discovering the new digital assistant within the group of digital assistants. In some examples, the operations further include determining that the digital assistant has been removed from the group of digital assistants enabled for simultaneous execution on the MAD. Here, the multi-assistant interface performs a warm word arbitration routine in response to determining that the digital assistant has been removed from the group of digital assistants. In some embodiments, for a corresponding digital assistant of a digital assistant within the group of digital assistants, the operations further include determining whether to add or remove a warm word from the respective active sets of warm words associated with the corresponding digital assistant. In these embodiments, the multi-assistant interface performs a warm word arbitration routine in response to determining whether to add or remove a warm word from the respective sets of warm words.

[0012] In some examples, the operations further include determining a change in the ambient context, and the multi-assistant interface executes a warm word arbitration routine in response to determining the change in the ambient context. In some embodiments, the operations further include obtaining effective warm word constraints. In these embodiments, the effective warm word constraints include at least one of the following: the availability of memory and computing resources on the MAD for detecting warm words, the computational requirements for enabling each warm word in each active set of warm words associated with each digital assistant in the group of digital assistants, an acceptable false acceptance rate tolerance, or an acceptable false rejection rate tolerance. Here, the number of warm words in the final set of warm words enabled for detection by the MAD is based on the obtained active warm word constraints. In some examples, executing the warm word arbitration routine includes identifying any shared warm words corresponding to warm words present in at least two of the active sets of warm words, and determining the final set of warm words is based on assigning a higher priority to warm words identified as shared warm words.

[0013] In some embodiments, executing the warm word arbitration routine includes determining the frequency of detection of the warm word by MAD for each warm word in each active set of warm words for each digital assistant in the group of digital assistants. In these embodiments, determining the final set of warm words is based on the determined frequency of detection of each warm word in each active set of warm words for each digital assistant in the group of digital assistants. In some examples, executing the warm word arbitration routine includes determining the time when the warm word was most recently detected by MAD for each warm word in each active set of warm words for each digital assistant in the group of digital assistants. Here, determining the final set of warm words is based on the determined time when each warm word in each active set of warm words for each digital assistant in the group of digital assistants was most recently detected.

[0014] In some embodiments, the operations further include receiving a voice command instructing the MAD to enable the first digital assistant and the second digital assistant to run simultaneously on the MAD. Here, the voice command is spoken by a user of the MAD and captured by the MAD in the streaming audio. In these embodiments, after receiving the voice command, the operations further include enabling the first digital assistant and the second digital assistant to run simultaneously on the MAD, wherein the group of digital assistants includes the first digital assistant and the second digital assistant. In some examples, the operations further include receiving a multi-assistant configuration request from a software application running on the MAD or on another device that communicates with the MAD to enable the first digital assistant and the second digital assistant to run simultaneously on the MAD. In these examples, after receiving the multi-assistant configuration request, the operations also include enabling the first digital assistant and the second digital assistant to run simultaneously on the MAD, wherein the group of digital assistants includes the first digital assistant and the second digital assistant.

[0015] The details of one or more embodiments of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will become apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]

[0016] [Figure 1A] FIG. 1 is a schematic diagram of an exemplary system including a user using multiple assistant warm words to control long-term actions. [Figure 1B] FIG. 1 is a schematic diagram of an exemplary system including a user using multiple assistant warm words to control long-term actions. [Figure 1C] FIG. 1 is a schematic diagram of an exemplary system including a user using multiple assistant warm words to control long-term actions. [Figure 2] FIG. 1 is a schematic diagram of a multi-assistant warm word detection process. [Figure 3] 1 is a flowchart of an exemplary arrangement of operations for a method of using multiple assistant warm words. [Figure 4] FIG. 1 is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION

[0017] Like reference symbols in the various drawings refer to like elements.

[0018] A user's interaction with an Assistant-enabled device is designed to be primarily, but not exclusively, through voice input. As a result, the Assistant-enabled device needs some way to identify when any given utterance in the surrounding environment is directed at the device, as opposed to being directed at an individual in the environment or originating from a non-human source (e.g., a television or music player). One way to achieve this is through the use of hot words, reserved by agreement among users in the environment as predetermined words or words to be spoken to attract the device's attention. In the exemplary environment, the hot word used to attract the Assistant's attention is the words "OK computer." As a result, whenever the words "OK computer" are spoken, they are picked up by the microphone and transmitted to a hot word detector, which performs speech understanding techniques to determine whether the hot word was spoken and, if so, waits for a subsequent command or query. Thus, utterances directed to an Assistant-enabled device take the general form [hot word][query], where the "hot word" in this example is "OK computer" and the "query" can be any question, command, declaration, or other request that can be voice-recognized, analyzed, and acted upon by the system, either alone or in cooperation with a server over a network.

[0019] When a user gives several hotword-based commands to an Assistant-enabled device, such as a mobile phone or smart speaker, the user's interaction with the phone or speaker can be awkward. The user may say, "OK computer, play the assignments playlist." The phone or speaker may start playing the first song in the playlist. The user may want to move on to the next song and say, "OK computer, next." To move on to another song, the user may say, "OK computer, next" again. To alleviate the need to keep repeating hotwords before speaking a command, the Assistant-enabled device may be configured to recognize / detect a narrow set of hot phrases or warm words that directly trigger the respective action. In this example, the warm word "next" serves the dual purpose of hotword and command, such that instead of saying "OK computer, next," the user can simply say "next" to invoke the Assistant-enabled device and trigger it to perform the respective action.

[0020] An assistant-enabled device may include multiple digital assistants, each with a set of active warm words to control long-term operations. As used herein, long-term operations refer to applications or events that the digital assistant runs for an extended period of time and that can be controlled by the user while the application or event is ongoing. For example, if a digital assistant for a clock application sets a timer for 30 minutes, the timer is a long-term operation from the time the timer is set until the timer ends, or until the resulting alert is acknowledged after the timer ends. In this case, a warm word such as "stop" may be active so that the user can stop the timer by simply saying "stop" without first speaking a hot word. Similarly, a command to instruct a digital assistant for a music application to play music from a streaming music service is a long-term operation while the digital assistant is streaming music from the streaming music service through a playback device. In this case, the active set of warm words to control the playback of music the digital assistant is streaming through a playback device may be "stop," "pause," "volume up," "volume down," "next," "previous," etc. Long-term actions may include multi-step dialogue queries such as "make a restaurant reservation," with different sets of warm words active depending on the given stage of the multi-step dialogue. For example, an assistant-enabled device may prompt a user to select from a list of restaurants, which may activate a set of warm words, each including a respective identifier (e.g., restaurant name or number in the list), to select a restaurant from the list and complete the action of making a reservation for that restaurant.

[0021] One challenge with devices that enable multiple digital assistants is limiting the number of words / phrases enabled simultaneously so that quality and efficiency do not degrade. For example, the number of false positives, which refers to when an assistant-enabled device incorrectly detects / recognizes one of the active words, increases significantly as the number of warm words enabled simultaneously increases. Furthermore, multiple digital assistants may support the same words / phrases, creating redundancy and consuming resources that could be allocated to supporting additional words / phrases and / or larger models that can detect words / phrases with lower false acceptance or rejection rates.

[0022] Embodiments herein are directed to enabling a set of one or more warm words associated with an ongoing long-term operation and at least one digital assistant of a multi-assistant-enabled device. That is, the warm words enabled for detection by the multi-assistant-enabled device are associated with a high probability of being spoken by a user after an initial command to control the long-term operation and may also be used to control at least one digital assistant to perform other operations different from the ongoing long-term operation. In this way, while the assistant-enabled device is performing a long-term operation commanded by the user, the user may speak any of the enabled warm words to trigger a respective action performed by at least one of the digital assistants to control the long-term operation. The warm word detector and speaker identification may operate on the multi-assistant-enabled device and have low power consumption.

[0023] 1A-1C illustrate example systems 100, 100a-c that activate warm words 112 for actions on one or more digital assistants 105, 105a-n, which control long-term behavior associated with an initial warm word 112 spoken by a user 102 in an initial command to control the long-term behavior. Briefly, as described in more detail below, a multi-assistant device (MAD) 104 simultaneously runs a group of digital assistants 105, 105a-n. The digital assistants 105 of the MAD 104 begin playing music 122 in response to an utterance 106, "OK computer, play music," spoken by a user 102. Additionally, the multi-assistant interface 210 (FIG. 2) of the MAD 104 performs an arbitration routine to validate a final set 113 of warm words 112 for detection by the MAD 104. While digital assistant 105b is performing a long-running operation with music 122 as playback audio from speaker 18, MAD 104 can detect / recognize the warm word 112 (FIG. 1C) spoken by user 102 as an action to control the long-running operation, such as a command to stop the playback audio of music 122. Advantageously, the arbitration routine may identify digital assistants 105 with duplicate warm words, and MAD 104 enables only a single instance of the duplicate warm word 112, thereby conserving resources of MAD 104.

[0024] Systems 100a-100c include MAD 104 executing a group of digital assistants 105, 105a-n, enabled for simultaneous execution on MAD 104, with which user 102 may interact via voice. Each digital assistant 105 in the group of digital assistants 105 may include a respective hotword detector 108, speech recognizer 116, and natural language understanding (NLU) module 124. Additionally or alternatively, MAD 104 includes an assistant client, which may be a standalone application on an operating system or may form all or part of the operating system (e.g., executing a dedicated hotword detector 108, speech recognizer 116, and / or NLU module 124). While the group of digital assistants 105 execute simultaneously on MAD 104, the digital assistants 105 may all be linked together or otherwise associated with one another in one or more data structures. For example, each of the digital assistants 105 may be registered to the same user account, the same set of user accounts, a particular structure, and / or all assigned to a particular structure in the device topology representation. The device topology representation may include a corresponding unique identifier for each digital assistant 105 and, optionally, corresponding unique identifiers of other digital assistants 105 that may interact through the digital assistant. Additionally, the device topology representation may specify assistant attributes associated with each digital assistant 105. The attributes of a given digital assistant 105 may indicate, for example, one or more input and / or output modalities supported by each digital assistant 105, the processing capabilities of each digital assistant 105, the manufacturer, model, and / or unique identifier (e.g., serial number) of each digital assistant 105 (based on which processing capabilities may be determined), and / or other attributes.

[0025] In the illustrated example, the MAD 104 corresponds to a smart speaker with which the user 102 may interact. However, the MAD 104 may include other computing devices, such as, but not limited to, a smartphone, tablet, smart display, desktop / laptop, smartwatch, smart glasses / headset, smart appliance, headphones, or vehicle infotainment device. The MAD 104 includes data processing hardware 10 and memory hardware 12 that stores instructions that, when executed on the data processing hardware 10, cause the data processing hardware 10 to perform operations. The MAD 104 includes an array of one or more microphones 16 configured to capture sounds, such as voices, directed toward the MAD 104. The MAD 104 may also include or communicate with an audio output device (e.g., speaker) 18 that may output audio, such as music 122 and / or synthesized voice, from at least one of the digital assistants 105. Additionally, MAD 104 may include or communicate with one or more cameras 19 configured to capture images within the environment and output image data.

[0026] In some configurations, the MAD 104 communicates with a user device 50 associated with the user 102. In the illustrated example, the user device 50 includes a smartphone with which the user 102 may interact. However, the user device 50 may include other computing devices, such as, but not limited to, a smartwatch, a smart display, smart glasses, a smartphone, a smart headset, a tablet, a smart appliance, headphones, a computing device, a smart speaker, or other assistant-enabled devices. The user device 50 may include at least one microphone 52 present on the user device 50 that communicates with the MAD 104. In these configurations, the user device 50 may also communicate with one or more microphones 16 present on the MAD 104. Additionally, the user 102 may control and / or configure the MAD 104 and interact with the digital assistant 105 using an interface 200, such as a graphical user interface (GUI) 300 rendered for display on the screen of the user device 50.

[0027] 1A shows a user 102 speaking the utterance 106, "OK computer, play music," near a MAD 104. A microphone 16 of the MAD 104 receives the utterance 106 and processes audio data 402 corresponding to the utterance 106. Initial processing of the audio data 402 may include filtering the audio data 402 and converting the audio data 402 from an analog signal to a digital signal. Once the MAD 104 processes the audio data 402, the MAD may store the audio data 402 in a buffer in the memory hardware 12 for further processing. Using the audio data 402 in the buffer, the MAD 104 may use a hotword detector 108 (e.g., a hotword detector 108 of one of the digital assistants 105) to detect whether the audio data 402 contains a hotword. The hotword detector 108 is configured to identify hotwords contained in the audio data 402 without performing voice recognition on the audio data 402.

[0028] In some implementations, the hot word detector 108 is configured to identify hot words in an early portion of the utterance 106. In this example, the hot word detector 108 may determine that the utterance 106, "OK computer, play music," includes the hot word 110, "OK computer," if the hot word detector 108 detects acoustic features in the audio data 402 that are characteristic of the hot word 110. The acoustic features may be Mel-Frequency Cepstral Coefficients (MFCCs), which are a representation of the short-term power spectrum of the utterance 106, or may be Mel-scale filter bank energy of the utterance 106. For example, the hot word detector 108 may detect that the utterance 106, "OK computer, play music," includes the hot word 110, "OK computer," based on generating MFCCs from the audio data 402 and classifying the MFCCs as including MFCCs similar to MFCCs that are characteristic of the hot word "OK computer" stored in the hot word model of the hot word detector 108. As another example, the hot word detector 108 may detect that the utterance 106, "OK computer, play music," contains the hot word 110, "OK computer," based on generating mel-scale filter bank energies from the audio data 402 and classifying the mel-scale filter bank energies as containing mel-scale filter bank energies similar to mel-scale filter bank energies characteristic of the hot word "OK computer" stored in the hot word model of the hot word detector 108.

[0029] When the hot word detector 108 determines that the audio data 402 corresponding to the utterance 106 includes the hot word 110, the MAD 104 may trigger a wake-up process to begin speech recognition on the audio data 402 corresponding to the utterance 106. For example, FIG. 2 shows the MAD 105 including a speech recognizer 116 (e.g., a speech recognizer 116 of one of the digital assistants 105) that uses an automatic speech recognition model 117 that may perform speech recognition or semantic interpretation on the audio data 402 corresponding to the utterance 106. The speech recognizer 116 may perform speech recognition on the portion of the audio data 402 that follows the hot word 110. In this example, the speech recognizer 116 may identify the words "play music" in the command 118.

[0030] In some examples, the MAD 104 is configured to communicate with a remote system 130 over the network 120. The remote system 130 may include remote resources, such as remote data processing hardware 132 (e.g., a remote server or CPU) and / or remote memory hardware 134 (e.g., a remote database or other storage hardware). The MAD 104 may utilize the remote resources to perform various functions related to speech processing and / or synthesis playback communication. In some implementations, the speech recognizer 116 is located on the remote system 130 in addition to or instead of the MAD 104. When the hotword detector 108 triggers the MAD 104 to wake up in response to detecting a hotword 110 in an utterance 106, the MAD 104 may transmit initial audio data 402 corresponding to the utterance 106 to the remote system 130 over the network 120. Here, the MAD 104 may transmit a portion of the initial audio data 402 including the hotword 110 for the remote system 130 to confirm the presence of the hotword 110. Alternatively, the MAD 104 may send only the portion of the initial audio data 402 corresponding to the portion of the utterance 106 after the hot word 110 to the remote system 130, which executes the speech recognizer 116 to perform speech recognition and returns a transcription of the initial audio data 402 to the MAD 104.

[0031] 1A-2, the MAD 104 may further include an NLU module 124 (e.g., an NLU module 124 of one of the digital assistants 105) that performs semantic interpretation on the utterance 106 to identify queries / commands directed to the MAD 104. Specifically, the NLU module 124 identifies words in the utterance 106 identified by the speech recognizer 116 and performs semantic interpretation to identify any voice commands in the utterance 106. The NLU module 124 of the MAD 104 (and / or the remote system 130) may identify the words "play music" as a command specifying a long-running action of the digital assistant 105 (i.e., playing music 122). In the illustrated example, the digital assistant 105 running on the MAD 104 begins to perform the long-running action of playing music 122 as playback audio (e.g., track 1) from the speaker 18 of the MAD 104. The digital assistant 105 may stream music 122 from a streaming service (not shown), or the digital assistant 105 may instruct the MAD 104 to play music stored on the MAD 104. An exemplary long-term operation includes music playback, although long-term operations may also include other types of media playback, such as videos, podcasts, and / or audiobooks. Long-term operations may also include home automation (e.g., adjusting light levels, controlling the thermostat, etc.).

[0032] 1A-2, MAD 104 (and / or server 130) may include a multi-assistant interface 210 that executes assistant detector 220, warm word enabler 230, and warm word detector 240. Assistant detector 220 may be configured to discover one or more active digital assistants 105 running simultaneously on MAD 104. Based on a change in one of the group of assistant devices 105 discovered by assistant detector 220, the warm words 112 detected by warm word detector 240, and / or the context 212 in the environment of the user and / or user device 50, warm word enabler 230 is triggered to execute a warm word arbitration routine that determines which warm words 112 to include in the final set 113 of warm words 112 based on the processing constraints of MAD 104. The warm word detector 240 receives the final set 113 of warm words 112 and may detect warm words 112 from the final set 113 in the streaming audio captured by the MAD 104 without performing speech recognition on the captured audio, as described in more detail below.

[0033] The MAD 104 may notify the user 102 (e.g., Barb) who spoke the utterance 106 that a long-running action is being performed. For example, the digital assistant 105 may generate a synthesized voice 123 (FIG. 1A) for audible output from the speaker 18 of the MAD 104 stating, "Barb, it's okay to talk to the music playback controls instead of saying 'OK computer.'" In an additional example, the digital assistant 105 provides a notification to the user device 50 associated with the user 102 (e.g., Barb) to inform the user 102 which warm word 112 is currently valid for controlling the long-running action.

[0034] 2, while the digital assistant 105 is running, the MAD 104 uses the assistant detector 220 to discover the digital assistant 105 for a group of digital assistants 105 enabled for simultaneous execution on the MAD 104. For example, one or more active digital assistants 105 may be discovered based on which long-term operation the MAD 104 is currently performing. In another example, when the MAD 104 receives a user command specifying a long-term operation to be performed by the corresponding digital assistant 105, one or more active digital assistants 105 are added to the group of digital assistants 105. The user command specifying the long-term operation may include a user input instruction via any one of touch, voice, gesture, gaze, and / or an input device (e.g., a mouse or stylus) to interact with the MAD 104. Additionally or alternatively, the user command specifying a long-term operation for the corresponding digital assistant 105 may trigger the discovery of one or more additional digital assistants 105 that interface / interact with the corresponding digital assistant 105. In another example, the assistant detector 220 monitors the activity of applications running on the MAD 104 to identify active applications that include a corresponding digital assistant 105.

[0035] In some implementations, assistant detector 220 maintains a list of previous digital assistants 105 in a group of digital assistants 105 enabled for concurrent execution on MAD 104. Here, the list of previous digital assistants 105 may refer to a list of digital assistants 105 discovered by assistant detector 220 that is not the most recent (i.e., latest) discovery of the assistant device 105. In this example, after discovering one or more digital assistants 105 in the group of digital assistants 105, assistant detector 220 may determine that the list of previous digital assistants 105 does not include the same digital assistant 105 associated with the current state of MAD 104. In other words, assistant detector 220 may determine that the digital assistants 105 in the list of previous digital assistants 105 are different from the digital assistant 105 currently running on MAD 104. This change between the list of previous digital assistants 105 and the list of current digital assistants 105 triggers enabler 230 to execute a warm word reconciliation routine to add or remove warm words 112 from the final set 113 of warm words 112 enabled for detection by MAD 104 based on the current digital assistant 105. For example, if a new digital assistant 105 starts running on MAD 104, assistant detector 220 may discover that the new digital assistant 105 in the group of digital assistants 105 is enabled for simultaneous execution on MAD 105 by comparing the list of current digital assistants 105 that includes the digital assistant 105 with a previous list of digital assistants 105 that does not include the digital assistant 105. In response, warm word enabler 230 triggers execution of a warm word reconciliation routine to add / update any warm words 112 to the final set 113 of warm words 112.Conversely, if the digital assistant 105 stops running on the MAD 104, the assistant detector 220 may discover that the digital assistant 105 has been removed from the group of digital assistants 105 by comparing the list of previous digital assistants 105 that included the digital assistant 105 with the list of current digital assistants 105 that do not include the digital assistant 105. In response, the warm word enabler 230 triggers execution of a warm word reconciliation routine to add / update any warm words 112 to the final set 113 of warm words 112. However, in some examples, if the assistant detector 220 determines that the list of the previous digital assistant 105 is the same as the list of the current digital assistant 105, the assistant detector 220 may not send the list of the current digital assistant 105 to update the final set 113 of warm words 112.

[0036] In some implementations, MAD 104 identifies warm words 112 to be added or removed from each active set 111 of warm words 112 for one or more of the digital assistants 105 in the group of digital assistants 105. For example, user 102 may say "pause" to pause playback of music 122. While the music is paused, MAD 104 may determine that the warm word 122 "pause" is to be removed from each active set 111 of warm words 112 associated with the digital assistant 105 controlling the music playback because the user 102 is unlikely to say this warm word. In other words, once the user says "pause," digital assistant 105 may designate the warm word 112 "pause" as an inactive warm word 112 because the user 102 is unlikely to repeat the warm word 112 "pause." In response to determining that the warm word 112 "pause" is to be removed from the active set 111 of warm words 112, the MAD 104 executes a warm word arbitration routine (via the warm word enabler 230) to update the final set 113 of warm words 112. Continuing the example, the MAD 104 may determine that the warm word 112 "play" is to be added to the active set 111 of warm words 112 in response to the user 102 speaking "pause" because it is likely that the user 102 will resume playing the music 122 within a threshold time. In response to determining that the warm word 112 "play" is to be added to the active set 111 of warm words 112, the MAD executes a warm word arbitration routine to update the final set 113 of warm words 112. In other embodiments, the warm word reconciliation routine is triggered based on the addition and deletion of warm words 112 in the respective active sets 111 of warm words 112 for the corresponding digital assistant 105 (e.g., based on configuration changes requested by the user).

[0037] In some implementations, the warm word enabler 230 executes the warm word arbitration routine in response to determining a change in the ambient context 212. For example, the ambient context 212 may include background noise in the user's 102's environment, and an increase in volume may constitute a change in the ambient context 212 that triggers the warm word arbitration routine to update / modify the final set 113 of warm words 112 to include a more capable warm word model 114 for one or more of the valid warm words 112 that more accurately detects warm words 112 in a noisy environment. In another example, the ambient context 212 may include lighting in the MAD 104's environment. Here, the camera 19 of the MAD 104 may capture image data indicating that the lighting is dimmer (e.g., the user 102 is starting their bedtime routine) as a change in the ambient context 212 that triggers the warm word arbitration routine to enable a warm word 112 to control the alarm clock digital assistant 105. Similarly, changes in ambient context 212 may include various time periods, and the warm word arbitration routine is triggered to enable / disable warm words 112 based on the time period. In these examples, MAD 104 may automatically execute the warm word arbitration routine to update the valid warm words 112 based on the daily routine of user 102 of MAD 104 corresponding to individual time periods. In yet a further example, ambient context 218 may include proximity to other assistant-enabled devices. For example, if MAD 104 is a vehicle system, user 102 may instruct vehicle digital assistant 105 to turn on the climate control when the vehicle entering a threshold proximity range of another assistant-enabled device (e.g., a smart thermostat in the user's 102 home) causes a change in ambient context 218 that triggers MAD 104 to execute the warm word arbitration routine (e.g., update final set 113 to include the warm word 112 corresponding to the smart thermostat).

[0038] 2, for each active digital assistant 105 in the group of active digital assistants 105 discovered by the assistant detector 220, the warm word enabler 230 may receive a respective active set 111, 111a-n of warm words 112, each of which specifies a respective action to be performed by the respective digital assistant 105. In other words, receiving the respective active set 111 of warm words 112 associated with the digital assistant 105 includes receiving the respective active set 111 of warm words 112 associated with the corresponding digital assistant 105 in response to the corresponding digital assistant 105 performing a long-term action. Based on the respective active set 111 of warm words 112 associated with each digital assistant 105 in the group of digital assistants 105, the warm word enabler 230 performs a warm word arbitration routine to enable a final set 113 of warm words 112 for detection by the MAD 104. The final set 113 of warm words 112 may include warm words selected from the active sets 111 of warm words 112 associated with one or more of the digital assistants 105 in the group of digital assistants 105.

[0039] In some embodiments, for each warm word 112 in the active set 111 of warm words 112, the multi-assistant user interface 210 further receives a respective warm word model 114 configured to detect the corresponding warm word 112 in streaming audio without performing speech recognition. For example, the MAD 104 (and / or the server 130) further includes one or more warm word models 114. Here, the warm word models 114 may be stored in the memory hardware 12 of the MAD 104 or in remote memory hardware 134 on the server 130. If stored on the server 130, the MAD 104 may obtain the warm word model 114 for the corresponding warm word 112 and request the server 130 to provide the obtained warm word model 114 so that the MAD 104 can enable the warm word model 114 (via the warm word enabler 230). An active warm word model 114 running on the warm word detector 240 of the MAD 104 may detect utterances of the corresponding warm word 112 in the streaming audio captured by the MAD 104 without performing speech recognition on the captured audio. Furthermore, a single warm word model 114 may be able to detect all of the warm words 112 in the streaming audio. Thus, a warm word model 114 may detect one or more warm words 112, and a different warm word model 114 may detect one or more other warm words 112.

[0040] In some configurations, the warm word enabler 230 receives code associated with a long-running application (e.g., a music application running in the foreground or background of the MAD 104) to identify any warm words 112 and associated warm word models 114 that the application's developer wants users 102 to be able to speak to interact with the application and respective actions for each warm word 112. In other examples, the warm word enabler 230 receives, via a warm word application programming interface (API) running on the MAD 104, a respective warm word model 114 for at least one warm word 112 in a respective active set 111 of warm words 112, configured to detect the corresponding warm word 112 in streaming audio without performing speech recognition. The warm words 112 in the registry may also be related to follow-up queries that users 102 (or typical users) tend to issue following a given query, e.g., "OK computer, next track."

[0041] In additional embodiments, by enabling the final set 113 of warm words 112 via the warm word enabler 230, the MAD 104 runs the speech recognizer 116 in a low-power and low-fidelity state. Here, the speech recognizer 116 is constrained or biased to recognize only one or more warm words 112 that are enabled when spoken in the speech captured by the MAD 104. Because the speech recognizer 116 recognizes only a limited number of terms / phrases, the number of parameters for the speech recognizer 116 can be significantly reduced, thereby reducing the memory requirements and the number of calculations required to recognize warm words in speech. Therefore, the low-power and low-fidelity characteristics of the speech recognizer 116 may be suitable for execution on a digital signal processor (DSP). In these embodiments, the speech recognizer 116 running on the MAD 104 may recognize utterances 147 of valid warm words 112 in the streaming audio captured by the MAD 104 instead of using the warm word model 114.

[0042] As described above, the warm word enabler 230 performs a warm word reconciliation routine to identify which warm words 112 can be enabled, what quality of warm word models 114 to use, and / or what error tolerance is appropriate for the MAD 104, based on the resource constraints of the MAD 104. For example, the warm word reconciliation routine may assign the highest priority to warm words 112 shared across multiple digital assistants 105 in a group of digital assistants 105, and the remaining warm words 112 are ranked based on metadata and / or usage history of the warm words 112. For example, the warm word enabler 230 identifies any shared warm words 112 corresponding to warm words 112 present in at least two active sets 111 of warm words 112, and the warm word reconciliation routine assigns a higher priority to the warm words 112 identified as shared warm words 112 by enabling them. Additionally or alternatively, the warm word reconciliation routine assigns a higher priority to similar warm words 112, where warm words 112 that sound similar and are associated with the same intent may be merged into a single warm word model 114. Alternatively, warm words 112 that sound similar and are associated with different intents may have separate warm word models 114 to limit errors in warm word 112 detection.

[0043] In some embodiments, for each warm word 112 in the active set 111 of warm words 112 for each digital assistant 105, a priority score for the warm word 112 can be determined based on the characteristics of the corresponding digital assistant 105. The priority score can also be determined based on the user's characteristics and / or the embedding of the warm word. The priority score can be determined taking into account warm words that may be semantically related (e.g., "play" and "pause") and / or phonetically similar and may be more relevant to a particular digital assistant 105, for example, based on the user's preferred digital assistant 105. In some embodiments, the priority score can be determined based on the output of a machine learning model.

[0044] In some examples, the warm word arbitration routine executed by the warm word enabler 230 of the multi-assistant interface 210 may obtain valid warm word constraints when determining the final set 113 of warm words 112 to enable for detection by the MAD 104. Here, the number of warm words 112 in the final set 113 of warm words 112 enabled for detection by the MAD 104 is based on the obtained valid warm word constraints. For example, the valid warm word constraints may include the resource availability of the memory hardware 12 and data processing hardware 10 on the MAD 104 for detecting warm words 112, or the computational requirements for enabling each warm word 112 in each active set 111 of warm words 112 associated with each digital assistant 105 in the group of digital assistants. Additionally or alternatively, the valid warm word constraints may include a tolerance range for an acceptable false acceptance rate or a tolerance range for an acceptable false rejection rate. In these examples, the valid warm word constraints define the boundaries within which the warm word reconciliation routine must operate in order to make the final set 113 of warm words 112 valid.

[0045] The warm word enabler 230 may maintain a log of identification information and / or timestamps when each warm word 112 is detected by the warm word detector 240 to generate additional parameters for the warm word reconciliation routine. For example, the warm word enabler 230 may determine the frequency of detection of the warm word 112 by the MAD 104 when executing the warm word reconciliation routine. In these examples, the warm word enabler 230 may determine the frequency of detection of each warm word 112 in each active set 111 of warm words 112 of each digital assistant 105 in a group of digital assistants 105, and the final set 113 of warm words 112 is based on the determined detection frequency. Here, the determined detection frequency defines an affinity for the warm word 112 for each digital assistant 105. In other words, each digital assistant 105 may have a different affinity for each warm word 112 based on the frequency of detection of the warm word 112 by each digital assistant 105. Similarly, the warm word arbitration routine may include, for each warm word 112 in each active set 111 of warm words 112 for each digital assistant 105, the time the warm word 112 was most recently detected by the MAD 104. Here, determining the final set 113 of warm words 112 is based on the determined time each warm word 112 in each active set 111 of warm words 112 for each digital assistant 105 was most recently detected. In these examples, recently detected warm words 112 may be assigned a higher priority during the warm word arbitration routine than warm words 112 that have not been detected recently and / or frequently.

[0046] 1A , in the case of a long-term operation of playing music 122, the assistant detector 220 detects / determines that the digital assistant 105 for the music application is active. The warm word enabler 230 receives an active set 111a of warm words 112 to control the long-term operation of playing music, the active set 111a including the warm words 112 "stop," "next," "previous," "turn up the volume," and "turn down the volume," each associated with a respective action for controlling the playback of music 122 from the speaker 18 of the MAD 104. Based on receiving the active set 111a of warm words 112, the warm word enabler executes a warm word arbitration routine to enable the warm words 112 "stop," "next," "previous," "turn up the volume," and "turn down the volume" while the digital assistant 105 is performing the long-term operation, and deactivates these warm words 112 when the long-term operation ends.

[0047] 1B, while the digital assistant 105 running on the MAD 104 is performing a long-duration operation of playing music 122, the MAD 104 receives an utterance 146 spoken by the user 102, "Set the timer for 30 minutes." Here, running the timer 125 is a long-duration operation from the time the timer is set to until the timer expires, or until the resulting alert is acknowledged after the timer expires. Based on the long-duration operation of running the timer 125, the assistant detector 220 detects / determines that the digital assistant 105 for the clock application is active, which causes a change in the current assistant 105 (i.e., triggers the warm word enabler 230 to execute a warm word reconciliation routine). In response, the warm word enabler 230 receives a respective active set 111b of warm words 112, each of which specifies a respective action to be performed by the digital assistant (e.g., the clock digital assistant). In this case, for a long-term operation that operates the timer 125, the respective active set 111b of warm words 112 received by the warm word enabler 230 includes the warm words 112 "reset," "lap," "stop," and "add time," each associated with a respective action for controlling the operation that operates the timer 125 on the MAD 104.

[0048] Based on the active set 111b of received warm words 112 and the active set 111a of received warm words 112, the warm word enabler 230 executes a warm word reconciliation routine to enable a final set 113 of warm words 112 for detection by the MAD 104 by selecting warm words 112 from the respective active sets 111a, 111b of warm words 112. Here, the warm word reconciliation routine identifies that the active sets 111a and 111b each contain the warm word 112 "stop." Rather than allowing detection of two separate warm word models 114 for the warm word 112 "stop," the warm word reconciliation routine selects only one warm word model 114 for the warm word 112 "stop." In some implementations, the warm word arbitration routine determines that the MAD 104 has sufficient capacity to enable a higher quality (i.e., additional parameters, reduced latency, and / or increased sensitivity) implementation of the warm word model 114 for the warm word 112 "stop." As shown in the example, the multi-assistant interface 210 enables a final set 113 of warm words 112 including the warm words 112 "next," "previous," "volume up," "volume down," "stop," "reset," "lap," and "add time."

[0049] 1C , while the MAD 104 is performing a long-duration operation of playing music 122 and running a timer 125, the user 102 speaks an utterance 147 that includes a warm word 112 from a final set 113 of warm words 112 that are enabled for detection by the MAD 104. In the illustrated example, the user 102 speaks the enabled warm word 112, "stop." Without performing speech recognition on the captured audio, the MAD 104 may apply the warm word model 114 for the final set 113 of warm words 112 to identify whether the utterance 147 includes any of the warm words 112 in the final set 113. The final set 113 of warm words 112 may be "next," "previous," "volume up," "volume down," "stop," "reset," "lap," and "add time." The MAD 104 compares the audio data 402 corresponding to the utterance 147 with the valid warm word models 114 corresponding to the valid warm words 112 “next,” “previous,” “volume up,” “volume down,” “stop,” “reset,” “rap,” and “add time,” and determines that the warm word model 114 enabled for the warm word 112 “stop” detects the warm word 112 “stop” in the utterance 147 without performing speech recognition on the audio data 402.

[0050] Because MAD 104 detects the shared warm word 112 "stop" in the audio data 402, MAD 104 may perform query interpretation using NLU module 124 to obtain additional context for MAD 104 and determine which long-term action (and corresponding digital assistant 105) the user 102 is referring to. For example, NLU module 124 may determine, based on the context that a timer is running, that the user 102 is more likely to want to stop the long-term action of running timer 125 rather than wanting to stop the long-term action of playing music 122. Alternatively, the context may indicate that the user 102 has moved out of the room away from MAD 104, and NLU module 120 may determine that the user 102 is more likely to want to stop music 122 because the user 102 is no longer in the vicinity of MAD 104. In other words, downstream actions to control long-term behavior may depend on the context determined by the NLU module 124 when determining which actions to perform on the active shared warm words 112 for performing actions by multiple digital assistants 105.

[0051] In some embodiments, the user 102 issues a voice command spoken near the MAD 104 and captured by the MAD 104 in the streaming audio. Here, the voice command spoken by the user 102 commands the MAD 104 to enable the first digital assistant 105 and the second digital assistant 105 to run simultaneously on the MAD 104. In these embodiments, the MAD 104 may include an assistant client (e.g., a standalone application on an operating system), and the first digital assistant 105 includes an assistant client already running on the MAD 104 and activates the second digital assistant 105 in response to receiving the voice command. In other words, after receiving the voice command, the first digital assistant 105 and the second digital assistant 105 are enabled to run simultaneously with each other on the MAD 104. Here, the group of digital assistants 105 includes the first digital assistant 105 and the second digital assistant 105.

[0052] In other embodiments, MAD 104 receives a multi-assistant configuration request from a software application running on MAD 104 or another assistant device (e.g., user device 50) that communicates with MAD 104 to enable the first digital assistant 105 and the second digital assistant 105 to run simultaneously on MAD 104. In these embodiments, the software application may correspond to an assistant client as a standalone operation on the operating system, and the user 102 may configure MAD 104 to run the first digital assistant 105 and the second digital assistant 105 to run simultaneously. For example, the user device 50 running the software application may submit a request to MAD 104 to run the first digital assistant 105 and the second digital assistant 105 simultaneously. In another example, a home automation software application running on MAD 104 is configured to include a first digital assistant 105 (e.g., a browser digital assistant 105) and a second digital assistant 105 (e.g., a music digital assistant 105). In these embodiments, after receiving a multi-assistant configuration request, MAD 104 enables the first digital assistant 105 and the second digital assistant 105 to run simultaneously with each other, and the group of digital assistants 105 includes the first digital assistant 105 and the second digital assistant 105.

[0053] 3 is a flowchart of an exemplary arrangement of operations of a method 300 for detecting warm words for multiple digital assistants. At operation 302, for each digital assistant 105 in a group of digital assistants 105, 105a-n enabled for simultaneous execution on a multi-assistant device (MAD) 104, the method 300 includes receiving a respective active set 111, 111a-n of warm words 112, each of which specifies a respective action to be performed by each digital assistant 105. At operation 304, based on the respective active sets 111 of warm words 112 associated with each digital assistant 105 in the group of digital assistants 105, the method 300 also includes, by the multi-assistant interface 210 running on the MAD 104, executing a warm word arbitration routine to enable a final set 113 of warm words 114 for detection by the MAD 104. Each corresponding warm word 112 in the final set 113 of warm words 112 enabled for detection by MAD 104 is selected from a respective active set 111 of warm words 112 for at least one digital assistant 105 in the group of digital assistants 105.

[0054] While the final set 113 of warm words 112 is enabled for detection by the MAD 104, the method 300 includes, at operation 306, receiving audio data 402 corresponding to the utterance 146 captured by the MAD 104. At operation 308, the method 300 also includes detecting warm words 112 from the final set 113 of warm words 112 in the audio data 402. The method 300 also includes, at operation 310, instructing a digital assistant 105 associated with the detected warm word 112 from the group of digital assistants 105 to perform a respective action specified by the detected warm word 112.

[0055] 4 is a schematic diagram of an exemplary computing device 400 that may be used to implement the systems and methods described herein. Computing device 400 is intended to represent various types of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown here, their connections and relationships, and their functions are meant to be exemplary only and are not meant to limit the implementation of the invention(s) described and / or claimed herein.

[0056] Computing device 400 includes processor 410, memory 420, storage device 430, high-speed interface / controller 440 connecting to memory 420 and high-speed expansion port 450, and low-speed interface / controller 460 connecting to low-speed bus 470 and storage device 430. Components 410, 420, 430, 440, 450, and 460 are interconnected using various buses and may be mounted on a common motherboard or otherwise mounted as desired. Processor 410 (e.g., data processing hardware 12 and / or remote data processing hardware 132 of FIG. 1A ) may process instructions for execution within computing device 400, including instructions stored in memory 420 or storage device 430, for displaying graphical information for a graphical user interface (GUI) on an external input / output device, such as display 480 coupled to high-speed interface 440. In other implementations, multiple processors and / or multiple buses may be used as needed, along with multiple memories and types of memory. Also, multiple computing devices 400 may be connected (eg, as a bank of servers, a group of blade servers, or a multi-processor system) with each device providing a portion of the required operations.

[0057] Memory 420 (e.g., memory hardware 12 and / or remote memory hardware 134 of FIG. 1A) stores information non-temporarily within computing device 400. Memory 420 may be a computer-readable medium, volatile memory unit(s), or non-volatile memory unit(s). Non-transient memory 420 may be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by computing device 400. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.

[0058] Storage device 430 can provide mass storage for computing device 400. In some embodiments, storage device 430 is a computer-readable medium. In various different implementations, storage device 430 may be a device array including a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or a storage area network or other configuration of devices. In further embodiments, a computer program product is tangibly embodied on an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as memory 420, storage device 430, or memory on processor 410.

[0059] The high-speed controller 440 manages bandwidth-intensive operations of the computing device 400, while the low-speed controller 460 manages less bandwidth-intensive operations. This allocation of roles is merely exemplary. In some implementations, the high-speed controller 440 is coupled to memory 420, a display 480 (e.g., via a graphics processor or accelerator), and a high-speed expansion port 450 that can accept various expansion cards (not shown). In some implementations, the low-speed controller 460 is coupled to a storage device 430 and a low-speed expansion port 490. The low-speed expansion port 490 may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet) and may be coupled to one or more input / output devices such as a keyboard, a pointing device, a scanner, or, for example, a network device such as a switch or a router via a network adapter.

[0060] The computing device 400, as shown, may be implemented in many different forms. For example, it may be implemented as a standard server 400a, or multiple times in a group of such servers 400a, as a laptop computer 400b, or as part of a rack server system 400c.

[0061] Various implementations of the systems and techniques described herein may be realized in digital electronic and / or optical circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementation in one or more computer programs executable and / or interpretable by a programmable system including at least one programmable processor, which may be special purpose or general purpose, coupled to receive data and instructions from and transmit data and instructions to a storage system, at least one input device, and at least one output device.

[0062] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform a task. In some embodiments, a software application may be referred to as an "application," "app," or "program." Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.

[0063] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in a high-level procedural programming language and / or an object-oriented programming language and / or an assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0064] The processes and logic flows described herein may be implemented by one or more programmable processors, also referred to as data processing hardware, executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be implemented by special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). Processors suitable for executing computer programs include, by way of example, both general-purpose and special purpose microprocessors and any one or more processors of any kind of digital computer. Generally, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or is operably coupled to receive data from or transmit data to them, or both. A computer need not, however, have such devices. Computer-readable media suitable for storing computer program instructions and data include all types of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0065] To interact with a user, one or more aspects of the present disclosure may be implemented on a computer having a display device, such as a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen, for displaying information to the user, and optionally a keyboard and pointing device, such as a mouse or trackball, through which the user can provide input to the computer. Other types of devices may also be used to provide interaction with the user; for example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user may be received in any form, including acoustic, speech, or tactile input. Additionally, the computer may interact with the user by sending and receiving documents to devices used by the user, for example, by sending a web page to a web browser on the user's client device in response to a request received from the web browser.

[0066] Although several embodiments have been described, it will be understood that various modifications can be made without departing from the spirit and scope of the present disclosure. Accordingly, other embodiments are within the scope of the following claims.

Claims

1. A computer-implemented method (300) that, when executed on data processing hardware (10), causes the data processing hardware to perform an operation, the operation comprising: For each digital assistant (105) in a group of digital assistants (105) enabled for simultaneous execution on a multi-assistant device (MAD (104)), receive a respective active set (111) of warm words (112), each specifying a respective action that each digital assistant (105) will perform; and executing a warm word arbitration routine by a multi-assistant interface (200) running on the MAD (104) to enable a final set (113) of warm words (112) for detection by the MAD (104) based on the respective active sets (111) of the warm words (112) associated with each digital assistant (105) in the group of digital assistants (105), wherein each corresponding warm word (112) in the final set (113) of warm words (112) enabled for detection by the MAD (104) is selected from the respective active sets (111) of the warm words (112) for at least one digital assistant (105) in the group of digital assistants (105); While the final set (113) of warm words (112) is being validated for detection by the MAD (104), The operations further include receiving audio data (402) corresponding to speech captured by the MAD (104); detecting a warm word (112) from the final set (113) of warm words (112) in the audio data (402); Instructing the digital assistant (105) associated with the detected warm word (112) from the group of digital assistants (105) to perform the respective action specified by the detected warm word (112); A computer-implemented method (300) comprising:

2. 2. The computer-implemented method of claim 1, wherein receiving a respective active set (111) of the warm words (112) associated with the digital assistant (105) includes receiving, for at least one warm word (112) in the respective active set (111) of the warm words (112), via a warm word application programming interface (API) running on the MAD (104), a respective warm word model configured to detect the corresponding warm word (112) in streaming audio without performing speech recognition.

3. The operation further includes, for a corresponding digital assistant (105) of the digital assistant (105) in the group of digital assistants (105), Receiving a user command (118) specifying a long-term action to be performed by the corresponding digital assistant (105); Executing the long-term action specified by the user command (118) via the corresponding digital assistant (105); Including, 3. The computer-implemented method (300) of claim 1 or 2, wherein receiving an active set (111) of each of the warm words (112) associated with the digital assistant (105) includes receiving an active set (111) of each of the warm words (112) associated with the corresponding digital assistant (105) in response to the corresponding digital assistant (105) performing the long-term operation.

4. 4. The computer-implemented method of claim 3, wherein each warm word in each active set of warm words is associated with a respective action for controlling the long-term operation performed by the corresponding digital assistant.

5. The operation further comprises: Discovering a new digital assistant (105) within a group of digital assistants (105) enabled for concurrent execution on the MAD (104); The multi-assistant interface (200) executes the warm word arbitration routine in response to discovering the new digital assistant (105) within the group of digital assistants (105). A computer-implemented method (300) according to any one of claims 1 to 4.

6. The operation further comprises: determining that a digital assistant (105) has been removed from a group of the digital assistants (105) enabled for concurrent execution on the MAD (104); The multi-assistant interface (200) executes the warm word arbitration routine in response to determining that the digital assistant (105) has been removed from the group of digital assistants (105). A computer-implemented method (300) according to any one of claims 1 to 5.

7. The operation further includes, for a corresponding digital assistant (105) of the digital assistant (105) in the group of digital assistants (105), determining whether to add or remove a warm word (112) from the active set (111) of each of the warm words (112) associated with the corresponding digital assistant (105); The computer-implemented method (300) of any one of claims 1 to 6, wherein the multi-assistant interface (200) executes the warm word reconciliation routine in response to determining the addition of the warm word (112) or the deletion of the warm word (112) in the respective set of warm words (112).

8. The operation further comprises: determining a change in the surrounding context (212); The computer-implemented method (300) of any one of claims 1 to 7, wherein the multi-assistant interface (200) executes the warm word reconciliation routine in response to determining a change in the ambient context (212).

9. The operation further comprises: obtaining a valid warm word constraint, the valid warm word constraint comprising: the availability of memory and computing resources on the MAD (104) for warm word (112) detection; The computational requirements for validating each warm word (112) in the active set (111) of each warm word (112) associated with each digital assistant (105) in the group of digital assistants (105); an acceptable range of false acceptance rates, or an acceptable false rejection rate tolerance; 9. The computer-implemented method of claim 1, wherein the number of warm words in the final set of warm words enabled for detection by the MAD is based on the obtained valid warm word constraints.

10. executing the warm word arbitration routine identifying any shared warm words (112) corresponding to warm words (112) present in at least two of the active sets (111) of warm words (112); 10. The computer-implemented method (300) of claim 1, wherein determining the final set (113) of warm words (112) is based on assigning higher priorities to warm words (112) identified as shared warm words (112).

11. executing the warm word arbitration routine For each warm word (112) in each active set (111) of warm words (112) for each digital assistant (105) in the group of digital assistants (105), determining the frequency of detection of the warm word (112) by the MAD (104); The computer-implemented method (300) of any one of claims 1 to 10, wherein determining the final set (113) of warm words (112) is based on the determined detection frequency of each warm word (112) in each active set (111) of warm words (112) for each digital assistant (105) in the group of digital assistants (105).

12. executing the warm word arbitration routine For each warm word (112) in each active set (111) of the warm words (112) for each digital assistant (105) in the group of digital assistants (105), determining the time when the warm word (112) was most recently detected by the MAD (104); The computer-implemented method (300) of any one of claims 1 to 11, wherein determining the final set (113) of warm words (112) is based on the determined time when each warm word (112) in each active set (111) of warm words (112) for each digital assistant (105) in the group of digital assistants (105) was most recently detected.

13. The operation further comprises: receiving a voice command (118) commanding the MAD (104) to enable a first digital assistant (105) and a second digital assistant (105) to run simultaneously on the MAD (104), the voice command (118) being spoken by a user of the MAD (104) and captured by the MAD (104) in streaming audio; The computer-implemented method (300) of any one of claims 1 to 12, wherein the operations further include enabling the first digital assistant (105) and the second digital assistant (105) to run simultaneously with each other on the MAD (104) after receiving the voice command (118), and the group of digital assistants (105) includes the first digital assistant (105) and the second digital assistant (105).

14. The operation further comprises: Receiving a multi-assistant configuration request from a software application running on the MAD (104) or on another device that communicates with the MAD (104) to enable a first digital assistant (105) and a second digital assistant (105) to run simultaneously on the MAD (104); After receiving the multi-assistant configuration request, enabling the first digital assistant (105) and the second digital assistant (105) to run simultaneously with each other on the MAD (104), wherein the group of digital assistants (105) includes the first digital assistant (105) and the second digital assistant (105). A computer-implemented method (300) described in any one of claims 1 to 14.

15. A system (100), comprising: data processing hardware (10); and memory hardware (12) in communication with the data processing hardware (10), the memory hardware (12) storing instructions that, when executed on the data processing hardware (10), cause the data processing hardware (10) to perform operations, the operations including: For each digital assistant (105) in a group of digital assistants (105) enabled for simultaneous execution on a multi-assistant device (MAD (104)), receive a respective active set (111) of warm words (112), each specifying a respective action that each digital assistant (105) will perform; and executing a warm word arbitration routine by a multi-assistant interface (200) running on the MAD (104) to enable a final set (113) of warm words (112) for detection by the MAD (104) based on the respective active sets (111) of the warm words (112) associated with each digital assistant (105) in the group of digital assistants (105), wherein each corresponding warm word (112) in the final set (113) of warm words (112) enabled for detection by the MAD (104) is selected from the respective active sets (111) of the warm words (112) for at least one digital assistant (105) in the group of digital assistants (105); While the final set (113) of warm words (112) is being validated for detection by the MAD (104), The operations further include receiving audio data (402) corresponding to speech captured by the MAD (104); detecting a warm word (112) from the final set (113) of warm words (112) in the audio data (402); Instructing the digital assistant (105) associated with the detected warm word (112) from the group of digital assistants (105) to perform the respective action specified by the detected warm word (112); A system (100) comprising:

16. 16. The system (100) of claim 15, wherein receiving a respective active set (111) of the warm words (112) associated with the digital assistant (105) includes receiving, for at least one warm word (112) in the respective active set (111) of the warm words (112), via a warm word application programming interface (API) running on the MAD (104), a respective warm word model configured to detect the corresponding warm word (112) in streaming audio without performing speech recognition.

17. The operation further includes, for a corresponding digital assistant (105) of the digital assistant (105) in the group of digital assistants (105), Receiving a user command (118) specifying a long-term action to be performed by the corresponding digital assistant (105); and executing the long-term action specified by the user command (118) via the corresponding digital assistant (105); The system (100) of claim 15 or 16, wherein receiving an active set (111) of each of the warm words (112) associated with the digital assistant (105) includes receiving an active set (111) of each of the warm words (112) associated with the corresponding digital assistant (105) in response to the corresponding digital assistant (105) performing the long-term operation.

18. 18. The system (100) of claim 17, wherein each warm word (112) in each active set (111) of warm words (112) is associated with a respective action for controlling the long-term operation performed by the corresponding digital assistant (105).

19. The operation further comprises: Discovering a new digital assistant (105) within a group of digital assistants (105) enabled for concurrent execution on the MAD (104); The multi-assistant interface (200) executes the warm word arbitration routine in response to discovering the new digital assistant (105) within the group of digital assistants (105). The system (100) of any one of claims 15 to 18.

20. The operation further comprises: determining that a digital assistant (105) has been removed from a group of the digital assistants (105) enabled for concurrent execution on the MAD (104); The multi-assistant interface (200) executes the warm word arbitration routine in response to determining that the digital assistant (105) has been removed from the group of digital assistants (105). The system (100) described in any one of claims 15 to 19.

21. The operation further includes, for a corresponding digital assistant (105) of the digital assistant (105) in the group of digital assistants (105), determining whether to add or remove a warm word (112) from the active set (111) of each of the warm words (112) associated with the corresponding digital assistant (105); The system (100) of any one of claims 15 to 20, wherein the multi-assistant interface (200) executes the warm word reconciliation routine in response to determining the addition of the warm word (112) or the deletion of the warm word (112) in each set of the warm words (112).

22. The operation further comprises: determining a change in the surrounding context (212); The system (100) of any one of claims 15 to 21, wherein the multi-assistant interface (200) executes the warm word reconciliation routine in response to determining a change in the ambient context (212).

23. The operation further comprises: obtaining a valid warm word constraint, the valid warm word constraint comprising: the availability of memory and computing resources on the MAD (104) for warm word (112) detection; The computational requirements for validating each warm word (112) in the active set (111) of each warm word (112) associated with each digital assistant (105) in the group of digital assistants (105); an acceptable range of false acceptance rates, or an acceptable false rejection rate tolerance; 23. The system (100) of claim 15, wherein the number of warm words (112) in the final set (113) of warm words (112) enabled for detection by the MAD (104) is based on the obtained valid warm word constraints.

24. executing the warm word arbitration routine identifying any shared warm words (112) corresponding to warm words (112) present in at least two of the active sets (111) of warm words (112); 24. The system (100) of claim 15, wherein determining the final set (113) of warm words (112) is based on assigning higher priorities to warm words (112) identified as shared warm words (112).

25. executing the warm word arbitration routine For each warm word (112) in each active set (111) of warm words (112) for each digital assistant (105) in the group of digital assistants (105), determining the frequency of detection of the warm word (112) by the MAD (104); The system (100) described in any one of claims 15 to 24, wherein determining the final set (113) of warm words (112) is based on the determined detection frequency of each warm word (112) in each active set (111) of the warm words (112) for each digital assistant (105) in the group of digital assistants (105).

26. executing the warm word arbitration routine For each warm word (112) in each active set (111) of the warm words (112) for each digital assistant (105) in the group of digital assistants (105), determining the time when the warm word (112) was most recently detected by the MAD (104); The system (100) described in any one of claims 15 to 25, wherein determining the final set (113) of warm words (112) is based on the determined time when each warm word (112) in each active set (111) of the warm words (112) for each digital assistant (105) in the group of digital assistants (105) was most recently detected.

27. The operation further comprises: receiving a voice command (118) commanding the MAD (104) to enable a first digital assistant (105) and a second digital assistant (105) to run simultaneously on the MAD (104), the voice command (118) being spoken by a user of the MAD (104) and captured by the MAD (104) in streaming audio; The system (100) of any one of claims 15 to 26, wherein the operations further include enabling the first digital assistant (105) and the second digital assistant (105) to run simultaneously with each other on the MAD (104) after receiving the voice command (118), and the group of digital assistants (105) includes the first digital assistant (105) and the second digital assistant (105).

28. The operation further comprises: Receiving a multi-assistant configuration request from a software application running on the MAD (104) or on another device that communicates with the MAD (104) to enable a first digital assistant (105) and a second digital assistant (105) to run simultaneously on the MAD (104); After receiving the multi-assistant configuration request, the system (100) includes enabling the first digital assistant (105) and the second digital assistant (105) to run simultaneously with each other on the MAD (104), wherein the group of digital assistants (105) includes the first digital assistant (105) and the second digital assistant (105).

Citation Information

Patent Citations

  • Method and system for autocompletion for languages having ideographs and phonetic characters

    JP2012108959A

  • Document classification system, control method of document classification system, and control program of document classification system

    JP2016027510A

  • Contextual Hot Words

    JP2020503568A

  • Network microphone device with command keyword adjustments

    JP2022536765A

  • Crowdsourced on-boarding of digital assistant operations

    US20180336049A1