Multi-user warm words
The method and system optimize warm word detection in Assistant-enabled devices by managing multiple users' interactions through user detection, arbitration, and low-power recognition, addressing inefficiencies and false positives.
Patent Information
- Application Number
- JP2025528765
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-17
- Filing Date
- 2023-11-07
- Publication Date
- 2025-11-28
AI Technical Summary
Existing Assistant-enabled devices struggle to efficiently manage multiple users' interactions due to conflicting commands and computational limitations, leading to increased false positives and inefficiencies in warm word detection.
A method and system for detecting multiple users in an environment, obtaining respective sets of active warm words, and executing a warm word arbitration routine to select a final set of warm words for detection, considering computational resources, user preferences, and tolerance rates, while enabling speaker verification and using low-power speech recognition.
Enhances the ability of Assistant-enabled devices to accurately and efficiently manage multiple users' interactions by optimizing warm word detection, reducing false positives, and ensuring that only authorized users' commands are executed.
Smart Images

Figure 2025538477000001_ABST
Abstract
Description
[Technical Field]
[0001] TECHNICAL FIELD This disclosure relates to multi-user warm words. [Background technology]
[0002] The way users interact with assistant-enabled devices is primarily, but not exclusively, designed through voice input. For example, a user may ask the device to perform an action, including media playback (e.g., music or podcasts), and the device responds by starting to play audio that matches the user's criteria. In situations where a device (e.g., a smart speaker) is widely shared by multiple users in an environment, the device may need to accommodate multiple actions requested by users that may conflict with each other. Summary of the Invention
[0003] One aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations, the operations including detecting the presence of multiple users in an environment of an assistant-enabled device (AED), the AED running a digital assistant, and obtaining, for each of the multiple users, a respective set of active warm words, each of which specifies a respective action to be performed by the digital assistant. The operations also include executing a warm word arbitration routine to enable a final set of warm words for detection by the AED based on the respective set of active warm words for each of the multiple users. The final set of warm words enabled for detection by the AED includes warm words selected from the respective set of active warm words for at least one of the multiple users detected in the environment of the AED. The operations also include receiving audio data corresponding to speech captured by the AED while the final set of warm words is enabled for detection by the AED, detecting warm words from the final set of warm words in the audio data, and instructing the digital assistant to perform the respective actions specified by the detected warm words.
[0004] Embodiments of the present disclosure may include one or more of the following optional features. In some embodiments, detecting the presence of multiple users within the environment of the AED includes detecting the presence of at least one of the multiple users within the environment based on proximity information of the AED to a user device associated with at least one of the multiple users. In additional embodiments, detecting the presence of multiple users within the environment of the AED includes receiving image data corresponding to a scene of the environment and detecting the presence of at least one of the multiple users within the environment based on the image data. In some examples, detecting the presence of multiple users within the environment of the AED includes detecting the presence of the corresponding users within the environment of the AED based on receiving audio data characterizing voice queries issued by the corresponding users and directed to the digital assistant, performing speaker identification on the audio data to identify the corresponding users who issued the voice queries, and determining that the identified corresponding users who issued the voice queries are present within the environment of the AED. In these examples, the voice query issued by the corresponding user includes a command that causes the digital assistant to perform a long-lasting action specified by the command, and obtaining the set of respective active warm words includes, in response to the digital assistant performing the long-lasting action specified by the command, adding one or more warm words that specify a respective action for controlling the long-lasting action to the set of respective active warm words of the corresponding user that issued the voice query.
[0005] In some embodiments, the operations also include detecting the presence of a new user within the environment of the AED, and the warm word arbitration routine is executed in response to detecting the presence of the new user within the environment. The operations may optionally include determining that one of the multiple users is no longer present within the environment of the AED, and the warm word arbitration routine is executed in response to determining that one of the multiple users is no longer present within the environment. In some examples, the operations also include determining, for one of the multiple users present within the environment of the AED, to add a new warm word to or remove one of the warm words from a respective set of active warm words, and the warm word arbitration routine is executed in response to determining, for one of the multiple users present within the environment of the AED, to add a new warm word to or remove one of the warm words from a respective set of active warm words. In some embodiments, the operations also include determining a change in ambient context of the AED, and the warm word arbitration routine is executed in response to determining the change in ambient context.
[0006] In some examples, executing the warm word arbitration routine includes obtaining enabled warm word constraints and determining, based on the enabled warm word constraints, a number of warm words to enable in the final set of warm words for detection by the AED. Here, the enabled warm word constraints include at least one of the following: availability of memory and computing resources of the AED for detecting warm words, computational requirements for enabling each warm word in each active warm word set for each user of a plurality of users present in the AED's environment, an acceptable false acceptance rate tolerance, or an acceptable false rejection rate tolerance. In some embodiments, executing the warm word arbitration routine includes, for each corresponding user of the plurality of users, ranking the warm words in each active warm word set from highest priority to lowest priority based on the warm word prioritization signal, and enabling, for each corresponding user of the plurality of users, a final set of warm words for detection by the AED based on the ranking of the warm words in each active warm word set. In some examples, performing the warm word arbitration routine includes identifying, for each of at least two of the plurality of users, any shared warm words that correspond to warm words present in the respective sets of active warm words, and determining that the final set of warm words is based on assigning a higher priority to warm words identified as shared warm words for inclusion in the final set of warm words.
[0007] In some embodiments, executing the warm word reconciliation routine includes determining a warm word affinity score for each user and validating a final set of warm words for detection by the AED based on the warm word affinity score determined for each corresponding user, wherein the warm word affinity score determined for each corresponding user may be based on at least one of the frequency of use of the warm word by the corresponding user, the frequency of interaction between the corresponding user and the digital assistant, the duration of the corresponding user's presence in the AED's environment, the user's proximity to the AED, or the current user context.
[0008] In some examples, the operations further include: determining that the warm words detected in the audio data corresponding to the utterance captured by the user device include a speaker-specific warm word selected from the set of active warm words for each of one or more users among the multiple users, so that the digital assistant performs the respective action specified by the speaker-specific warm word only if the speaker-specific warm word is spoken by one of the one or more users corresponding to the set of active warm words; and performing speaker verification on the audio data to determine that the utterance was spoken by one of the one or more users corresponding to the set of active warm words from which the detected warm word was selected, based on determining that the warm words detected in the audio data corresponding to the utterance captured by the user device include the speaker-specific warm word. Instructing the digital assistant to perform the respective action specified by the detected warm word is based on speaker verification performed on the audio data.
[0009] In some embodiments, the final set of warm words may be enabled for detection by activating and running a respective warm word model on the Assistant-enabled device for each warm word in the final set of warm words, and detecting warm words from the final set of warm words in the audio data may include detecting warm words in the audio data using the respective activated warm word models without performing speech recognition on the audio data. In these embodiments, detecting warm words in the audio data may include extracting audio features of the audio data, processing the extracted audio features using the respective activated warm word models to generate warm word confidence scores, and determining that the audio data corresponding to the utterance contains warm words if the warm word confidence scores satisfy a warm word confidence threshold. The final set of warm words may be enabled for detection by running a speech recognizer on the AED, the speech recognizer being biased to recognize warm words in the final set of warm words, and detecting warm words from the final set of warm words in the audio data may include recognizing repeated warm words in the audio data using the speech recognizer running on the AED.
[0010] Another aspect of the present disclosure provides a system including data processing hardware and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed by the data processing hardware, cause the data processing hardware to perform operations, including detecting the presence of multiple users in an environment of an assistant-enabled device (AED), where the AED runs a digital assistant, and the operations further include obtaining, for each of the multiple users, a respective set of active warm words, each of which specifies a respective action to be performed by the digital assistant. The operations also include executing a warm word arbitration routine to enable a final set of warm words for detection by the AED based on the respective set of active warm words for each of the multiple users. The final set of warm words enabled for detection by the AED includes warm words selected from the respective set of active warm words for at least one of the multiple users detected in the environment of the AED. The operations also include receiving audio data corresponding to speech captured by the AED while the final set of warm words is enabled for detection by the AED, detecting warm words from the final set of warm words in the audio data, and instructing the digital assistant to perform respective actions specified by the detected warm words.
[0011] This aspect may include one or more of the following optional features. In some implementations, detecting the presence of multiple users within the environment of the AED includes detecting that at least one of the multiple users is present within the environment based on proximity information of the AED to a user device associated with at least one of the multiple users. In additional implementations, detecting the presence of multiple users within the environment of the AED includes receiving image data corresponding to a scene of the environment and detecting the presence of at least one of the multiple users within the environment based on the image data. In some examples, detecting the presence of multiple users within the environment of the AED includes detecting the presence of the corresponding users within the environment of the AED based on receiving audio data characterizing voice queries issued by the corresponding users and directed to the digital assistant, performing speaker identification on the audio data to identify the corresponding users who issued the voice queries, and determining that the identified corresponding users who issued the voice queries are present within the environment of the AED. In these examples, the voice query issued by the corresponding user includes a command that causes the digital assistant to perform a long-lasting action specified by the command, and obtaining the set of respective active warm words includes, in response to the digital assistant performing the long-lasting action specified by the command, adding one or more warm words that specify a respective action for controlling the long-lasting action to the set of respective active warm words of the corresponding user that issued the voice query.
[0012] In some embodiments, the operations also include detecting the presence of a new user within the environment of the AED, and the warm word arbitration routine is executed in response to detecting the presence of the new user within the environment. The operations may optionally include determining that one of the multiple users is no longer present within the environment of the AED, and the warm word arbitration routine is executed in response to determining that one of the multiple users is no longer present within the environment. In some examples, the operations also include determining, for one of the multiple users present within the environment of the AED, to add a new warm word to or remove one of the warm words from a respective set of active warm words, and the warm word arbitration routine is executed in response to determining, for one of the multiple users present within the environment of the AED, to add a new warm word to or remove one of the warm words from a respective set of active warm words. In some embodiments, the operations also include determining a change in ambient context of the AED, and the warm word arbitration routine is executed in response to determining the change in ambient context.
[0013] In some examples, executing the warm word arbitration routine includes obtaining enabled warm word constraints and determining, based on the enabled warm word constraints, a number of warm words to enable in the final set of warm words for detection by the AED. Here, the enabled warm word constraints include at least one of the following: availability of memory and computing resources of the AED for warm word detection, computational requirements for enabling each warm word in each active warm word set for each user of a plurality of users present in the AED's environment, an acceptable false acceptance rate tolerance, or an acceptable false rejection rate tolerance. In some embodiments, executing the warm word arbitration routine includes, for each corresponding user of the plurality of users, ranking the warm words in each active warm word set from highest priority to lowest priority based on the warm word prioritization signal, and enabling, for each corresponding user of the plurality of users, a final set of warm words for detection by the AED based on the ranking of the warm words in each active warm word set. In some examples, performing the warm word arbitration routine includes identifying, for each of at least two of the plurality of users, any shared warm words that correspond to warm words present in the respective sets of active warm words, and determining that the final set of warm words is based on assigning a higher priority to warm words identified as shared warm words for inclusion in the final set of warm words.
[0014] In some embodiments, executing the warm word reconciliation routine includes determining a warm word affinity score for each user and validating a final set of warm words for detection by the AED based on the warm word affinity score determined for each corresponding user, wherein the warm word affinity score determined for each corresponding user may be based on at least one of the frequency of use of the warm word by the corresponding user, the frequency of interaction between the corresponding user and the digital assistant, the duration of the corresponding user's presence in the AED's environment, the user's proximity to the AED, or the current user context.
[0015] In some examples, the operations further include: determining that the warm words detected in the audio data corresponding to the utterance captured by the user device include a speaker-specific warm word selected from the set of active warm words for each of one or more users among the multiple users, so that the digital assistant performs the respective action specified by the speaker-specific warm word only if the speaker-specific warm word is spoken by one of the one or more users corresponding to the set of active warm words; and performing speaker verification on the audio data to determine that the utterance was spoken by one of the one or more users corresponding to the set of active warm words from which the detected warm word was selected, based on determining that the warm words detected in the audio data corresponding to the utterance captured by the user device include the speaker-specific warm word. Instructing the digital assistant to perform the respective action specified by the detected warm word is based on speaker verification performed on the audio data.
[0016] In some embodiments, the final set of warm words may be enabled for detection by activating and running a respective warm word model on the assistant-enabled device for each warm word in the final set of warm words, and detecting warm words from the final set of warm words in the audio data may include detecting warm words in the audio data without performing speech recognition on the audio data using the respective activated warm word models. In these embodiments, detecting warm words in the audio data may include extracting audio features of the audio data, processing the extracted audio features using the respective activated warm word models to generate warm word confidence scores, and determining that the audio data corresponding to the utterance contains the warm word if the warm word confidence score satisfies a warm word confidence threshold. The final set of warm words may be enabled for detection by running a speech recognizer on the AED, the speech recognizer being biased to recognize warm words in the final set of warm words, and detecting warm words from the final set of warm words in the audio data may include recognizing repeated warm words in the audio data using the speech recognizer running on the AED.
[0017] The details of one or more embodiments of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will become apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]
[0018] [Figure 1A] FIG. 1 is a schematic diagram of an example system including multiple users with respective sets of active warm words, each specifying a respective action for the digital assistant to perform. [Figure 1B]FIG. 1 is a schematic diagram of an example system including multiple users with respective sets of active warm words, each specifying a respective action for the digital assistant to perform. [Figure 1C] FIG. 1 is a schematic diagram of an example system including multiple users with respective sets of active warm words, each specifying a respective action for the digital assistant to perform. [Figure 2] 1 is an exemplary data store for storing registered user data. [Figure 3] A and B are exemplary graphical user interfaces (GUIs) that are rendered on the screen of a user device. [Figure 4] FIG. 10 is a schematic diagram of an example warm word arbitration process for validating a final set of warm words for detection on an Assistant-enabled device. [Figure 5] FIG. 1 is a schematic diagram of a speaker identification process. [Figure 6] 10 is a flowchart of an example configuration of operations for a method of enabling a final set of warm words for detection at an assistant-enabled device when the presence of multiple users is detected within the environment of the assistant-enabled device. [Figure 7] FIG. 1 is a schematic diagram of an example computing device that may be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION
[0019] Like reference symbols in the various drawings indicate like elements.
[0020] The way a user interacts with an Assistant-enabled device is primarily, but not exclusively, designed through voice input. For example, a user may ask the device to perform an action, including media playback (e.g., music or podcasts), and the device responds by starting to play audio that matches the user's criteria. As a result, an Assistant-enabled device needs some way to identify when any given utterance in the surrounding environment is directed at the device, as opposed to being directed at an individual in the environment or originating from a non-human source (e.g., a television or music player). One way to achieve this is through the use of hot words, reserved by agreement among users in the environment as the predetermined word(s) to be spoken to attract the device's attention. In an exemplary environment, the hot word used to attract the Assistant's attention is the words "OK, computer." As a result, whenever the words "OK, computer" are spoken, they are picked up by the microphone and transmitted to a hot word detector, which performs speech modeling techniques to determine if the hot word was spoken and, if so, waits for the next command or query. Thus, utterances directed to an Assistant-enabled device take the general form [hot word][query], where the "hot word" in this example is "OK computer" and the "query" can be any question, command, declaration, or other request that can be voice-recognized, analyzed, and acted upon by the system, either alone or in conjunction with a server over a network.
[0021] When a user provides several sequences of hotword-based commands to an Assistant-enabled device, such as a mobile phone or smart speaker, the user's interaction with the phone or speaker can be awkward. The user may say, "OK, computer, play the assignments playlist." The phone or speaker may start playing the first song in the playlist. The user wants to skip to the next song and can say, "OK, computer, next." To skip to another song, the user can say, "OK, computer, next" again. To alleviate the need to continually repeat hotwords before speaking commands, the Assistant-enabled device may be configured to recognize / detect a narrow set of hot phrases or warm words that directly trigger the respective action. In this example, the warm word "next" serves two purposes: as a hotword and as a command; therefore, instead of saying "OK, computer, next," the user can simply say "next" to invoke the Assistant-enabled device to trigger the respective action. Other non-limiting warm words and hot phrases may include "what's the weather?", "set a timer," "turn the volume up," and "turn the volume down."
[0022] A set of warm words can be activated to control long-lasting actions. As used herein, a long-lasting action refers to an application or event that the digital assistant runs for an extended period of time and that can be controlled by the user while the application or event is ongoing. For example, if the digital assistant sets a timer for 30 minutes, the timer is a long-lasting action from the time the timer is set to until the timer ends or the resulting alert is acknowledged after the timer ends. In this example, a warm word such as "stop the timer" is activated, allowing the user to stop the timer by simply saying "stop the timer" without first speaking a hot word. Similarly, a command to instruct the digital assistant to play music from a streaming music service is a long-lasting action while the digital assistant is streaming music from the streaming music service through a playback device. In this example, the active set of warm words can be "pause," "pause music," "volume up," "volume down," "next," "previous," etc., to control the playback of music that the digital assistant is streaming through the playback device. A long-running action may involve a multi-step dialogue query such as "make a restaurant reservation," with different sets of warm words active depending on the given stage of the multi-step dialogue. For example, a digital assistant may prompt a user to select from a list of restaurants, which may activate a set of warm words each containing a respective identifier for selecting a restaurant from the list (e.g., the name or number of a restaurant on the list) to complete the action of making a reservation for that restaurant.
[0023] One challenge with warm words is limiting the number of simultaneously active words / phrases to avoid a decrease in quality and efficiency. For example, the number of false positives, which indicates when an Assistant-enabled device incorrectly detects / recognizes one of the active words, increases significantly with the number of simultaneously active warm words. Furthermore, a user who seeds a command to initiate a long-running action cannot prevent others from speaking an active warm word to control the long-running action.
[0024] Because warm words and / or hot phrases can be active (e.g., always on) for extended periods of time, there can often be limitations on how many different words / phrases can be enabled for detection by an assistant-enabled device (AED) at any given time. For example, a computational budget based on computing resource constraints due to the AED's processing power and memory availability can affect the number of different warm words that can be enabled at any given time. Because models for recognizing / detecting warm words and / or hot phrases typically run on a digital signal processor (DSP), careful consideration of the computational budget is especially important for battery-powered devices. Another challenge with warm words is limiting the number of words / phrases that are enabled simultaneously so that quality and efficiency are not compromised. For example, the number of false positives (i.e., also known as the "false alarm rate"), which indicates when an AED incorrectly detects / recognizes one of the warm words, increases significantly when a large number of warm words are simultaneously enabled for detection by the AED.
[0025] Because the total number of different warm words that can be enabled for detection is typically limited by one or more of the factors described above, challenges arise for shared AEDs intended to serve multiple users (e.g., members of a household sharing an Assistant-enabled AED) because different users may have different preferences and may each rely on a different set of warm words to be active. For example, one user may want to issue commands to control music playback on the AED, while another user sharing the same AED may want to issue messaging commands to facilitate communication of messages between that user and some remote recipient. Furthermore, some users may have different tolerance preferences regarding the consequences of detecting a false acceptance (and / or false rejection) of a warm word. For example, some users may tolerate excessive triggering of "stop," while others may not. Ideally, an AED aims to enable many warm words in each of its active warm word sets at any given time, based on the AED's computational budget.
[0026] Embodiments herein are directed to detecting the presence of multiple users within the environment of the AED and obtaining respective sets of active warm words, each of which specifies a respective action to be performed by the digital assistant. Based on the respective sets of active warm words obtained for each user, embodiments further direct executing a warm word arbitration routine to validate a final set of warm words for detection by the AED, whereby the final set of warm words includes warm words selected from the respective sets of active warm words for at least one user of the multiple users detected within the environment of the AED.
[0027] As used herein, "active" warm words include warm words that a user prefers to enable the AED for detection / recognition without the user having to speak a predetermined hot word to "wake up" the AED, while the warm words in the final set of warm words enabled for detection are those that the AED can assertively detect / recognize. Thus, to the extent the AED's computational budget permits, the warm word arbitration routine aims to include all of the active warm words in the final set of warm words that will be enabled for detection by the AED. Otherwise, if the total number of allowable warm words in the final set of warm words is less than the total number of active warm words from each active warm word set, the warm word arbitration routine performs the task of selecting active warm words for inclusion in the final set of warm words that are ranked with a higher priority than those not selected for inclusion in the final set of warm words. As a result, there may be situations and scenarios in which a warm word found in a given user's set of active warm words may not ultimately be included in the final set of warm words and therefore may not be enabled for detection by the AED, resulting in the unselected warm word not being recognized / detected when spoken in an utterance unless the utterance also includes a predefined hot word (e.g., "hey, computer"). Thus, the warm word arbitration routine aims to dynamically select warm words for inclusion in the final set of warm words on an ongoing basis in order to maximize the total number of allowable warm words in the final set of warm words enabled for detection by the AED.
[0028] 1A-1C illustrate an exemplary system 100 for enabling a final set of warm words 112F for detection by an assistant-enabled device (AED) 104, where the final set of warm words 112F includes warm words 112 selected from a respective active set of warm words 112A for at least one user among multiple users detected within the environment of the AED 104. The AED 104 may run one or more digital assistants 105, and each warm word 112 (whether present in one of the active set and / or final set of warm words 112A, 112F) specifies a respective action for at least one of the one or more digital assistants 105 to execute. The user 102 can interact with the digital assistant 105 via voice. For simplicity, examples herein illustrate a state in which the AED 104 runs a single digital assistant 105. However, the present disclosure is not limited to the number of digital assistants 105, and the AED 104 can run any combination of digital assistants simultaneously or individually at any given time.
[0029] In some implementations, for each warm word 112 in the final set of warm words 112F, the AED 104 further receives a respective warm word model 330 configured to detect the corresponding warm word 112 in streaming audio without performing speech recognition. For example, the AED 104 (and / or the server 130) further includes one or more warm word models 330, where the warm word models 330 may be stored in the memory hardware 12 of the AED 104 or in the remote memory hardware 134 of the server 130. If stored in the server 130, the AED 104 may request the server 130 to retrieve the warm word model 330 for the corresponding warm word 112 and provide the retrieved warm word model 330 so that the AED 104 (via the warm word reconciliation routine 401) can validate the warm word model 330. An active warm word model 330 running on the AED 104 may detect utterances of the corresponding warm words 112 in the streaming audio captured by the AED 104 without performing speech recognition of the captured audio. Furthermore, a single warm model 330 may be able to detect all of the warm words 112 from the final set of warm words 112F in the streaming audio.
[0030] In some configurations, the AED 104 receives code associated with an application loaded on the AED 104 (e.g., a music application running in the foreground or background of the AED 104) to identify any warm words 112 that the application's developer wants users 102 to be able to speak to interact with the application, along with associated warm word models 330 and actions corresponding to each warm word 112. In other examples, the AED 104 receives, for at least one warm word 112 in a respective set of active warm words 112A for at least one user 102, a respective warm word model 330 configured to detect the corresponding warm word 112 in streaming audio without performing speech recognition, via a warm word application programming interface (API) running on the AED 104. The warm words 112 in the registry may also relate to follow-up queries that users 102 (or typical users) tend to issue following a given query, such as "Ok computer, play my song playlist."
[0031] In additional embodiments, enabling the final warm word set 112F causes the AED 104 to run the speech recognizer 116 in a low-power and low-fidelity state. Here, the speech recognizer 116 is constrained or biased to recognize only the warm words in the final warm word set 112F when spoken in the speech captured by the AED 104. Because the speech recognizer 116 recognizes only a limited number of terms / phrases, the number of parameters of the speech recognizer 116 can be significantly reduced, thereby reducing memory requirements and the number of calculations required to recognize active warm words in the speech. Therefore, the low-power and low-fidelity characteristics of the speech recognizer 116 may be suitable for implementation on a digital signal processor (DSP). In these embodiments, the speech recognizer 116 running on the AED 104 may recognize the utterance 106 of the warm words 112 in the streaming audio captured by the AED 104 instead of using the warm word model 330. In some examples, the detection of a warm word 112 by a corresponding warm word model 330 is confirmed by a speech recognizer 116 that performs speech recognition on the audio data.
[0032] In the illustrated example, the AED 104 includes a smart speaker. However, the AED 104 can include other computing devices, such as, but not limited to, a smartphone, tablet, smart display, headset, desktop / laptop, smartwatch, smart appliance, headphones, other wearables, or vehicle infotainment devices. The AED 104 includes data processing hardware 10 and memory hardware 12 that stores instructions that, when executed on the data processing hardware 10, cause the data processing hardware 10 to perform operations. The AED 104 includes an array of one or more microphones 16 configured to capture sounds, such as speech, directed at the AED 104. The AED 104 may also include or communicate with an audio output device (e.g., a speaker) 18 that can output audio, such as music 122 and / or synthesized speech, from the digital assistant 105. In some configurations, the AED 104 also includes or is in communication with a display 13 configured to display content from various sources. Additionally, the AED 104 may include or be in communication with one or more cameras 19 configured to capture images within the environment and output image data 412 (FIG. 4).
[0033] In some configurations, the AED 104 communicates with multiple user devices 50 associated with multiple users 102. In the illustrated example, second and third users 102b, 102c each include a respective user device 50, including a smartphone, with which the respective users 102 may interact. However, the user devices 50 may include other computing devices, such as, but not limited to, a smartwatch, a smart display, smart glasses, a smartphone, smart glasses / headset, a tablet, a smart appliance, headphones, a computing device, a smart speaker, or other assistant-enabled devices. Each user device 50 may include at least one microphone 52 present on the user device 50 and in communication with the AED 104. In these configurations, the user device 50 may also communicate with one or more microphones 16 present on the AED 104. Additionally, the multiple users 102 may control and / or configure the AED 104 and may interact with the digital assistant 105 using an interface 200, such as a graphical user interface (GUI) 200, rendered for display on the respective screens of each user device 50.
[0034] 1A-1C and 4, during execution of the digital assistant 105, the AED 104 detects multiple users 102a-102c in the environment using the user detector 410 and obtains a respective set of active warm words 112A for each user 102 detected in the environment. Based on the respective set of active warm words 112A for each of the multiple users 102 detected in the environment of the AED 104, a warm word selector 400 operating on the AED 104 selects a final set of warm words 112F to enable simultaneous detection at the AED 104. For example, the warm word selector 400 receives proximity information 54 (FIG. 4) about the location of each of the multiple users 102a-102c relative to the AED 104 by the user detector 410. In some implementations, for one or more of the users 102 having a respective user device 50, each user device 50 can broadcast proximity information 54 receivable by the user detector 410, which the AED 104 uses to determine the proximity of each user device 50 to the AED 104. The proximity information 54 from each user device 50 may include a wireless communication signal, such as WiFi, Bluetooth, or ultrasound, and the signal strength of the wireless communication signal received by the user detector 410 may correlate the proximity (e.g., distance) of the user device 50 to the AED 104. The proximity information 54 received from each user device 50 may include a device identifier 50 that uniquely identifies the device 50 and can be used by the user detector 410 to resolve the identity of the user 102 using conventional techniques.
[0035] In further embodiments, the user detector 410 automatically detects one or more of the multiple users 102 of the environment by receiving image data 412 corresponding to a scene of the environment and acquired by the camera 19. Here, the user detector 410 detects the multiple users 102 based on the received image data 312. The user detector 410 may anonymously detect the users 102 based on the image data 412 without uniquely identifying the users 102. In some embodiments, the user detector 410 performs facial recognition on the received image data 312 and attempts to uniquely identify each user 102 based on the performed facial recognition. In these embodiments, each user 102 explicitly grants the digital assistant 105 the privilege to perform facial recognition, and each user 102 has the option to revoke the granted privilege at any time.
[0036] Similarly, the user detector 410 may detect multiple users 102 in an environment by performing speaker identification (FIG. 5) and resolving the identities of the users 102 within the environment. Here, the user detector 410 may detect one or more of the multiple users 102 based on received audio data 502 associated with an issued query, and the user detector 410 may continue to detect the user 102 who issued the query (and associated audio data 502) for a threshold time after the user 102 speaks. Notably, the user detector 510 may use any combination of techniques to obtain results that can be correlated to detect the number of different users 102 within the AED 102's environment. Each user device 50 may broadcast a set of current active warm words 112A for the associated user 102, which is received by the warm word selector (and warm word reconciliation routine 401). Similarly, the set of active warm words 112A for one or more of the detected users 102 may be obtained from profile information associated with the detected users once the identities of the users 102 have been resolved.
[0037] In some implementations, the user detector 410 resolves the identities of each of multiple users. In some scenarios, the user 102 is identified as a registered user 200 of the AED 104 and digital assistant 105 who is authorized to access or control various functions of the AED 104. The AED 104 may have multiple different registered users 200, each with a registered user account that indicates specific permissions or rights regarding functions of the AED 104. For example, the AED 104 may operate in a multi-user environment, such as a home with multiple family members, whereby each family member corresponds to a registered user 200 with permissions to access a different respective set of resources. Illustratively, a mother named Barb speaking the command "play a song playlist" would result in the digital assistant 105 streaming music from a rock song playlist associated with the mother. This is in contrast to a different song playlist created and associated with another registered user 200 in the home, such as a teenage daughter, whose playlist includes pop music.
[0038] FIG. 2 illustrates an exemplary data store storing enrollment user data / information for each of multiple enrolled users 200a-n of the AED 104. Here, each enrolled user 200 of the AED 104 may undertake a voice enrollment process to obtain a respective enrollment speaker vector 154 from audio samples of multiple enrollment phrases spoken by the enrolled user 200. For example, a speaker identification model 510 (FIG. 5) may generate one or more enrollment speaker vectors 154 from audio samples of enrollment phrases spoken by each enrolled user 200, which may be combined, e.g., averaged or otherwise accumulated, to form the respective enrollment speaker vector 154. One or more of the enrolled users 200 may perform a voice enrollment process using the AED 104, with the microphone 16 capturing audio samples of these users speaking enrollment utterances, from which the speaker identification model 510 generates the respective enrollment speaker vector 154. The model 510 may execute on the AED 104, the server 120, or a combination thereof. Additionally, one or more of the enrolled users 200 may enroll with the AED 104 by providing authorization and authentication credentials to an existing user account on the AED 104, where the existing user account may store the enrollment speaker vector 154 obtained from a previous voice enrollment process, with other devices also linked to the user account.
[0039] In some examples, the enrollment speaker vector 154 of an enrolled user 200 includes a text-dependent enrollment speaker vector. For example, the text-dependent enrollment speaker vector may be extracted from one or more audio samples of each enrolled user 200 speaking a predetermined term, such as a hot word 110 (e.g., "OK, computer") used to invoke the AED 104 to wake up from a sleep state. In other examples, the enrollment speaker vector 154 of an enrolled user 200 is obtained from one or more audio samples of each enrolled user 200 speaking phrases of different lengths with different terms / words, and is text-independent. In these examples, the text-independent enrollment speaker vector may be obtained over time from audio samples obtained from speech interactions the user 102 has with the AED 104 or other devices linked to the same account.
[0040] Additionally, the AED 104 (and / or the server 120) may optionally store one or more other text-dependent speaker vectors 158 extracted from one or more audio samples of each enrolled user 200 who speaks a particular term or phrase. For example, the enrolled user 200 a may include a respective text-dependent speaker vector 158 for each of one or more warm words 112 that, when detected by the AED 104, may be spoken to cause the AED 104 to perform a respective action to control long-term behavior or to execute some other command. Thus, the text-dependent speaker vector 158 for each enrolled user 200 represents the speech characteristics of each enrolled user 200 who speaks a particular warm word 112. The text-dependent speaker vector 154 stored for each enrolled user 200 associated with a particular warm word 112 may be used to verify each enrolled user 200 who speaks the particular warm word 112 to instruct the AED 104 to perform an action to control long-term behavior.
[0041] FIG. 2 also shows that the AED 104 (and / or server 120) stores warm word preferences 212 for each registered user 200. The warm word preferences 212 may include a list of warm words that the user 102 selects for inclusion in their respective active warm word set 112A. Some of the warm words in each active warm word set 112A may be activity-based, such that the warm words are only “active” depending on context information indicating a current activity associated with the user. For example, the activity-based warm words may include music playback settings (stop, pause, volume up, volume down, etc.) that are only included in the set of active warm words while the AED 104 is streaming and playing music. Here, the context information indicates the current activity of performing the long-term action of streaming music for playback from the AED, and thus the activity-based warm words, each specifying a respective action of the digital assistant 105, cause the long-term action of music playback from the AED 104 to be activated. In other examples, activity-based warm words may include warm words that each specify a respective action for the digital assistant 105 to perform based on a context related to the current application 107 with which the user 102 is currently interacting. For example, the user 102 may interact with a cooking application 107 running on a user device 50 associated with the user 102, and the activity-based warm words may relate to actions that need to be performed for a recipe conveyed by the cooking application (e.g., preheat, set timer) and / or actions 107 to control the cooking application (e.g., next screen) by speaking, allowing the user to navigate the cooking application hands-free.The warm word preference warm word list may include preferred warm words that the user 102 selects for inclusion in their respective active warm word sets 112A, regardless of any activities the user 102 is engaged in. Here, preferred warm words are those included in the active warm word set 112A that the user 102 wants to speak without speaking a predefined hot word to cause the AED 102 to perform the respective action specified by the warm word. Some preferred warm words may be time-sensitive, such that the user 102 defines the period / hour / day during which the warm word is active and should be included in their respective active warm word sets. For example, between 9:00 and 10:00 PM when the user 102 is getting ready for bed, the user 102 may want to speak the warm word "set alarm" to enable the user 102 to set the alarm the next morning without first having to speak a predefined hot word (e.g., hey, computer).
[0042] The warm word preferences 212 stored for each registered user may further include a tolerance range for an acceptable false acceptance rate and / or a tolerance range for an acceptable false rejection rate for that user's active warm word set. These tolerance ranges may determine how sensitive the resulting warm word model is to detecting the presence of warm words in speech. These tolerance ranges may be included in the effective warm word constraints 332 received by the warm word arbitration routine 401 when validating the final warm word set 112F for detection by the AED 104.
[0043] FIG. 1A shows a user 102 speaking a first utterance 106, 106a near an AED 104: "Ok computer, play the song playlist." The microphone 16 of the AED 104 receives the utterance 106 and processes audio data 502 ( FIGS. 4 and 5 ) corresponding to the utterance 106a. Initial processing of the audio data 402 may include filtering the audio data 402 and converting the audio data 502 from an analog signal to a digital signal. Once the AED 104 processes the audio data 502, the AED may store the audio data 502 in a buffer in the memory hardware 12 for further processing. With the audio data 502 in the buffer, the AED 104 may use a hotword detector 108 to detect whether the audio data 402 contains a hotword. The hotword detector 108 is configured to identify hotwords contained in the audio data 502 without performing speech recognition on the audio data 502.
[0044] The hotword detector 108 is configured to identify hotwords in an early portion of the utterance 106. In this example, the hotword detector 108 may determine that the utterance 106, "Ok, computer, play my music playlist," includes the hotword 110, "Ok, computer," if the hotword detector 108 detects acoustic features of the audio data 402 that are characteristic of the hotword 110. The acoustic features may be Mel-Frequency Cepstral Coefficients (MFCCs), which are a representation of the short-term power spectrum of the utterance 106, or may be Mel-scale filter bank energy of the first utterance 106. For example, the hotword detector 108 may detect that the utterance 106, "Ok, computer, play my music playlist," includes the hotword 110, "Ok, computer," based on generating MFCCs from the audio data 402 and classifying the MFCCs as including MFCCs similar to MFCCs that are characteristic of the hotword "Ok, computer" stored in a hotword model of the hotword detector 108. As another example, the hot word detector 108 may detect that the utterance 106 "Ok, computer, play music" contains the hot word 110 "Ok, computer" based on generating mel-scale filter bank energy from the audio data 402 and classifying the mel-scale filter bank energy as containing mel-scale filter bank energy similar to mel-scale filter bank energy characteristic of the hot word "Ok, computer" stored in the hot word model of the hot word detector 108.
[0045] When the hot word detector 108 determines that the audio data 402 corresponding to the utterance 106 includes the hot word 110, the AED 104 may trigger a wake-up process to begin speech recognition on the audio data 402 corresponding to the utterance 106. For example, a speech recognizer 116 executing on the AED 104 may perform speech recognition or semantic interpretation on the audio data 402 corresponding to the utterance 106. The speech recognizer 116 may perform speech recognition on the portion of the audio data 402 that follows the hot word 110. In this example, the speech recognizer 116 may identify the words "play my song playlist" as the command 118 that follows the hot word 110.
[0046] In some implementations, the voice recognizer 116 is located on the server 120 in addition to or instead of the AED 104. When the hot word detector 108 triggers the AED 104 to wake up in response to detecting the hot word 110 in the utterance 106, the AED 104 may transmit audio data 402 corresponding to the utterance 106 to the server 120 over the network 132. The AED 104 may transmit the portion of the audio data 402 including the hot word 110 for the server 120 to verify the presence of the hot word 110. Alternatively, the AED 104 may transmit only the portion of the audio data 402 corresponding to the portion of the utterance 106 after the hot word 110 to the server 120. The server 120 executes the voice recognizer 116 to perform voice recognition and returns a transcription of the audio data 402 to the AED 104. The AED 104 then identifies the words in the utterance 106, and the AED 104 performs semantic interpretation to identify any voice commands. The AED 104 (and / or server 120) may identify a command that causes the digital assistant 105 to perform the "play music" long-running action. In the illustrated example, the digital assistant 105 begins performing the long-running action of playing music 122 as playback audio from the speaker 18 of the AED 104. The digital assistant 105 may stream the music 122 from a streaming service (not shown), or the digital assistant 105 may instruct the AED 104 to play music stored on the AED 104.
[0047] The AED 104 (and / or server 120) may include an action identifier 124 configured to identify one or more long-term actions that the digital assistant 105 is currently performing. For each long-term action that the digital assistant 105 is currently performing, the warm word selector 400, via the warm word reconciliation routine 401, can select one or more corresponding warm word sets 112, each associated with a respective action for controlling the long-term action, for inclusion in the final warm word set 112F. In some examples, the warm word selector 400 accesses the warm word preferences 212 from registered user data / information (e.g., stored in memory hardware 12) of the registered user 200 of the AED 104, or from another registry or table that associates the identified long-term action with one or more corresponding warm word sets 112 that are highly correlated with the long-term action. For example, if the long action corresponds to a function of setting a timer, the associated set of one or more warm words 112 available for the warm word selector 126 to activate includes the warm word 112 "stop timer" to instruct the digital assistant 105 to stop the timer. Similarly, for the long action of "call [contact]," the associated set of warm words 112 includes the warm word(s) 112 of "hang up" and / or "end call" to end an ongoing call. In the illustrated example, for the long action of playing music 122, the associated set of one or more warm words 112 available for the warm word selector 126 to activate includes the warm words 112 "next," "pause," "previous," "volume up," and "volume down," each associated with a respective action for controlling the playback of music 122 from the speaker 18 of the AED 104.Thus, the warm word selector 400 can determine to include these warm words 112 in the final set of warm words 112F enabled for detection by the AED 104 while the digital assistant 105 is performing a long operation, and can disable these warm words 112 when the long operation ends. Similarly, the warm word arbitration routine 401 can enable / disable different warm words 112 depending on the state of the ongoing long operation. For example, if the user speaks "pause" to pause playback of music 122, the warm word arbitration routine 401 can add the warm word 112 for "play" to the final set of warm words 112F that resume playback of music 122. In some configurations, instead of accessing a registry or warm word preferences 212 from the registered user data / information of the registered user 200, the warm word selector 400 examines the code associated with a long-running application (e.g., a music application running in the foreground or background of the AED 104) and identifies any warm words 112 that the application's developer speaks to the user 102 to interact with the application and the respective actions for each warm word 112.
[0048] In some implementations, after adding warm words 112 correlated with long-term actions for inclusion in the final set of warm words 112F enabled for detection, the digital assistant 105 associates these warm words 112 only with the user 102 who spoke the utterance 106 that includes a command 118 for the digital assistant 105 to perform the long-term action. That is, the digital assistant 105 configures some of the warm words 112 in the final set of warm words 112F to be speaker-specific, so that they depend on the speaking voice of the particular user 102 who provided the initial command 118 to initiate the long-term action. As will become apparent, the warm words' 112 dependence on the speaking voice of a particular user 102a causes the AED 104 (e.g., via the digital assistant 105) to only perform each action specified by one of the warm words 112 when the warm word is spoken by a particular user 102, thereby suppressing performance of each action (or at least requiring approval from a particular user 102) when the warm word 112 is spoken by a different speaker.
[0049] As shown in FIG. 5 , in some examples, the user detector 410 resolves the identity of the user 102 who spoke the utterance 106 by executing a speaker identification process 500. The speaker identification process 500 may execute on the data processing hardware 12 of the AED 104. The process 500 may also execute on the server 120. The speaker identification process 500 identifies the user 102 who spoke the utterance 106 by first extracting a first speaker identification vector 511 representing characteristics of the utterance 106 from audio data 502 corresponding to the utterance 106 spoken by the user 102. Here, the speaker identification process 500 may execute a speaker identification model 510 configured to receive the audio data 502 as input and generate the first speaker identification vector 511 as output. The speaker identification model 510 may be a neural network model trained to output the speaker identification vector 511 under machine or human supervision. The speaker identification vector 511 output by the speaker identification model 510 may include an N-dimensional vector having values corresponding to speech features of the utterance 106 associated with the user 102. In some examples, the speaker identification vector 511 is a d-vector.
[0050] Once the first speaker identification vector 511 is output from the model 510, the speaker identification process 500 determines whether the extracted speaker identification vector 511 matches any of the enrollment speaker vectors 154 stored in the AED 104 (e.g., in memory hardware 12) for the enrolled users 200a-n (FIG. 2) of the AED 104. As described above with reference to FIG. 2, the speaker identification model 510 may generate the enrollment speaker vector 154 for the enrolled user 200 during the voice enrollment process. Each enrollment speaker vector 154 may be used as a reference vector corresponding to a voiceprint or unique identifier representative of the voice characteristics of the respective enrolled user 200.
[0051] In some implementations, the speaker identification process 500 uses a comparator 520 that compares the first speaker identification vector 511 with a respective enrollment speaker vector 154 associated with each enrolled user 200a-n of the AED 104. Here, the comparator 520 may generate a score for each comparison that indicates the likelihood that the utterance 106 corresponds to the identity of the respective enrolled user 200, and when the score meets a threshold, the identity is accepted. When the score does not meet the threshold, the comparator may reject the identity. In some implementations, the comparator 520 calculates a respective cosine distance between the first speaker identification vector 511 and each enrollment speaker vector 154 and determines that the first speaker identification vector 511 matches one of the enrollment speaker vectors 154 when the respective cosine distance meets a cosine distance threshold.
[0052] In some examples, the first speaker identification vector 511 is a text-dependent speaker identification vector extracted from a portion of the audio data that includes the hot word 110, and each enrollment speaker vector 154 is also text-dependent on the same hot word 110. Using a text-dependent speaker vector can improve accuracy in determining whether the first speaker identification vector 511 matches any of the enrollment speaker vectors 154. In other examples, the first speaker identification vector 511 is a text-independent speaker identification vector extracted from the entire audio data that includes both the hot word 110 and the command 118, or from a portion of the audio data that includes the command 118.
[0053] When the speaker identification process 500 determines that the first speaker identification vector 511 matches one of the enrollment speaker vectors 154, the process 500 identifies the user 102 who spoke the utterance 106 as the respective enrolled user 200 associated with one of the enrollment speaker vectors 154 that matches the extracted speaker identification vector 511. In the illustrated example, the comparator 520 determines a match based on the respective cosine distances between the first speaker identification vector 511 and the enrollment speaker vector 154 associated with the first enrolled user 200a satisfying a cosine distance threshold. In some scenarios, the comparator 520 identifies the user 102 as the respective first enrolled user 200a associated with the enrollment speaker vector 154 having the shortest respective cosine distance from the first speaker identification vector 511 if the shortest respective cosine distance also meets the cosine distance threshold.
[0054] Conversely, when the speaker identification process 500 determines that the first speaker identification vector 511 does not match any of the enrollment speaker vectors 154, the process 500 may identify the user 102 who spoke the utterance 106 as a guest user of the AED 104. Accordingly, the user detector 410 may associate the set of one or more activated warm words 112 with the guest user and may use the first speaker identification vector 511 as a reference speaker vector representing the speech characteristics of the guest user's voice. In some examples, the guest user may enroll with the AED 104, and the AED 104 may store the first speaker identification vector 511 as the enrollment speaker vector 154 for each of the newly enrolled users.
[0055] 1A , the AED 104 notifies the identified user 102 (e.g., Verb) that warm words 112 are enabled and associated with a set 112A of respective active warm words for controlling long-term actions that the user 102 can speak of any of the warm words 112 to instruct the AED 104 to perform a respective action for controlling long-term actions. For example, the digital assistant 105 may generate synthesized speech 123 for audible output from the speaker 18 of the AED 104 stating, "Verb, you can now control music playback with your voice without having to say, 'OK, computer.'" In an additional example, the digital assistant 105 can provide a notification to a user device 50 (e.g., a smartphone) linked to the identified user's user account to inform the identified user 102 (e.g., Verb) which warm words 112 are currently active for controlling long-term actions.
[0056] As shown in FIG. 3A , a graphical user interface (GUI) 300, 300 a executing on the user device 50 may display valid warm words 112 and associated respective actions for controlling long-term actions. In particular, the GUI 300 a of FIG. 3A illustrates a screen display related to controlling a long-term action initiated by the user 102 a and, therefore, presents only warm words and associated actions related to controlling an ongoing long-term action. Accordingly, additional warm words may be present in each activated warm word set 112A for the user 102 a. For example, FIG. 1A also illustrates the warm words 112 “Call [contact],” “Turn off the lights,” and “Turn on the lights” included in each active warm word set 112A for the user 102 a, but are not associated with the long-term action of playing a playlist of Barb’s songs. Each warm word itself may serve as a descriptor identifying the respective action. 3A presents an exemplary GUI 300a displayed on the screen of the user device 50 that informs the user 102 which warm words 112 are active for the user 102 to speak to control long-term actions, and which warm words 112N are not enabled or are simply inactive / disabled and therefore unavailable to control long-term actions when spoken by the user 102. Specifically, the GUI 300a renders the active warm words 112 "next," "pause," "previous," "volume up," and "volume down," as well as the inactive warm word 112N "play." If the user 102 pauses music playback, the "play" warm word can be the active warm word 112 and the "pause" warm word can be the inactive warm word 112N. Each warm word 112 is associated with a respective action for controlling the playback of music 122 from the speaker 18 of the AED 104.
[0057] Additionally, the GUI 300a may render to display an identifier of the long action (e.g., "Playing Track 1"), an identifier of the AED 104 (e.g., smart speaker) currently performing the long action, and / or the identity of the active user 102 (e.g., Verb) that initiated the long action. In some implementations, the identity of the active user 102 includes an image 304 of the active user 102. Thus, by identifying the active user 102 and the active warm words 112, the GUI 300a identifies the active user 102 as the "controller" of the long action, who can speak any of the active warm words 112 displayed in the GUI 300a to perform a respective action to control the long action. As noted above, the warm words in the set of active warm words 112 associated with controlling the long action may optionally depend on the voice spoken by Verb 102, since Verb 102 seeded the initial command 118 "play music" to initiate the long action. By making the set of active warm words 112 dependent on the voice spoken by Barb 102, the AED 104 (e.g., via the digital assistant 105) will only perform the respective action associated with one of the warm words 112 when the warm word 112 is spoken by Barb 102, and will suppress the performance of the respective action (or at least require approval from Barb 102) when the warm word 112 is spoken by a different speaker.
[0058] The user device 50 may also render graphical elements 302 for display in the GUI 300a to perform each action associated with each active warm word 112 to play the music 122 from the speaker 18 of the AED 104. In the example shown, the graphical elements 302 are associated with playback controls for long duration actions to play the music 122, which, when selected, cause the device 50 to perform the respective action. For example, the graphical elements 302 may include playback controls for performing the action "next" associated with the warm word 112, the action "pause" associated with the warm word 112, the action "previous" associated with the warm word 112, the action "increase volume" associated with the warm word 112, and the action "decrease volume" associated with the warm word 112. The GUI 300a may receive user input instructions via any one of touch, speech, gesture, gaze, and / or an input device (e.g., a mouse or style) to control the playback of the music 122 from the speaker 18 of the AED 104. For example, the user 102 may present a user input indication indicating a selection of a "Next" control (e.g., by touching a graphical button in the GUI 300a that universally represents "Next") to cause the AED 104 to perform an action to advance to the next song in a playlist associated with the music 122.
[0059] Referring to FIG. 3B, a GUI 300, 300b executing on the user device 50 (or AED 104) of a second user 102b (e.g., Jim) displays a warm word configuration screen, allowing the user 102b to select which warm words 112 to add to the user's 102b's respective active warm word sets 112A, 112Ab (FIGS. 1A-1C). In the illustrated example, the GUI 300b displays a list of available warm words 112 as respective graphical elements, allowing the user to select to add or remove corresponding available warm words from each active warm word set 112A that the user wants the AED 104 to detect and ultimately perform the respective action specified by the corresponding warm word when spoken by the user 102b. The GUI 300b allows the user 102b to select activity-based warm words that the user 102b wants to activate by adding them to the respective active warm word sets 112A when the user 102b is performing a particular activity. Here, a particular activity may be identified based on the current application running on the user device 50 or AED 104 with which the user 102b is interacting. For example, the warm word setting screen presented by GUI 300b may present a list of applications loaded on the user device 50 and, for each corresponding application, list available activity-based warm words that the user 102 may select to add to the respective set of active warm words 112A when the user 102 is interacting with the corresponding application. A non-exhaustive list of applications, each with available activity-based warm words, may include a cooking application and a music player application. In the illustrated example, when the user 102 is interacting with the cooking application, the user may select the activity-based warm words "preheat oven" and "set timer," which may be added to the respective set of active warm words 112Ab for the user 102b.1A , a user may view a recipe for roasting boneless chicken in a cooking app, and as the user views the recipe creation steps, the activity-based warm word “preheat oven” may be added to the respective set of active warm words 112Ab. Here, when the user 102b, the AED 104, detects the warm word “preheat oven” captured by the AED 104 from audio data characterizing speech spoken by the user 102b, the digital assistant 105 prompts the smart oven to turn on and preheat. In particular, as described in more detail below, in order for the AED 104 to detect the warm word in the audio data and perform the respective action, the warm word reconciliation routine 401 must affirmatively select the warm word “preheat oven” to be enabled as one of the final set of warm words 112F.
[0060] 3B, FIG. 1B shows that once the smart oven's temperature is successfully preheated, the activity warm word "set timer" may be added to the respective set of active warm words 112Ab. Here, the AED 104 may receive an ambient context signal 440 from the smart oven indicating that the smart oven is at the specified preheat temperature, thereby triggering the AED 104 to add the activity-based warm word "set timer" to, and remove the activity-based warm word "even preheat" from, the respective set of active warm words 112Ab for the user 102b. In particular, as described in more detail below, the addition and removal of the warm words "set timer" and "preheat oven" from the set 112Ab of active warm words for the second user 102b may trigger the warm word selector 400 to re-execute the warm word arbitration routine 401 to determine whether to add the warm word "set timer" to the set 112F of final warm words enabled for detection by the AED 104. In the example shown in FIG. 1B, the AED 104 outputs a notification indicating that the smart oven has successfully preheated. For example, the digital assistant 105 may generate a synthesized voice 123 for audible output from the speaker 18 of the AED 104 stating, "The oven temperature has reached 350 degrees." In a further example, the digital assistant 105 notifies the identified user 102b (e.g., Jim) by providing a notification to a user device 50 (e.g., a smartphone) linked to the identified user's user account when the warm word "preheat the oven" is added by the warm word selector 400 to the final set 112F of warm words enabled for detection by the AED 104.
[0061] The GUI 300b in FIG. 3B also shows options for users to select other activity-based warm words for activities related to the smart doorbell and incoming calls. For example, as shown in FIG. 1B, the smart doorbell may be assigned activity-based warm words such as “show camera,” “ignore,” and “speak” under visitor presence conditions, such as when a visitor rings the smart doorbell and / or when the smart doorbell detects the presence of a visitor in proximity to the smart doorbell. Under one or more visitor presence conditions, these activity-based warm words may be added to the respective active warm word set 112A for each user 102 that selects to include these activity-based warm words, for example, by using the warm word configuration screen presented by the GUI 300b. The smart doorbell may transmit an ambient context signal 440 (FIG. 4) to the AED 104 indicating the occurrence of one or more visitor presence conditions, which may cause the AED 104 to sound a visitor notification 125, such as a doorbell chime, for audible output from the speaker 18 of the AED 104. The AED 104 may provide the visitor notification 125 as a synthesized voice in addition to or instead of the doorbell chime. At the same time, the ambient context signal 440 may trigger the warm word selector 400 to execute a warm word arbitration routine 401 that decides to add the activity-based warm words "show camera," "ignore," and "speak" to the final set 112F of warm words enabled for detection by the AED 104.Thus, in the example of FIG. 1B , any of users 102a-c can say (without speaking a defined hotword) "Show camera" to display visitor image data captured by the smart doorbell's camera or a camera near the smart doorbell on a screen in communication with the AED 104, select "Ignore" to dismiss the visitor notification (optionally having the smart doorbell relay a pre-recorded message to the visitor), or select "Speak" to open the microphone in communication with the smart doorbell and microphone 16 of the AED 104 (and / or microphone 52 of the user device 50) to provide intercom communication functionality between the user 102 and the visitor.
[0062] Continuing with FIG. 3B , the warm word settings screen also shows a list of preferred warm words that the user 102b can select for inclusion in their respective active warm word sets 112A, regardless of activity. In the illustrated example, the user 102b selects the warm words “Call [contact],” “Play music,” and “Turn off the lights,” “Turn on the lights” as preferred warm words to add to their active warm word sets 112A, 112Ab. The list of preferred warm words may be input in advance or may be based on commonly used voice commands that the digital assistant 105 prompts for the warm words and learns over time. Some of the preferred warm words may be custom words or phrases that the user assigns to respective actions (e.g., routines) that are performed when spoken by the user. In the illustrated example, the warm words "What's the weather like" and "Set an alarm" are listed as preferred warm words, but are shown as inactive warm words 112N because user 102b did not choose to add these warm words to their respective list of active warm words 112A. The lists of preferred and activity-based warm words selected for inclusion in each set of active warm words 112A may be stored in each warm word preference 212 for each registered user 200.
[0063] 1A-1C and 4, user devices associated with users 102 detected in the environment by user detector 410 may push / broadcast their respective active warm word sets 112A periodically and / or whenever a warm word is added or removed from the active set. Here, the user devices may push / broadcast their respective active warm word sets 112A from their respective user devices 50 in communication with AED 50. Additionally or alternatively, warm word selector 400 may obtain the active warm word sets 112A by accessing warm word preferences 212 if user detector 410 uniquely identifies the user as one of registered users 200. In some examples, the warm word selector 400 determines when a warm word is added to a different user's set of active warm words 112A in response to receiving an ambient context signal 440 indicating that a particular activity is currently in progress and that the particular activity is assigned to one or more activity-based warm words communicated in the warm word preferences 212. The warm word selector 400 can obtain the set of active warm words 112 using other techniques, such as prompting a detected user to provide a set of active warm words.
[0064] 4 , in some embodiments, the warm word selector 400 executes a warm word arbitration routine 401 to validate a final set of warm words 112F for detection by the AED 104 based on the set of active warm words 112A obtained for each user 102 from among the multiple users 102 detected by the user detector 401. The warm word selector 400 may execute the warm word arbitration routine 401 periodically at predetermined intervals and / or in response to a particular event / condition. In some examples, the warm word arbitration routine 401 executes in response to determining whether to add a new warm word to or remove one of the warm words from the respective set of active warm words 112A for one of multiple different users present in the environment. 1B shows a first user 102a (e.g., Barb) near an AED 104 uttering a second utterance 106, 106b: "Ok, computer, stop the music when I get in the car and start my commute routine." The microphone 16 of the AED 104 receives the utterance 106b, processes the audio data 502, and detects a hot word 110 using a hot word detector 108, thereby waking the AED 104 and performing speech recognition to recognize a command to perform the action of stopping the music and starting Barb's commute routine after she gets in the car. As a result of the command 118 causing the digital assistant 105 to stop its extended operation of playing music 122 from Barb's playlist, the music playback warm words pause, next song, previous song, volume up, and volume down have now been removed from Barb's respective active warm word set 112Aa. In this example, the warm word selector 400 may receive the set of active warm words 112Aa for each verb, now updated to omit the music play warm word, thereby triggering execution of the warm word arbitration routine 401 to update the final set of warm words 112F enabled for detection by the AED.Similarly, adding the activity-based warm words “show camera,” “ignore,” and “speak” to the respective active warm word sets 112Aa-Ac for each of the users 102a-c may trigger the warm word selector 400 to execute the warm word reconciliation routine 401.
[0065] In some examples, the warm word reconciliation routine 401 executes in response to the user detector 410 detecting the presence of a new user within the environment of the AED 104. In particular, as the AED obtains a respective set of active warm words 112A for any detected new users, the warm word reconciliation routine 401 executes to determine whether the final set of warm words 112F should be updated to include any of the active warm words for the new user. By the same concept, updating the final set of warm words 112F may include removing some warm words from the final set to make space (i.e., in terms of processing / memory capacity and / or error tolerance thresholds) for any new active warm words for the new user.
[0066] Additionally or alternatively, the warm word reconciliation routine 401 may execute in response to the AED 104 determining that one of the users previously detected is no longer present within the environment of the AED 104. For example, FIG. 1C shows that the first user 102a is no longer present in the environment. The user detector 410 may continually update a list of users currently detected in the environment of the AED 104, thereby enabling the AED 104 to ascertain when a user is no longer present.
[0067] In some implementations, the AED 104 determines a change in ambient context, and the warm word reconciliation routine 401 executes in response to determining the change in ambient context. In some examples, the warm word selector 400 receives an ambient context signal 440 from a source indicating a change in ambient context. For example, a smart doorbell may provide an ambient context signal 440 indicating the occurrence of one or more visitor presence conditions, or an oven may provide an ambient context signal 440 when the oven reaches preheat temperature. An incoming call to one or more registered users 200 of the AED 104 may also serve as an ambient context signal 440, potentially causing the warm word reconciliation routine 401 to update the final warm word set 112F to include a new warm word associated with the incoming phone call event, such as, but not limited to, a warm word such as "answer" or "ignore," which, when spoken, causes the AED 104 to answer or ignore the call.
[0068] Execution of the warm word reconciliation routine 401 may include obtaining valid warm word constraints 430 and determining, based on the valid warm word constraints 430, the number of warm words to enable in the final warm word set 112F for detection by the AED 104. The warm word constraints 430 may enable the warm word reconciliation routine 401 to determine the warm word capacity of the AED 104 at any given time. The warm word constraints 430 may include at least one of the availability of memory and computing resources of the AED 104 for warm word detection, the computational requirements for enabling each warm word in each active warm word set 112A for each user of multiple users present in the AED's environment, an acceptable false acceptance rate tolerance, or an acceptable false rejection rate tolerance. The acceptable false acceptance and rejection rate tolerances may be based on global settings, device settings, or may be user-defined for each user. In some examples, the tolerance range of acceptable false accept and reject rates may be dynamic in case fading in warm word detection sensitivity occurs when the user is less active or away from the device.
[0069] The availability of memory and computing resources may take into account whether the AED 104 is battery powered, and if so, the current capacity of the battery. The availability of memory and computing resources may also take into account the current warm words in the final set of warm words, the applications currently running on the AED 104, as well as processing power and storage capacity.
[0070] The computational requirements may determine the memory and processing resources needed to execute each warm word model associated with each warm word in the set of active warm words 112A. For example, the computational requirements may include the warm word model size / parameters for each warm word, as well as whether any of the warm words are speaker-specific, which requires performance of the speaker ID process 500. Additionally or alternatively, the computational requirements may determine the memory and processing resources needed to run the speech recognizer 116 in a low-power mode sufficient to only recognize utterances of the warm words. In particular, a low tolerance for an acceptable false acceptance rate may require more processing of each warm word model to detect the warm word than a high tolerance for an acceptable false acceptance rate.
[0071] In some embodiments, executing the warm word arbitration routine 401 includes, for each corresponding user 102 of the multiple users 102, ranking the warm words in each active warm word set 112A from highest priority to lowest priority based on the warm word prioritization signal 413, and enabling a final warm word set 112F for detection by the AED 104 based on the ranking of the warm words in each active warm word set 112A for each corresponding user 102. Here, the warm word prioritization signal 413 may include at least one of: a frequency of use of each warm word in each active warm word set 112A by the corresponding user; a current state of the AED 104; an ambient context of the AED 104; or co-presence information indicating previous warm word use and / or actions performed by the AED when the corresponding user was previously present in an environment with one or more combinations of other users of the multiple users. The current state of the AED may include any long-term actions that the digital assistant is currently performing on the AED 104. For example, if the digital assistant is playing music at the highest volume setting, the prioritization signal 413 may cause the warm word arbitration routine 401 to rank the warm words "play" and "turn up the volume" in the set of active warm words 112A with a lower priority because the user is less likely to speak these warm words. The ambient context may include activity recognition performed by the AED based on one or more signals. For example, FIG. 1A shows that a third user 102c is watching a movie with headphones on and has enabled a do-not-disturb mode on their user device 50.The routine 401 can determine the user context of the user 103c based on image data 412 received from the camera 19 (or other camera), indicating that the user is wearing headphones and looking away from the AED 104, and / or based on a communication signal received from the user device 50 (or headphones) indicating that a "do not disturb" mode is detected on the user device. If enabled, the headphones may also communicate a signal indicating that the headphones are currently being worn and that audio content is playing. The user context of the user 102c may additionally or alternatively be ascertained based on a received signal indicating that a movie is in progress. As a result, the routine 401 can rank the active warm word sets 112A of the user 103c with a lower priority than the active warm word sets 112Aa, 112Ab of the first and second users 102a, 102b. However, as shown in FIG. 1B, the third user 103c may allow visitor alert notifications while the do-not-disturb mode is enabled, such that an ambient context signal 440 indicating a visitor is present at the doorbell causes the digital assistant 105 to output a visitor alert audibly through the user's headphones and / or visually by displaying a graphic on the user device 50 or the television the user 103c is currently watching, notifying the user 102c of the visitor's presence. Thus, while under the conditions of FIG. 1A, each active warm word set 112Ac is ranked with a low priority and is not included in the final warm word set (except for "Call [contact]," which is also active for the other users 102a, 102b), the visitor alert provided to the third user 102c in FIG. 1B may change the current user context of the user 102c and now indicate that the third user 102c may be interested in interacting with the digital assistant 105 to learn more about the doorbell visitor.As shown in FIG. 1B , the warm words “show camera,” “ignore,” and “speak” are included and ranked highest in the active warm word sets 112Aa-Ac of all users 102a-c and are ultimately selected for inclusion in the final warm word set 112F enabled for detection by the AED 104.
[0072] The co-presence information can indicate to what extent a corresponding user utilizes a warm word in the presence of other users. For example, a first user 102a and a second user 102b may each frequently speak the warm word "play music" in each other's presence, whereas the first user 102a rarely speaks the warm word "play music" in the presence of a third user 102c.
[0073] The arbitration routine 401 may select higher priority warm words from the ranked active warm word sets for all of the users 102 for inclusion in the final warm word set 112F to enable detection by the AED 104. As previously described, the number of warm words included in the final warm word set 112F is limited by the valid warm word constraint 430. In some examples, to optimize the performance of the arbitration routine 401, the arbitration routine 401 considers only warm words from the top N warm words in each ranked active warm word set 112A per user. The value of N may be fixed or variable for different users. In examples implementing a variable value of N among different sets of users, the routine 401 may determine a warm word affinity score for each corresponding user 102 from among the multiple users and enable the final warm word set based on the warm word affinity score. The warm word affinity score for each corresponding user 102 may be based on the frequency of use of the warm word by the corresponding user 102 and / or the frequency of interaction between the corresponding user 102 and the digital assistant 105. Here, the frequency of use of a warm word may indicate how often the corresponding user 102 uses the warm word when interacting with the digital assistant 105, while the frequency of interaction may indicate how often the corresponding user 102 generally interacts with the digital assistant 105. The frequency of use of a warm word and / or the frequency of interaction may be further constrained by the current period, particular days of the week, and / or time of day that the presence of the corresponding user 102 is detected in the environment of the AED 104. For example, a user may frequently ask "What's the weather like?" at the same time every morning. The warm word affinity score may additionally or alternatively be based on at least one of the period of time that the corresponding user 102 is present in the environment of the AED 104, the proximity of the corresponding user 102 to the AED 104, or the current user context.For example, depending on the corresponding user's proximity information 54, the arbitration routine 401 may assign a higher priority to warm words in their respective active warm word sets for users who are closer to the AED 104 than for other users because closer users are more likely to interact with the AED than more distant users.
[0074] 1A and 1B, the third user 103c is wearing headphones and watching a movie, and the "do not disturb" mode is enabled on the user device 50. In this case, the current context of the user 102c indicates that the user 103c is unlikely to interact with the digital assistant 103, and the proximity information 58 indicates that the third user 103c is relatively far from the AED 104 compared to the other users. Therefore, the warm word arbitration routine 401 may determine a lower affinity score for the third user 103c than for the other users 103a and 103b. Furthermore, the affinity score of the first user 102a may further increase if the first user 102c subsequently uses one of the music playback warm words after issuing the command 118 "play the song playlist" in the first utterance 106a. Similarly, the affinity score for a second user 102b may increase if the user 102b frequently uses the warm word "preheat oven" while viewing recipes in the cooking application 107, as shown in FIG. 1A. By the same concept, in FIG. 1B, the warm word "set timer" is ranked higher in the respective set of active warm words after the oven reaches preheat temperature because the user 102b is now more likely to place the food in the oven and set the timer for the time the food will bake in the oven. In FIG. 1B, the warm word "preheat oven" is now ranked lowest in the respective set of active warm words 112Ab for a second user 102c.
[0075] To illustrate how the combination of current user context and proximity affects the warm word affinity score of the third user 102b, the current user context and proximity of the third user 102c in FIG. 1A associate the third user 102c with a low affinity score, while the visitor notification 125 event in FIG. 1B associates the second user 102b with a slightly higher affinity score. However, in FIG. 1C, the proximity information 58 now indicates that the third user 102c is approaching the AED 104, and the current user context now indicates that the "do not disturb" mode has been disabled and the user is no longer wearing headphones. As a result, the arbitration routine 401 may increase the affinity score of the third user 102c because the user 102c is more likely to interact with the digital assistant 105 compared to the example shown in FIG. 1B, and particularly the example shown in FIG. 1A.
[0076] In addition to the duration of presence, the arbitration routine 401 may further predict the expected likelihood of the user's presence in the future and determine the affinity score based on the expected likelihood of the user's presence. For example, if the action identifier 124 determines that the transcription output from the speech recognizer 116 for the second utterance 106b includes the command 118, "Once I get in my car, I'll stop the music and start my commute routine," the arbitration routine 401 may predict that the user's 102a's presence in the environment is likely to end. As a result, the arbitration routine may begin gradually decreasing the affinity score of the first user 102a until the first user 102a is no longer detected, as shown in FIG. 1C .
[0077] Continuing with reference to FIGS. 1A-1C and 4, the warm word arbitration routine 401 can increase the score of warm words shared by two or more active warm word sets 112A because these warm words are more likely to be spoken. These shared warm words can be ranked higher within their respective active warm word sets 112A so that they are not accidentally omitted from the top N active warm words, as described above. Furthermore, related warm word models for detecting these shared warm words across multiple users may be merged together to limit the number of similar warm word detection models running on the AED 104. Furthermore, warm words shared by different users can cause the digital assistant 105 to perform different actions depending on which user speaks the warm word 112. For example, in the example of FIG. 1C, when the second user 102b utters "play music," which is effective for detection in the final warm word set, the digital assistant 105 may stream music from a first music player indicated as a preferred music player for the second user 102b. In contrast, when the third user 102c utters the same warm word "play music," the digital assistant 105 may stream music from a different second music player preferred by the third user 102c. In particular, the speaker identification process 500 of FIG. 5 may be performed on the audio data 502 characterizing the utterance of the warm word to uniquely identify the user who uttered the warm word, and appropriate action (e.g., select which music player to stream music on) may be performed by the digital assistant 105. In this example, the warm word reconciliation routine 401 identifies that each of the active warm word sets 112Ab, 112Ac each includes the warm word 112 "play music."Rather than enabling detection of two separate warm word models 330 for the warm word 112 "play music," the warm word reconciliation routine selects only one warm word model 330 for the warm word 112 "play music." In some embodiments, the warm word reconciliation routine 401 determines that the AED 104 has sufficient capacity to enable execution of a higher quality (i.e., additional parameters, additional behaviors, reduced latency, and / or increased sensitivity) warm word model 330 for the warm word 112 "play music." Optionally, the warm word reconciliation routine 401 determines that the AED 104 has sufficient capacity to enable execution of a different architecture, such as transitioning from the warm word model 330 to an ASR model.
[0078] 1C , the first user 102a is no longer detected, and the affinity score for the third user 102c is further amplified because the third user 102c approaches the AED 104 and is no longer wearing headphones. Additionally, the first user's 102a's song playlist has stopped playing. As a result, the final set 112F of warm words for which detection is enabled includes the warm words “tell a joke,” “set a timer,” “play music,” “call [contact],” and “turn off the lights.” The warm word “tell a joke” may be speaker-specific for the third user 102c and may be selected from the respective set of active warm words 112Ac for the third user 102c, while the warm word “set a timer” may be speaker-specific for the second user 102b and may be selected from the respective set of active warm words 112Ab for the second user 102b. The warm words "play music," "call [contact]," and "turn off the lights" are shared warm words that can be uttered by either user 102b, 102c and, when detected in streaming audio by the AED 104 (e.g., using an appropriate warm word model 330 or speech recognizer 116), cause the digital assistant 105 to perform the respective action specified by the warm word without including a predetermined hot word (e.g., "Ok computer"). In an example, a third user 102c utters a third utterance 106, 106c that includes the warm word 112 "tell me a joke" from the final set of warm words 112F enabled for detection by the AED 104. Without performing speech recognition on the captured audio, the AED 104 can apply the warm word model 330 to the final set of warm words 112F to identify whether the utterance 106c includes any of the warm words included in the final set of warm words 112F.The AED 104 compares the audio data 502 corresponding to the utterance 106c with the enabled warm word models 330 corresponding to the warm words 112 “Tell a joke,” “Set a timer,” “Play music,” “Call [contact],” and “Turn off the lights,” and determines that the enabled warm word model 330 for the warm word 112 “Tell a joke” detects the warm word 112 “Tell a joke” in the utterance 102c without performing speech recognition on the audio data 402. Most notably, the user speaks the utterance 106c of the warm word 112 without appending a predetermined hot word 110 to the utterance 106c, and the appropriate warm word model 330 detects the presence of the warm word 112 in the audio data and triggers the AED 104 to invoke the digital assistant 105 to perform the action prescribed by the warm word (e.g., retrieve and play a joke). In some examples, the speech recognizer 116 operates in a low-power mode to recognize only the presence of warm words 112 in the final set of warm words 112F. In these examples, the speech recognizer 116 can recognize when one of the warm words in the final set of warm words 112F is not spoken in the presence of a predetermined hot word and invoke the digital assistant to perform the respective action. In examples, the digital assistant 105 retrieves jokes from a search engine, an on-device application, or other source and audibly outputs the joke as synthesized speech 129. The digital assistant 105 may additionally or alternatively output the joke as a text representation displayed on the screen 13 of the AED 104 or another screen in communication with the AED 104.
[0079] FIG. 6 includes a flowchart of an exemplary arrangement of operations of a method 600 for enabling a final warm word set 112F for detection by an assistant-enabled device (AED) 104 when multiple users 102 are present in the AED's 104 environment. Operations performed by the method 600 may be described with reference to FIGS. 1-5. As used throughout this disclosure, the "environment" of an AED 104 refers to users within proximity to interact with the AED via speech and, optionally, other means. Thus, the "environment" may include the same room or even residence as the AED, a vehicle cabin, the exterior of a vehicle, a doctor's office where the AED is located, a business where the AED is located, and within sufficient proximity to the AED for the AED's microphone to capture audio. The data processing hardware 12 of the AED 104 executes instructions stored in the memory hardware 14, which cause the AED to perform operations. Optionally, a server 130 may execute some or all of the operations.
[0080] At operation 602, the method 600 includes detecting the presence of multiple users 102 within an environment of the AED 104. The AED 104 runs a digital assistant 105. The AED 104 may run multiple digital assistants simultaneously in some configurations. At operation 604, for each user 102 of the multiple users 102, the method also includes a respective set of active warm words 112A, each specifying a respective action for the digital assistant 105 to perform.
[0081] In operation 606, based on the respective active warm word sets for each of the plurality of users, the method 600 includes executing the warm word reconciliation routine 401 to enable a final set of warm words 112F to be detected by the AED 104, where the final set of warm words 112F enabled for detection by the AED 104 includes warm words 112 selected from the respective active warm word sets 112A for at least one user 102 of the plurality of users 102 detected within the environment of the AED 104.
[0082] In operation 608, the final set of warm words is valid for detection by the AED 104, and the method includes receiving audio data 502 corresponding to the utterance 106 captured by the AED 104, detecting warm words 112 from the final set of warm words 112F in the audio data 502, and instructing the digital assistant to perform the respective actions specified by the detected warm words.
[0083] In some examples, the final set of warm words 112F is enabled for detection by activating a respective warm word model 330 to run on the AED for each warm word in the final set of warm words 112F. Here, the method uses each activated warm word model 330 to detect warm words in the audio data 502 without performing speech recognition on the audio data 502. More specifically, detecting warm words in the voice data may include extracting voice features in the audio data 502, using each activated warm word model 330 to process the extracted voice features to generate a warm word confidence score, and determining that the voice data corresponding to the utterance contains the warm word if the warm word confidence score satisfies a warm word confidence threshold. In some examples, the warm word confidence threshold may be identified for a corresponding enrolled user 200 by accessing a tolerance range of an acceptable false acceptance rate and / or a tolerance range of an acceptable false rejection rate stored in the warm word preferences 212 for the enrolled user 200. In a further example, if the final warm word set 112F contains two different warm words that are phonetically similar, the warm word confidence threshold for detecting each of these warm words may be increased to reduce the tendency for false acceptance, in which the warm word model 330 mistakenly detects a warm word in the audio instead of the phonetically similar warm word that was actually spoken.
[0084] 7 is a schematic diagram of an exemplary computing device 700 that can be used to implement the systems and methods described herein. Computing device 700 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown, their connections and relationships, and their functions are for illustrative purposes only and are not intended to limit the implementation of the invention as described and / or claimed herein.
[0085] Computing device 700 includes a processor 710, memory 720, a storage device 730, a high-speed interface / controller 740 connecting to memory 720 and a high-speed expansion port 750, and a low-speed interface / controller 760 connecting to a low-speed bus 770 and storage device 730. Each of the components 710, 720, 730, 740, 750, and 760 are interconnected using various buses and may be mounted on a common motherboard or otherwise as desired. Processor 710 (e.g., data processing hardware 10, 132 of FIG. 1 ) processes instructions for execution in computing device 700, including instructions stored in memory 720 or storage device 730, and can display graphical information for a graphical user interface (GUI) on an external input / output device, such as a display 780 connected to high-speed interface 740. In other embodiments, multiple processors and / or multiple buses may be used, along with multiple memories and multiple types of memories, as desired. Additionally, multiple computing devices 700 may be connected, each providing a portion of the required operations (eg, as a bank of servers, a group of blade servers, or a multi-processor system).
[0086] The memory 720 stores information non-temporarily within the computing device 700. The memory 720 (e.g., memory hardware 12, 134 in FIG. 1) may be a computer-readable medium, a volatile memory unit(s), or a non-volatile memory unit(s). The non-transient memory 720 may be a physical device used to temporarily or permanently store programs (e.g., instruction sequences) or data (e.g., program state information) for use by the computing device 700. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase-change memory (PCM), and disk or tape.
[0087] The storage device 730 can provide mass storage for the computing device 700. In some implementations, the storage device 730 is a computer-readable medium. In various different implementations, the storage device 730 can be a series of devices including a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or a storage area network or other configuration of devices. In further implementations, a computer program product is tangibly embodied on an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as memory 720, the storage device 730, or memory on the processor 710.
[0088] The high-speed controller 740 manages bandwidth-intensive operations of the computing device 700, and the low-speed controller 760 manages less bandwidth-intensive operations. This allocation of roles is merely exemplary. In some implementations, the high-speed controller 740 is coupled to memory 720, a display 780 (e.g., via a graphics processor or accelerator), and a high-speed expansion port 750 that can accept various expansion cards (not shown). In some implementations, the low-speed controller 760 is coupled to a storage device 730 and a low-speed expansion port 790. The low-speed expansion port 790 may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet) and may be coupled to one or more input / output devices such as a keyboard, pointing device, scanner, etc., or to a network device such as a switch or router, for example, via a network adapter.
[0089] As shown, computing device 700 can be implemented in a number of different forms. For example, it may be implemented as a standard server 700a, or multiple times within a group of such servers 700a, as a laptop computer 700b, or as part of a rack server system 700c.
[0090] Various implementations of the systems and techniques described herein may be realized in digital electronic and / or optical circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementation in one or more computer programs executable and / or interpretable by a programmable system including at least one programmable processor. Such programmable processor may be specialized or general-purpose and may be coupled to receive and transmit data and instructions from a storage system, at least one input device, and at least one output device.
[0091] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform a task. In some examples, a software application may be referred to as an "application," "app," or "program." Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0092] Non-transitory memory may be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by a computing device. Non-transitory memory may be volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.
[0093] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language, and / or assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0094] The processes and logic flows described herein may be implemented by one or more programmable processors, also referred to as data processing hardware, executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be implemented by special-purpose logic circuitry, such as an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit). Processors suitable for executing computer programs include, for example, both general-purpose and special-purpose processors, as well as any one or more processors of any kind of digital computer. Generally, a processor receives instructions and data from a read-only memory, a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or is operably connected to receive data from or transmit data to them, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0095] To interact with a user, one or more aspects of the present disclosure can be implemented in a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, or a touch screen, for displaying information to the user, and optionally a keyboard and pointing device, such as a mouse or trackball, through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user. For example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user can be received in any form, such as acoustic input, speech input, or tactile input. Furthermore, a computer can interact with a user by sending documents to and receiving documents from a device used by the user, for example, by sending a web page to a web browser on the user's client device in response to a request received from the web browser.
[0096] Although several embodiments have been described, it will be understood that various modifications can be made without departing from the spirit and scope of the present disclosure. Accordingly, other embodiments are within the scope of the following claims.
Claims
1. A computer-implemented method (600) that, when executed by data processing hardware (710), causes the data processing hardware (710) to perform operations, the operations comprising: detecting the presence of multiple users in an environment of an assistant-enabled device (AED (104)), the AED (104) running a digital assistant (105), the operations further comprising: For each user of the plurality of users, obtain a respective set of active warm words (112), each of which specifies a respective action to be performed by the digital assistant (105); and executing a warm word arbitration routine (401) to enable a final set of warm words (112) for detection by the AED (104) based on the respective set of active warm words (112) for each of the plurality of users, wherein the final set of warm words (112) enabled for detection by the AED (104) includes warm words (112) selected from the respective set of active warm words (112) for at least one user of the plurality of users detected in the environment of the AED (104), the operations further comprising: While the final warm word set (112) is enabled for detection by the AED (104), receiving audio data (402) corresponding to speech (106) captured by the AED (104); detecting warm words (112) from the final set of warm words (112) in the audio data (402); and instructing the digital assistant (105) to perform the respective action specified by the detected warm word (112).
2. 10. The computer-implemented method of claim 1, wherein detecting the presence of the plurality of users within the environment of the AED comprises detecting the presence of the at least one of the plurality of users within the environment based on proximity information of the AED to a user device associated with at least one of the plurality of users.
3. Detecting the presence of the plurality of users within the environment of the AED (104) includes: receiving image data (312) corresponding to a scene of the environment; and detecting the presence of at least one of the plurality of users in the environment based on the image data.
4. Detecting the presence of the plurality of users in the environment of the AED (104) includes: detecting the presence of a corresponding user (102) in the environment of the AED (104); Receiving voice data characterizing a voice query issued by the corresponding user (102) and directed to the digital assistant (105); performing speaker identification on the voice data to identify the corresponding user (102) who issued the voice query; and determining that the identified corresponding user (102) who issued the voice query is present within the environment of the AED (104).
5. the voice query issued by the corresponding user (102) includes a command (118) that causes the digital assistant (105) to perform a long-term action specified by the command (118); 5. The computer-implemented method of claim 4, wherein obtaining the respective set of active warm words includes adding one or more warm words specifying respective actions for controlling the long-term behavior to the respective set of active warm words of the corresponding user who issued the voice query in response to the digital assistant performing the long-term behavior specified by the command.
6. The operation is detecting the presence of a new user within the environment of the AED (104); 6. The computer-implemented method of claim 1, wherein the warm word reconciliation routine is executed in response to detecting the presence of the new user in the environment.
7. The operation is determining that one of the plurality of users is no longer present within the environment of the AED (104); 7. The computer-implemented method of claim 1, wherein the warm word reconciliation routine is executed in response to determining that the one of the plurality of users is no longer present in the environment.
8. The operation is determining, for one of the plurality of users present in the environment of the AED (104), adding a new warm word (112) to the respective set of active warm words (112) or deleting one of the warm words (112) therefrom; 8. The computer-implemented method of claim 1, wherein the warm word reconciliation routine is executed in response to determining, for the one of the plurality of users present in the environment of the AED, the addition of the new warm word to the respective set of active warm words or the removal of the one of the warm words from the respective set of active warm words.
9. The operation is determining a change in the ambient context of the AED (104); 9. The computer-implemented method (600) of claim 1, wherein the warm word reconciliation routine (401) is executed in response to determining the change in surrounding context.
10. Executing the warm word arbitration routine (401) includes: obtaining an enabled warm word constraint (430), the enabled warm word constraint (430) comprising: the availability of memory and computing resources of the AED (104) for the detection of warm words (112); the computational requirements for validating each warm word (112) of the respective set of active warm words (112) for each user of the plurality of users present in the environment of the AED (104); an acceptable range of false acceptance rates, or and at least one of an acceptable false rejection rate tolerance range, and said performing further comprises:
10. The computer-implemented method of claim 1, further comprising determining a number of warm words to enable in the final set of warm words for detection by the AED based on the enabled warm word constraints.
11. Executing the warm word arbitration routine (401) includes: and for each corresponding user (102) of the plurality of users, ranking the warm words (112) in the respective set of active warm words (112) from highest priority to lowest priority based on a warm word prioritization signal (413), wherein the warm word prioritization signal (413) comprises: the frequency of use of each warm word (112) in the respective set of active warm words (112) by the corresponding user (102); the current state of the AED (104); the ambient context of the AED (104); or and co-presence information indicative of previous warm word (112) uses and / or actions performed by the AED (104) when the corresponding user (102) was previously present in the environment with one or more combinations of others of the plurality of users, wherein the performing further comprises:
11. The computer-implemented method of claim 1, further comprising: for each corresponding user of the plurality of users, validating the final set of warm words for detection by the AED based on the ranking of the warm words in the respective set of active warm words.
12. Executing the warm word arbitration routine (401) includes: determining a warm word affinity score for each corresponding user (102) of the plurality of users; and validating the final set of warm words for detection by the AED based on the warm word affinity score determined for each corresponding user.
13. The warm word affinity score determined for each corresponding user (102) is: the frequency of use of the warm words (112) by the corresponding users (102); The frequency of interactions between the corresponding user (102) and the digital assistant (105); the duration of the presence of the corresponding user (102) within the environment of the AED (104); the proximity of the user to the AED (104); or 13. The computer-implemented method (600) of claim 12, based on at least one of: a current user context;
14. Executing the warm word arbitration routine (401) includes: identifying, for each of at least two of the plurality of users, any shared warm words (112) corresponding to warm words (112) present in the respective sets of active warm words (112); determining the final set of warm words based on assigning higher priorities to warm words identified as shared warm words for inclusion in the final set of warm words.
15. The operation is Determining that the warm words (112) detected in the audio data (402) corresponding to the utterance (106) captured by the user device (50) include a speaker-specific warm word (112) selected from the respective active warm word sets (112) for each of the one or more users, such that the digital assistant (105) performs the respective action specified by the speaker-specific warm word (112) only when the speaker-specific warm word (112) is spoken by one or more users among the plurality of users corresponding to the respective active warm word sets (112); and based on determining that the warm words (112) detected in the audio data (402) corresponding to the utterance (106) captured by the user device (50) include speaker-specific warm words (112), performing speaker verification on the audio data (402) to determine that the utterance (106) was spoken by one of the one or more users corresponding to the respective set of active warm words (112) from which the detected warm words (112) were selected; Instructing the digital assistant (105) to perform the respective action specified by the detected warm word (112) is based on the speaker verification performed on the audio data (402). A computer-implemented method (600) according to any one of claims 1 to 14.
16. the final set of warm words (112) is enabled for detection by activating and running on the assistant-enabled device, for each warm word (112) in the final set of warm words (112), a respective warm word model (330); 16. The computer-implemented method of claim 1, wherein detecting the warm words from the final set of warm words in the audio data comprises using the activated respective warm word models to detect the warm words in the audio data without performing speech recognition on the audio data.
17. Detecting the warm word (112) in the audio data (402) comprises: extracting audio features from the audio data (402); generating a warm word confidence score by processing the extracted audio features using each of the activated warm word models (330); and determining that the audio data corresponding to the utterance contains the warm word if the warm word confidence score satisfies a warm word confidence threshold.
18. the final set of warm words (112) is enabled for detection by running a voice recognizer (116) on the AED (104), the voice recognizer (116) being biased to recognize the warm words (112) in the final set of warm words (112); 18. The computer-implemented method of claim 1, wherein detecting the warm words from the final set of warm words in the audio data comprises recognizing the repeated warm words in the audio data using the speech recognizer implemented on the AED.
19. A system (100), comprising: Data processing hardware (710); and memory hardware (720) in communication with the data processing hardware (710), the memory hardware (720) storing instructions that, when executed by the data processing hardware (710), cause the data processing hardware (710) to perform operations, the operations including: detecting the presence of multiple users in an environment of an assistant-enabled device (AED (104)), the AED (104) running a digital assistant (105), the operations further comprising: For each user of the plurality of users, obtain a respective set of active warm words (112), each of which specifies a respective action to be performed by the digital assistant (105); and executing a warm word arbitration routine (401) to enable a final set of warm words (112) for detection by the AED (104) based on the respective set of active warm words (112) for each user of the plurality of users, wherein the final set of warm words (112) enabled for detection by the AED (104) includes warm words (112) selected from the respective set of active warm words (112) for at least one user of the plurality of users detected in the environment of the AED (104), the operations further comprising: While the final warm word set (112) is enabled for detection by the AED (104), receiving audio data (402) corresponding to speech (106) captured by the AED (104); detecting warm words (112) from the final set of warm words (112) in the audio data (402); and instructing the digital assistant (105) to perform the respective action specified by the detected warm word (112).
20. 20. The system of claim 19, wherein detecting the presence of the plurality of users within the environment of the AED includes detecting the presence of the at least one of the plurality of users within the environment based on proximity information of the AED to a user device associated with at least one of the plurality of users.
21. Detecting the presence of the plurality of users within the environment of the AED (104) includes: receiving image data (312) corresponding to a scene of the environment; and detecting the presence of at least one of the plurality of users in the environment based on the image data.
22. Detecting the presence of the plurality of users in the environment of the AED (104) includes: detecting the presence of a corresponding user (102) in the environment of the AED (104); Receiving voice data characterizing a voice query issued by the corresponding user (102) and directed to the digital assistant (105); performing speaker identification on the voice data to identify the corresponding user (102) who issued the voice query; and determining that the identified corresponding user (102) who issued the voice query is present within the environment of the AED (104).
23. the voice query issued by the corresponding user (102) includes a command (118) that causes the digital assistant (105) to perform a long-term action specified by the command (118); 23. The system (100) of claim 22, wherein obtaining the respective set of active warm words (112) includes, in response to the digital assistant (105) performing the long-term action specified by the command (118), adding one or more warm words (112) specifying respective actions for controlling the long-term action to the respective set of active warm words (112) of the corresponding user (102) who issued the voice query.
24. The operation is detecting the presence of a new user within the environment of the AED (104); 24. The system (100) of any of claims 19 to 23, wherein the warm word reconciliation routine (401) is executed in response to detecting the presence of the new user in the environment.
25. The operation is determining that one of the plurality of users is no longer present within the environment of the AED (104); 25. The system (100) of any of claims 19 to 24, wherein the warm word reconciliation routine (401) is executed in response to determining that the one of the plurality of users is no longer present in the environment.
26. The operation is determining, for one of the plurality of users present in the environment of the AED (104), adding a new warm word (112) to the respective set of active warm words (112) or deleting one of the warm words (112) therefrom; 26. The system of claim 19, wherein the warm word reconciliation routine is executed in response to determining, for the one of the plurality of users present in the environment of the AED, the addition of the new warm word to the respective set of active warm words or the removal of the one of the warm words from the respective set of active warm words.
27. The operation is determining a change in the ambient context of the AED (104); 27. The system (100) of any of claims 19 to 26, wherein the warm word reconciliation routine (401) is executed in response to determining the change in ambient context.
28. Executing the warm word arbitration routine (401) includes: obtaining an enabled warm word constraint (430), the enabled warm word constraint (430) comprising: the availability of memory and computing resources of the AED (104) for the detection of warm words (112); the computational requirements for validating each warm word (112) of the respective set of active warm words (112) for each user of the plurality of users present in the environment of the AED (104); an acceptable range of false acceptance rates, or and at least one of an acceptable false rejection rate tolerance range, and said performing further comprises:
28. The system (100) of claim 19, further comprising determining a number of warm words (112) to enable in the final set of warm words (112) for detection by the AED (104) based on the enabled warm word constraints (430).
29. Executing the warm word arbitration routine (401) includes: and for each corresponding user (102) of the plurality of users, ranking the warm words (112) in the respective set of active warm words (112) from highest priority to lowest priority based on a warm word prioritization signal (413), wherein the warm word prioritization signal (413) comprises: the frequency of use of each warm word (112) in the respective set of active warm words (112) by the corresponding user (102); the current state of the AED (104); the ambient context of the AED (104); or and co-presence information indicative of previous warm word (112) uses and / or actions performed by the AED (104) when the corresponding user (102) was previously present in the environment with one or more combinations of others of the plurality of users, wherein the performing further comprises:
29. The system (100) of claim 19, further comprising: for each corresponding user (102) of the plurality of users, validating the final set of warm words (112) for detection by the AED (104) based on the ranking of the warm words (112) in the respective set of active warm words (112).
30. Executing the warm word arbitration routine (401) includes: determining a warm word affinity score for each corresponding user (102) of the plurality of users; and validating the final set of warm words for detection by the AED based on the warm word affinity score determined for each corresponding user.
31. The warm word affinity score determined for each corresponding user (102) is: the frequency of use of the warm words (112) by the corresponding users (102); The frequency of interactions between the corresponding user (102) and the digital assistant (105); the duration of the presence of the corresponding user (102) within the environment of the AED (104); the proximity of the user to the AED (104); or 31. The system (100) of claim 30, based on at least one of: a current user context.
32. Executing the warm word arbitration routine (401) includes: identifying, for each of at least two of the plurality of users, any shared warm words (112) corresponding to warm words (112) present in the respective sets of active warm words (112); and determining the final set of warm words based on assigning a higher priority to warm words identified as shared warm words for inclusion in the final set of warm words.
33. The operation is Determining that the warm words (112) detected in the audio data (402) corresponding to the utterance (106) captured by the user device (50) include a speaker-specific warm word (112) selected from the respective active warm word sets (112) for each of the one or more users, such that the digital assistant (105) performs the respective action specified by the speaker-specific warm word (112) only when the speaker-specific warm word (112) is spoken by one or more users among the plurality of users corresponding to the respective active warm word sets (112); and based on determining that the warm words (112) detected in the audio data (402) corresponding to the utterance (106) captured by the user device (50) include speaker-specific warm words (112), performing speaker verification on the audio data (402) to determine that the utterance (106) was spoken by one of the one or more users corresponding to the respective set of active warm words (112) from which the detected warm words (112) were selected; 33. The system (100) of any of claims 19 to 32, wherein instructing the digital assistant (105) to perform the respective action specified by the detected warm word (112) is based on the speaker verification performed on the audio data (402).
34. the final set of warm words (112) is enabled for detection by activating and running on the assistant-enabled device, for each warm word (112) in the final set of warm words (112), a respective warm word model (330); 34. The system of claim 19, wherein detecting the warm word from the final set of warm words in the audio data includes using the activated respective warm word models to detect the warm word in the audio data without performing speech recognition on the audio data.
35. Detecting the warm word in the audio data includes: extracting audio features from the audio data (402); generating a warm word confidence score by processing the extracted audio features using each of the activated warm word models (330); and determining that the audio data corresponding to the utterance contains the warm word if the warm word confidence score satisfies a warm word confidence threshold.
36. the final set of warm words (112) is enabled for detection by running a voice recognizer (116) on the AED (104), the voice recognizer (116) being biased to recognize the warm words (112) in the final set of warm words (112); 36. The system of claim 19, wherein detecting the warm words from the final set of warm words in the audio data comprises recognizing the repeated warm words in the audio data using the speech recognizer implemented in the AED.
Citation Information
Patent Citations
Speaker Dependent Follow Up Actions And Warm Words
US20220189465A1