Multi-user warm word
By detecting multiple users in an environment supporting assistant device and obtaining their active set of warm words, and executing warm words arbitration routines, the problem of warm words management and coordination in a multi-user environment is solved, and the accurate detection and execution of user commands is achieved.
Patent Information
- Application Number
- CN202380078640.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-17
- Filing Date
- 2023-11-07
- Publication Date
- 2025-06-10
AI Technical Summary
In a multi-user shared support assistant device environment, it is difficult for the prior art to effectively manage and coordinate the warm word collection of multiple users, resulting in the device that may accidentally detect or ignore user commands.
The final warm word set is enabled by detecting the presence of multiple users in an environment supporting the assistant device, obtaining the warm word active set of each user, and performing the warm word arbitration routine for the digital assistant to detect and perform the corresponding actions.
It realizes effective management and coordination of warm words in a multi-user environment, improves the device's accurate detection and execution ability of user commands, and reduces the occurrence of false detection and neglect.
Smart Images

Figure CN120129893A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to multi-user warm words. Background Art
[0002] If not exclusively, the ways in which users interact with the device of the support assistant are mainly designed with the help of voice input. For example, a user may request the device to perform an action including media playback (e.g., music or podcasts), where the device responds by initiating the playback of audio that matches the user's criteria. In the case where the device (e.g., a smart speaker) is shared by multiple users in the environment, the device may need to handle multiple actions that may compete with each other requested by the users. Summary of the Invention
[0003] One aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations including: detecting the presence of multiple users within the environment of a device (AED) that supports an assistant of a digital assistant; and for each of the multiple users, obtaining a respective active set of warm words that each specify a respective action for the digital assistant to perform. Based on the respective active sets of warm words for each of the multiple users, the operations further include: performing a warm word arbitration routine to enable a final set of warm words for detection by the AED. The final set of warm words enabled for detection by the AED includes warm words selected from the respective active sets of warm words for at least one of the multiple users detected within the environment of the AED. When the final set of warm words is enabled for detection by the AED, the operations further include: receiving audio data corresponding to the utterance captured by the AED; detecting a warm word from the final set of warm words in the audio data; and instructing the digital assistant to perform the respective action specified by the detected warm word.
[0004] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, detecting the presence of multiple users within the environment of an AED includes: detecting the presence of at least one of the multiple users within the environment based on proximity information of the AED relative to a user device associated with at least one of the multiple users. In additional implementations, detecting the presence of multiple users within the environment of an AED includes: receiving image data corresponding to a scene of the environment; and detecting the presence of at least one of the multiple users within the environment based on the image data. In some examples, detecting the presence of multiple users within the environment of an AED includes: detecting the presence of a corresponding user within the environment of the AED based on voice data received that characterizes a voice query issued by the corresponding user and directed to a digital assistant; performing speaker identification on the voice data to identify the corresponding user who issued the voice query; and determining that the identified corresponding user who issued the voice query is present within the environment of the AED. In these examples, the voice query issued by the corresponding user includes a command that causes the digital assistant to perform a long-term operation specified by the command, and obtaining the corresponding active set of warm words includes: in response to the digital assistant performing the long-term operation specified by the command, adding one or more warm words that specify corresponding actions for controlling the long-term operation to the corresponding active set of warm words for the corresponding user who issued the voice query.
[0005] In some implementations, the operation further includes: detecting the presence of a new user within the environment of the AED, wherein the warm word arbitration routine is executed in response to detecting the presence of a new user within the environment. The operation may optionally include: determining that one of the multiple users is no longer present within the environment of the AED, wherein the warm word arbitration routine is executed in response to determining that one of the multiple users is no longer present within the environment. In some examples, the operation further includes: determining to add a new warm word to the corresponding active set of warm words for one of the multiple users present within the environment of the AED or remove one of the warm words from the active set of warm words, wherein the warm word arbitration routine is executed in response to determining to add a new warm word to the corresponding active set of warm words for one of the multiple users present within the environment of the AED or remove one of the warm words from the active set of warm words. In some implementations, the operation further includes: determining that the environmental context of the AED has changed, wherein the warm word arbitration routine is executed in response to determining that the environmental context has changed.
[0006] In some examples, performing a warm word arbitration routine includes: obtaining enabled warm word constraints; determining, based on the enabled warm word constraints, the number of warm words to be enabled for detection by an AED in a final set of warm words. Here, the enabled warm word constraints include at least one of the following: availability of memory and computing resources on the AED for detecting warm words, computational requirements for each warm word in a respective active set of warm words enabled for each of a plurality of users present in the environment of the AED, acceptable false acceptance rate tolerance, or acceptable false rejection rate tolerance. In some implementations, performing the warm word arbitration routine includes: for each respective user of the plurality of users, sorting the warm words in the respective active set of warm words from highest priority to lowest priority based on a warm word prioritization signal; and enabling the final set of warm words for detection by the AED based on the sorting of the warm words in the respective active set of warm words for each respective user of the plurality of users. In some examples, performing the warm word arbitration routine includes: identifying any shared warm words corresponding to warm words in the respective active sets of warm words for at least two of the plurality of users; and determining the final set of warm words based on assigning a higher priority to the warm words identified as shared for inclusion in the final set of warm words.
[0007] In some implementations, performing the warm word arbitration routine includes: determining a warm word affinity score for each user; and enabling the final set of warm words for detection by the AED based on the warm word affinity scores determined for each respective user. Here, the warm word affinity scores determined for each respective user can be based on at least one of the following: frequency of use of warm words by the respective user; frequency of interaction between the respective user and the digital assistant; duration of presence of the respective user in the environment of the AED; proximity of the user to the AED; or current user context.
[0008] In some examples, the operation further includes: determining that a warm word detected in audio data corresponding to an utterance captured by a user device includes a speaker-specific warm word selected from the respective active sets of warm words for one or more of the plurality of users, such that the digital assistant performs the respective action specified by the speaker-specific warm word only when the speaker-specific warm word is spoken by any of the one or more users corresponding to the respective active sets of warm words; and performing speaker verification on the audio data based on determining that the warm word detected in the audio data corresponding to the utterance captured by the user device includes a speaker-specific warm word to determine that the utterance is spoken by one of the one or more users corresponding to the respective active sets of warm words from which the detected warm word was selected. Here, instructing the digital assistant to perform the respective action specified by the detected warm word is based on the speaker verification performed on the audio data.
[0009] In some implementations, the final set of warm words is enabled for detection by activating the respective warm word models for each warm word in the final set of warm words to run on a device supporting the assistant, and detecting a warm word from the final set of warm words in the audio data includes: detecting the warm word in the audio data using the activated respective warm word model without performing speech recognition on the audio data. In these implementations, detecting a warm word in the audio data may include: extracting audio features of the audio data; using the activated respective warm word model to generate a warm word confidence score by processing the extracted audio features; and when the warm word confidence score meets a warm word confidence threshold, determining that the audio data corresponding to the utterance includes the warm word. The final set of warm words may be enabled for detection by executing a speech recognizer on an AED, the speech recognizer being biased towards recognizing warm words in the final set of warm words, and detecting a warm word from the final set of warm words in the audio data may include: using the speech recognizer executed on the AED to recognize repeated warm words in the audio data.
[0010] Another aspect of the present disclosure provides a system that includes: data processing hardware and memory hardware communicatively coupled to the data processing hardware. The memory hardware stores instructions that, when executed by the data processing hardware, cause the data processing hardware to perform operations that include: detecting the presence of a plurality of users within an environment of a device (AED) that supports a digital assistant; and for each user of the plurality of users, obtaining a respective active set of warm words that each specify a respective action for the digital assistant to perform. Based on the respective active sets of warm words for each user of the plurality of users, the operations further include: executing a warm word arbitration routine to enable a final set of warm words for detection by the AED. The final set of warm words enabled for detection by the AED includes warm words selected from the respective active sets of warm words for at least one user of the plurality of users detected within the environment of the AED. When the final set of warm words is enabled for detection by the AED, the operations further include: receiving audio data corresponding to an utterance captured by the AED; detecting a warm word from the final set of warm words in the audio data; and instructing the digital assistant to perform the respective action specified by the detected warm word.
[0011] This aspect may include one or more of the following optional features. In some implementations, detecting the presence of multiple users within the environment of the AED includes: detecting the presence of at least one of the multiple users within the environment based on proximity information of the AED relative to a user device associated with at least one of the multiple users. In additional implementations, detecting the presence of multiple users within the environment of the AED includes: receiving image data corresponding to the scene of the environment; and detecting the presence of at least one of the multiple users within the environment based on the image data. In some examples, detecting the presence of multiple users within the environment of the AED includes: detecting the presence of a corresponding user within the environment of the AED based on voice data received that characterizes a voice query issued by the corresponding user and directed to the digital assistant; performing speaker identification on the voice data to identify the corresponding user who issued the voice query; and determining that the identified corresponding user who issued the voice query is present within the environment of the AED. In these examples, the voice query issued by the corresponding user includes a command that causes the digital assistant to perform a long-term operation specified by the command, and obtaining the corresponding active set of warm words includes: in response to the digital assistant performing the long-term operation specified by the command, adding one or more warm words that specify the corresponding actions for controlling the long-term operation to the corresponding active set of warm words for the corresponding user who issued the voice query.
[0012] In some implementations, the operation further includes: detecting the presence of a new user within the environment of the AED, wherein the warm word arbitration routine is executed in response to detecting the presence of a new user within the environment. The operation may optionally include: determining that one of the multiple users is no longer present within the environment of the AED, wherein the warm word arbitration routine is executed in response to determining that one of the multiple users is no longer present within the environment. In some examples, the operation further includes: determining to add a new warm word to the corresponding active set of warm words for one of the multiple users present within the environment of the AED or remove one of the warm words from the active set of warm words, wherein the warm word arbitration routine is executed in response to determining to add a new warm word to the corresponding active set of warm words for one of the multiple users present within the environment of the AED or remove one of the warm words from the active set of warm words. In some implementations, the operation further includes: determining that the environmental scenario of the AED has changed, wherein the warm word arbitration routine is executed in response to determining that the environmental scenario has changed.
[0013] In some examples, performing a warm word arbitration routine includes: obtaining enabled warm word constraints; determining, based on the enabled warm word constraints, the number of warm words to be enabled for detection by an AED in a final set of warm words. Here, the enabled warm word constraints include at least one of the following: availability of memory and computing resources on the AED for detecting warm words, computational requirements for each warm word in a respective active set of warm words enabled for each of a plurality of users present in the environment of the AED, an acceptable false acceptance rate tolerance, or an acceptable false rejection rate tolerance. In some implementations, performing a warm word arbitration routine includes: for each respective user of the plurality of users, sorting the warm words in the respective active set of warm words from highest priority to lowest priority based on a warm word prioritization signal; and enabling a final set of warm words for detection by the AED based on the sorting of the warm words in the respective active set of warm words for each respective user of the plurality of users. In some examples, performing a warm word arbitration routine includes: identifying any shared warm words corresponding to warm words in the respective active sets of warm words for at least two of the plurality of users; and determining the final set of warm words based on assigning a higher priority to the warm words identified as shared for inclusion in the final set of warm words.
[0014] In some implementations, performing a warm word arbitration routine includes: determining a warm word affinity score for each user; and enabling a final set of warm words for detection by the AED based on the warm word affinity scores determined for each respective user. Here, the warm word affinity scores determined for each respective user can be based on at least one of the following: frequency of use of warm words by the respective user; frequency of interaction between the respective user and the digital assistant; duration of presence of the respective user in the environment of the AED; proximity of the user to the AED; or current user context.
[0015] In some examples, the operation further includes: determining that a warm word detected in audio data corresponding to a spoken utterance captured by a user device includes a speaker-specific warm word selected from the respective active sets of warm words for one or more of the plurality of users, such that the digital assistant performs a respective action specified by the speaker-specific warm word only when the speaker-specific warm word is spoken by any of the one or more users corresponding to the respective active sets of warm words; and performing speaker verification on the audio data based on determining that the warm word detected in the audio data corresponding to the spoken utterance captured by the user device includes a speaker-specific warm word to determine that the spoken utterance is spoken by one of the one or more users corresponding to the respective active sets of warm words from which the detected warm word was selected. Here, instructing the digital assistant to perform the respective action specified by the detected warm word is based on the speaker verification performed on the audio data.
[0016] In some implementations, the final set of warm words is enabled for detection by activating a respective warm word model for each warm word in the final set of warm words to run on a device supporting the assistant, and detecting a warm word from the final set of warm words in audio data includes: detecting the warm word in the audio data using the activated respective warm word model without performing speech recognition on the audio data. In these implementations, detecting a warm word in audio data may include: extracting audio features of the audio data; using the activated respective warm word model to generate a warm word confidence score by processing the extracted audio features; and when the warm word confidence score meets a warm word confidence threshold, determining that the audio data corresponding to the utterance includes the warm word. The final set of warm words may be enabled for detection by performing a speech recognizer on an AED, the speech recognizer being biased towards recognizing warm words in the final set of warm words, and detecting a warm word from the final set of warm words in audio data may include: using the speech recognizer performed on the AED to recognize repeated warm words in the audio data.
[0017] Details of one or more implementations of the present disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figures 1A to 1C is a schematic diagram of an example system including multiple users having respective active sets of warm words each specifying a respective action for a digital assistant to perform.
[0019] Figure 2 is an example data repository storing registered user data.
[0020] Figure 3A and Figure 3B is an example graphical user interface (GUI) rendered on a screen of a user device.
[0021] Figure 4 is a schematic diagram of an example warm word arbitration process for enabling a final set of warm words for detection on a device supporting the assistant.
[0022] Figure 5 is a schematic diagram of a speaker recognition process.
[0023] Figure 6 is a flowchart of an example operational arrangement of a method for enabling a final set of warm words for detection on a device supporting the assistant when the presence of multiple users is detected in the environment of the device supporting the assistant.
[0024] Figure 7 is a schematic diagram of an example computing device that may be used to implement the systems and methods described herein.
[0025] In the various figures, like reference numerals indicate like elements. Detailed Description
[0026] If not exclusively, the manner in which a user interacts with a support assistant device is designed mainly by means of voice input. For example, a user may request that the device perform an action including media playback (e.g., music or podcasts), where the device responds by initiating the playback of audio that matches the user's criteria. Thus, the support assistant device must have some way of discerning when any given utterance in the surrounding environment is directed at the device rather than at an individual in the environment, or originates from a non-human source (e.g., a television or music player). One way to achieve this is to use a hotword, which, according to a convention among users in the environment, is a predetermined word reserved to be spoken to arouse the device's attention. In an example environment, the hotword used to arouse the assistant's attention is "OK computer". Thus, whenever the word "OK computer" is spoken, the word is picked up by a microphone and conveyed to a hotword detector, which performs speech modeling techniques to determine whether the hotword has been spoken, and if so, waits for a subsequent command or query. Thus, an utterance directed at the support assistant device takes the general form [HOTWORD] [QUERY], where "HOTWORD" is "OK computer" in this example, and "QUERY" can be any question, command, announcement, or other request that can be speech-recognized, parsed, and acted upon by the system either alone or in combination with a server via a network.
[0027] In cases where a user provides a series of several hotword-based commands to a support assistant's device such as a phone or a smart speaker, the user's interaction with the phone or speaker can become clumsy. The user might say, "Ok computer, play my homework playlist". The phone or speaker might start playing the first song on that playlist. The user might want to advance to the next song and say, "Ok computer, next". To advance to yet another song, the user might say, "Ok computer, next" again. To alleviate the need to constantly repeat the hotword before stating a command, the support assistant's device can be configured to recognize / detect a narrow set of hot phrases or warm words to directly trigger corresponding actions. In this example, the warm word "next" serves the dual purpose of hotword and command, such that the user can simply say "next" to evoke the support assistant's device to trigger the execution of the corresponding action, rather than saying "Ok computer, next". Other non-limiting warm words and hot phrases can include "what’s the weather", "set a timer", "volume up", and "volume down".
[0028] The set of warm words can be active for controlling long - term operations. As used herein, a long - term operation refers to an application or event that a digital assistant performs for an extended duration and an application or event that can be controlled by a user while the application or event is in progress. For example, when a digital assistant sets a timer for 30 minutes, the timer is a long - term operation from the time the timer is set until the timer expires or until a resulting alert is acknowledged after the timer expires. In such a case, warm words such as "stop timer" can be active to allow the user to stop the timer by simply saying "stop timer" without first saying a hot word. Similarly, when the digital assistant is streaming music from a streaming music service via a playback device, a command to instruct the digital assistant to play music from the streaming music service is a long - term operation. In such a case, the active set of warm words can be "pause", "pause music", "volume up", "volume down", "next", "previous", etc. for controlling the playback of the music that the digital assistant is streaming via the playback device. Long - term operations can include multi - step conversation queries such as "book a restaurant", in which different sets of warm words will be active depending on the given stage of the multi - step conversation. For example, the digital assistant can prompt the user to select from a list of restaurants, and a set of warm words each including a corresponding identifier (e.g., the name or number of a restaurant in the list) can become active for selecting a restaurant from the list and completing the action of booking a reservation for that restaurant.
[0029] One challenge with warm words is to limit the number of words / phrases that are active simultaneously so as not to degrade quality and efficiency. For example, the larger the number of warm words that are active simultaneously, the significantly greater the number of false positives indicating when a device supporting the assistant incorrectly detects / identifies one of the active words. Additionally, a user who seeds a command to initiate a long - running operation cannot prevent others from saying the active warm words for controlling that long - running operation.
[0030] Since warm words and / or hot phrases can be active for an extended period of time (e.g., always on), there are typically limitations on how many different words / phrases can be enabled for detection by an assistant-enabled device (AED) at any given time. For example, a computational budget based on computational resource constraints due to processing power and memory availability on the AED may affect how many different warm words can be enabled at any given time. Consideration of the computational budget is particularly important for battery-powered devices because models for identifying / detecting warm words and / or hot phrases typically run on a digital signal processor (DSP). Another challenge with warm words is to limit the number of words / phrases enabled simultaneously so as not to degrade quality and efficiency. For example, the larger the number of warm words enabled simultaneously for detection by the AED, the greater the number of false positives (i.e., also referred to as the "false alarm rate") indicating when the AED incorrectly detects / identifies one of the warm words.
[0031] Since the total number of different warm words that can be enabled for detection is typically limited due to one or more of the factors mentioned above, difficulties arise on a shared AED intended to serve multiple users (e.g., family members sharing an assistant-enabled device) because different users can have different preferences in which different sets of warm words can be active. For example, one user may wish to issue commands for controlling music playback on the AED, while another user sharing the same AED may wish to issue messaging commands for facilitating communication of messages between that user and a remote recipient. Additionally, some users among the users may have different tolerance preferences regarding the consequences of false acceptance (and / or false rejection) detection of warm words. For example, some users may tolerate over-triggering of "stop", while others may not. Ideally, the AED is designed to enable as many warm words as possible in each active set of warm words based on the computational budget of the AED at any given time.
[0032] Implementations herein involve: detecting the presence of multiple users within the environment of the AED, and obtaining corresponding active sets of warm words each specifying a respective action for the digital assistant to perform. Based on the corresponding active sets of warm words obtained for each user, the implementations further involve: performing a warm word arbitration routine to enable a final set of warm words for detection by the AED, where the final set of warm words includes warm words selected from the corresponding active sets of warm words for at least one of the multiple users detected within the environment of the AED.
[0033] As used herein, "active" warm words include warm words that a user would prefer to have the AED enabled to detect / recognize without the user having to say a predefined hot word to "wake up" the AED, and the warm words in the final set of warm words enabled for detection are definitely enabled for the AED to detect / recognize. Thus, as long as the computational budget of the AED permits, the warm word arbitration routine will aim to include all active warm words in the final set of warm words enabled for detection by the AED. Otherwise, when the total number of permitted warm words in the final set of warm words is less than the total number of active warm words from the corresponding active set of warm words, the task of the warm word arbitration routine is to select those active warm words for inclusion in the final set of warm words, those active warm words being ranked with a higher priority than those warm words not selected for inclusion in the final set of warm words. Thus, there can be the following situations and scenarios: A warm word found in the active set of warm words for a given user may ultimately not be included in the final set of warm words and thus not be enabled for detection by the AED, such that an unselected warm word will not be recognized / detected when spoken in a discourse unless the discourse also includes a predefined hot word (e.g., "Hey Computer"). Thus, the warm word arbitration routine aims to continuously and dynamically select warm words for inclusion in the final set of warm words in order to maximize the total number of permitted warm words in the final set of warm words enabled for detection by the AED.
[0034] Figures 1A to 1C Example system 100 illustrating enabling a final set 112F of warm words for detection by an assistant-supporting device (AED) 104, where the final set 112F of warm words includes warm words 112 selected from corresponding active sets 112A of warm words for at least one of a plurality of users detected within the environment of AED 104. AED 104 may execute one or more digital assistants 105, and each warm word 112 (whether present in one of the active set 112A and / or final set 112F of warm words) specifies a corresponding action for at least one of the one or more digital assistants 105 to perform. User 102 may interact with digital assistant 105 via voice. For simplicity, the examples herein depict AED 104 executing a single digital assistant 105. However, the present disclosure is not limited to the number of digital assistants 105 that AED 104 is capable of executing at any given time, such that AED 104 may execute any combination of digital assistants concurrently or individually at any given time.
[0035] In some implementations, for each warm word 112 in the final set 112F of warm words, the AED 104 additionally receives a corresponding warm word model 330, which is configured to detect the corresponding warm word 112 in the streaming audio without performing speech recognition. For example, the AED 104 (and / or the server 130) additionally includes one or more warm word models 330. Here, the warm word model 330 can be stored on the memory hardware 12 of the AED 104 or on the remote memory hardware 134 of the server 130. If stored on the server 130, the AED 104 can request the server 130 to retrieve the warm word model 330 for the corresponding warm word 112 and provide the retrieved warm word model 330, so that the AED 104 (via the warm word arbitration routine 401) can enable the warm word model 330. The enabled warm word model 330 running on the AED 104 can detect the utterance of the corresponding warm word 112 in the captured audio without performing speech recognition on the streaming audio captured by the AED 104. Further, a single warm word model 330 can be capable of detecting all warm words 112 from the final set 112F of warm words in the streaming audio.
[0036] In some configurations, the AED 104 receives code associated with an application loaded on the AED 104 (e.g., a music application running in the foreground or background of the AED 104) to identify any warm words 112 that the developer of the application wants the user 102 to be able to say to interact with the application, the associated warm word models 330, and the corresponding actions for each warm word 112. In other examples, the AED 104 receives a corresponding warm word model 330 for at least one warm word 112 in the corresponding active set 112A of warm words for at least one user 102 via a warm word application programming interface (API) executed on the AED 104, which is configured to detect the corresponding warm word 112 in the streaming audio without performing speech recognition. The warm words 112 in the registry can also be related to the subsequent queries that the user 102 (or a typical user) tends to issue after a given query (e.g., "Ok computer, play my music playlist").
[0037] In additional implementations, enabling the final set 112F of warm words causes the AED 104 to execute the speech recognizer 116 in a low-power and low-fidelity state. Here, when spoken in the utterances captured by the AED 104, the speech recognizer 116 is constrained or biased to recognize only the warm words in the final set 112F of warm words. Since the speech recognizer 116 recognizes only a limited number of terms / phrases, the number of parameters of the speech recognizer 116 can be significantly reduced, thereby reducing the memory requirements and the amount of computation needed to recognize the active warm words in speech. Thus, the low-power and low-fidelity characteristics of the speech recognizer 116 can be suitable for execution on a digital signal processor (DSP). In these implementations, instead of using the warm word model 330, the speech recognizer 116 executed on the AED 104 can recognize the utterances 106 of the warm words 112 in the streaming audio captured by the AED 104. In some examples, the detection of the warm words 112 by the corresponding warm word model 330 is confirmed by the speech recognizer 116 that performs speech recognition on the audio data.
[0038] In the example shown, the AED 104 includes a smart speaker. However, the AED 104 can include other computing devices, such as but not limited to smartphones, tablets, smart displays, head-mounted devices, desktop / laptop computers, smartwatches, smart appliances, headphones, other wearable devices, or vehicle infotainment devices. The AED 104 includes data processing hardware 10 and memory hardware 12 that stores instructions that, when executed on the data processing hardware 10, cause the data processing hardware 10 to perform operations. The AED 104 includes an array of one or more microphones 16 that are configured to capture acoustic sounds, such as speech directed at the AED 104. The AED 104 can also include or communicate with an audio output device (e.g., a speaker) 18 that can output audio, such as music 122 and / or synthetic speech from the digital assistant 105. In some configurations, the AED 104 also includes or communicates with a display 13 that is configured to display content from various sources. Additionally, the AED 104 can include or communicate with one or more cameras 19 that are configured to capture images within the environment and output image data 412 ( Figure 4 ).
[0039] In some configurations, the AED 104 communicates with a plurality of user devices 50 associated with a plurality of users 102. In the example shown, the second user 102b and the third user 102c each include a respective user device 50 that includes a smart phone with which the respective user 102 can interact. However, the user device 50 can include other computing devices such as, but not limited to, smart watches, smart displays, smart glasses, smart phones, smart glasses / head-mounted devices, tablets, smart appliances, head-mounted earphones, computing devices, smart speakers, or another assistant-enabled device. Each user device 50 can include at least one microphone 52 that resides on the user device 50 and communicates with the AED 104. In these configurations, the user device 50 can also communicate with one or more microphones 16 that reside on the AED 104. Additionally, the plurality of users 102 can control and / or configure the AED 104 and interact with the digital assistant 105 using an interface 200 such as a graphical user interface (GUI) 200 that is rendered for display on a respective screen of each user device 50.
[0040] Continuing to refer Figures 1A to 1C and Figure 4 , during the execution of the digital assistant 105, the AED 104 uses the user detector 410 to detect the plurality of users 102a–102c in the environment and obtains a respective active set 112A of warm words for each user 102 detected in the environment. Based on the respective active sets 112A of warm words for each of the plurality of users 102 detected in the environment of the AED 104, the warm word selector 400 running on the AED 104 selects a final set 112F of warm words to enable for simultaneous detection on the AED 104. For example, the warm word selector 400 receives proximity information 54 ( Figure 4 ) of the position of each of the plurality of users 102a–102c relative to the AED 104 via the user detector 410. In some implementations, for one or more of the users 102 having a respective user device 50, the respective user device 50 is capable of broadcasting proximity information 54 that can be received by the user detector 410, and the AED 104 uses the proximity information to determine the proximity of each user device 50 relative to the AED 104. The proximity information 54 from each user device 50 can include a wireless communication signal (such as WiFi, Bluetooth, or ultrasonic), where the signal strength of the wireless communication signal received by the user detector 410 can be related to the proximity (e.g., distance) of the user device 50 relative to the AED 104. The proximity information 54 received from each user device 50 can include a device identifier 50 that uniquely identifies the device 50 and can be used by the user detector 410 to resolve the identity of the user 102 using conventional techniques.
[0041] In additional implementations, the user detector 410 automatically detects one or more of the multiple users 102 in the environment by receiving image data 412 corresponding to the scene of the environment and obtained by the camera 19. Here, the user detector 410 detects the multiple users 102 based on the received image data 312. The user detector 410 can anonymously detect the users 102 based on the image data 412 without uniquely identifying the users 102. In some implementations, the user detector 410 performs face recognition on the received image data 312 and attempts to uniquely identify each user 102 based on the performed face recognition. In these implementations, each user 102 explicitly grants the digital assistant 105 the privilege to perform face recognition, and each user 102 has an option to revoke the granted privilege at any time.
[0042] Similarly, the user detector 410 can detect multiple users 102 in the environment by performing speaker recognition ( Figure 5 ) to solve for the identity of the users 102 within the environment. Here, the user detector 410 can detect one or more of the multiple users 102 based on the received audio data 502 associated with the issued query, and the user detector 410 can continue to detect the user 102 that issued the query (and associated audio data 502) for a threshold amount of time after the user 102 has spoken. It is noted that the user detector 510 can use any combination of techniques to obtain results that can be related to the number of different users 102 detected within the environment of the AED 102. Each user device 50 can broadcast the current active set 112A of wake words for the associated user 102 received by the wake word selector (and wake word arbitration routine 401). Similarly, once the identity of the user 102 is resolved, the active set 112A of wake words for one or more of the detected users 102 can be retrieved from the profile information associated with the detected users.
[0043] In some implementations, the user detector 410 resolves the identity of each of the multiple users. In some scenarios, the user 102 is identified as the registered user 200 of the AED 104, and the registered user is authorized to access or control various functions of the AED 104 and the digital assistant 105. The AED 104 may have multiple different registered users 200 each having a registered user account, and the registered user account indicates specific permissions or rights regarding the functionality of the AED 104. For example, the AED 104 may operate in a multi-user environment (such as a family with multiple family members), where each family member corresponds to a registered user 200 having permissions to access different respective resource sets. By way of illustration, a mother named Barb saying the command "play my music playlist" will cause the digital assistant 105 to stream music from the rock music playlist associated with the mother rather than from a different music playlist created by another registered user 200 of the family (such as a teenage daughter whose playlist includes pop music) and associated with that registered user.
[0044] Figure 2 An example data repository showing the registered user data / information of each of the multiple registered users 200a - 200n of the AED 104 is shown. Here, each registered user 200 of the AED 104 may perform a voice registration process to obtain a corresponding registered speaker vector 154 from audio samples of multiple registration phrases spoken by the registered user 200. For example, the speaker discriminant model 510 ( Figure 5 ) may generate one or more registered speaker vectors 154 from audio samples of the registration phrases spoken by each registered user 200, and the one or more registered speaker vectors may be combined (e.g., averaged or otherwise accumulated) to form the corresponding registered speaker vector 154. One or more of the registered users 200 may use the AED 104 to perform the voice registration process, where the microphone 16 captures audio samples of these users speaking the registration utterances, and the speaker discriminant model 510 generates the corresponding registered speaker vector 154 from the audio samples. The model 510 may be executed on the AED 104, the server 120, or a combination thereof. Additionally, one or more of the registered users 200 may register with the AED 104 by providing authorization and authentication credentials to an existing user account having the AED 104. Here, the existing user account may store the registered speaker vector 154 obtained from a previous voice registration process with another device also linked to the user account.
[0045] In some examples, the enrolled speaker vector 154 of the enrolled user 200 includes a text-related enrolled speaker vector. For example, the text-related enrolled speaker vector can be extracted from one or more audio samples of the corresponding enrolled user 200 who utters a pre-determined term, such as a hotword 110 (e.g., "Ok computer") for invoking the AED 104 to wake up from the dormant state. In other examples, the enrolled speaker vector 154 for the enrolled user 200 is text-independent and is obtained from one or more audio samples of the corresponding enrolled user 200 who utters phrases with different terms / words and different lengths. In these examples, the text-independent enrolled speaker vector can be obtained over time from audio samples that are obtained from the voice interaction of the user 102 with the AED 104 or other devices linked to the same account.
[0046] Additionally, the AED 104 (and / or the server 120) may optionally store one or more other text-related speaker vectors 158 each extracted from one or more audio samples of the corresponding enrolled user 200 who utters a specific term or phrase. For example, the enrolled user 200a may include a corresponding text-related speaker vector 158 for each of one or more warm words 112, which when enabled for detection by the AED 104 can be uttered to cause the AED 104 to perform a corresponding action for controlling long-term operations or execute some other command. Thus, the text-related speaker vector 158 for the corresponding enrolled user 200 represents the voice characteristics of the corresponding enrolled user 200 who utters the specific warm word 112. The text-related speaker vector 154 associated with the specific warm word 112 stored for the corresponding enrolled user 200 can be used to authenticate the corresponding enrolled user 200 who utters the specific warm word 112 to command the AED 104 to perform an action for controlling long-term operations.
[0047] Figure 2Also shown is that the AED 104 (and / or the server 120) stores warm word preferences 212 for the corresponding registered user 200. The warm word preferences 212 can include a list of warm words selected by the user 102 for inclusion in the respective active set 112A of warm words. Some of the warm words in the respective active set 112A of warm words can be activity-based, such that the warm words are “active” only based on the context information indicating the current activity associated with the user. For example, the activity-based warm words can include music playback settings (e.g., stop, pause, volume up, volume down, etc.), which are included in the active set of warm words only when the AED 104 is streaming music for playback. Here, the context information will indicate the current activity of performing the long-term operation of streaming music from the AED for playback, and thus will cause the activation of the activity-based warm words that each specify the corresponding actions for the digital assistant 105 to control the music playback from the AED 104. In another example, the activity-based warm words can include warm words that each specify the corresponding actions for the digital assistant 105 to perform based on the context related to the current application 107 that the user 102 is currently interacting with. For example, the user 102 may be interacting with a cooking application 107 executing on the user device 50 associated with the user 102, and the activity-based warm words can be related to the actions required to be performed for the recipe communicated by the cooking application (e.g., preheat the oven, set the timer) and / or the actions for controlling the cooking application 107 by voice (e.g., next screen) such that the user can navigate the cooking application in a hands-free manner. The list of warm words in the warm word preferences can include preferred warm words selected by the user 102 for inclusion in the respective active set 112A of warm words, which are not related to any activity that the user 102 is performing. Here, the preferred warm words are those warm words included in the active set 112A of warm words that the user 102 expects to say without saying the predefined hot word to cause the AED 102 to perform the corresponding actions specified by the warm words. Some of the preferred warm words can be time-sensitive such that the user 102 defines the period / time / date during which the warm word should be active and is included in the respective active set of warm words, such that the user 102. For example, between 9 pm and 10 pm, when the user 102 is getting ready to sleep, the user 102 may want to say the warm word “set alarm” to allow the user 102 to set his / her alarm for the next morning without having to first say the predefined hot word (e.g., HeyComputer).
[0048] The warm word preference 212 stored for each respective registered user may further include an acceptable false acceptance rate tolerance and / or an acceptable false rejection rate tolerance for the active set of warm words for him / her. These tolerances may specify the sensitivity of the resulting warm word model for detecting the presence of warm words in speech. When the final set 112F of warm words is enabled for detection by the AED 104, these tolerances may be included within the enabled warm word constraints 332 received by the warm word arbitration routine 401.
[0049] Figure 1A FIG. shows user 102 speaking first utterances 106, 106a near the AED 104: "Ok computer, play my music playlist". The microphone 16 of the AED 104 receives the utterance 106 and processes the audio data 502 corresponding to the utterance 106a ( Figure 4 and Figure 5 ). Initial processing of the audio data 402 may involve filtering the audio data 402 and converting the audio data 502 from an analog signal to a digital signal. When the AED 104 processes the audio data 502, the AED may store the audio data 502 in a buffer of the memory hardware 12 for additional processing. With the audio data 502 in the buffer, the AED 104 may use the hot word detector 108 to detect whether the audio data 402 includes a hot word. The hot word detector 108 is configured to identify hot words included in the audio data 502 without performing speech recognition on the audio data 502.
[0050] The hotword detector 108 is configured to identify hotwords in an initial portion of the utterance 106. In this example, if the hotword detector 108 detects an acoustic feature that is characteristic of the hotword 110 in the audio data 402, the hotword detector 108 may determine that the utterance 106 "Ok computer, play my music playlist" includes the hotword 110 "ok computer". The acoustic feature may be a Mel-frequency cepstral coefficient (MFCC) that is a representation of a short-term power spectrum of the utterance 106, or may be a Mel-scale filter bank energy of the utterance 106. For example, based on generating MFCCs from the audio data 402 and classifying that the MFCCs include MFCCs that are similar to MFCCs that are characteristic of the hotword "okcomputer" as stored in the hotword model of the hotword detector 108, the hotword detector 108 may detect that the utterance 106 "Ok computer, play music" includes the hotword 110 "ok computer". As another example, based on generating Mel-scale filter bank energies from the audio data 402 and determining that the Mel-scale filter bank energies include Mel-scale filter bank energies that are similar to the Mel-scale filter bank energies that are characteristic of the hot word “ok computer” as stored in the hot word model of the hot word detector 108, the hot word detector 108 can detect that the utterance 106 “Ok computer, play music” includes the hot word 110 “ok computer”.
[0051] When the hotword detector 108 determines that the audio data 402 corresponding to the utterance 106 includes the hotword 110, the AED 104 can trigger a wake-up process to initiate speech recognition of the audio data 402 corresponding to the utterance 106. For example, a speech recognizer 116 running on the AED 104 can perform speech recognition or semantic interpretation on the audio data 402 corresponding to the utterance 106. The speech recognizer 116 can perform speech recognition on the portion of the audio data 402 following the hotword 110. In this example, the speech recognizer 116 can identify the words "play my music playlist" as a command 118 following the hotword 110.
[0052] In some implementations, the speech recognizer 116 is located on the server 120, as a supplement to or replacement for the AED 104. When the hot word detector 108 triggers the AED 104 to wake up in response to detecting the hot word 110 in the utterance 106, the AED 104 can transmit the audio data 402 corresponding to the utterance 106 to the server 120 via the network 132. The AED 104 can transmit the portion of the audio data 402 that includes the hot word 110 for the server 120 to confirm the presence of the hot word 110. Alternatively, the AED 104 can transmit only the portion of the audio data 402 corresponding to the portion of the utterance 106 that comes after the hot word 110 to the server 120. The server 120 executes the speech recognizer 116 to perform speech recognition and returns a transcription of the audio data 402 to the AED 104. Subsequently, the AED 104 identifies the words in the utterance 106, and the AED 104 performs semantic interpretation and identifies any voice commands. The AED 104 (and / or the server 120) can identify a command to have the digital assistant 105 perform a long-term operation of "play music". In the example shown, the digital assistant 105 begins performing the long-term operation of playing the music 122 as playback audio from the speaker 18 of the AED 104. The digital assistant 105 can stream the music 122 from a streaming service (not shown), or the digital assistant 105 can instruct the AED 104 to play music stored on the AED 104.
[0053] The AED 104 (and / or the server 120) may include an operation recognizer 124 configured to recognize one or more long-running operations that the digital assistant 105 is currently performing. For each long-running operation that the digital assistant 105 is currently performing, the hotword selector 400 may select, via the hotword arbitration routine 401, a corresponding set of one or more hotwords 112 each associated with a respective action for controlling the long-running operation for inclusion in the final set of hotwords 112F. In some examples, the hotword selector 400 accesses the hotword preferences 212 from the registered user data / information of the registered user 200 of the AED 104 (e.g., stored on the memory hardware 12) or another registry or table associating the identified long-running operation with a corresponding set of one or more hotwords 112 highly relevant to the long-running operation. For example, if the long-running operation corresponds to the set timer function, the associated set of one or more hotwords 112 available for the hotword selector 126 to activate includes the hotword 112 "stop timer" for instructing the digital assistant 105 to stop the timer. Similarly, for the long-running operation of "Call [contact name]", the associated set of hotwords 112 includes the hotwords 112 "hang up" and / or "end call" for ending an ongoing call. In the example shown, for the long-running operation of playing music 122, the associated set of one or more hotwords 112 available for the hotword selector 126 to activate includes the hotwords 112 "next", "pause", "previous", "volume up", and "volume down" each associated with a respective action for controlling the playback of the music 122 from the speaker 18 of the AED 104. Thus, the hotword selector 400 may determine to include these hotwords 112 in the final set of hotwords 112F enabled for detection by the AED 104 while the digital assistant 105 is performing the long-running operation and may disable these hotwords 112 once the long-running operation ends. Similarly, the hotword arbitration routine 401 may enable / disable different hotwords 112 based on the status of the ongoing long-running operation. For example, if the user says "pause" to pause the playback of the music 122, the hotword arbitration routine 401 may add the hotword 112 for "play" to the final set of hotwords 112F to resume the playback of the music 122. In some configurations, instead of accessing the registry or hotword preferences 212 from the registered user data / information of the registered user 200, the hotword selector 400 examines the code associated with the application of the long-running operation (e.g., a music application running in the foreground or background of the AED 104) to identify any hotwords 112 that the developer of the application intended for the user 102 to be able to say to interact with the application and the corresponding actions for each hotword 112.
[0054] In some implementations, after adding warm words 112 related to long-term operations for inclusion in the final set 112F of warm words enabled for detection, the digital assistant 105 associates these warm words 112 only with the user 102 who utters a speech 106 having a command 118 for the digital assistant 105 to perform a long-term operation. That is, the digital assistant 105 configures a portion of the warm words 112 in the final set 112F of warm words to be speaker-specific, such that they depend on the speech voice of the particular user 102 who provides the initial command 118 to initiate the long-term operation. As will become apparent, by making the warm words 112 depend on the speech voice of a particular user 102a, the AED 104 (e.g., via the digital assistant 105) will perform the corresponding action specified by the warm word only when one of the warm words 112 is uttered by the particular user 102, and thus suppress the execution of the corresponding action (or at least require approval from the particular user 102) when the warm word 112 is uttered by a different speaker.
[0055] Reference Figure 5 , in some examples, the user detector 410 resolves the identity of the user 102 who utters the speech 106 by performing a speaker recognition process 500. The speaker recognition process 500 can be executed on the data processing hardware 12 of the AED 104. The process 500 can also be executed on the server 120. The speaker recognition process 500 identifies the user 102 who utters the speech 106 by first extracting a first speaker discriminant vector 511 representing the characteristics of the speech 106 from the audio data 502 corresponding to the speech 106 uttered by the user 102. Here, the speaker recognition process 500 can execute a speaker discriminant model 510, which is configured to receive the audio data 502 as input and generate the first speaker discriminant vector 511 as output. The speaker discriminant model 510 can be a neural network model trained under machine or human supervision to output the speaker discriminant vector 511. The speaker discriminant vector 511 output by the speaker discriminant model 510 can include an N-dimensional vector having values corresponding to the speech characteristics of the speech 106 associated with the user 102. In some examples, the speaker discriminant vector 511 is a d-vector.
[0056] Once the first speaker discriminant vector 511 is output from the model 510, the speaker recognition process 500 determines whether the extracted speaker discriminant vector 511 matches any of the registered speaker vectors 154 stored on the AED 104 (e.g., in the memory hardware 12) for the registered users 200a–200n of the AED 104 ( Figure 2 )). As referenced above Figure 2As described, the speaker discriminator model 510 can generate an enrollment speaker vector 154 for the enrolled user 200 during the voice enrollment process. Each enrollment speaker vector 154 can be used as a reference vector corresponding to a voiceprint or unique identifier representing the characteristics of the voice of the corresponding enrolled user 200.
[0057] In some implementations, the speaker identification process 500 uses a comparator 520 that compares the first speaker discriminator vector 511 with the corresponding enrollment speaker vectors 154 associated with each of the enrolled users 200a–200n of the AED 104. Here, the comparator 520 can generate a score for each comparison that indicates the likelihood that the utterance 106 corresponds to the identity of the corresponding enrolled user 200, and accepts the identity when the score meets a threshold. When the score does not meet the threshold, the comparator can reject the identity. In some implementations, the comparator 520 calculates the corresponding cosine distance between the first speaker discriminator vector 511 and each enrollment speaker vector 154, and determines that the first speaker discriminator vector 511 matches one of the enrollment speaker vectors in the enrollment speaker vector 154 when the corresponding cosine distance meets the cosine distance threshold.
[0058] In some examples, the first speaker discriminator vector 511 is a text-dependent speaker discriminator vector extracted from a portion of the audio data that includes the hotword 110, and each enrollment speaker vector 154 is also text-dependent on the same hotword 110. The use of text-dependent speaker vectors can improve the accuracy in determining whether the first speaker discriminator vector 511 matches any of the enrollment speaker vectors in the enrollment speaker vector 154. In other examples, the first speaker discriminator vector 511 is a text-independent speaker discriminator vector extracted from the entire audio data that includes both the hotword 110 and the command 118 or from a portion of the audio data that includes the command 118.
[0059] When the speaker recognition process 500 determines that the first speaker discriminant vector 511 matches one of the enrolled speaker vectors 154 in the enrolled speaker vector 154, the process 500 identifies the user 102 who uttered the utterance 106 as the corresponding enrolled user 200 associated with the enrolled speaker vector in the enrolled speaker vector 154 that matches the extracted speaker discriminant vector 511. In the example shown, the comparator 520 determines a match based on the corresponding cosine distance between the first speaker discriminant vector 511 and the enrolled speaker vector 154 associated with the first enrolled user 200a satisfying the cosine distance threshold. In some scenarios, the comparator 520 identifies the user 102 as the corresponding first enrolled user 200a associated with the enrolled speaker vector 154 having the shortest corresponding cosine distance from the first speaker discriminant vector 511, provided that the shortest corresponding cosine distance also satisfies the cosine distance threshold.
[0060] Conversely, when the speaker recognition process 500 determines that the first speaker discriminant vector 511 does not match any of the enrolled speaker vectors 154 in the enrolled speaker vector 154, the process 500 may identify the user 102 who uttered the utterance 106 as a guest user of the AED 104. Thus, the user detector 410 may associate the set of activated one or more warm words 112 with the guest user and use the first speaker discriminant vector 511 as a reference speaker vector representing the voice characteristics of the voice of the guest user. In some cases, the guest user may enroll with the AED 104, and the AED 104 may store the first speaker discriminant vector 511 as the corresponding enrolled speaker vector 154 for the newly enrolled user.
[0061] In Figure 1A the example shown, the AED 104 notifies the identified user 102 (e.g., Barb) associated with the corresponding active set 112A of warm words for controlling long-term operations that the warm words 112 are enabled and the user 102 may speak any of the warm words 112 to instruct the AED 104 to perform the corresponding actions for controlling long-term operations. For example, the digital assistant 105 may generate a synthetic voice 123 for audible output from the speaker 18 of the AED 104, the synthetic voice stating: "Barb, you may speak music playback controls without saying 'Ok computer'". In an additional example, the digital assistant 105 may provide a notification to the user device 50 (e.g., a smart phone) linked to the user account of the identified user to inform the identified user 102 (e.g., Barb) of which warm words 112 are currently active for controlling long-term operations.
[0062] Reference Figure 3A , on the user device 50, the graphical user interfaces (GUIs) 300, 300a that are executed can display the enabled warm words 112 and the associated corresponding actions for controlling long-term operations. It is worth noting that Figure 3A the GUI 300a depicts the presentation of a screen related to controlling a long-term operation initiated by the user 102a, and thus, only presents the warm words and the associated actions related to the control of the ongoing long-term operation. Therefore, there may be additional warm words in the corresponding activated set 112A of warm words for the user 102. For example, Figure 1A also shown are the warm words 112 "Call [Contact]", "Turn off lights", and "Turn on lights" included in the corresponding active set 112A of warm words for the user 102a, although the warm words are not related to the long-term operation of the playback of Barb's music playlist. Each warm word itself can be used as a descriptor for identifying the corresponding action. Figure 3A An example GUI 300a is provided that is displayed on the screen of the user device 50. The example GUI is used to inform the user 102 which warm words 112 are active for the user 102 to speak to control long-term operations, and which warm words 112N are not enabled or are merely inactive / disabled and thus not available for controlling long-term operations when spoken by the user 102. Specifically, the GUI 300a renders the active warm words 112 "next", "pause", "previous", "volume up", and "volume down" and the unenabled warm word 112N "play". If the user 102 pauses the music playback, the warm word for "play" can become an active warm word 112, and the warm word for "pause" can become an inactive warm word 112N. Each warm word 112 is associated with a corresponding action for controlling the playback of the music 122 from the speaker 18 of the AED 104.
[0063] Additionally, the GUI 300a can render an identifier of a long-term operation (e.g., "Playing Track 1"), an identifier of the AED 104 that is currently performing the long-term operation (e.g., a smart speaker), and / or the identity of the active user 102 who initiated the long-term operation (e.g., Barb) for display. In some implementations, the identity of the active user 102 includes an image 304 of the active user 102. Thus, by identifying the active user 102 and the active hotword 112, the GUI 300a reveals the active user 102 as the "controller" of the long-term operation, and the "controller" can speak any of the active hotwords 112 displayed in the GUI 300a to perform the corresponding action for controlling the long-term operation. As mentioned above, the hotwords in the active set of hotwords 112 that are relevant to controlling the long-term operation can optionally depend on the speaking voice of Barb 102, since Barb 102 issued the initial command 118 "play music" to initiate the long-term operation. By making the active set of hotwords 112 depend on the speaking voice of Barb 102, the AED 104 (e.g., via the digital assistant 105) will perform the corresponding action associated with that hotword 112 only when one of the hotwords 112 is spoken by Barb 102, and will inhibit the execution of the corresponding action (or at least require approval from Barb 102) when the hotword 112 is spoken by a different speaker.
[0064] The user device 50 may also render a graphical element 302 for display in the GUI 300a, the rendered graphical element being used to perform a corresponding action associated with the corresponding active warm word 112 to play back music 122 from the speaker 18 of the AED 104. In the example shown, the graphical element 302 is associated with playback controls for the long-term operation of playing back music 122, which, when selected, causes the device 50 to perform a corresponding action. For example, the graphical element 302 may include playback controls for: performing an action associated with the warm word 112 "next", performing an action associated with the warm word 112 "pause", performing an action associated with the warm word 112 "previous", performing an action associated with the warm word 112 "volume up", and performing an action associated with the warm word 112 "volume down". The GUI 300a may receive a user input indication via any one of touch, voice, gesture, gaze, and / or an input device (e.g., a mouse or stylus) to control the playback of music 122 from the speaker 18 of the AED 104. For example, the user 102 may provide a user input indication that selects the "next" control (e.g., by touching a graphical button in the GUI 300a that generally represents "next") to cause the AED 104 to perform the action of advancing to the next song in the playlist associated with the music 122.
[0065] Reference Figure 3B , on the user devices 50 (or AED 104) of a second user 102b (e.g., Jim), the GUIs 300, 300b that are executed display a warm word configurator screen that allows the user 102b to select which warm words 112 to add to the respective active sets 112A, 112Ab of the warm words of the user 102b ( Figures 1A to 1C)。In the example shown, the GUI 300b presents a list of available warm words 112 that are displayed as corresponding graphical elements that can be selected to add or remove a corresponding available warm word that the user wants the AED 104 to detect and ultimately execute a corresponding action when spoken by the user 102b from the corresponding active set 112A of warm words, the corresponding action being specified by the corresponding warm word. The GUI 300b provides the ability to select activity-based warm words that the user 102b wants to activate by adding them to the corresponding active set 112A of warm words while the user 102b is performing a particular activity. Here, the particular activity can be ascertained based on the current application being executed on the user device 50 or the AED 104 with which the user 102b is interacting. For example, the warm word configurator screen presented by the GUI 300b can provide a list of applications loaded on the user device 50 and list, for each corresponding application, the available activity-based warm words that the user 102 can select to add to the corresponding active set 112A of warm words when the user 102 is interacting with the corresponding application. A non-exhaustive list of applications that each have available activity-based warm words can include a cooking application and a music player application. In the example shown, when the user 102 is interacting with the cooking application, the user can select the activity-based warm words "pre-heat oven" and "set timer" that can be added to the corresponding active set 112Ab of warm words for the user 102b. For example, in Figure 1A the user may be viewing a recipe for roasting a boneless chicken in a cooking app, and when the user is viewing the recipe preparation steps, the activity-based warm word "pre-heat oven" can be added to the corresponding active set 112Ab of warm words. Here, when the AED 104 detects the warm word "pre-heat oven" in the audio data that characterizes the utterance spoken by the user 102b and captured by the AED 104, the digital assistant 105 instructs the smart oven to turn on and pre-heat at the time of the user 102b. It is noted that, as described in more detail below, the warm word arbitration routine 401 must positively select to enable the warm word "pre-heat oven" as one of the warm words in the final set 112F of warm words in order to permit the AED 104 to detect the warm word in the audio data and execute the corresponding action.
[0066] Continuing to refer to Figure 3B the GUI 300b, Figure 1BOnce the temperature of the smart oven is successfully preheated, the active warm word "set timer" can be added to the corresponding active set 112Ab of warm words. Here, the AED 104 can receive from the smart oven an ambient context signal 440 indicating that the smart oven is at a specified preheat temperature, thereby triggering the AED 104 to add the activity-based warm word "set timer" to the corresponding active set 112Ab of warm words for the user 102b and remove the activity-based warm word "pre-heat even" from the corresponding active set of warm words. It should be noted that, as described in more detail below, adding the warm words "set timer" and "pre-heat oven" to the corresponding active set 112Ab of warm words for the second user 102b and removing the warm word from the corresponding active set of warm words can trigger the warm word selector 400 to execute the warm word arbitration routine 401 again to determine whether to add the warm word "set timer" to the final set 112F of warm words enabled for detection by the AED 104. In Figure 1B In the example shown, once the smart oven is successfully preheated, the AED 104 outputs an indicated notification. For example, the digital assistant 105 can generate a synthetic voice 123 for audible output from the speaker 18 of the AED 104, the synthetic voice stating "Oven Reached 350-deg (oven reached 350 degrees)". In an additional example, when the warm word "pre-heat oven" has been added by the warm word selector 400 to the final set 112F of warm words enabled for detection by the AED 104, the digital assistant 105 can provide a notification to the user device 50 (e.g., a smart phone) linked to the user account of the identified user to inform the identified user 102b (e.g., Jim).
[0067] Figure 3B The GUI 300b also shows options for the user to select other activity-based warm words for activities related to the smart doorbell and incoming calls. For example, as Figure 1BAs shown, the smart doorbell can be designated with activity-based warm words, such as "Show Cam", "Ignore", and "Talk", under visitor presence conditions (such as when a visitor rings the smart doorbell and / or the smart doorbell detects the presence of a visitor approaching the smart doorbell). Under one or more of the visitor presence conditions, these activity-based warm words can be added to the corresponding active set 112A of warm words for each user 102, and the corresponding active set of warm words is selected, for example, via using the warm word configurator screen presented by the GUI 300b to include these activity-based warm words. The smart doorbell can send an ambient context signal 440 indicating the occurrence of one or more visitor presence conditions to the AED 104 ( Figure 4 ), where the AED 104 can emit a visitor notification 125, such as a doorbell ring, for audible output from the speaker 18 of the AED 104. The AED 104 can provide the visitor notification 125 as synthesized speech, as a supplement to and in place of the doorbell ring. Meanwhile, the ambient context signal 440 can trigger the warm word selector 400 to execute a warm word arbitration routine 401 that determines to add the activity-based warm words "ShowCam", "Ignore", and "Talk" to the final set 112F of warm words enabled for detection by the AED 104. Thus, in the Figure 1B example, any of the users 102a - 102c can (without saying a predefined hot word) say "Show Cam" to cause the screen communicating with the AED 104 to display image data of the visitor captured by the camera of the smart doorbell or a camera near the smart doorbell, say "Ignore" to select to dismiss the visitor notification (and optionally, cause the smart doorbell to broadcast a pre-recorded message to the visitor), or say "Talk" to turn on the microphone communicating with the smart doorbell and the microphone 16 of the AED 104 (and / or the microphone 52 of the user device 50), thereby providing an intercom communication capability between the user 102 and the visitor.
[0068] Continuing to refer to Figure 3B, the warm word configurator screen also shows a list of preferred warm words that user 102b can select for inclusion in the respective active set 112A of warm words, which are not related to the activity. In the example shown, user 102b selects the warm words "Call [Contact]", "Play Music", "Turn off lights", and "Turn on lights" as preferred warm words to add to the active sets 112A, 112Ab of warm words. The list of preferred warm words can be pre-populated or based on common voice commands learned over time that are elevated to warm words by digital assistant 105. Some of the preferred warm words can be custom words or phrases that the user assigns to the corresponding actions (e.g., routines) to be performed when spoken by the user. In the example shown, the warm words "What’s the weather" and "Set Alarm" are listed as preferred warm words but are indicated as inactive warm words 112N because user 102b did not select to add these warm words to the corresponding active list 112A of warm words. The list of preferred warm words and activity-based warm words selected for inclusion in the respective active set 112A of warm words can be saved in the respective warm word preferences 212 for each registered user 200.
[0069] Refer back to Figures 1A to 1C and Figure 4 , the associated user device of user 102 detected by user detector 410 in the environment can periodically and / or whenever a warm word is added to or removed from the active set, push / broadcast the respective active set 112A of warm words of user 102. Here, the user device can push / broadcast the respective active set 112A of warm words from the corresponding user device 50 that communicates with AED 50. Additionally or alternatively, when user detector 410 uniquely identifies the user as one of the registered users 200, warm word selector 400 can obtain the active set 112A of warm words by accessing warm word preferences 212. In some examples, warm word selector 400 determines when to add a warm word to the active set 112A of warm words for different users in response to receiving an environmental context signal 440 that indicates a particular activity is currently ongoing and the particular activity is assigned to one or more activity-based warm words communicated in warm word preferences 212. Warm word selector 400 can use other techniques (such as prompting the detected user to provide his / her active set of warm words) to obtain the active set of warm words 112.
[0070] Refer to Figure 4In some implementations, the warm word selector 400 executes a warm word arbitration routine 401 for enabling a final set 112F of warm words for detection by the AED 104 based on an active set 112A of warm words obtained for each user 102 among a plurality of users 102 detected by the user detector 401. The warm word selector 400 may execute the warm word arbitration routine 401 periodically during specified intervals and / or in response to a particular event / condition. In some examples, the warm word arbitration routine 401 is executed in response to determining to add a new warm word to a corresponding active set 112A of warm words for one of a plurality of different users present in the environment or to remove one of the warm words from the active set of warm words. For illustration, Figure 1B A first user 102a (e.g., Barb) is shown speaking a second utterance 106, 106b near the AED 104: "Ok computer, stop music and start commute routine when I get in the car". The microphone 16 of the AED 104 receives the utterance 106b and processes the audio data 502 to detect the hot word 110 using the hot word detector 108, thereby causing the AED 104 to wake up and perform speech recognition to recognize the command to stop the music and perform the action of starting Barb's commute routine once Barb is in her car. Since the command 118 causes the digital assistant 105 to stop the long-standing operation of playing back the music 122 from Barb's playlist, the music playback warm words pause, next song, previous song, volume up, and volume down have now been removed from the corresponding active set 112Aa for Barb's warm words. In this example, the warm word selector 400 receives the corresponding active set 112Aa of warm words for Barb that has been updated to now omit the music playback warm word, thereby triggering the execution of the warm word arbitration routine 401 to potentially update the final set 112F of warm words enabled for detection by the AED. Similarly, adding the activity-based warm words "Show Cam", "Ignore", and "Talk" to the corresponding active sets 112Aa-112Ac of warm words for each of the users 102a-102c can trigger the warm word selector 400 to execute the warm word arbitration routine 401.
[0071] In some examples, the warm word arbitration routine 401 executes in response to the user detector 410 detecting the presence of a new user within the environment of the AED 104. Notably, since the AED obtains the corresponding active set 112A of warm words for any newly detected user, the warm word arbitration routine 401 executes to determine whether the final set 112F of warm words should be updated to include any of the active warm words from the active warm words for the new user. Similarly, updating the final set 112F of warm words can include removing some warm words from the final set to make room (i.e., in terms of processing / memory capabilities and / or error acceptance thresholds) for any new active warm words from the new active warm words for the new user.
[0072] Additionally or alternatively, the warm word arbitration routine 401 can execute in response to the AED 104 determining that one of the previously detected users no longer exists within the environment of the AED 104. For example, Figure 1C illustrates that the first user 102a no longer exists in the environment. The user detector 410 can continuously update the list of users currently detected within the environment of the AED 104, thereby allowing the AED 104 to ascertain when a user no longer exists.
[0073] In some implementations, the AED 104 determines that the environmental context has changed, and the warm word arbitration routine 401 executes in response to determining that the environmental context has changed. In some examples, the warm word selector 400 receives an environmental context signal 440 from a source indicating that the environmental context has changed. For example, a smart doorbell can provide an environmental context signal 440 indicating the occurrence of one or more visitor presence conditions, and an oven can provide an environmental context signal 440 when the oven has reached its preheated temperature. An incoming call to one or more registered users 200 of the AED 104 can also be used as the environmental context signal 440, which causes the warm word arbitration routine 401 to potentially update the final set 112F of warm words to include new warm words related to the incoming phone call event, such as but not limited to warm words like "answer" or "ignore" that, when spoken, will cause the AED 104 to answer or ignore the incoming call.
[0074] Execution of the warm word arbitration routine 401 can include: obtaining an enabled warm word constraint 430; and determining, based on the enabled warm word constraint 430, the number of warm words to be enabled in the final set of warm words 112F for detection by the AED 104. The warm word constraint 430 can permit the warm word arbitration routine 401 to determine the warm word capacity of the AED 104 at any given time. The warm word constraint 430 can include at least one of the following: memory and computational resource availability of the AED 104 for detecting warm words, computational requirements for each warm word in the respective active set 112A of warm words enabled for each of a plurality of users present in the environment of the AED, acceptable false acceptance rate tolerance or acceptable false rejection rate tolerance. The acceptable false acceptance rate tolerance and the acceptable false rejection rate tolerance can be based on global settings, device settings, or user-defined on a per-user basis. In some examples, the acceptable false acceptance rate tolerance and the acceptable false rejection rate tolerance can be dynamic, where a decrease in warm word detection sensitivity occurs when the user is less active or farther from the device.
[0075] The memory and computational resource availability can consider whether the AED 104 is battery-powered and, if so, the current capacity of the battery. The memory and computational resource availability can further consider the current warm words in the final set of warm words, in addition to processing power and storage capacity, and consider the applications currently running on the AED 104.
[0076] The computational requirements can determine the memory and processing resources required to execute the respective warm word models associated with each warm word in the active set 112A of warm words. For example, the computational requirements can include the warm word model size / parameters for each warm word and whether any of the warm words are speaker-specific and require execution of the speaker identification process 500. Additionally or alternatively, the computational requirements can determine the memory and processing resources required to execute the speech recognizer 116 in a low-power mode sufficient to only discern the utterance of a warm word. It is noted that a lower acceptable false acceptance rate tolerance may require more processing of the respective warm word models used to detect warm words compared to a higher acceptable false acceptance rate tolerance.
[0077] In some implementations, the execution of the warm word arbitration routine 401 includes: for each corresponding user 102 among multiple users 102, sorting the warm words in the corresponding active set of warm words from the highest priority to the lowest priority based on the warm word priority sorting signal 413; and enabling the final set of warm words 112F for detection by the AED 104 based on the sorting of the warm words in the corresponding active set 112A of warm words for each corresponding user 102. Here, the warm word priority sorting signal 413 may include at least one of the following: the usage frequency of each warm word in the corresponding active set 112A of warm words for the corresponding user, the current state of the AED 104, the environmental context of the AED 104, or coexistence information indicating previous warm word usage and / or operations performed by the AED when the corresponding user was previously present in the environment in combination with one or more of the other users among the multiple users. The current state of the AED may include any long-term operation that the digital assistant is currently performing on the AED 104. For example, if the digital assistant is playing music at the highest volume setting, the priority sorting signal 413 may cause the warm word arbitration routine 401 to sort the warm words "play" and "volume up" in the active set 112A of warm words with a low priority because the user is less likely to say these warm words. The environmental context may include activity recognition performed by the AED based on one or more signals. For example, Figure 1A depicts a third user 102c watching a movie using his / her headset and having enabled the do not disturb mode on his / her user device 50. The routine 401 may determine the user context of the user 103c based on the image data 412 received from the camera 19 (or another camera) indicating that the user is wearing the headset and has shifted his / her gaze away from the AED 104 and / or the communication signal received from the user device 50 (or the headset) indicating that the do not disturb mode is enabled on the user device. If enabled, the headset may also communicate a signal indicating that the headset is currently being worn and is playing back audio content. The user context of the user 102c may additionally or alternatively be ascertained based on the received signal indicating that the movie is in progress. Thus, the routine 401 may sort the corresponding active set 112A of warm words for the user 103c with a lower priority than the active sets 112Aa, 112Ab of warm words for the first user 102a and the second user 102b. However, as Figure 1B shown, the third user 103c may permit a visitor alert notification when the do not disturb mode is enabled, such that the environmental context signal 440 indicating a visitor at the doorbell causes the digital assistant 105 to audibly output the visitor alert from the user's headset and / or visually display the visitor alert by showing a graphic on the user device 50 or the television that the user 103c is currently viewing to inform the user 102c of the visitor. Thus, although inFigure 1A Under the conditions, the corresponding active set 112Ac of warm words is sorted with a low priority and not included in the final set of warm words (except for "Call [Contact]" which is also active for other users 102a, 102b), but Figure 1B the visitor alert sent to the third user 102c in can change the current user context of the user 102c to indicate that the third user 102c may be interested in interacting with the digital assistant 105 to find out information about the visitor at the doorbell. As Figure 1B shown, the warm words "Show Cam", "ignore", and "Talk" are included in the active sets 112Aa - 112Ac of warm words for all users 102a - 102c and are sorted highest, and are finally selected for inclusion in the final set 112F of warm words enabled for detection by the AED 104.
[0078] Coexistence information can indicate how warm words are utilized by the corresponding users in the presence of other users. For example, the first user 102a and the second user 102b may each often say the warm word "play music" when in each other's presence, while the first user 102a rarely says the warm word "play music" when the third user 102c is present.
[0079] The arbitration routine 401 can select higher-priority warm words from the sorted active set of warm words for all users 102 for inclusion in the final set 112F of warm words to be enabled for detection by the AED 104. As previously mentioned, the number of warm words included in the final set 112F of warm words is limited by the enabled warm word constraint 430. In some examples, to optimize the execution of the arbitration routine 401, the arbitration routine 401 only considers the warm words that are the top N warm words sorted for each user from the respective active sets 112A of warm words. The value of N can be fixed or can be variable for different users. In an example where the value of N is made variable among different sets of users, routine 401 can determine a warm word affinity score for each corresponding user 102 among multiple users and enable the final set of warm words based on the warm word affinity score. The warm word affinity score for each corresponding user 102 can be based on the frequency of warm word usage by the corresponding user 102 and / or the frequency of interaction between the corresponding user 102 and the digital assistant 105. Here, the frequency of warm word usage can indicate how frequently the corresponding user 102 uses warm words when interacting with the digital assistant 105, while the frequency of interaction can indicate how frequently the corresponding user 102 generally interacts with the digital assistant 105. The frequency of warm word usage and / or interaction can be further constrained by the current duration of detecting the presence of the corresponding user 102 within the environment of the AED 104, a particular day-of-week date, and / or the time of day. For example, a user may frequently ask "What’s the weather" at approximately the same time every morning. The warm word affinity score can additionally or alternatively be based on at least one of the following: the duration of the presence of the corresponding user 102 within the environment of the AED 104, the proximity of the corresponding user 102 relative to the AED 104, or the current user context. For example, the proximity information 54 for a corresponding user can cause the arbitration routine 401 to assign higher priority to the warm words in the respective active sets of warm words for a user who is closer to the AED 104 than other users, because a closer user may be more likely to interact with the AED compared to a more distant user.
[0080] depicting a third user 103c wearing his / her headset and watching a movie and having enabled the do not disturb mode on his / her user device 50 Figure 1A and Figure 1BIn the example, the warm word arbitration routine 401 can determine a lower affinity score for the third user 103c than for other users 103a, 103b because the current context of user 102c indicates a low likelihood that user 103c will interact with the digital assistant 103, and the proximity information 58 indicates that the third user 103c is located relatively far from the AED 104 compared to other users. Additionally, if the first user 102c subsequently uses one of the music playback warm words after issuing the command 118 "play my music playlist" in the first utterance 106a, the affinity score for the first user 102a can be further increased. Similarly, if user 102b frequently utilizes the warm word "pre-heat oven" while viewing a recipe in the cooking application 107 as shown in Figure 1A the affinity score for the second user 102b can be increased. Similarly, after the oven reaches the preheated temperature, the warm word "Set Timer" is ranked higher in the corresponding active set of warm words in Figure 1B because there is a higher likelihood that user 102b will place the dish in the oven and set the timer for the duration that the dish will be baked in the oven. In Figure 1B the "preheat oven" warm word is now ranked lowest in the corresponding active set 112Ab of warm words for the second user 102c.
[0081] To illustrate how the combination of the current user context and proximity affects the warm word affinity score for the third user 102b, the current user context and affinity for the third user 102c in Figure 1A associate the third user 102c with a low affinity score, while the event of the visitor notification 125 in Figure 1B associates the second user 102b with a slightly higher affinity score. However, in Figure 1C the proximity information 58 now indicates that the third user 102c has moved closer to the AED 104, and the current user context now indicates that the do not disturb mode is disabled and the user is no longer wearing his / her headset. Therefore, the arbitration routine 401 can increase the affinity score for the third user 102c because there is a higher likelihood that user 102c will interact with the digital assistant 105 compared to the example depicted in Figure 1B and especially the example depicted in Figure 1A .
[0082] In addition to the duration of the presence, the arbitration routine 401 can further predict the expected likelihood of a future user presence and determine an affinity score based on the expected likelihood of the user presence. For example, when the action recognizer 124 determines that the transcription of the second utterance 106b output from the speech recognizer 116 includes the command 118 “… stop music and start commute routine when I get in the car”, the arbitration routine 401 can predict that the presence of the user 102a in the environment will likely end. Thus, the arbitration routine can start gradually decreasing the affinity score for the first user 102a until the first user 102a is no longer detected as shown in Figure 1C until the first user 102a is no longer detected.
[0083] Continuing to refer to Figures 1A to 1C and Figure 4 , the warm word arbitration routine 401 can boost the scores for warm words shared by two or more active sets in the active set 112A of warm words because there is an increased likelihood that these warm words will be spoken. These shared warm words can be ranked higher in the respective active sets 112A of warm words so that they are not inadvertently omitted from the top N active warm words as discussed above. Additionally, the associated warm word models for detecting these shared warm words among multiple users can be merged with each other to limit the number of similar warm word detection models executed on the AED 104. Furthermore, warm words shared by different users can cause the digital assistant 105 to perform different actions depending on which user speaks the warm word 112. For example, in the Figure 1C example, the second user 102b speaking “play music” which is enabled for detection in the final set of warm words can cause the digital assistant 105 to stream music from the first music player designated as the preferred music player for the second user 102b. In contrast, the third user 102c speaking the same warm word “play music” can cause the digital assistant 105 to stream music from a different second music player preferred by the third user 102c. Notably, Figure 5The speaker identification process 500 can be performed on the audio data 502 representing the utterance of the warm word, so as to uniquely identify the user who utters the warm word, such that appropriate actions (e.g., which music player to select for streaming music) can be performed by the digital assistant 105. In this example, the warm word arbitration routine 401 identifies that the respective active sets 112Ab, 112Ac of warm words each include the warm word 112 "play music". Instead of enabling the detection of two separate warm word models 330 for the warm word 112 "play music", the warm word arbitration routine selects only one warm word model 330 for the warm word 112 "play music". In some implementations, the warm word arbitration routine 401 determines that the AED 104 has sufficient capacity and is capable of performing a higher quality (i.e., additional parameters, additional operations, reduced latency, and / or increased sensitivity) warm word model 330 for the warm word 112 "play music". Optionally, the warm arbitration routine 401 determines that the AED 104 has sufficient capacity and is capable of performing a different architecture, such as moving from the warm word model 330 to the ASR model.
[0084] Reference Figure 1C, the first user 102a is no longer detected, and since the third user 102c moves closer to the AED 104 and no longer wears the headset, the affinity score for the third user 102c is further improved. Additionally, the playback of the music playlist of the first user 102a has stopped. Thus, the final set 112F of warm words enabled for detection includes the warm words "tell me a joke", "set a timer", "play music", "call [contact]", and "turn off lights". The warm word "tell me a joke" can be speaker-specific to the third user 102c and selected from the corresponding active set 112Ac of warm words for the third user 102c, while the warm word "set a timer" can be speaker-specific to the second user 102b and selected from the corresponding active set 112Ab of warm words for the second user 102b. The warm words "play music", "call [contact]", and "turn off lights" are shared warm words that can be spoken by either of the users 102b, 102c to cause the corresponding actions specified by the warm words to be performed by the digital assistant 105 when detected in the streaming audio by the AED 104 (e.g., using the appropriate warm word model 330 or the speech recognizer 116) without including a pre-determined hot word (e.g., Ok Computer). In this example, the third user 102c speaks the third utterances 106, 106c, which include the warm word 112 "tell me a joke" from the final set 112F of warm words enabled for detection by the AED 104. Without performing speech recognition on the captured audio, the AED 104 can apply the warm word model 330 for the final set 112F of warm words to identify whether the utterance 106c includes any of the warm words included in the final set 112F of warm words. The AED 104 compares the audio data 502 corresponding to the utterance 106c with the enabled warm word models 330 corresponding to the warm words 112 "tell me a joke", "set a timer", "play music", "call [contact]", and "turn off lights", and determines without performing speech recognition on the audio data 402 that the warm word model 330 enabled for the warm word 112 "tell me a joke" detects the warm word 112 "tell me a joke" in the utterance 102c.Most notably, the user utters utterance 106c without prefixing the utterance 106c with a predetermined hot word 110, such that the appropriate warm word model 330 detects the presence of warm word 112 in the audio data and triggers the AED 104 to evoke the digital assistant 105 to perform the action specified by the warm word (e.g., retrieve and play a joke). In some examples, the speech recognizer 116 operates in a low-power mode that only recognizes the presence of warm words 112 in the final set 112F of warm words. In these examples, the speech recognizer 116 can recognize when one of the warm words in the final set 112F of warm words is uttered without the presence of a predefined hot word and evoke the digital assistant to perform the corresponding action. In this example, the digital assistant 105 retrieves a joke from a search engine, a device application, or other source and audibly outputs the joke as synthesized speech 129. The digital assistant 105 can additionally or alternatively output the joke as a text representation displayed on the screen 13 of the AED 104 or on another screen in communication with the AED 104.
[0085] Figure 6 Flowchart of an exemplary operational arrangement of method 600 that includes enabling a final set 112F of warm words for detection by an AED 104 when multiple users 102 are present in an environment of the assistant-enabled device (AED) 104. The operations performed by method 600 may be described with reference to FIGS. 1 through Figure 5 As used throughout this disclosure, the "environment" of the AED 104 refers to the proximity within which a user is for interacting with the AED by voice and optionally also by other means. Thus, the "environment" can include the same room as the AED or even a home, a vehicle cabin, the exterior of a vehicle, a doctor's office where the AED is located, a business establishment where the AED is located, and within the proximity of the AED sufficient for the AED's microphone to capture audio. The data processing hardware 12 of the AED 104 can execute instructions stored on the memory hardware 14 that cause the AED to perform operations. Optionally, the server 130 can perform some or all of the operations.
[0086] At operation 602, method 600 includes: detecting the presence of multiple users 102 within the environment of the AED 104. The AED 104 executes the digital assistant 105. In some configurations, the AED 104 can execute multiple digital assistants simultaneously. At operation 604, for each user 102 of the multiple users 102, the method further includes: respective active warm word sets 112A that each specify a corresponding action for the digital assistant 105 to perform.
[0087] At operation 606, based on the respective active sets of warm words for each of a plurality of users, method 600 includes: performing a warm word arbitration routine 401 to enable a final set 112F of warm words for detection by AED 104. Here, the final set 112F of warm words enabled for detection by AED 104 includes warm words 112 selected from the respective active sets 112A of warm words for at least one user 102 of the plurality of users 102 detected within the environment of AED 104.
[0088] At operation 608, when the final set of warm words is enabled for detection by AED 104, the method includes: receiving audio data 502 corresponding to the utterance 106 captured by AED 104; detecting warm words 112 from the final set 112F of warm words in the audio data 502; and instructing the digital assistant to perform the respective actions specified by the detected warm words.
[0089] In some examples, the final set 112F of warm words is enabled for detection by activating respective warm word models 330 for each warm word in the final set 112F to run on the AED. Here, the method uses the activated respective warm word models 330 to detect warm words in the audio data 502 without performing speech recognition on the audio data 502. More specifically, detecting warm words in the audio data may include: extracting audio features in the audio data 502; generating a warm word confidence score by processing the extracted audio features using the activated respective warm word models 330; and determining that the audio data corresponding to the utterance includes a warm word when the warm word confidence score meets a warm word confidence threshold. In some examples, the warm word confidence threshold can be ascertained for the registered user 200 by accessing an acceptable false acceptance rate tolerance and / or an acceptable false rejection rate tolerance stored in the warm word preference 212 of the corresponding registered user 200. In additional examples, when the final set 112F of warm words includes two different warm words with similar pronunciations, the warm word confidence threshold for detecting each of these warm words can be increased to reduce the propensity for false acceptance where the warm word model 330 incorrectly detects a warm word in the audio instead of the actually spoken similar - pronunciation warm word.
[0090] Figure 7 is a schematic diagram of an example computing device 700 that can be used to implement the systems and methods described in this document. Computing device 700 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown here, their connections and relationships, and their functions are intended to be exemplary only and are not intended to limit the implementation of the invention described and / or claimed in this document.
[0091] The computing device 700 includes a processor 710, a memory 720, a storage device 730, a high-speed interface / controller 740 connected to the memory 720 and a high-speed expansion port 750, and a low-speed interface / controller 760 connected to a low-speed bus 770 and the storage device 730. Each of the components 710, 720, 730, 740, 750, and 760 is interconnected using various buses and can be mounted on a common motherboard or otherwise as appropriate. The processor 710 (e.g., the data processing hardware 10, 132 of FIG. 1) can process instructions for execution within the computing device 700, including instructions stored in the memory 720 or on the storage device 730, to display graphical information of a graphical user interface (GUI) on an external input / output device such as a display 780 coupled to the high-speed interface 740. In other implementations, multiple processors and / or multiple buses and multiple memories and multiple types of memories can be used as appropriate. Additionally, multiple computing devices 700 can be connected, where each device provides a portion of the necessary operations (e.g., as a server group, blade server cluster, or multi-processor system).
[0092] The memory 720 stores information non-transitorily within the computing device 700. The memory 720 (e.g., the memory hardware 12, 134 of FIG. 1) can be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. The non-transitory memory 720 can be a physical device for storing programs (e.g., instruction sequences) or data (e.g., program state information) on a temporary or permanent basis for use by the computing device 700. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., commonly used for firmware such as a boot program). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and magnetic disk or tape.
[0093] The storage device 730 can provide large-capacity storage for the computing device 700. In some implementations, the storage device 730 is a computer-readable medium. In various different implementations, the storage device 730 can be a floppy disk device, a hard disk device, an optical disk device, or a magnetic tape device, a flash memory, or other similar solid-state memory devices, or an array of devices (including devices in a storage area network or other configurations). In additional implementations, the computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable medium or a machine-readable medium, such as the memory 720, the storage device 730, or the memory on the processor 710.
[0094] The high-speed controller 740 manages the bandwidth-intensive operations of the computing device 700, while the low-speed controller 760 manages the lower-bandwidth-intensive operations. This division of responsibilities is merely exemplary. In some implementations, the high-speed controller 740 is coupled to the memory 720, the display 780 (e.g., via a graphics processor or accelerator), and a high-speed expansion port 750 that can accept various expansion cards (not shown). In some implementations, the low-speed controller 760 is coupled to the storage device 730 and a low-speed expansion port 790. The low-speed expansion port 790, which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner; or to a networking device, such as a switch or a router, via a network adapter.
[0095] The computing device 700 can be implemented in many different forms, as shown in the figure. For example, the computing device can be implemented as a standard server 700a or implemented multiple times in a group of such servers 700a, as a laptop computer 700b, or as part of a rack server system 700c.
[0096] Various implementations of the systems and techniques described herein can be implemented in digital electronic and / or optical circuitry, integrated circuit systems, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in one or more computer programs executable and / or interpretable on a programmable system that includes at least one programmable processor, which can be dedicated or general-purpose and can be coupled to receive data and instructions from a storage system, at least one input device, and at least one output device and to transmit data and instructions to a storage system, at least one input device, and at least one output device.
[0097] A software application (i.e., software resource) can refer to computer software that causes a computing device to perform tasks. In some examples, a software application can be referred to as an "application", "app", or "program". Example applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0098] A non-transitory memory can be a physical device for storing programs (e.g., sequences of instructions) or data (e.g., program state information) for use by a computing device on a temporary or permanent basis. A non-transitory memory can be a volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., commonly used for firmware such as a boot program). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and magnetic disks or tapes.
[0099] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages and / or assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, device, and / or apparatus (e.g., a disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives the machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0100] The processes and logical flows described in this specification can be performed by one or more programmable processors (also referred to as data processing hardware) that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by special purpose logic circuitry, such as an FPGA (field programmable gate array) or ASIC (application specific integrated circuit). By way of example, processors suitable for the execution of a computer program include both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices for storing data (e.g., magnetic disks, magneto - optical disks, or optical disks), or be operatively coupled to receive data from one or more mass storage devices or transfer data to one or more mass storage devices or both. However, a computer need not have such devices. Computer - readable media suitable for storing computer program instructions and data include all forms of non - volatile memory, media and memory devices, by way of example including semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto - optical disks, and CD ROM and DVD - ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0101] In order to provide for interaction with a user, one or more aspects of the present disclosure may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube), an LCD (liquid crystal display) monitor, or a touch screen) for displaying information to the user and, optionally, a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide for interaction with the user; for example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input received from the user may be in any form, including acoustic input, speech input, or tactile input. Additionally, a computer may interact with the user by sending documents to and receiving documents from a device used by the user; for example, by sending a web page to a web browser on a client device of the user in response to a request received from the web browser.
[0102] A variety of implementations have been described. However, it should be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims.
Claims
1. A computer-implemented method (600) that, when executed by data processing hardware (710), causes the data processing hardware (710) to perform operations including: Detecting the presence of a plurality of users within an environment of an assistant-enabled device AED (104) that executes a digital assistant (105); For each user of the plurality of users, obtaining a respective active set of warm words (112) that each specify a respective action for the digital assistant (105) to perform; Based on the respective active sets of warm words (112) for each user of the plurality of users, performing a warm word arbitration routine (401) to enable a final set of warm words (112) for detection by the AED (104), the final set of warm words (112) enabled for detection by the AED (104) including warm words (112) selected from the respective active sets of warm words (112) for at least one user of the plurality of users detected within the environment of the AED (104); And When the final set of warm words (112) is enabled for detection by the AED (104): Receiving audio data (402) corresponding to an utterance (106) captured by the AED (104); Detecting warm words (112) from the final set of warm words (112) in the audio data (402); and Instructing the digital assistant (105) to perform the respective action specified by the detected warm words (112).
2. The computer-implemented method (600) according to claim 1, wherein detecting the presence of the plurality of users within the environment of the AED (104) Comprises: Detecting the presence of at least one user of the plurality of users within the environment based on proximity information (54) of the AED (104) relative to a user device (50) associated with at least one user of the plurality of users.
3. The computer-implemented method (600) according to claim 1, wherein detecting the presence of the plurality of users within the environment of the AED (104) Comprises: Receiving image data (312) corresponding to a scene of the environment; And Detecting the presence of at least one user of the plurality of users within the environment based on the image data (312).
4. The computer-implemented method (600) according to claim 1, wherein detecting the presence of a corresponding user (102) within the environment of the AED (104) comprises detecting the presence of the corresponding user (102) within the environment of the AED (104) based on: Receiving voice data representing a voice query issued by the corresponding user (102) and directed to the digital assistant (105); Performing speaker recognition on the voice data to identify the corresponding user (102) that issued the voice query; and Determine that the identified corresponding user (102) who issued the voice query is present within the environment of the AED (104).
5. The computer-implemented method (600) according to claim 4, wherein: the voice query issued by the corresponding user (102) includes a command (118), the command causing the digital assistant (105) to perform a long-term operation specified by the command (118); and obtaining the corresponding active set of warm words (112) includes: in response to the digital assistant (105) performing the long-term operation specified by the command (118), adding one or more warm words (112) specifying corresponding actions for controlling the long-term operation to the corresponding active set of warm words (112) for the corresponding user (102) who issued the voice query.
6. The computer-implemented method (600) according to any one of claims 1 to 5, wherein the operation further includes: detecting the presence of a new user within the environment of the AED (104), wherein the warm word arbitration routine (401) is executed in response to detecting the presence of the new user within the environment.
7. The computer-implemented method (600) according to any one of claims 1 to 6, wherein the operation further includes: determining that one of the plurality of users is no longer present within the environment of the AED (104), wherein the warm word arbitration routine (401) is executed in response to determining that one of the plurality of users is no longer present within the environment.
8. The computer-implemented method (600) according to any one of claims 1 to 7, wherein the operation further includes: determining to add a new warm word (112) to the corresponding active set of warm words (112) for one of the plurality of users present within the environment of the AED (104) or remove one of the warm words (112) from the corresponding active set of warm words, wherein the warm word arbitration routine (401) is executed in response to determining to add the new warm word (112) to the corresponding active set of warm words (112) for one of the plurality of users present within the environment of the AED (104) or remove one of the warm words (112) from the corresponding active set of warm words.
9. The computer-implemented method (600) according to any one of claims 1 to 8, wherein the operation further includes: determining that the environmental context of the AED (104) has changed, wherein the warm word arbitration routine (401) is executed in response to determining that the environmental context has changed.
10. The computer-implemented method (600) according to any one of claims 1 to 9, wherein executing the warm word arbitration routine (401) includes: obtaining enabled warm word constraints (430), the enabled warm word constraints (430) including at least one of the following: Memory and computational resource availability on the AED (104) for detecting warm words (112); Computational requirements for each warm word (112) in the respective active set of warm words (112) enabled for each of the multiple users present in the environment of the AED (104); Acceptable false acceptance rate tolerance; Or Acceptable false rejection rate tolerance; And Determine the number of warm words (112) to be enabled in the final set of warm words (112) for detection by the AED (104) based on the enabled warm word constraints (430).
11. The computer-implemented method (600) according to any one of claims 1 to 10, wherein performing the warm word arbitration routine (401) Comprises: For each corresponding user (102) of the multiple users, sort the warm words (112) in the respective active set of warm words (112) from highest priority to lowest priority based on a warm word prioritization signal (413), the warm word prioritization signal (413) comprising at least one of the following: The usage frequency of each warm word (112) in the respective active set of warm words (112) by the corresponding user (102); The current state of the AED (104); The environmental context of the AED (104); or Coexistence information indicating previous warm word (112) usage and / or operations performed by the AED (104) when the corresponding user (102) was previously present in the environment with one or more combinations of other users among the multiple users; And Enable the final set of warm words (112) for detection by the AED (104) based on the sorting of the warm words (112) in the respective active set of warm words (112) for each corresponding user (102) of the multiple users.
12. The computer-implemented method (600) according to any one of claims 1 to 11, wherein performing the warm word arbitration routine (401) Comprises: For each corresponding user (102) of the multiple users, determine a warm word affinity score; And Enable the final set of warm words (112) for detection by the AED (104) based on the warm word affinity scores determined for each corresponding user (102).
13. The computer-implemented method (600) according to claim 12, wherein the warm word affinity scores determined for each corresponding user (102) are based on at least one of the following: The frequency of warm word (112) usage by the corresponding user (102); The frequency of interaction between the corresponding user (102) and the digital assistant (105); The duration of the presence of the corresponding user (102) in the environment of the AED (104); The proximity of the user to the AED (104); or The current user context.
14. The computer-implemented method (600) according to any one of claims 1 to 13, wherein performing the warm word arbitration routine (401) comprises: identifying any shared warm words (112) corresponding to the warm words (112) present in the respective active sets of warm words (112) for each of at least two of the plurality of users; and determining that the final set of warm words (112) is based on assigning a higher priority to the warm words (112) identified as shared for inclusion in the final set of warm words (112).
15. The computer-implemented method (600) according to any one of claims 1 to 14, wherein the operation further comprises: determining that the warm words (112) detected in the audio data (402) corresponding to the utterance (106) captured by the user device (50) include speaker-specific warm words (112) selected from the respective active sets of warm words (112) for each of one or more of the plurality of users, such that the digital assistant (105) performs the respective actions specified by the speaker-specific warm words (112) only when the speaker-specific warm words (112) are spoken by any of the one or more users corresponding to the respective active sets of warm words (112); and performing speaker verification on the audio data (402) based on determining that the warm words (112) detected in the audio data (402) corresponding to the utterance (106) captured by the user device (50) include speaker-specific warm words (112) to determine that the utterance (106) is spoken by one of the one or more users corresponding to the respective active sets of warm words (112) from which the detected warm words (112) were selected, wherein instructing the digital assistant (105) to perform the respective actions specified by the detected warm words (112) is based on the speaker verification performed on the audio data (402).
16. The computer-implemented method (600) according to any one of claims 1 to 15, wherein: the final set of warm words (112) is enabled for detection by activating a respective warm word model (330) for each warm word (112) in the final set of warm words (112) to run on the assistant-enabled device; and detecting the warm words (112) from the final set of warm words (112) in the audio data (402) comprises: using the activated respective warm word models (330) to detect the warm words (112) in the audio data (402) without performing speech recognition on the audio data (402).
17. The computer-implemented method (600) according to claim 16, wherein detecting the warm words (112) in the audio data (402) comprises: extracting audio features of the audio data (402); Use the activated corresponding hotword model (330) to generate a hotword confidence score by processing the extracted audio features; and When the hotword confidence score meets the hotword (112) confidence threshold, determine that the audio data (402) corresponding to the utterance (106) includes the hotword (112).
18. The computer-implemented method (600) according to any one of claims 1 to 17, wherein: The final set of hotwords (112) is enabled for detection by executing a speech recognizer (116) on the AED (104), and the speech recognizer (116) is biased towards recognizing the hotwords (112) in the final set of hotwords (112); and Detecting the hotwords (112) from the final set of hotwords (112) in the audio data (402) includes: using the speech recognizer (116) executed on the AED (104) to recognize repeated hotwords (112) in the audio data (402).
19. A system (100), comprising: Data processing hardware (710); and Memory hardware (720) communicatively coupled to the data processing hardware (710), the memory hardware (720) storing instructions that, when executed on the data processing hardware (710), cause the data processing hardware (710) to perform operations including: Detect the presence of multiple users within the environment of an assistant-supporting device AED (104) that executes a digital assistant (105); For each of the multiple users, obtain a corresponding active set of hotwords (112) that each specify a corresponding action for the digital assistant (105) to perform; Based on the corresponding active sets of hotwords (112) for each of the multiple users, execute a hotword arbitration routine (401) to enable a final set of hotwords (112) for detection by the AED (104), and the final set of hotwords (112) enabled for detection by the AED (104) includes hotwords (112) selected from the corresponding active sets of hotwords (112) for at least one of the multiple users detected within the environment of the AED (104); and When the final set of hotwords (112) is enabled for detection by the AED (104): Receive audio data (402) corresponding to an utterance (106) captured by the AED (104); Detect hotwords (112) from the final set of hotwords (112) in the audio data (402); and Instruct the digital assistant (105) to perform the corresponding action specified by the detected hotwords (112).
20. The system (100) according to claim 19, wherein detecting the presence of the multiple users within the environment of the AED (104) includes: Detect the presence of at least one user among the plurality of users in the environment based on proximity information (54) of the AED (104) relative to a user device (50) associated with at least one user among the plurality of users.
21. The system (100) according to claim 19, wherein detecting the presence of the plurality of users in the environment of the AED (104) comprises: receiving image data (312) corresponding to a scene of the environment; and detecting the presence of at least one user among the plurality of users in the environment based on the image data (312).
22. The system (100) according to claim 19, wherein detecting the presence of the plurality of users in the environment of the AED (104) comprises detecting the presence of a corresponding user (102) in the environment of the AED (104) based on: receiving voice data characterizing a voice query issued by the corresponding user (102) and directed to the digital assistant (105); performing speaker recognition on the voice data to identify the corresponding user (102) who issued the voice query; and determining that the identified corresponding user (102) who issued the voice query is present in the environment of the AED (104).
23. The system (100) according to claim 22, wherein: the voice query issued by the corresponding user (102) includes a command (118), the command causing the digital assistant (105) to perform a long-term operation specified by the command (118); and obtaining the corresponding active set of warm words (112) includes: in response to the digital assistant (105) performing the long-term operation specified by the command (118), adding one or more warm words (112) specifying corresponding actions for controlling the long-term operation to the corresponding active set of warm words (112) for the corresponding user (102) who issued the voice query.
24. The system (100) according to any one of claims 19 to 23, wherein the operation further comprises: detecting the presence of a new user in the environment of the AED (104), wherein the warm word arbitration routine (401) is executed in response to detecting the presence of the new user in the environment.
25. The system (100) according to any one of claims 19 to 24, wherein the operation further comprises: determining that a user among the plurality of users is no longer present in the environment of the AED (104), wherein the warm word arbitration routine (401) is executed in response to determining that a user among the plurality of users is no longer present in the environment.
26. The system (100) according to any one of claims 19 to 25, wherein the operation further comprises: Determine whether to add a new warm word (112) to the corresponding active set of warm words (112) for a user among the multiple users existing in the environment of the AED (104), or remove one of the warm words (112) from the corresponding active set of warm words, wherein the warm word arbitration routine (401) is executed in response to determining to add the new warm word (112) to the corresponding active set of warm words (112) for the user among the multiple users existing in the environment of the AED (104), or remove one of the warm words (112) from the corresponding active set of warm words.
27. The system (100) according to any one of claims 19 to 26, wherein the operation further comprises: Determine a change in the environmental context of the AED (104), wherein the warm word arbitration routine (401) is executed in response to determining the change in the environmental context.
28. The system (100) according to any one of claims 19 to 27, wherein executing the warm word arbitration routine (401) comprises: Obtain enabled warm word constraints (430), the enabled warm word constraints (430) including at least one of the following: Availability of memory and computing resources on the AED (104) for detecting warm words (112); Computing requirements for each warm word (112) in the corresponding active set of warm words (112) enabled for each user among the multiple users existing in the environment of the AED (104); Acceptable false acceptance rate tolerance; or Acceptable false rejection rate tolerance; and Based on the enabled warm word constraints (430), determine the number of warm words (112) to be enabled in the final set of warm words for detection by the AED (104).
29. The system (100) according to any one of claims 19 to 28, wherein executing the warm word arbitration routine (401) comprises: For each corresponding user (102) among the multiple users, sort the warm words (112) in the corresponding active set of warm words (112) from the highest priority to the lowest priority based on a warm word priority sorting signal (413), the warm word priority sorting signal (413) including at least one of the following: The usage frequency of each warm word (112) in the corresponding active set of warm words (112) by the corresponding user (102); The current state of the AED (104); The environmental context of the AED (104); or Coexistence information indicating previous warm word (112) usage and / or operations performed by the AED (104) when the corresponding user (102) was previously present in the environment together with one or more combinations of other users among the multiple users; and Enable the final set of warm words (112) for detection by the AED (104) based on the ranking of the warm words (112) in the respective active set of warm words (112) for each corresponding user (102) among the multiple users.
30. The system (100) according to any one of claims 19 to 29, wherein performing the warm word arbitration routine (401) comprises: For each corresponding user (102) among the multiple users, determine a warm word affinity score; and Enable the final set of warm words (112) for detection by the AED (104) based on the warm word affinity scores determined for each corresponding user (102).
31. The system (100) according to claim 30, wherein the warm word affinity scores determined for each corresponding user (102) are based on at least one of the following: The frequency of use of the warm words (112) by the corresponding user (102); The frequency of interaction between the corresponding user (102) and the digital assistant (105); The duration of the presence of the corresponding user (102) in the environment of the AED (104); The proximity of the user to the AED (104); or The current user context.
32. The system (100) according to any one of claims 19 to 31, wherein performing the warm word arbitration routine (401) comprises: Identify any shared warm words (112) corresponding to the warm words (112) present in the respective active sets of warm words (112) for at least two users among the multiple users; and Determine that the final set of warm words (112) is based on assigning a higher priority to the warm words (112) identified as shared for inclusion in the final set of warm words (112).
33. The system (100) according to any one of claims 19 to 32, wherein the operation further comprises: Determine that the warm words (112) detected in the audio data (402) corresponding to the utterance (106) captured by the user device (50) include speaker-specific warm words (112) selected from the respective active sets of warm words (112) for each of one or more users among the multiple users, such that the digital assistant (105) performs the respective actions specified by the speaker-specific warm words (112) only when the speaker-specific warm words (112) are spoken by any of the one or more users corresponding to the respective active sets of warm words (112); and Based on determining that the warm word (112) detected in the audio data (402) corresponding to the utterance (106) captured by the user device (50) includes a speaker-specific warm word (112), speaker verification is performed on the audio data (402) to determine that the utterance (106) is spoken by one of the one or more users corresponding to the respective active set of warm words (112) from which the detected warm word (112) was selected. Wherein instructing the digital assistant (105) to perform the respective action specified by the detected warm word (112) is based on the speaker verification performed on the audio data (402).
34. The system (100) according to any one of claims 19 to 33, wherein: The final set of warm words (112) is enabled for detection by activating the respective warm word model (330) for each warm word (112) in the final set of warm words (112) to run on the assistant-enabled device; and Detecting the warm word (112) from the final set of warm words (112) in the audio data (402) includes: using the activated respective warm word model (330) to detect the warm word (112) in the audio data (402) without performing speech recognition on the audio data (402).
35. The system (100) according to claim 34, wherein detecting the warm word in the audio data includes: Extracting audio features of the audio data (402); Using the activated respective warm word model (330) to generate a warm word confidence score by processing the extracted audio features; and When the warm word confidence score meets the warm word (112) confidence threshold, determining that the audio data (402) corresponding to the utterance (106) includes the warm word (112).
36. The system (100) according to any one of claims 19 to 35, wherein: The final set of warm words (112) is enabled for detection by performing a speech recognizer (116) on the AED (104), the speech recognizer (116) being biased towards recognizing the warm word (112) in the final set of warm words (112); and Detecting the warm word (112) from the final set of warm words (112) in the audio data (402) includes: using the speech recognizer (116) performed on the AED (104) to recognize the repeated warm word (112) in the audio data (402).