Voice query handling in environment with multiple users

By implementing the control method of cyclic mode and cyclic queue in a multi-user environment, the problem of multi-user competition query actions is solved, and the user experience is improved and the device is reasonably allocated.

CN120077433APending Publication Date: 2025-05-30GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380074718.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-06
Filing Date
2023-09-12
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In a multi-user environment, the support assistant device needs to handle query actions that may compete between multiple users, resulting in improper execution of the actions or poor user experience.

Method used

By executing a loop mode on the data processing hardware, the digital assistant controls the action execution of the query command based on the loop queue, ensuring that each user has the opportunity to issue a query and prevent the execution of the repeated action by the speaker identification.

Benefits of technology

It effectively avoids the situation where a single user exclusively occupies the device, ensures fairness among multiple users and reasonable allocation of devices, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077433A_ABST
    Figure CN120077433A_ABST
Patent Text Reader

Abstract

The method (500) includes receiving a first query (116) issued by a first user, the first query including a command (111) for causing a digital assistant (105) to perform a first action, and enabling a loop mode (350) to control execution of the action. The method further includes, while performing the first action, receiving audio data (402) corresponding to a second query (146) including a command for performing a second action, performing speaker recognition on the audio data, determining that the second query is spoken by the first user, blocking performing the second action, and prompting at least another user to issue a query. The method further includes receiving a third query (148) issued by a second user, the third query including a command for causing the digital assistant to perform a third action, and performing execution of the third action when the digital assistant completes performing the first action.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to handling voice queries in an environment with multiple users. Background Art

[0002] If not exclusively, the way users interact with an assistant-enabled device is mainly designed with voice input. For example, a user can ask the device to perform an action including media playback (e.g., music or podcasts), where the device responds by initiating the playback of audio that matches the user's criteria. In instances where the device (e.g., a smart speaker) is shared by multiple users in an environment, the device may need to handle multiple actions that may compete with each other requested by the users. Summary of the Invention

[0003] One aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations including detecting multiple users in the environment of an assistant-enabled device (AED). The operations further include receiving a first query issued by a first user among the multiple users. The first query includes a command for causing a digital assistant to perform a first action. The operations further include enabling a round-robin mode that, when enabled, causes the digital assistant to control the execution of actions of query commands after the first query based on a round-robin queue. Here, the round-robin queue includes the multiple users detected in the environment of the AED. When the digital assistant is performing the first action and when the round-robin mode is enabled, the operations further include receiving audio data corresponding to a second query spoken by one of the multiple users and captured by the AED. Here, the second query includes a command for causing the digital assistant to perform a second action. The operations also include performing speaker recognition on the audio data corresponding to the second query to determine that the second query is spoken by the first user who issued the first query. Based on determining that the second query is spoken by the first user who issued the first query, the operations further include preventing the digital assistant from performing the second action and prompting at least another user different from the first user among the multiple users detected in the environment to have an opportunity to issue a query. The operations further include receiving a third query issued by a second user among the multiple users detected in the environment, the third query includes a command for causing the digital assistant to perform a third action, and when the digital assistant finishes performing the first action, the operations further include performing the execution of the third action.

[0004] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, prompting at least the other user among the plurality of users detected in the environment includes providing user-selectable options as an output of a user interface of a user device associated with the second user, the user-selectable options, when selected, issuing a third query from the second user to perform the third action. Here, receiving the third query issued by the second user is based on receiving a user input indication indicating a selection of the user-selectable option. In these implementations, providing the user-selectable options as an output from the user interface may include displaying the user-selectable options as graphical elements on a screen of a user device associated with the second user via the user interface. The graphical elements prompt the second user with an opportunity to issue the third query. Additionally or alternatively, providing the user-selectable options as an output from the user interface includes providing the user-selectable options as an audible output from a speaker communicating with the data processing hardware via the user interface. The audible output prompts the second user with an opportunity to issue the third query.

[0005] In some examples, the first query further includes a constraint on subsequent queries, the constraint including one of an action category, a time limit for an action, a time limit for the loop pattern, or a threshold number of actions per query. In these examples, the operation further includes determining that the second query does not violate the constraint in the first query, and updating the loop queue to include the second query spoken by the first user. In these examples, the operation may further include detecting that the first user has left the environment of the device of the support assistant, and updating the loop queue to remove the second query spoken by the first user.

[0006] In some implementations, detecting the plurality of users within the environment of the device of the support assistant includes detecting at least one of the plurality of users based on proximity information of a user device associated with at least one of the plurality of users. Additionally or alternatively, detecting the plurality of users within the environment of the device of the support assistant includes receiving image data corresponding to a scene of the environment, and detecting at least one of the plurality of users based on the image data. In other implementations, detecting the plurality of users within the environment of the device of the support assistant includes receiving a list indicating each of the plurality of users to be added to the loop queue. In some examples, the loop queue includes, for each corresponding user among the plurality of users detected in the environment, the identity of the corresponding user and a query count of queries received from the corresponding user. In some implementations, receiving the first query issued by the first user includes receiving a user input indication indicating a user intent to issue the first query from a user device associated with the first user.

[0007] In some examples, receiving the first query issued by the first user includes receiving initial audio data corresponding to the first query issued by the first user and captured by the device of the support assistant. In these examples, after receiving the initial audio data corresponding to the first query issued by the first user, the operation further includes performing speaker recognition on the initial audio data to identify the first user who issued the first query. The speaker recognition includes extracting a first speaker discrimination vector representing the characteristics of the first query issued by the first user from the initial audio data corresponding to the first query issued by the first user, and determining that the extracted speaker discrimination vector matches any of the registered speaker vectors stored on the device of the support assistant. Each registered speaker vector is associated with a different corresponding registered user of the device of the support assistant. When the first speaker discrimination vector matches one of the registered speaker vectors, the operation further includes identifying the first user who issued the first query as the corresponding registered user associated with the one of the registered speaker vectors that matches the extracted speaker discrimination vector. In some implementations, performing speaker recognition on the audio data corresponding to the second query to determine that the second query was spoken by the first user who issued the first query includes extracting a second speaker discrimination vector representing the characteristics of the second query from the audio data corresponding to the second query, and determining that the second extracted speaker discrimination vector matches the reference speaker vector of the first user.

[0008] Another aspect of the present disclosure provides a system that includes data processing hardware and memory hardware that communicates with the data processing hardware. The memory hardware stores instructions that, when executed by the data processing hardware, cause the data processing hardware to perform operations that include detecting a plurality of users in an environment of an assistant-supporting device (AED). The operations further include receiving a first query issued by a first user among the plurality of users. The first query includes a command for causing a digital assistant to perform a first action. The operations further include enabling a loop mode that, when enabled, causes the digital assistant to control the execution of actions of query commands after the first query based on a loop queue. Here, the loop queue includes the plurality of users detected within the environment of the AED. When the digital assistant is performing the first action and when the loop mode is enabled, the operations further include receiving audio data corresponding to a second query spoken by one of the plurality of users and captured by the AED. Here, the second query includes a command for causing the digital assistant to perform a second action. The operations also include performing speaker recognition on the audio data corresponding to the second query to determine that the second query was spoken by the first user who issued the first query. Based on determining that the second query was spoken by the first user who issued the first query, the operations also include preventing the digital assistant from performing the second action and prompting at least one other user among the plurality of users detected in the environment and different from the first user to have an opportunity to issue a query. The operations further include receiving a third query issued by a second user among the plurality of users detected in the environment, the third query including a command for causing the digital assistant to perform a third action, and when the digital assistant finishes performing the first action, the operations further include performing the execution of the third action.

[0009] This aspect may include one or more of the following optional features. In some implementations, prompting at least the other user among the plurality of users detected in the environment includes providing user-selectable options as an output of a user interface of a user device associated with the second user, the user-selectable options, when selected, issuing a third query from the second user to perform the third action. Here, receiving the third query issued by the second user is based on receiving a user input indication indicating a selection of the user-selectable option. In these implementations, providing the user-selectable options as an output from the user interface may include displaying the user-selectable options as graphical elements on a screen of the user device associated with the second user via the user interface, the graphical elements prompting the second user with an opportunity to issue the third query. Additionally or alternatively, providing the user-selectable options as an output from the user interface includes providing the user-selectable options as an audible output from a speaker communicating with the data processing hardware via the user interface. The audible output prompts the second user with an opportunity to issue the third query.

[0010] In some examples, the first query further includes a constraint on subsequent queries, the constraint including one of an action category, a time limit for the action, a time limit for the loop pattern, or a threshold number of actions per query. In these examples, the operation further includes determining that the second query does not violate the constraint in the first query, and updating the loop queue to include the second query spoken by the first user. In these examples, the operation may further include detecting that the first user has left the environment of the device of the support assistant, and updating the loop queue to remove the second query spoken by the first user.

[0011] In some implementations, detecting the plurality of users within the environment of the device of the support assistant includes detecting at least one of the plurality of users based on proximity information of a user device associated with at least one of the plurality of users. Additionally or alternatively, detecting the plurality of users within the environment of the device of the support assistant includes receiving image data corresponding to a scene of the environment, and detecting at least one of the plurality of users based on the image data. In other implementations, detecting the plurality of users within the environment of the device of the support assistant includes receiving a list indicating each of the plurality of users to be added to the loop queue. In some examples, the loop queue includes, for each corresponding user among the plurality of users detected in the environment, the identity of the corresponding user and a query count of queries received from the corresponding user. In some implementations, receiving the first query issued by the first user includes receiving a user input indication indicating a user intent to issue the first query from a user device associated with the first user.

[0012] In some examples, receiving the first query issued by the first user includes receiving initial audio data corresponding to the first query issued by the first user and captured by the device of the support assistant. In these examples, after receiving the initial audio data corresponding to the first query issued by the first user, the operation further includes performing speaker identification on the initial audio data to identify the first user who issued the first query. The speaker identification includes extracting a first speaker discrimination vector representing the characteristics of the first query issued by the first user from the initial audio data corresponding to the first query issued by the first user, and determining that the extracted speaker discrimination vector matches any registered speaker vector stored on the device of the support assistant. Each registered speaker vector is associated with a different corresponding registered user of the device of the support assistant. When the first speaker discrimination vector matches one of the registered speaker vectors, the operation further includes identifying the first user who issued the first query as the corresponding registered user associated with the one of the registered speaker vectors that matches the extracted speaker discrimination vector. In some implementations, performing speaker identification on the audio data corresponding to the second query to determine that the second query was spoken by the first user who issued the first query includes extracting a second speaker discrimination vector representing the characteristics of the second query from the audio data corresponding to the second query, and determining that the second extracted speaker discrimination vector matches the reference speaker vector of the first user.

[0013] Details of one or more implementations of the present disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1A – Figure 1C is a schematic diagram of an example system including multiple users who control the device of the support assistant.

[0015] Figure 2A – Figure 2C is an example GUI rendered on the screen of a user device to display a circular queue.

[0016] Figure 3 is a schematic diagram of a query handling process.

[0017] Figure 4A is a schematic diagram of a speaker identification process.

[0018] Figure 4B is a schematic diagram of a speaker verification process.

[0019] Figure 5A flowchart of an example operational arrangement of a method for handling voice queries in an environment with multiple users.

[0020] Figure 6 A schematic diagram of an example computing device that can be used to implement the systems and methods described herein.

[0021] In the various figures, the same reference symbols indicate the same elements. Detailed Description

[0022] If not exclusively, the manner in which a user interacts with a support assistant device is primarily designed by means of voice input. For example, a user can request that the device perform an action including media playback (e.g., music or podcasts), where the device responds by initiating the playback of audio that matches the user's criteria. In instances where the device (e.g., a smart speaker) is shared by multiple users in an environment, the device may need to handle multiple actions that may compete with each other as requested by the users. In cases where one or more of the multiple users issue multiple individual requests to the device, the device can respond to the individual requests in an ordered manner while actively prompting users who have not yet made requests to submit requests. By ensuring that each user in the environment has an opportunity to submit a request to the device, instances where individual users monopolize the device are avoided.

[0023] Once the device receives a request from a user, it can implement one or more scheduling policies to resolve the individual request. For example, the device can implement a round-robin scheduling policy, where each user's first request is resolved individually in order before resolving subsequent requests from each user. Alternatively, the device can implement a scheduling policy based on the length of the requested action, the importance of the requested action, or the priority of the requested action. For example, the device can grant priority to requests submitted by the host of an event. Similarly, the device can implement a scheduling policy that divides users into multiple queues, where each queue in the multiple queues corresponds to a type of action (e.g., playing audio, controlling lights, etc.).

[0024] In addition to controlling playback actions to prevent a single user from monopolizing a playlist, the device can also control other types of media, such as podcasts and videos. Similarly, the device can prevent a single user from monopolizing the device with multiple follow-up questions by prompting other users present to submit questions before answering the multiple follow-up questions. This can additionally extend to controlling various aspects of the home connected to the device. For example, the host of a party can seek to control the lighting level or the type of music being played during the party to ensure a soothing atmosphere. The host can say "only allow lighting below 60% and jazz music". For the duration of the party, the device can prevent or limit the extent to which other participants at the party can adjust the lighting level and prohibit participants from requesting music outside the jazz music genre.

[0025] In addition to restricting individuals, the device can also operate to be more inclusive of the individuals present in the home. For example, the device can assist the individuals in the environment in creating a shopping list, thereby ensuring that all individuals are given the opportunity to add items to the shopping list. For example, in response to an initial individual issuing a query to add multiple items to the shopping list, the device can prompt the individuals to add items to the shopping list. Similarly, the device can proactively prompt individuals to participate in events, such as setting the morning alarm. Here, the device can suggest that an individual request an alarm in response to receiving a request to set an alarm for a different individual / receiving a request to set an alarm from a different individual. Additionally, the device can ensure that individuals are not excluded from the interaction by attracting / prompting an individual who has not spoken recently to join the interaction between the device and another individual near the individual.

[0026] Figure 1A – Figure 1C An example system 100a–c for handling queries in an environment with multiple users 102, 102a–n using a scheduling strategy that employs a cyclic pattern is shown. The cyclic pattern actively attracts the multiple users 102 detected in the environment based on a cyclic queue 350. Briefly, and as described in more detail below, including a query handler 300 ( Figure 3)The support assistant device (AED) 104 detects multiple users 102, 102a–c in the environment and starts playing music 122 in response to receiving a first query 106 “Ok computer, play Glory, and let’s stick to pop music tonight” issued by user 102a. While the AED 104 is performing the action of playing music 122 as playback audio from the speaker 18, the AED 104 receives a second query 146 “Play Cold Water next” spoken by the same user 102a ( Figure 1B ). Because the query handler 300 detects / identifies that there are other users 102b, 102c in the environment, the query handler 300 prompts one or more of the other users 102b, 102c to issue queries to be added to the loop queue 350 before performing the action associated with the second query 146 issued by user 102a.

[0027] Systems 100a–100c include an AED 104 that executes a digital assistant 105 with which multiple users 102 can interact by issuing queries that include commands for performing actions. In the example shown, the AED 104 corresponds to a smart speaker with which multiple users 102 can interact. However, the AED 104 can include other computing devices such as, but not limited to, smart phones, tablet computers, smart displays, desktop / laptop computers, smart watches, smart glasses / head-mounted devices, smart appliances, head-mounted headphones, or vehicle infotainment devices. The AED 104 includes data processing hardware 10 and memory hardware 12 that stores instructions that, when executed on the data processing hardware 10, cause the data processing hardware 10 to perform operations. The AED 104 includes an array of one or more microphones 16 that are configured to capture acoustic sounds such as speech directed at the AED 104. The AED 104 can also include or communicate with an audio output device (e.g., a speaker) 18 that can output audio such as music 122 and / or synthesized speech from the digital assistant 105. Additionally, the AED 104 can include or communicate with one or more cameras 19 that are configured to capture images within the environment and output image data 312 ( Figure 3 ).

[0028] In some configurations, the AED 104 communicates with a plurality of user devices 50, 50a–n associated with a plurality of users 102. In the example shown, each of the plurality of user devices 50a–c includes a smart phone with which the respective user 102 can interact. However, the user device 50 can include other computing devices such as, but not limited to, smart watches, smart displays, smart glasses, smart phones, smart glasses / head-mounted devices, tablet computers, smart appliances, headphones, computing devices, smart speakers, or another assistant-enabled device. Each of the plurality of user devices 50a–n can include at least one microphone 52, 52a–n resident on the user device 50 that communicates with the AED 104. In these configurations, the user device 50 can also communicate with one or more microphones 16 resident on the AED 104. Additionally, the plurality of users 102 can control and / or configure the AED 104 and interact with the digital assistant 105 using an interface 200 such as a graphical user interface (GUI) 200, 200a–n rendered to be displayed on a respective screen of each user device 50.

[0029] As Figure 1A – Figure 1C and Figure 3 shown, the digital assistant 105 implementing the query processor 300 uses a circular queue 350 to manage queries issued by the plurality of users 102. In some implementations, the query processor 300 maintains a log of the detected users 102 in the environment and one or more queries received from each of the users 102 in the circular queue 350 and ensures that the digital assistant 105 performs actions associated with the received queries in an ordered manner according to the circular queue 350. In this sense, the query processor 300 maintains the order of the queries received from the plurality of users 102 within the circular queue 350 and prevents a single user 102 from monopolizing the digital assistant 105 with consecutive queries. Additionally, the query processor 300 can facilitate adding additional and / or conflicting queries from the plurality of users 102 to the circular queue 350.

[0030] In some examples, the query processor 300 identifies metadata in the issued queries that specifies the order of actions recorded within the circular queue 350. Here, the metadata for an issued query to “play” a song can include the length or tempo of the song, where the query processor 300 can perform performances of consecutive short songs but separate two long songs with other songs in between to avoid performing consecutive long songs. Similarly, the query 300 can group music of similar genres or tempos together within the constraints of the circular queue 350 to provide a smoother transition between actions.

[0031] The GUI 200 of each corresponding user device 50 associated with a user 102 can display a circular queue 350 that includes the detected identity of the user 102 and queries associated with each corresponding user 102. For example, the circular queue 350 includes, for each user 102 among a plurality of users 102 detected in the environment, the identity of the corresponding user 102 and a query count for the queries received from the corresponding user 102. The query count can refer to the number of songs that the user 102 has submitted, where each song corresponds to an entry in the circular queue 350, or can refer to the number of actions that the user 102 has submitted, where each entry in the circular queue 350 corresponds to an action that can include multiple songs. In some configurations, the AED 104 includes a screen and renders the GUI 200 to display the circular queue 350 on the screen. For example, the AED 104 can include a smart display, a tablet computer, or a smart TV within the environment. Figure 2A – Figure 2C Example GUIs 200b1–200b3 are provided that are displayed on the screen of the user device 50b associated with the user 102b to inform the user 102b of the status of the circular queue 350. Additionally, the GUI 200 can be rendered to display an identifier of the current action (e.g., "Play Glory"), an identifier of the AED 104 that is currently performing the action (e.g., a smart speaker), and / or the identity of the user 102a who initiated the current action being performed by the digital assistant 105 (e.g., Barb). As mentioned above, by enabling the circular mode in which the query handler 300 manages the circular queue 350, when the AED 104 (e.g., via the query handler 300) determines that Barb 102a has issued a first action recorded in the circular queue 350, the AED (e.g., via the digital assistant 105) will block the execution of a second action (or at least require approval from other users 102b, 102c).

[0032] Continuing to refer Figure 1A – Figure 1C and Figure 3 and, during the execution of the digital assistant 105, the AED 104 uses the user detector 310 of the query handler 300 to detect multiple users 102a–c in the environment. For example, the query handler 300 receives proximity information 54 of each of the multiple users 102a–c relative to the position of the AED 104 via the user detector 310 ( Figure 3)。In some implementations, each user device 50a–c of multiple users 102 broadcasts proximity information 54 that can be received by the user detector 310, and the AED 104 can use this proximity information to determine the proximity of each user device 50 relative to the AED 104. The proximity information 54 from each user device 50 can include a wireless communication signal, such as WiFi, Bluetooth, or ultrasonic, where the signal strength of the wireless communication signal received by the user detector 310 can be related to the proximity (e.g., distance) of the user device 50 relative to the AED 104. In implementations where the user 102 does not have a user device 50 or has a user device 50 that does not share proximity information 54, the user detector 310 can detect the user 102 based on an explicit input (e.g., a visitor list) 313 received from the user 102a that issued the first query 106. For example, the user detector 310 receives a visitor list 313 from a seed user 102 (e.g., user 102a), and the visitor list indicates each user 102 among the multiple users 102 to be added to the circular queue 350. In other implementations, the user detector 310 automatically detects the multiple users 102 in the environment by receiving image data 312 corresponding to the scene of the environment and obtained by the camera 19. Here, the user detector 310 detects the multiple users 102 based on the received image data 312. Similarly, the user detector 310 can detect the multiple users 102 in the environment by performing speaker recognition ( Figure 4A and Figure 4B ) to resolve the identity of the user 102 who issued a query by speaking within the environment. Here, the user detector 310 detects the multiple users 102 based on the received audio data 402 associated with the issued query, and the user detector 310 can continue to detect the user 102 who issued the query (and the associated audio data 402) for a threshold amount of time after the user 102 has spoken.

[0033] In some implementations, the user detector 310 maintains a list of previous users 314 present in the environment of the AED 104. Here, the list of previous users 314 can refer to a list of users 102 previously detected by the user detector 310, which can include one or more users not in the most recent (i.e., latest) detection of the user 102 and / or the most recent detection of the user 102 can include one or more additional users not in the list of previous users 314. In this example, after receiving the proximity information 54, the image data 312, and / or the visitor list, the user detector 310 can determine that the list of previous users 314 does not include the same users 102 associated with the current state of the environment (i.e., the list of current users 316). In other words, the user detector 310 can determine that the users 102 in the list of previous users 314 are different from the users 102 in the list of current users 316. This change between the list of previous users 314 and the list of current users 316 triggers the query handler 300 to update the circular queue 350 to add entries to the circular queue 350 or remove entries from the circular queue 350 based on the list of current users 316, and generate an update 352 to be displayed in the GUI 300 of each user device 50 associated with the detected user 102. For example, if user 102a leaves the environment, the user detector 310 can detect that user 102a has left by comparing the list of previous users 314 that includes user 102a with the list of current users 314 that does not include user 102a. In response, the query handler 300 updates the circular queue 350 to remove any queries issued by user 102a and generates an update to the circular queue 350. However, in some examples, when the user detector 310 determines that the list of previous users 314 is the same as the list of detected current users 316, the user detector 310 may not send the list of current users 316 to update the circular queue 350.

[0034] In some examples, when there is a difference between the list of the previous user 314 and the list of the current user 316, the user detector 310 outputs only the list of the current user 316 (thereby triggering the query handler to generate an update 352 to the circular queue 350). For example, the user detector 310 may be configured with a change threshold, and when the difference detected between the list of the previous user 314 and the list of the current user 316 meets the change threshold (e.g., exceeds the threshold), the user detector 310 outputs the list of the current user 316 to the query handler 300. The threshold may be zero, where the slightest difference between the list of the previous user 314 and the list of the current user 316 detected by the indication determiner 210 (e.g., as long as user 102 enters or leaves the environment) may trigger the query handler 300 to update the circular queue 350. Conversely, the threshold may be higher than zero to prevent unnecessary updates to the circular queue 350, as a type of queue interruption sensitivity mechanism. For example, the change threshold may be time-based (e.g., a certain amount of time), where if user 102 temporarily leaves the environment (e.g., goes to a different room) but returns within the threshold amount of time, the AED 104 does not update the circular queue 350. Here, when the difference between the list of the previous user 314 and the list of the current user 316 continues to reach the threshold amount of time, the user detector 310 outputs only the list of the current user 316, thereby updating the circular queue (350).

[0035] Referring again to Figure 1A – Figure 1C , the user detector 310 can identify the two users via proximity information 54 received from the respective user devices 50a, 50b of user 102a (e.g., Barb) and user 102b (e.g., Jeff). In this example, the user detector 310 may not detect user 102c (e.g., user 102c chooses not to share proximity information 54 from user device 50c) by detecting user device 50c. However, the user detector 310 still detects the presence of user 102c (e.g., via voice data, image data, and / or input from other users 102a, 102b) and outputs the list of the current user 316. Thereafter, the query handler 300 adds the list of the current user 316 including the detected users 102a–c to the circular queue 350 that records the queries issued by users 102a–c

[0036] Continuing Figure 1AIn the example, user 102a among multiple users 102a–c is shown issuing a first query 106 “Ok computer, play Glory, and let’s stick to pop music tonight” near AED 104. Here, the first query 106 issued by user 102a is spoken by user 102a and includes initial audio data 402 corresponding to the first query 106 ( Figure 3 ). The first query 106 may further include a user input indication indicating a user intent to issue the first query via any one of touch, voice, gesture, gaze, and / or an input device (e.g., a mouse or a stylus) for interacting with AED 104. Optionally, based on receiving the initial audio data 402 corresponding to the first query 106, the query processor 300 resolves the identity of the speaker of the first query 106 by performing a speaker recognition process 400a ( Figure 4A ) on the audio data 402 and determining that the first query 106 is issued by user 102a. In other implementations, user 102a issues the first query 106 without speaking. In these implementations, user 102a issues the first query 106 via a user device 50a associated with user 102a (e.g., entering text corresponding to the first query 106 into a GUI 200a displayed on the screen of user device 50a, selecting the first query 106 displayed on the screen of user device 50a, etc.). Here, AED 104 may resolve the identity of user 102 who issued the first query 106 by identifying the user device 50a associated with user 102a.

[0037] The microphone 16 of AED 104 receives the first query 106 and processes the initial audio data 402 corresponding to the first query 106. The initial processing of the audio data 402 may involve filtering the audio data 402 and converting the audio data 402 from an analog signal to a digital signal. When AED 104 processes the audio data 402, AED may store the audio data 402 in a buffer of the memory hardware 12 for additional processing. In the case where there is audio data 402 in the buffer, AED 104 may use a hotword detector 108 to detect whether the audio data 402 includes a hotword. The hotword detector 108 is configured to identify hotwords included in the audio data 402 without performing speech recognition on the audio data 402.

[0038] In some implementations, the hotword detector 108 is configured to identify a hotword in the initial portion of the first query 106. In this example, if the hotword detector 108 detects acoustic features that are characteristic of the hotword 110 in the audio data 402, the hotword detector 108 may determine that the first query 106 “Ok computer, play Glory, and let’s stick to pop music tonight” includes the hotword 110 “ok computer”. The acoustic features may be Mel Frequency Cepstral Coefficients (MFCCs), which are representations of the short-term power spectrum of the first query 106, or may be the Mel-scale filter bank energies of the first query 106. For example, the hotword detector 108 may detect that the first query 106 “Ok computer, play Glory, and let’s stick to pop music tonight” includes the hotword 110 “ok computer” based on generating MFCCs from the audio data 402 and classifying the MFCCs as including MFCCs similar to those characteristic of the hotword “ok computer” stored in the hotword model of the hotword detector 108. As another example, the hotword detector 108 may detect that the first query 106 “Ok computer, play Glory, and let’s stick to pop music tonight” includes the hotword 110 “ok computer” based on generating Mel-scale filter bank energies from the audio data 402 and classifying the Mel-scale filter bank energies as including Mel-scale filter bank energies similar to those characteristic of the hotword “ok computer” stored in the hotword model of the hotword detector 108.

[0039] When the hotword detector 108 determines that the initial audio data 402 corresponding to the first query 106 includes the hotword 110, the AED 104 may trigger a wake-up process to initiate speech recognition of the audio data 402 corresponding to the first query 106. For example, Figure 3 It is shown that the AED 104 includes a speech recognizer 170 that employs an automatic speech recognition model 172, which may perform speech recognition or semantic interpretation on the audio data 402 corresponding to the first query 106. The speech recognizer 170 may perform speech recognition on the portion of the audio data 402 that follows the hotword 110. In this example, the speech recognizer 170 may identify the words “play Glory, and let’s stick to pop music tonight” in the first query 106.

[0040] In some examples, the AED 104 is configured to communicate with the remote system 130 via the network 120. The remote system 130 may include remote resources such as remote data processing hardware 132 (e.g., a remote server or CPU) and / or remote memory hardware 134 (e.g., a remote database or other storage hardware). The query processor 300 may be executed on the remote system 130 as a supplement to or an alternative to the AED 104. The AED 104 may utilize the remote resources to perform various functions related to speech processing and / or synthesized playback communication. In some implementations, the speech recognizer 170 is located on the remote system 130 as a supplement to or an alternative to the AED 104. After the hotword detector 108 triggers the AED 104 to wake up in response to detecting the hotword 110 in the first query 106, the AED 104 may send the initial audio data 402 corresponding to the first query 106 to the remote system 130 via the network 120. Here, the AED 104 may send the portion of the initial audio data 402 that includes the hotword 110 for the remote system 130 to confirm the presence of the hotword 110. Alternatively, the AED 104 may only send the portion of the initial audio data 402 corresponding to the portion of the utterance 106 after the hotword 110 to the remote system 130, where the remote system 130 executes the speech recognizer 170 to perform speech recognition and returns the transcription of the initial audio data 402 to the AED 104.

[0041] Continuing to refer Figure 1A – Figure 1C and Figure 3 , the query processor 300 may further include a natural language understanding (NLU) module 320 that performs semantic interpretation on the first query 106 to identify queries / commands directed to the AED 104. Specifically, the NLU module 320 identifies the words recognized by the speech recognizer 170 in the first query 106 and performs semantic interpretation to identify any speech commands in the first query 106. The NLU module 320 of the AED 104 (and / or the remote system 130) may identify the word "playGlory" as a command 111 for performing a first action (i.e., playing the music 122) and the word "and let’s stickto pop music tonight" as a constraint 113 for restricting the action of the query command after the first query 106. In Figure 1AIn the example shown, digital assistant 105 begins to perform a first action of playing music 122 as playback audio (e.g., track #1) from speaker 18 of AED 104. Digital assistant 105 may stream music 122 from a streaming service (not shown), or digital assistant 105 may instruct AED 104 to play music stored on AED 104. Additionally, query processor 300 enables a loop mode that includes loop queue 350. The loop mode, when enabled, causes digital assistant 105 to control the execution of actions of query commands following first query 106 based on loop queue 350. Here, loop queue 350 includes a list of current users 316, the list of current users including multiple users 102a–c detected by user detector 310 within the environment of AED 104. In this example, in response to receiving constraint 113 in first query 106, query processor 300 restricts the actions of digital assistant 105 to play only a certain genre of music (i.e., pop music).

[0042] In Figure 1A the example shown, also referring to Figure 3 , query processor 300 enables a loop mode that includes loop queue 350 and adds constraint 113 to action constraint data store 330. Query processor 300 maintains a record of action constraints 332 in action constraint data store 330 (e.g., stored on memory hardware 12), and query processor 300 may restrict which actions are added to loop queue 350 for digital assistant 105 to execute based on action constraints 332. For example, query processor 300 may first verify that command 111 and / or constraint 113 in first query 106 do not conflict with any action constraints 332, and then execute the first action associated with command 111 and add the action “play Glory” to the loop queue( Figure 2A ).

[0043] The AED 104 can notify a user 102a (e.g., Barb) who issued the first query 106 that the round-robin mode is enabled to use the round-robin queue 350 to control subsequent queries (e.g., queue subsequent queries). For example, the digital assistant 105 can generate synthesized speech 123 to be audibly output from the speaker 18 of the AED 104, the synthesized speech stating "Barb, roundrobin mode is now enabled to control queries". In an additional example, the digital assistant 105 provides a notification (e.g., update 352) to the user device 50a associated with the user 102a (e.g., Barb) to inform the user 102a of the entries in the round-robin queue 350 and / or any action constraints 232 stored in the action constraint data repository 330.

[0044] Reference Figure 2A – Figure 2C , a graphical user interface (GUI) 200 executed on the user device 50 can display the round-robin queue 350, which includes the detected users 102 and the queries associated with each detected user 102. As used herein, the GUI 200 can receive user input indications via any of touch, voice, gesture, gaze, and / or an input device (e.g., a mouse or stylus) for interacting with the round-robin queue 350 (via the digital assistant 105). Each listed action can be used as a descriptor that identifies the corresponding command 111 for causing the digital assistant 105 to perform the action. Figure 2AAn example GUI 200b1 is provided that is displayed on the screen of user device 50b to inform user 102b of the content in circular queue 350. Specifically, GUI 200b1 renders circular queue 350 that includes users 102a (Barb), 102b (Jeff), and 102c (Tina) in a repeating order. Circular queue 350 additionally includes placeholders for queries for each of users 102a–c. As shown, the first entry for user 102a (Barb) includes corresponding action 111 to “Play Glory,” while subsequent entries for user 102a and other users 102b, 102c (Jeff, Tina) are empty. In other words, circular queue 350 includes, for each user 102 among multiple detected users 102a–c, the identity of user 102 (e.g., Barb, Jeff, Tina) and a query count that includes queries received from the corresponding user 102. As shown, user 102a (Barb) has a single query count corresponding to the action “Play Glory,” while users 102b (Jeff) and 102c (Tina) have no queries in circular queue 350.

[0045] Additionally, GUI 200 may be rendered to display an identifier of the current action (e.g., “Play Glory”), an identifier of the AED 104 that is currently performing the action (e.g., smart speaker), an indicator that the loop mode is enabled, and / or the identity of user 102a that issued the first query 106 (e.g., Barb). In an implementation where first query 106 includes constraint 113, GUI 200 renders an identifier of constraint 113 (e.g., popular music) for display. Thus, user 102b may consult user device 50b to view the current action of digital assistant 105 and any constraints 113 that will limit queries issued by user 102.

[0046] Reference Figure 4A and Figure 4B, in some implementations, the AED 104 (or the remote system 130 communicating with the AED 104) further includes an example data repository 430 that stores the registered user data / information of each of the multiple registered users 432a–n of the AED 104. Here, each registered user 432 of the AED 104 can perform a voice registration process to obtain a corresponding registered speaker vector 154 from audio samples of multiple registration phrases spoken by the registered user 432. For example, the speaker discrimination model 410 can generate one or more registered speaker vectors 154 from the audio samples of the registration phrases spoken by each registered user 432, and these audio samples can be combined, such as averaged or otherwise accumulated, to form the corresponding registered speaker vector 154. One or more of the registered users 432 can use the AED 104 to perform the voice registration process, where the microphone 16 captures the audio samples of these users speaking the registration utterances, and the speaker discrimination model 410 generates the corresponding registered speaker vectors 154 from these audio samples. The model 410 can be executed on the AED 104, the remote system 130, or a combination thereof. Additionally, one or more of the registered users 432 can register with the AED 104 by providing authorization and authentication credentials to an existing user account of the AED 104. Here, the existing user account can store the registered speaker vectors 154 obtained from a previous voice registration process with another device that is also linked to the user account.

[0047] In some examples, the registered speaker vector 154 of the registered user 432 includes a text-dependent registered speaker vector. For example, a text-dependent registered speaker vector can be extracted from one or more audio samples of the corresponding registered user 432 who speaks a predetermined term, such as a hot word 110 (e.g., "Ok computer") for invoking the AED 104 to wake up from a sleep state. In other examples, the registered speaker vector 154 of the registered user 432 is a text-independent registered speaker vector obtained from one or more audio samples of the corresponding registered user 102 who speaks phrases with different terms / words and different lengths. In these examples, the text-independent registered speaker vectors can be obtained from the audio samples over time, and these audio samples are obtained from the voice interactions between the user 102 and the AED 104 or other devices linked to the same account.

[0048] Reference Figure 4A, the speaker recognition process 400a identifies the user 102a (e.g., Barb) who uttered the first query 106 by first extracting a first speaker discrimination vector 411 that represents the characteristics of the first query 106 uttered by the user 102a from the initial audio data 402 corresponding to the first query 106. Here, the speaker recognition process 400a can execute a speaker discrimination model 410 that is configured to receive the audio data 402 corresponding to the second query 146 as input and generate the first speaker discrimination vector 411 as output. The speaker discrimination model 410 can be a neural network model that is trained under machine or human supervision to output the speaker discrimination vector 411. The speaker discrimination vector 411 output by the speaker discrimination model 410 can include an N-dimensional vector that has values corresponding to the speech characteristics of the first query 106 associated with the user 102a. In some examples, the speaker discrimination vector 411 is a d-vector. In some examples, the first speaker discrimination vector 411 includes a set of speaker discrimination vectors, each of which is associated with a different user who is also authorized to control the AED 104. For example, in addition to the user 102a who uttered the first query 106, other authorized users can include other individuals who were present when the user 102a uttered the first query 106 that issued a command 111 for performing the first action and / or individuals added / designated by the user 102a as being authorized.

[0049] Once the first speaker discrimination vector 411 is output from the model 410, the speaker recognition process 400a determines whether the extracted speaker discrimination vector 411 matches any of the registered speaker vectors 154 stored on the AED 104 (e.g., stored in the memory hardware 12) of the registered users 432a–n of the AED 104. As described above, the speaker discrimination model 410 can generate the registered speaker vectors 154 of the registered users 200 during the voice registration process. Each registered speaker vector 154 can be used as a reference vector 155 corresponding to a voiceprint or unique identifier that represents the characteristics of the voice of the corresponding registered user 432.

[0050] In some implementations, the speaker recognition process 400a uses a comparator 420 that compares a first speaker discrimination vector 411 with corresponding registered speaker vectors 154 associated with each registered user 432a–n of the AED 104. Here, the comparator 420 can generate a score for each comparison that indicates the likelihood that the initial audio data 402 corresponding to the first query 106 corresponds to the identity of the corresponding registered user 432, and when the score meets a threshold, that identity is accepted. When the score does not meet the threshold, the comparator 420 can reject the identity of the speaker who issued the first query 106. In some implementations, the comparator 420 calculates a corresponding cosine distance between the first speaker discrimination vector 411 and each registered speaker vector 154, and determines that the first speaker discrimination vector 411 matches one of the registered speaker vectors in the registered speaker vectors 154 when the corresponding cosine distance meets a cosine distance threshold.

[0051] In some examples, the first speaker discrimination vector 411 is a text-related speaker discrimination vector extracted from a portion of one or more words corresponding to the first query 106, and each registered speaker vector 154 is also a text-related registered speaker vector on the same one or more words. The use of text-related speaker vectors can improve the accuracy in determining whether the first speaker discrimination vector 411 matches any of the registered speaker vectors in the registered speaker vectors 154. In other examples, the first speaker discrimination vector 411 is a text-independent speaker discrimination vector extracted from the entire initial audio data 402 corresponding to the first query 106.

[0052] When the speaker recognition process 400a determines that the first speaker discrimination vector 411 matches one of the registered speaker vectors in the registered speaker vectors 154, the process 400a identifies the user 102a who issued the first query 106 as the corresponding registered user 432a associated with the registered speaker vector in the registered speaker vectors 154 that matches the extracted speaker discrimination vector 411. In the example shown, the comparator 420 determines the match based on the corresponding cosine distance between the first speaker discrimination vector 411 and the registered speaker vector 154 associated with the registered user 432a meeting the cosine distance threshold. In some scenarios, the comparator 420 identifies the user 102a as the corresponding registered user 432a associated with the registered speaker vector 154 having the shortest corresponding cosine distance from the first speaker discrimination vector 411, provided that the shortest corresponding cosine distance also meets the cosine distance threshold.

[0053] Conversely, when the speaker recognition process 400a determines that the first speaker discrimination vector 411 does not match any of the registered speaker vectors in the registered speaker vector 154, the process 400a may identify the user 102a who uttered the utterance 106 as a guest user of the AED 104. Accordingly, the query processor 300 may add the guest user to the circular queue 350 and use the first speaker discrimination vector 411 as the reference speaker vector 155 representing the voice characteristics of the voice of the guest user. In some instances, the guest user may register with the AED 104, and the AED 104 may store the first speaker discrimination vector 411 as the corresponding registered speaker vector 154 of the new registered user.

[0054] Return to reference Figure 1B , when the digital assistant 105 plays music 122 as playback audio from the speaker 18 of the AED 104 and when the loop mode is enabled, the AED 104 receives a second query 146 that includes a command 118 for causing the digital assistant 105 to perform a second action. In the illustrated example, the user 102a who issued the first query 106 also issues a second query 146 "Play Cold Water next", which includes a command 118 for causing the digital assistant 105 to play a song (i.e., Cold Water) immediately after playing track #1 (i.e., Glory). Based on receiving the second query 146, the query processor 300 parses the identity of the speaker of the second query 146 by performing a speaker recognition process 400b ( Figure 4B ) on the audio data 402 corresponding to the second query 146 and determines that the second query 146 was issued by the user 102a who issued the first query 106. As referred to above Figure 4AIn the implementation described, in which the first query 106 issued by user 102a includes initial audio data 402 (e.g., the first query 106 is spoken by the first user 102a), the query processor 300 can first perform a speaker recognition process 400a on the initial audio data 402 corresponding to the first query 106 to identify the user 102a who issued the first query 106. The speaker recognition process 400b can be executed on the data processing hardware 12 of the AED 104. The speaker recognition process 400b can also be executed on the remote system 130. If the speaker verification process 400b for the audio data 402 corresponding to the second query 146 indicates that the second query 146 is spoken by a user 102 different from the user 102 who issued the first query 106, the AED 104 can continue to add the action associated with the second query 146 to the circular queue 350. Conversely, if the speaker verification process 400b for the audio data 402 corresponding to the second query 146 indicates that the second query 146 is spoken by the same user 102 who issued the first query 106, the AED 104 can block the execution of the second action (or at least require input from one or more other users 102 in the environment (e.g., in Figure 1B the

[0055] Referring again to Figure 4B and Figure 1B for an example, in response to receiving the second query 146, the AED 104 resolves the identity of the user 102 who spoke the second query 146 by performing the speaker recognition process 400b. The speaker recognition process 400b identifies the user 102a who spoke the first query 146 by first extracting a second speaker discrimination vector 412 representing the characteristics of the second query 146 from the audio data 402 corresponding to the first query 146 spoken by the user 102a. Here, the speaker verification process 400b can execute a speaker discrimination model 410 that is configured to receive the audio data 402 as input and generate the second speaker discrimination vector 412 as output. As discussed above in Figure 4A the speaker discrimination model 410 can be a neural network model trained under machine or human supervision to output the speaker discrimination vector 412. The second speaker discrimination vector 412 output by the speaker discrimination model 410 can include an N-dimensional vector having values corresponding to the speech characteristics of the utterance 146 associated with the user 102a. In some examples, the speaker discrimination vector 412 is a d-vector.

[0056] Once the second speaker discrimination vector 412 is output from the speaker discrimination model 410, the speaker verification process 400b determines whether the extracted speaker discrimination vector 412 matches the reference speaker vector 155 associated with the first registered user 432a stored on the AED 104 (e.g., stored in the memory hardware 12). The reference speaker vector 155 associated with the first registered user 432a may include the corresponding registered speaker vectors 154 associated with the first registered user 432a. As discussed above, the speaker discrimination model 410 may generate the registered speaker vectors 154 of the registered users 432 during the voice registration process. Each registered speaker vector 154 may be used as a reference vector corresponding to a voiceprint or unique identifier representing the characteristics of the voice of the corresponding registered user 432.

[0057] In some implementations, the speaker verification process 400b uses a comparator 420 that compares the second speaker discrimination vector 412 with the reference speaker vector 155 associated with the first registered user 432a among the registered users 432. Here, the comparator 420 may generate a comparison score that indicates the likelihood that the second query 146 corresponds to the identity of the first registered user 432a, and when the score meets a threshold, the identity is accepted. When the score does not meet the threshold, the comparator 420 may reject the identity. In some implementations, the comparator 420 calculates the corresponding cosine distance between the second speaker discrimination vector 412 and the reference speaker vector 155 associated with the first registered user 432a, and determines that the second speaker discrimination vector matches the reference speaker vector 155 when the corresponding cosine distance meets the cosine distance threshold.

[0058] When the speaker verification process 400b determines that the second speaker discrimination vector 412 matches the reference speaker vector 155 associated with the first registered user 432a, the process 400b identifies the user 102a who uttered the second query 146 as the first registered user 432a associated with the reference speaker vector 155. In the example shown, the comparator 420 determines the match based on the corresponding cosine distance between the second speaker discrimination vector 412 and the reference speaker vector 155 associated with the first registered user 432a meeting the cosine distance threshold. In some scenarios, the comparator 420 identifies the user 102a as the corresponding first registered user 432a associated with the reference speaker vector 155 having the shortest corresponding cosine distance from the second speaker discrimination vector 412, provided that the shortest corresponding cosine distance also meets the cosine distance threshold.

[0059] Refer to the above Figure 4A, in some implementations, the speaker recognition process 400a determines that the first speaker discrimination vector 411 does not match any of the registered speaker vectors in the registered speaker vector 154, and identifies the user 102a who uttered the first query 106 as a visitor user of the AED 104. Thus, the speaker verification process 400b can first determine whether the user 102a who uttered the first query 106 is identified as a registered user 432 or a visitor user by the speaker recognition process 400a. When the user 102a is a visitor user, the comparator 420 compares the second speaker discrimination vector 412 with the first speaker discrimination vector 411 obtained during the speaker recognition process 400a. Here, the first speaker discrimination vector 411 represents the characteristics of the first query 106 uttered by the visitor user 102a, and is therefore used as a reference vector to verify whether the second query 146 is also uttered by the visitor user 102a or another user 102. Here, the comparator 420 can generate a comparison score, which indicates the likelihood that the second query 146 corresponds to the identity of the visitor user 102a, and when the score meets the threshold, the identity is accepted. When the score does not meet the threshold, the comparator 420 can reject the identity of the visitor user who uttered the second query 146. In some implementations, the comparator 420 calculates the corresponding cosine distance between the first speaker discrimination vector 411 and the second speaker discrimination vector 412, and determines that the first speaker discrimination vector 411 matches the second speaker discrimination vector 412 when the corresponding cosine distance meets the cosine distance threshold.

[0060] Return reference Figure 1B , based on determining that the second query 146 is uttered by the user 102a who issued the first query 106, the query processor 300 (via the digital assistant 105) prevents the AED 104 from performing the second action, and instead prompts at least another user 102 different from the user 102a among the multiple users 102a–c detected in the environment to have the opportunity to issue a query. In other words, after determining that the user 102a is the issuer of the first query 106 and the speaker of the second query 146, the query processor 300 first verifies that other users 102 in the environment do not wish to issue a query before fulfilling the second query 146, thereby preventing the user 102a from monopolizing the AED 104. If other users 102 in the environment confirm that they do not wish to issue a query, then when the fulfillment of the first query 106 is completed, the digital assistant 105 can continue to fulfill the second query 146 issued by the user 102a.

[0061] In some implementations, before performing a second action included in a second query 146 issued by user 102a, the query handler 300 (via the digital assistant 105) prompts at least one other user 102 among a plurality of users 102 detected in the environment to provide a third query 148 to perform a third action including a corresponding command 119. In these implementations, prompting at least one other user 102 among the plurality of users 102 includes providing user-selectable options as output from the AED 104, which, when selected, issue the third query 148 to perform the third action including the corresponding command 119. For example, the digital assistant 105 may generate synthesized speech 123 to be audibly output from the speaker 18 of the AED 104 (or a speaker communicating with the data processing hardware (e.g., the speaker of the user device 50)), which prompts user 102b with the opportunity to issue the query "Jeff, would you like to select a song to play next?" As a response, user 102b (i.e., Jeff) among the plurality of users 102b, 102c different from user 102a is shown issuing the third query 148 "Yes, play Red next" near the AED 104. In an additional example, the digital assistant 105 may provide a notification to the user device 50b associated with user 102b (e.g., Jeff) to display the user-selectable options as a graphical element 210 on the screen of the user device 50b, which prompts user 102b with the opportunity to issue the third query 148. As a supplement or alternative to audibly prompting user 102b, as Figure 2B shown, the GUI 200b2 renders the graphical element 210 "Jeff, would you like to select a song to play next?" and "Yes" and "No" for display, which allows user 102b to issue the third query 148 (or choose to forgo the opportunity to issue the third query). Here, when the user device 50b receives a user input indication indicating a selection of one of the graphical elements 210 of the user-selectable options, the query handler 300 receives the third query 148.

[0062] Reference Figure 1B and Figure 3, in response to receiving the third query 148, the NLU module 320 executed on the AED 104 (and / or executed on the remote system 130) may recognize the phrase "play Red" as a command 119 for performing a third action (i.e., playing music 122). The query handler 300 may first determine whether the command 119 for performing the third action violates any action constraints 232 on the digital assistant 105 before updating the circular queue 350 to include the third action. As discussed above, the action constraints 332 may include the constraint 113 included in the first query 106, which restricts the genre of music (e.g., pop music). Here, the query handler 300 verifies that the action of playing "Red" in the third query 148 issued by the user 102b does not conflict with the constraint 113 that restricts the music genre to pop music, and then adds "Red" to the circular queue 350. In some implementations, the query handler 300 additionally verifies that the action of playing "Cold Water" included in the second query 146 does not conflict with the constraint 113 that restricts the music genre to pop music and / or any other action constraints 332 on the digital assistant 105, and then updates the circular queue 350 to include the command 118 in the second query 146 issued by the user 102a.

[0063] Although the examples mainly refer to actions of playing music to prevent a single user 102 from monopolizing the playlist, these actions can refer to any category of actions, including but not limited to search queries, control of assistant-supported devices (e.g., smart lights, smart thermostats), and playing other types of media (e.g., podcasts, videos, etc.). For example, the query handler 300 can help users 102 in the environment create a shopping list, thereby ensuring that all users 102 are given the opportunity to add items to the shopping list. For example, the AED 104 can prompt the second user 102 to add items in response to a query from the first user 102 to add multiple items to the shopping list, while still adding the multiple items requested by the first user 102 to the shopping list. Similarly, the query handler 300 can proactively prompt users 102 to include actions such as setting a morning alarm. Here, the AED 104 can suggest that the second user 102 request an alarm in response to receiving a request to set an alarm from the first user 102. In addition, the query handler 300 enables the AED 104 to ensure that users 102 are not excluded from the interaction by attracting / prompting users 102 who have not issued queries recently to join the interaction between the AED 104 and another user 102.

[0064] Additionally, the constraint 332 can include many restrictions on the action itself. For example, a time limit for the action can be appropriate, which limits how long the user 102 has a turn in each entry in the circular queue 350 (e.g., how many jokes the user 102 can request for each turn in the circular queue). The time limit for the circular pattern can control how long the event of applying the constraint 332 lasts. Similarly, the constraint 332 on the threshold number of actions per query can limit the number of actions that the user 102 can request for each issued query entry in the circular queue 350 (e.g., allowing the user 102 to request an entire album and / or playlist for each turn in the circular queue 350). Also, the threshold number of actions can include the total number of actions per user in the circular queue 350. For example, the user 102 can only submit 50 additional songs for the circular queue 350.

[0065] Continue Figure 1B and Figure 2C In the example of, after the query processor 300 verifies that the action of playing "Red" in the third query 148 issued by the user 102b does not conflict with the constraint 113 that restricts the music genre to pop music, the query processor 300 updates the circular queue 350 to add "Red" to the identity associated with the user 102b in the circular queue 350. As Figure 2C shown, the GUI 200b3 renders the circular queue 350 including the users 102a (Barb), 102b (Jeff), and 102c (Tina) in a repeating order. The action placeholder for Jeff has been updated to include the command 119, which is used to cause the digital assistant 105 to perform the action "play Red" once the AED 105 finishes executing the action (e.g., playing) "play Glory" in the command 111 in the first query 106. In other words, based on the circular queue 350, when the digital assistant 105 finishes executing the first action 111 in the first query 106 issued by the user 102a, it will execute the execution of the third action 119 in the third query 148 issued by the user 102b. Additionally, the query processor 300 updates the circular queue 350 to include the second query 146 issued by the user 102a. As shown, the action placeholder for the second item for Barb has been updated to include the command 118, which is used to cause the digital assistant 105 to perform the action "play Cold Water" once the AED 105 finishes executing the action (e.g., playing) "play Red" in the command 119 in the third query 148 and any additional intermediate queries submitted by users in the environment 102 different from the users 102a and 102b before the AED 105 finishes executing the actions associated with the third query 148.

[0066] Refer toFigure 1C AED 104 can notify a user 102a (e.g., Barb) who issues a first query 106 that, based on the circular queue 350, a second query 146 will be fulfilled after the fulfillment of the third query 148. For example, the digital assistant 105 can generate synthesized speech 123 to be audibly output from the speaker 18 of the AED 104, the synthesized speech stating "Barb, ColdWater will play after Jeff’s selection". In an additional example, the digital assistant 105 provides a notification (e.g., update 352) to a user device 50a associated with the user 102a (e.g., Barb) to inform the user 102a of the updated entry in the circular queue 350.

[0067] Figure 5 A flowchart of an example operational arrangement of a method 500 for handling queries in an environment having multiple users 102. At operation 502, the method 500 includes detecting multiple users 102 in an environment of an assistant-enabled device (AED) 104. The method 500 also includes receiving, at operation 504, a first query 106 issued by a first user 102a among the multiple users 102. The first query 106 includes a command 111 for causing the digital assistant 105 to perform a first action. The method 500 further includes enabling, at operation 506, a loop mode that, when enabled, causes the digital assistant 105 to control the execution of actions of query commands after the first query 106 based on the circular queue 350. Here, the circular queue 350 includes the multiple users 102 detected within the environment of the AED 104.

[0068] When the digital assistant 105 is performing the first action and when the loop mode is enabled, the method 500 further includes receiving, at operation 508, audio data 402 corresponding to a second query 146 spoken by one of the multiple users 102 and captured by the AED 104. Here, the second query 146 includes a command 118 for causing the digital assistant 105 to perform a second action. At operation 510, the method 500 also includes performing speaker recognition on the audio data 402 corresponding to the second query 146 to determine that the second query 146 was spoken by the first user 102a who issued the first query 106.

[0069] Based on determining that the second query 146 was spoken by the first user 102a who issued the first query 106, method 500 further includes, at operation 512, preventing the digital assistant 105 from performing the second action, and prompting at least one other user 102 different from the first user 102a among the plurality of users 102 detected in the environment with an opportunity to issue a query. At operation 514, method 500 includes receiving a third query 148 issued by a second user 102b among the plurality of users 102 detected in the environment, the third query 148 including a command 119 for causing the digital assistant 105 to perform a third action. When the digital assistant 105 finishes performing the first action, method 500 further includes, at operation 502, performing the execution of the third action.

[0070] Figure 6 FIG. 4 is a schematic diagram of an example computing device 600 that can be used to implement the systems and methods described in this document. Computing device 600 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown herein, their connections and relationships, and their functions are intended to be exemplary only and are not intended to limit the implementation of the invention described and / or claimed in this document.

[0071] The computing device 600 includes a processor 610, a memory 620, a storage device 630, a high-speed interface / controller 640 connected to the memory 620 and a high-speed expansion port 650, and a low-speed interface / controller 660 connected to a low-speed bus 670 and the storage device 630. Each of the components 610, 620, 630, 640, 650, and 660 is interconnected using various buses and can be mounted on a common motherboard or otherwise as appropriate. The processor 610 (e.g., the data processing hardware 10, 132 of FIG. 1) can process instructions for execution within the computing device 600, including instructions stored in the memory 620 or on the storage device 630, to display graphical information of a graphical user interface (GUI) on an external input / output device such as a display 680 coupled to the high-speed interface 640. In other implementations, multiple processors and / or multiple buses and multiple memories and multiple types of memory can be used as appropriate. Additionally, multiple computing devices 600 can be connected, where each device provides a portion of the necessary operations (e.g., as a server group, a blade server group, or a multi-processor system).

[0072] Memory 620 stores information non - transiently within computing device 600. Memory 620 (e.g., memory hardware 12, 134 of FIG. 1) can be a computer - readable medium, a volatile memory unit, or a non - volatile memory unit. The non - transient memory 620 can be a physical device for temporarily or permanently storing programs (e.g., sequences of instructions) or data (e.g., program state information) for use by computing device 600. Examples of non - volatile memory include, but are not limited to, flash memory and read - only memory (ROM) / programmable read - only memory (PROM) / erasable programmable read - only memory (EPROM) / electrically erasable programmable read - only memory (EEPROM) (e.g., commonly used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase - change memory (PCM), and magnetic disks or tapes.

[0073] Storage device 630 is capable of providing mass storage for computing device 600. In some implementations, storage device 630 is a computer - readable medium. In various different implementations, storage device 630 can be a floppy disk device, a hard disk device, an optical disk device, or a magnetic tape device, a flash memory, or other similar solid - state memory devices, or an array of devices (including devices in a storage area network or other configurations). In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer - or machine - readable medium, such as memory 620, storage device 630, or memory on processor 610.

[0074] High - speed controller 640 manages bandwidth - intensive operations of computing device 600, while low - speed controller 660 manages lower - bandwidth - intensive operations. Such a division of responsibilities is merely exemplary. In some implementations, high - speed controller 640 is coupled to memory 620, display 680 (e.g., via a graphics processor or accelerator), and a high - speed expansion port 650 that can accept various expansion cards (not shown). In some implementations, low - speed controller 660 is coupled to storage device 630 and low - speed expansion port 690. The low - speed expansion port 690, which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a networking device, such as a switch or a router, for example, via a network adapter.

[0075] Computing device 600 can be implemented in many different forms, as shown in the figure. For example, it can be implemented as a standard server 600a or multiple times as a group of such servers 600a, implemented as a laptop computer 600b, or implemented as part of a rack server system 600c.

[0076] Various implementations of the systems and techniques described herein can be implemented in digital electronic and / or optical circuitry, integrated circuit systems, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor which can be special purpose or general purpose and coupled to receive data and instructions from, and to send data and instructions to, a storage system, at least one input device, and at least one output device.

[0077] A software application (i.e., a software resource) can refer to computer software that causes a computing device to perform a task. In some examples, a software application can be referred to as an “application,” “app,” or “program.” Example applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.

[0078] A non-transitory memory can be a physical device for temporarily or permanently storing programs (e.g., sequences of instructions) or data (e.g., program state information) for use by a computing device. A non-transitory memory can be volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read only memory (ROM) / programmable read only memory (PROM) / erasable programmable read only memory (EPROM) / electrically erasable programmable read only memory (EEPROM) (e.g., typically used for firmware such as a boot program). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and magnetic disk or tape.

[0079] These computer programs (also referred to as programs, software, software applications or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages and / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, non-transitory computer-readable medium, device, and / or apparatus (e.g., a magnetic disk, optical disk, memory, programmable logic device (PLD)) that provides machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal that provides machine instructions and / or data to a programmable processor.

[0080] The processes and logical flows described in this specification can be performed by one or more programmable processors (also referred to as data processing hardware) that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by, or incorporated in, special purpose logic circuitry, such as an FPGA (field programmable gate array) or ASIC (application specific integrated circuit). By way of example, processors suitable for the execution of a computer program include both general and special purpose microprocessors, and any one or more processors of any type of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or operatively coupled to receive data from one or more mass storage devices or transfer data to one or more mass storage devices or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including by way of example semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0081] To provide for interaction with a user, one or more aspects of the present disclosure may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube), an LCD (liquid crystal display) monitor, or a touch screen) for displaying information to the user and, optionally, a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user may provide input to the computer. Other kinds of devices may also be used to provide for interaction with the user; for example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input received from the user may be in any form, including sound, speech, or tactile input. Additionally, the computer may interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a web page to a web browser on a client device of the user in response to a request received from the web browser.

[0082] A variety of implementations have been described. However, it should be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims.

Claims

1. A computer-implemented method (500) that, when executed by data processing hardware (610), causes the data processing hardware (610) to perform operations that include: detecting a plurality of users within an environment of an assistant-enabled device (104); receiving a first query (106) issued by a first user of the plurality of users, the first query (106) including a command (111) for causing a digital assistant (105) to perform a first action; enabling a loop mode that, when enabled, causes the digital assistant (105) to control the execution of actions of query commands subsequent to the first query (106) based on a loop queue (350) that includes the plurality of users detected within the environment of the assistant-enabled device (104); when the digital assistant (105) is performing the first action and when the loop mode is enabled: receiving audio data (402) corresponding to a second query (146) spoken by one of the plurality of users and captured by the assistant-enabled device (104), the second query (146) including a command (111) for causing the digital assistant (105) to perform a second action; performing speaker recognition on the audio data (402) corresponding to the second query (146) to determine that the second query (146) was spoken by the first user who issued the first query (106); based on determining that the second query (146) was spoken by the first user who issued the first query (106): preventing the digital assistant (105) from performing the second action; and prompting at least another user different from the first user among the plurality of users detected in the environment with an opportunity to issue a query; and receiving a third query (148) issued by a second user of the plurality of users detected in the environment, the third query (148) including a command (111) for causing the digital assistant (105) to perform a third action; and when the digital assistant (105) finishes performing the first action, performing the execution of the third action.

2. The method (500) of claim 1, wherein: prompting at least the another user among the plurality of users detected in the environment includes providing user-selectable options as an output of a user interface (200) of a user device (50) associated with the second user, the user-selectable options, when selected, issuing the third query (148) from the second user to perform the third action; and wherein receiving the third query (148) issued by the second user is based on receiving a user input indication indicating a selection of the user-selectable options.

3. The method (500) according to claim 2, wherein providing the user-selectable option as an output from the user interface (200) includes displaying, via the user interface (200), the user-selectable option as a graphical element (210) on a screen of the user device (50) associated with the second user, the graphical element (210) prompting the second user with an opportunity to issue the third query (148).

4. The method (500) according to claim 2, wherein providing the user-selectable option as an output from the user interface (200) includes providing, via the user interface (200), the user-selectable option as an audible output from a speaker (18) communicating with the data processing hardware (610), the audible output prompting the second user with an opportunity to issue the third query (148).

5. The method (500) according to any one of claims 1–4, wherein the first query (106) further includes a constraint on subsequent queries, the constraint including one of the following: Action category; Time limit for an action; Time limit for the loop pattern; or Threshold number of actions per query.

6. The method (500) according to claim 5, wherein the operation further includes: Determining that the second query (146) does not violate the constraint in the first query (106); and Updating the circular queue (350) to include the second query (146) spoken by the first user.

7. The method (500) according to claim 6, wherein the operation further includes: Detecting that the first user has left the environment of the device (104) of the support assistant; and Updating the circular queue (350) to remove the second query (146) spoken by the first user.

8. The method (500) according to any one of claims 1–7, wherein detecting the plurality of users within the environment of the device (104) of the support assistant includes detecting at least one of the plurality of users based on proximity information (54) of a user device (50) associated with at least one of the plurality of users.

9. The method (500) according to any one of claims 1–8, wherein detecting the plurality of users within the environment of the device (104) of the support assistant includes: Receiving image data (312) corresponding to a scene of the environment; and Detecting at least one of the plurality of users based on the image data (312).

10. The method (500) according to any one of claims 1–9, wherein detecting the plurality of users within the environment of the device (104) of the support assistant includes receiving a list indicating each user of the plurality of users to be added to the circular queue (350).

11. The method (500) according to any one of claims 1–10, wherein the circular queue (350) for each corresponding user of the plurality of users detected in the environment Comprising: The identity of the corresponding user; And The query count of the query received from the corresponding user.

12. The method (500) according to any one of claims 1–11, wherein receiving the first query (106) issued by the first user comprises receiving a user input indication from a user device (50) associated with the first user indicating a user intent to issue the first query (106).

13. The method (500) according to any one of claims 1–12, wherein receiving the first query (106) issued by the first user comprises receiving initial audio data (402) corresponding to the first query (106) issued by the first user and captured by the device (104) of the support assistant.

14. The method (500) according to claim 13, wherein the operation further comprises, after receiving the initial audio data (402) corresponding to the first query (106) issued by the first user, performing speaker recognition on the initial audio data (402) to identify the first user who issued the first query (106) by: Extracting a first speaker discrimination vector (411) representing the characteristics of the first query (106) issued by the first user from the initial audio data (402) corresponding to the first query (106) issued by the first user; Determining that the extracted speaker discrimination vector matches any of the registered speaker vectors (154) stored on the device (104) of the support assistant, each registered speaker vector (154) being associated with a different corresponding registered user of the device (104) of the support assistant; and Based on determining that the first speaker discrimination vector (411) matches one of the registered speaker vectors (154), identifying the first user who issued the first query (106) as the corresponding registered user associated with the registered speaker vector among the registered speaker vectors (154) that matches the extracted speaker discrimination vector.

15. The method (500) according to any one of claims 1–14, wherein performing speaker recognition on the audio data (402) corresponding to the second query (146) to determine that the second query (146) is spoken by the first user who issued the first query (106) Comprises: Extracting a second speaker discrimination vector (412) representing the characteristics of the second query (146) from the audio data (402) corresponding to the second query (146); and Determining that the second extracted speaker discrimination vector matches the reference speaker vector (155) of the first user.

16. A system (100), Comprising: Data processing hardware (610); And Memory hardware (620) that communicates with the data processing hardware (610), the memory hardware (620) storing instructions that, when executed on the data processing hardware (610), cause the data processing hardware (610) to perform operations, the operations including: Detecting a plurality of users within the environment of an assistant-enabled device (104); Receiving a first query (106) issued by a first user of the plurality of users, the first query (106) including a command (111) for causing a digital assistant (105) to perform a first action; Enabling a loop mode that, when enabled, causes the digital assistant (105) to control the execution of actions of query commands subsequent to the first query (106) based on a loop queue (350), the loop queue (350) including the plurality of users detected within the environment of the assistant-enabled device (104); When the digital assistant (105) is performing the first action and when the loop mode is enabled: Receiving audio data (402) corresponding to a second query (146) spoken by one of the plurality of users and captured by the assistant-enabled device (104), the second query (146) including a command (111) for causing the digital assistant (105) to perform a second action; Performing speaker recognition on the audio data (402) corresponding to the second query (146) to determine that the second query (146) was spoken by the first user who issued the first query (106); Based on determining that the second query (146) was spoken by the first user who issued the first query (106): Preventing the digital assistant (105) from performing the second action; and Prompting at least another user different from the first user among the plurality of users detected in the environment to have an opportunity to issue a query; and Receiving a third query (148) issued by a second user among the plurality of users detected in the environment, the third query (148) including a command (111) for causing the digital assistant (105) to perform a third action; and When the digital assistant (105) finishes performing the first action, performing the execution of the third action.

17. The system (100) according to claim 16, wherein: Prompting at least the other user among the plurality of users detected in the environment includes providing user-selectable options as an output of a user interface (200) of a user device (50) associated with the second user, the user-selectable options, when selected, issuing the third query (148) from the second user to perform the third action; and wherein receiving the third query (148) issued by the second user is based on receiving a user input indication indicating a selection of the user-selectable options.

18. The system (100) according to claim 17, wherein providing the user-selectable option as an output from the user interface (200) includes displaying the user-selectable option as a graphical element (210) via the user interface (200) on a screen of the user device (50) associated with the second user, the graphical element (210) prompting the second user with an opportunity to issue the third query (148).

19. The system (100) according to claim 17, wherein providing the user-selectable option as an output from the user interface (200) includes providing the user-selectable option as an audible output from a speaker (18) communicating with the data processing hardware (610) via the user interface (200), the audible output prompting the second user with an opportunity to issue the third query (148).

20. The system (100) according to any one of claims 16–19, wherein the first query (106) further includes a constraint on subsequent queries, the constraint including one of the following: Action category; Time limit for an action; Time limit for the loop pattern; or Threshold number of actions per query.

21. The system (100) according to claim 20, wherein the operation further includes: Determining that the second query (146) does not violate the constraint in the first query (106); and Updating the circular queue (350) to include the second query (146) spoken by the first user.

22. The system (100) according to claim 21, wherein the operation further includes: Detecting that the first user has left the environment of the support assistant device (104); and Updating the circular queue (350) to remove the second query (146) spoken by the first user.

23. The system (100) according to any one of claims 16–22, wherein detecting the plurality of users within the environment of the support assistant device (104) includes detecting at least one of the plurality of users based on proximity information (54) of a user device (50) associated with at least one of the plurality of users.

24. The system (100) according to any one of claims 16–23, wherein detecting the plurality of users within the environment of the support assistant device (104) includes: Receiving image data (312) corresponding to a scene of the environment; and Detecting at least one of the plurality of users based on the image data (312).

25. The system (100) according to any one of claims 16–24, wherein detecting the plurality of users within the environment of the support assistant device (104) includes receiving a list indicating each of the plurality of users to be added to the circular queue (350).

26. The system (100) according to any one of claims 16 - 25, wherein the circular queue (350) corresponds to each of the multiple users detected in the environment comprises: the identity of the corresponding user; and a query count of the queries received from the corresponding user.

27. The system (100) according to any one of claims 16 - 26, wherein receiving the first query (106) issued by the first user comprises receiving a user input indication from a user device (50) associated with the first user indicating a user intent to issue the first query (106).

28. The system (100) according to any one of claims 16 - 27, wherein receiving the first query (106) issued by the first user comprises receiving initial audio data (402) corresponding to the first query (106) issued by the first user and captured by the device (104) of the support assistant.

29. The system (100) according to claim 28, wherein the operation further comprises, after receiving the initial audio data (402) corresponding to the first query (106) issued by the first user, performing speaker recognition on the initial audio data (402) to identify the first user who issued the first query (106) by: extracting a first speaker discrimination vector (411) representing the characteristics of the first query (106) issued by the first user from the initial audio data (402) corresponding to the first query (106) issued by the first user; determining that the extracted speaker discrimination vector matches any of the registered speaker vectors (154) stored on the device (104) of the support assistant, each registered speaker vector (154) being associated with a different corresponding registered user of the device (104) of the support assistant; and based on determining that the first speaker discrimination vector (411) matches one of the registered speaker vectors (154), identifying the first user who issued the first query (106) as the corresponding registered user associated with the registered speaker vector among the registered speaker vectors (154) that matches the extracted speaker discrimination vector.

30. The system (100) according to any one of claims 16 - 29, wherein performing speaker recognition on the audio data (402) corresponding to the second query (146) to determine that the second query (146) is spoken by the first user who issued the first query (106) comprises: extracting a second speaker discrimination vector (412) representing the characteristics of the second query (146) from the audio data (402) corresponding to the second query (146); and determining that the second extracted speaker discrimination vector matches the reference speaker vector (155) of the first user.