Processing voice queries in an environment with multiple users

A round-robin scheduling strategy with speaker identification addresses the challenge of managing multiple user interactions in shared voice-activated devices, ensuring fair and orderly command execution.

JP7863260B2Active Publication Date: 2026-05-20GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
GOOGLE LLC
Filing Date
2023-09-12
Publication Date
2026-05-20

AI Technical Summary

Technical Problem

Existing voice-activated devices struggle to manage multiple user interactions effectively in shared environments, leading to monopolization of device resources by individual users.

Method used

Implementing a round-robin scheduling strategy and speaker identification to manage queries from multiple users, ensuring each user has an opportunity to interact while preventing monopolization.

Benefits of technology

Ensures fair and orderly execution of user commands, promoting inclusivity and preventing resource monopolization in multi-user environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007863260000001
    Figure 0007863260000001
  • Figure 0007863260000002
    Figure 0007863260000002
  • Figure 0007863260000003
    Figure 0007863260000003
Patent Text Reader

Abstract

The method (500) includes receiving a first query (116) issued by a first user, the first query including a command (111) for the digital assistant (105) to perform a first action, and further including enabling a round-robin mode (350) to control the performance of the action. The method also includes receiving audio data (402) corresponding to a second query (146) including a command to perform a second action while the first action is being performed, performing speaker identification on the audio data, determining that the second query was spoken by the first user, preventing the second action from being performed, and prompting at least another user to issue a query. The method further includes receiving a third query (148) issued by the second user, the third query including a command for the digital assistant to perform a third action, and further including performing the third action when the digital assistant completes performance of the first action.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to processing voice queries in an environment where there are multiple users.

Background Art

[0002] The way users interact with assistant-enabled devices is mainly designed by voice input, although not limited to this. For example, a user may ask a device to perform an action including media playback (e.g., music or podcast), and the device responds by starting to play audio that matches the user's criteria. In a situation where a device (e.g., a smart speaker) is widely shared by multiple users in an environment, the device may need to handle multiple actions requested by users who may compete with each other.

Summary of the Invention

[0003] One aspect of the present disclosure provides a method for execution on a computer that, when executed on data processing hardware, causes the data processing hardware to perform an operation that includes detecting multiple users in the environment of an assistant-enabled device (AED). The operation also includes receiving a first query issued by a first user of the multiple users. The first query includes a command for the digital assistant to perform a first action. The operation further includes enabling a round-robin mode, which, when enabled, causes the digital assistant to control the performance of actions commanded by queries following the first query, based on a round-robin queue, where the round-robin queue includes multiple users detected in the environment of the AED. While the digital assistant is performing the first action, and while the round-robin mode is enabled, the operation further includes receiving audio data corresponding to a second query spoken by one of the multiple users and captured by the AED, where the second query includes a command for the digital assistant to perform a second action. The operation also includes speaker identification on the audio data corresponding to the second query in order to determine that the second query was spoken by the first user who issued the first query. Based on the determination that the second query was spoken by the first user who issued the first query, the operation also includes preventing the digital assistant from performing the second action and prompting at least one other user among several users detected in the environment, who is different from the first user, to issue the query. The operation further includes receiving a third query issued by a second user among several users detected in the environment, the third query including a command for the digital assistant to perform the third action, and the operation further includes performing the third action once the digital assistant has completed performing the first action.

[0004] Embodiments of the present disclosure may include one or more of the following optional features. In some embodiments, prompting at least one other user among a plurality of users detected in the environment includes providing a user-selectable option as output from the user interface of a user device associated with a second user, which is selected to take a third action and issues a third query from the second user. Receiving the third query issued by the second user is based on receiving a user input instruction indicating a selection of the user-selectable option. In these embodiments, providing a user-selectable option as output from the user interface may include displaying the user-selectable option as a graphic element on the screen of a user device associated with a second user via the user interface. The graphic element prompts the second user to have the opportunity to issue a third query. Additionally or alternatively, providing a user-selectable option as output from the user interface may include providing a user-selectable option as an audible output from a speaker communicating with data processing hardware via the user interface. The audible output prompts the second user to have the opportunity to issue a third query.

[0005] In some examples, the first query further includes constraints for subsequent queries, the constraints including one of the following: the category of action, the time limit for the action, the time limit for round-robin mode, or the threshold number of actions per query. In these examples, the behavior further includes determining that the second query does not violate the constraints of the first query and updating the round-robin queue to include the second query spoken by the first user. In these examples, the behavior may further include detecting that the first user has left the environment of the assistant-enabled device and updating the round-robin queue to remove the second query spoken by the first user.

[0006] In some embodiments, detecting multiple users within an assistant-enabled device environment includes detecting at least one of the multiple users based on proximity information of a user device associated with at least one of the multiple users. Additionally or alternatively, detecting multiple users within an assistant-enabled device environment includes receiving image data corresponding to a scene in the environment and detecting at least one of the multiple users based on the image data. In other embodiments, detecting multiple users within an assistant-enabled device environment includes receiving a list indicating each of the multiple users to be added to a round-robin queue. In some examples, the round-robin queue includes, for each corresponding user among the multiple users detected in the environment, the identity of the corresponding user and a query count of queries received from the corresponding user. In some embodiments, receiving a first query issued by a first user includes receiving a user input instruction from a user device associated with the first user indicating the user's intent to issue the first query.

[0007] In some examples, receiving a first query issued by a first user involves receiving initial audio data corresponding to the first query issued by the first user and captured by the assistant-enabled device. In these examples, after receiving the initial audio data corresponding to the first query issued by the first user, the operation further includes performing speaker identification on the initial audio data to identify the first user who issued the first query. Speaker identification includes extracting a first speaker identification vector from the initial audio data corresponding to the first query issued by the first user that represents the features of the first query issued by the first user, and determining that the extracted speaker identification vector matches any registered speaker vector stored in the assistant-enabled device. Each registered speaker vector is associated with a different registered user of the assistant-enabled device. When the first speaker identification vector matches one of the registered speaker vectors, the operation also includes identifying the first user who issued the first query as the respective registered user associated with one of the registered speaker vectors that matches the extracted speaker identification vector. In some embodiments, speaker identification on audio data corresponding to a second query in order to determine that the second query was spoken by a first user who issued the first query includes extracting a second speaker identification vector representing the features of the second query from the audio data corresponding to the second query, and determining that the extracted second speaker identification vector matches the first user's reference speaker vector.

[0008] Other aspects of this disclosure provide a system including data processing hardware and memory hardware that communicates with the data processing hardware. The memory hardware, when executed by the data processing hardware, stores instructions that cause the data processing hardware to perform an operation which includes detecting multiple users in the environment of an automated external defibrillator (AED). The operation also includes receiving a first query issued by a first user of the multiple users. The first query includes a command for the digital assistant to perform a first action. The operation further includes enabling round-robin mode, which, when enabled, causes the digital assistant to control the performance of actions commanded by queries following the first query, based on a round-robin queue, where the round-robin queue includes multiple users detected in the environment of the AED. While the digital assistant is performing the first action and while round-robin mode is enabled, the operation further includes receiving audio data corresponding to a second query spoken by one of the multiple users and captured by the AED, where the second query includes a command for the digital assistant to perform a second action. The operation also includes speaker identification on the audio data corresponding to the second query in order to determine that the second query was spoken by the first user who issued the first query. Based on the determination that the second query was spoken by the first user who issued the first query, the operation also includes preventing the digital assistant from performing the second action and prompting at least one other user among several users detected in the environment, who is different from the first user, to have the opportunity to issue the query. The operation further includes receiving a third query issued by a second user among several users detected in the environment, the third query including a command for the digital assistant to perform the third action, and once the digital assistant has completed performing the first action, the operation further includes performing the third action.

[0009] This embodiment may include one or more of the following optional features. In some embodiments, prompting at least one other user among a plurality of users detected in the environment includes providing a user-selectable option as output from the user interface of a user device associated with a second user, which is selected to take a third action and issues a third query from the second user. Here, receiving the third query issued by the second user is based on receiving a user input instruction indicating a selection of the user-selectable option. In these embodiments, providing a user-selectable option as output from the user interface may include displaying the user-selectable option as a graphic element on the screen of a user device associated with a second user via the user interface, the graphic element prompting the second user to issue a third query. Additionally or alternatively, providing a user-selectable option as output from the user interface may include providing a user-selectable option as an audible output from a speaker communicating with data processing hardware via the user interface, the audible output prompting the second user to issue a third query.

[0010] In some examples, the first query further includes constraints for subsequent queries, the constraints including one of the following: the category of action, the time limit for the action, the time limit for round-robin mode, or the threshold number of actions per query. In these examples, the behavior further includes determining that the second query does not violate the constraints of the first query and updating the round-robin queue to include the second query spoken by the first user. In these examples, the behavior may further include detecting that the first user has left the environment of the assistant-enabled device and updating the round-robin queue to remove the second query spoken by the first user.

[0011] In some embodiments, detecting multiple users within an assistant-enabled device environment includes detecting at least one of the multiple users based on proximity information of a user device associated with at least one of the multiple users. Additionally or alternatively, detecting multiple users within an assistant-enabled device environment includes receiving image data corresponding to a scene in the environment and detecting at least one of the multiple users based on the image data. In other embodiments, detecting multiple users within an assistant-enabled device environment includes receiving a list indicating each of the multiple users to be added to a round-robin queue. In some examples, the round-robin queue includes, for each corresponding user among the multiple users detected in the environment, the identity of the corresponding user and a query count of queries received from the corresponding user. In some embodiments, receiving a first query issued by a first user includes receiving a user input instruction from a user device associated with the first user indicating the user's intent to issue the first query.

[0012] In some examples, receiving a first query issued by a first user involves receiving initial audio data corresponding to the first query issued by the first user and captured by the assistant-enabled device. In these examples, after receiving the initial audio data corresponding to the first query issued by the first user, the operation further includes performing speaker identification on the initial audio data to identify the first user who issued the first query. Speaker identification involves extracting a first speaker identification vector from the initial audio data corresponding to the first query issued by the first user, representing the features of the first query issued by the first user, and determining that the extracted speaker identification vector matches any registered speaker vector stored in the assistant-enabled device. Each registered speaker vector is associated with a different registered user of the assistant-enabled device. When the first speaker identification vector matches one of the registered speaker vectors, the operation also includes identifying the first user who issued the first query as the respective registered user associated with one of the registered speaker vectors that matches the extracted speaker identification vector. In some embodiments, speaker identification on audio data corresponding to a second query in order to determine that the second query was spoken by a first user who issued the first query includes extracting a second speaker identification vector representing the features of the second query from the audio data corresponding to the second query, and determining that the extracted second speaker identification vector matches the first user's reference speaker vector.

[0013] Details of one or more embodiments of this disclosure are described in the accompanying drawings and the following description. Other embodiments, features, and advantages will become apparent from the description and drawings, as well as from the claims. [Brief explanation of the drawing]

[0014] [Figure 1A] This is a schematic diagram of an exemplary system including multiple users controlling an assistant-enabled device. [Figure 1B]This is a schematic diagram of an exemplary system including multiple users controlling an assistant-enabled device. [Figure 1C] This is a schematic diagram of an exemplary system including multiple users controlling an assistant-enabled device. [Figure 2A] This is an exemplary GUI rendered on the user's device screen to display a round-robin queue. [Figure 2B] This is an exemplary GUI rendered on the user's device screen to display a round-robin queue. [Figure 2C] This is an exemplary GUI rendered on the user's device screen to display a round-robin queue. [Figure 3] This is a schematic diagram of the query processing process. [Figure 4A] This is a schematic diagram of the speaker identification process. [Figure 4B] This is a schematic diagram of the speaker verification process. [Figure 5] This is an illustrative flowchart of the operational setup for handling voice queries in an environment with multiple users. [Figure 6] This is a schematic diagram of an exemplary computing device that may be used to implement the systems and methods described herein. [Modes for carrying out the invention]

[0015] Similar reference symbols in various drawings indicate the same elements.

[0016] The way users interact with assistant-enabled devices is designed, though not limited, primarily through voice input. For example, a user might ask the device to perform an action, including media playback (e.g., music or podcasts), and the device responds by starting to play audio that matches the user's criteria. In situations where a device (e.g., a smart speaker) is widely shared by multiple users in an environment, the device may need to respond to multiple actions requested by users who may be competing with each other. If one or more users issue multiple individual requests to the device, the device can respond to the individual requests in an ordered manner while also proactively encouraging users who have not made requests to submit them. Ensuring that each user in the environment has the opportunity to submit requests to the device prevents situations where individual users monopolize the device.

[0017] When a device receives a request from a user, it may implement one or more scheduling strategies to handle individual requests. For example, a device may implement a round-robin scheduling strategy in which each user's first request is handled individually in order before subsequent requests from that user are handled. Alternatively, a device may implement a scheduling strategy based on the length of the requested action, the importance of the requested action, or the priority of the requested action. For example, a device may prioritize requests submitted by the event host. Similarly, a device may implement a scheduling strategy that divides users into multiple queues, each corresponding to a type of action (e.g., playing audio, controlling lighting, etc.).

[0018] In addition to controlling actions to prevent a single user from monopolizing playlists, the device may control other types of media such as podcasts and videos. Similarly, the device may prevent a single user from monopolizing the device with multiple related questions by prompting other existing users to submit questions before answering multiple related questions. Additionally, this may be extended to controlling aspects of the home connected to the device. For example, the host of a party may attempt to control the lighting level or the type of music being played during the party to ensure a calming atmosphere. The host may say, "Allow only lighting less than 60% and only jazz music." During the party, the device may prevent other party participants from adjusting the lighting level or limit the range within which it can be adjusted, and may prevent or limit bar participants from requesting music other than the jazz music genre.

[0019] In addition to restricting individuals, the device may operate to be more inclusive of the individuals present in the home. For example, the device may assist individuals in the environment in creating a shopping list, thereby ensuring that all individuals have the opportunity to add items to the shopping list. For example, the device may prompt an individual to add an item to the shopping list in response to a query issued by a first individual adding multiple items to the shopping list. Similarly, the device may actively prompt individuals to participate in events such as setting an alarm in the morning. Here, the device may propose that an individual request an alarm in response to receiving a request to set an alarm from / different individuals. Further, the device may ensure that an individual does not stay away from interactions by engaging / prompting an individual who has not recently spoken to participate in an interaction between the device and other individuals near that individual.

[0020] Figures 1A to 1C show exemplary systems 100a to 100c for processing queries using a scheduling strategy that employs a round-robin mode to actively engage multiple users 102 detected in the environment based on a round-robin queue 350, in an environment with multiple users 102, 102a to 102n. Briefly, as will be described in more detail below, an assistant-enabled device (AED) 104, including a query handler 300 (Figure 3), detects multiple users 102, 102a to 102c in the environment and, in response to receiving a first query 106 issued by user 102a, "Okay, computer, play Glory, let's have pop music all night," begins playing music 122. While AED 104 is playing music 122 as audio from speaker 18, AED 104 receives a second query 146 spoken by the same user 102a, "Next, play Cold Water" (Figure 1B). The query handler 300 detects / recognizes that other users 102b and 102c are present in the environment, and therefore prompts one or more of the other users 102b and 102c to issue a query to add to the round-robin queue 350 before performing any actions associated with the second query 146 issued by user 102a.

[0021] Systems 100a - 100c include an AED 104 that executes a digital assistant 105 through which multiple users 102 can interact by issuing queries including commands to perform actions. In the illustrated example, AED 104 corresponds to a smart speaker through which multiple users 102 can interact. However, AED 104 can include, but is not limited to, other computing devices such as smartphones, tablets, smart displays, desktops / laptops, smartwatches, smart glasses / headsets, smart appliances, headphones, or vehicle infotainment devices. AED 104 includes data processing hardware 10 and memory hardware 12 that stores instructions which, when executed by data processing hardware 10, cause the data processing hardware 10 to perform operations. AED 104 includes an array of one or more microphones 16 configured to capture acoustics such as utterances directed to AED 104. AED 104 may also include, or communicate with, an audio output device (e.g., a speaker) 18 that can output audio such as music 122 and / or synthesized utterances from digital assistant 105. Additionally, AED 104 may include, or communicate with, one or more cameras 19 configured to capture images in the environment and output image data 312 (FIG. 3).

[0022] In some configurations, the AED 104 communicates with multiple user devices 50, 50a-50n associated with multiple users 102. In the illustrated example, each user device 50 in the multiple user devices 50a-50c includes a smartphone that each user 102 can interact with. However, the user devices 50 may include, but are not limited to, other computing devices such as smartwatches, smart displays, smart glasses, smartphones, smart glasses / headsets, tablets, smart appliances, headphones, computing devices, smart speakers, or other assistant-enabled devices. Each user device 50 in the multiple user devices 50a-50n may include at least one microphone 52, 52a-52n present in the user device 50 that communicates with the AED 104. In these configurations, the user devices 50 may also communicate with one or more microphones 16 present in the AED 104. Additionally, multiple users 102 may control and / or configure the AED 104 and interact with the digital assistant 105 using a graphical user interface (GUI) 200, such as 200a to 200n, which is rendered for display on the respective screens of each user device 50.

[0023] As shown in Figures 1A to 1C and Figure 3, the digital assistant 105, which implements the query handler 300, manages queries issued by multiple users 102 using a round-robin queue 350. In some embodiments, the query handler 300 maintains a log of users 102 detected in the environment and one or more queries received from each of the users 102 in the round-robin queue 350 to ensure that the digital assistant 105 takes action associated with the received queries in an ordered manner according to the round-robin queue 350. In this sense, the query handler 300 maintains the order of queries received from multiple users 102 in the round-robin queue 350 and prevents a single user 102 from monopolizing the digital assistant 105 with consecutive queries. Furthermore, the query handler 300 can facilitate the addition of additional and / or conflicting queries from multiple users 102 to the round-robin queue 350.

[0024] In some examples, the query handler 300 identifies metadata of the issued query that defines the order of actions recorded in the round-robin queue 350. Here, the metadata of a query issued to "play" a song may include the length or tempo of the song, and the query handler 300 may perform the execution of consecutive short songs, but may insert other songs between two long songs to avoid performing consecutive long songs. Similarly, the query 300 may group similar musical genres or tempos together within the constraints of the round-robin queue 350 to provide smoother transitions between actions.

[0025] The GUI 200 of each user device 50 associated with user 102 may display a round-robin queue 350 containing the identity of the detected user 102 and the queries associated with each user 102. For example, the round-robin queue 350 may include, for each user 102 of multiple users 102 detected in the environment, the identity of the corresponding user 102 and a query count of queries received from the corresponding user 102. The query count may refer to the number of songs submitted by user 102, each song corresponding to an entry in the round-robin queue 350, or it may refer to the number of actions submitted by user 102, each entry in the round-robin queue 350 corresponding to an action that may contain multiple songs. In some configurations, AED 104 includes a screen and renders the GUI 200 to display the round-robin queue 350 on that screen. For example, AED 104 may include a smart display, tablet, or smart TV in the environment. Figures 2A to 2C provide exemplary GUIs 200b1 to 200b3 that are displayed on the screen of a user device 50b associated with user 102b to keep user 102b informed of the status of the round-robin queue 350. Additionally, the GUI 200 may be rendered to display an identifier for the current action (e.g., "Playing Glory"), an identifier for the AED 104 (e.g., smart speaker) currently performing the action, and / or the identity of user 102a (e.g., Barb) who initiated the current action, which is being performed by the digital assistant 105. As described above, by enabling round-robin mode, which includes a query handler 300 that manages the round-robin queue 350, when the AED 104 determines (e.g., by the query handler 300) that Barb 102a has already issued a first action recorded in the round-robin queue 350, the AED prevents the performance of a second action (e.g., via the digital assistant 105) (or at least requires approval from other users 102b, 102c).

[0026] Continuing to refer to Figures 1A-1C and Figure 3, during the execution of the digital assistant 105, the AED 104 detects multiple users 102a-102c in the environment using the user detector 310 of the query handler 300. For example, the query handler 300 receives proximity information 54 (Figure 3) for the location of each of the multiple users 102a-102c relative to the AED 104 via the user detector 310. In some embodiments, each user device 50a-50c of the multiple users 102 broadcasts proximity information 54 that can be received by the user detector 310, which the AED 104 can use to determine the proximity of each user device 50 to the AED 104. The proximity information 54 from each user device 50 may include wireless communication signals such as WiFi®, Bluetooth®, or ultrasound, and the signal strength of the wireless communication signals received by the user detector 310 may correlate with the proximity (e.g., distance) of the user device 50 to the AED 104. In embodiments where user 102 does not have a user device 50 or has a user device 50 that does not share proximity information 54, the user detector 310 may detect user 102 based on explicit input (e.g., guest list) 313 received from user 102a that issued the first query 106. For example, the user detector 310 receives the guest list 313 from a seed user 102 (e.g., user 102a) that represents each of the multiple users 102 and adds them to the round-robin queue 350. In other embodiments, the user detector 310 automatically detects multiple users 102 in the environment by receiving image data 312 that corresponds to the scene of the environment and is acquired by the camera 19. Here, the user detector 310 detects multiple users 102 based on the received image data 312. Similarly, the user detector 310 may detect multiple users 102 in the environment by performing speaker identification (Figures 4A and 4B) and resolving the identity of user 102 in the environment that issues queries by speaking.Here, the user detector 310 detects multiple users 102 based on the received audio data 402 associated with the issued query, and the user detector 310 may continue to detect users 102 who have issued a query (and associated audio data 402) for a threshold time after the user 102 has spoken.

[0027] In some embodiments, the user detector 310 maintains a list of previous users 314 present in the environment of the AED 104. Here, the list of previous users 314 may refer to a list of users 102 previously detected by the user detector 310, and the list may include one or more users not present in the most recent (i.e., latest) detection of user 102, and / or the latest detection of user 102 may include one or more additional users not present in the list of previous users 314. In this example, after receiving proximity information 54, image data 312, and / or the guest list, the user detector 310 may determine that the list of previous users 314 does not include the same user 102 associated with the current state of the environment (i.e., the list of current users 316). In other words, the user detector 310 may determine that user 102 in the list of previous users 314 is different from user 102 in the list of current users 316. This change between the previous list of users 314 and the current list of users 316 triggers the query handler 300 to update the round-robin queue 350 to add or remove entries based on the current list of users 316, generating an update 352 to be displayed in the GUI 300 for each user 50 associated with the detected user 102. For example, if user 102a leaves the environment, the user detector 310 may detect that user 102a has left by comparing the current list of users 314, which does not include user 102a, with the previous list of users 314, which includes user 102a. In response, the query handler 300 generates an update to the round-robin queue 350 by updating the round-robin queue 350 to remove any queries issued by user 102a. However, in some examples, when the user detector 310 determines that the previous list of users 314 is the same as the detected list of current users 316, the user detector 310 does not need to send the current list of users 316 to update the round-robin queue 350.

[0028] In some examples, when there is a difference between the list of previous users 314 and the list of current users 316, the user detector 310 outputs only the list of current users 316 (thus triggering the query handler to generate an update 352 to the round-robin queue 350). For example, the user detector 310 may consist of a change threshold, and when the difference detected between the list of previous users 314 and the list of current users 316 satisfies the change threshold (e.g., exceeds the threshold), the user detector 310 outputs the list of current users 316 to the query handler 300. The threshold may be zero, and a small difference between the list of previous users 314 and the list of current users 316 detected by the indicator decisioner 210 (e.g., as soon as user 102 enters or leaves the environment) may trigger the query handler 300 to update the round-robin queue 350. Conversely, as a type of queue interruption sensitivity mechanism, the threshold may be higher than zero to prevent unnecessary updates to the round-robin queue 350. For example, the change threshold may be temporary (e.g., a time unit), and if user 102 temporarily leaves the environment (e.g., goes to another room) but returns within the threshold time, AED 104 does not update the round-robin queue 350. Here, the user detector 310 updates the round-robin queue (350) by outputting only the list of the current user 316, when the difference between the list of the previous user 314 and the list of the current user 316 lasts for the threshold time.

[0029] Referring again to Figures 1A and 1C, the user detector 310 can identify user 102a (e.g., Barb) and user 102b (e.g., Jeff) by proximity information 54 received from their respective user devices 50a and 50b. In this example, the user detector 310 does not need to detect user 102c by detecting user device 50c (e.g., user 102c chooses not to share proximity information 54 from user device 50c). However, the user detector 310 still detects the presence of user 102c (e.g., by input from speech data, image data, and / or other users 102a and 102b) and outputs a list of current users 316. The query handler 300 then adds the list of current users 316, including the detected users 102a to 102c, to a round-robin queue 350 that records queries issued by users 102a to 102c.

[0030] Continuing with the example in Figure 1A, user 102a, one of several users 102a-102c, is shown issuing a first query 106, “Okay, computer, play Glory, let's have pop music all night,” near the AED 104. Here, the first query 106 issued by user 102a is spoken by user 102a and includes initial audio data 402 (Figure 3) corresponding to the first query 106. The first query 106 may further include a user input indication showing the user's intent to issue the first query via one of the following: touch, speech, gesture, gaze, and / or input device (e.g., mouse or stylus) in order to interact with the AED 104. Optionally, based on the reception of initial audio data 402 corresponding to the first query 106, the query handler 300 performs a speaker identification process 400a (Figure 4A) on the audio data 402 and resolves the speaker identity of the first query 106 by determining that the first query 106 was issued by user 102a. In other embodiments, user 102a issues the first query 106 without speaking. In these embodiments, user 102a issues the first query 106 via a user device 50a associated with user 102a (e.g., by typing text corresponding to the first query 106 into a GUI 200a displayed on the screen of user device 50a, or by selecting the first query 106 displayed on the screen of user device 50a). Here, the AED 104 can resolve the identity of user 102 who issued the first query 106 by recognizing the user device 50a associated with user 102a.

[0031] The microphone 16 of the AED 104 receives a first query 106 and processes the initial audio data 402 corresponding to the first query 106. Initial processing of the audio data 402 may include filtering the audio data 402 and converting the audio data 402 from an analog signal to a digital signal. Once the AED 104 has processed the audio data 402, it may store the audio data 402 in a buffer of the memory hardware 12 for further processing. With the audio data 402 in the buffer, the AED 104 may use a hotword detector 108 to detect whether the audio data 402 contains a hotword. The hotword detector 108 is configured to identify hotwords contained in the audio data 402 without performing speech recognition on the audio data 402.

[0032] In some embodiments, the hotword detector 108 is configured to identify a hotword in the initial part of the first query 106. In this example, the hotword detector 108 may determine that the first query 106, "Okay, computer, play Glory, let's have pop music all night," contains the hotword 110, "Okay, computer," if the hotword detector 108 detects an acoustic feature in the audio data 402 which is a feature of the hotword 110. The acoustic feature may be a Mel-frequency cepstrum coefficient (MFCC), which is a representation of the short-term power spectrum of the first query 106, or it may be the Mel-scale filter bank energy of the first utterance 106. For example, the hotword detector 108 may detect that the first query 106, "Okay, computer, play Glory, let's have pop music all night," contains the hotword 110, "Okay, computer," based on generating MFCCs from audio data 402 and classifying that the MFCCs contain MFCCs similar to the MFCCs that are features of the hotword "Okay, computer" stored in the hotword model of the hotword detector 108. As another example, the hotword detector 108 may detect that the first query 106, "Okay, computer, play Glory, let's have pop music all night," contains the hotword 110, "Okay, computer," based on generating Mel-scale filter bank energy from audio data 402 and classifying that the Mel-scale filter bank energy contains MFCCs similar to the Mel-scale filter bank energy that are features of the hotword "Okay, computer" stored in the hotword model of the hotword detector 108.

[0033] When the hotword detector 108 determines that the initial audio data 402 corresponding to the first query 106 contains the hotword 110, the AED 104 may trigger a wake-up process to begin speech recognition on the audio data 402 corresponding to the first query 106. For example, Figure 3 shows the AED 104 including a speech recognition device 170 that uses an automated speech recognition model 172 capable of performing speech recognition or semantic interpretation on the audio data 402 corresponding to the first query 106. The speech recognition device 170 may perform speech recognition on the portion of the audio data 402 following the hotword 110. In this example, the speech recognition device 170 may identify the words "Play Glory, let's have pop music all night" from the first query 106.

[0034] In some examples, the AED 104 is configured to communicate with a remote system 130 via a network 120. The remote system 130 may include remote resources such as remote data processing hardware 132 (e.g., a remote server or CPU) and / or remote memory hardware 134 (e.g., a remote database or other storage hardware). The query handler 300 may run on the remote system 130 in addition to or instead of the AED 104. The AED 104 may utilize the remote resources to perform various functions related to speech processing and / or synthesized playback communication. In some embodiments, the speech recognition device 170 is located on the remote system 130 in addition to or instead of the AED 104. When the hotword detector 108 triggers the AED 104 to wake up in response to detecting a hotword 110 in a first query 106, the AED 104 may send initial audio data 402 corresponding to the first query 106 to the remote system 130 via the network 120. Here, AED104 may send a portion of the initial audio data 402 containing the hotword 110 to the remote system 130 to confirm the presence of the hotword 110. Alternatively, AED104 may send only the portion of the initial audio data 402 corresponding to the portion of speech 106 after the hotword 110 to the remote system 130, and the remote system 130 will run the speech recognition device 170 to perform speech recognition and send a transcription of the initial audio data 402 back to AED104.

[0035] Continuing to refer to Figures 1A-1C and Figure 3, the query handler 300 may further include a natural language understanding (NLU) module 320 that performs semantic interpretation on the first query 106 to identify queries / commands directed to the AED 104. Specifically, the NLU module 320 identifies the words of the first query 106 identified by the speech recognition device 170 and performs semantic interpretation to identify any speech command of the first query 106. The NLU module 320 of the AED 104 (and / or remote system 130) identifies the word "play Glory" as a command 111 to perform a first action (i.e., play music 122) and identifies the word "let's have pop music all night" as a constraint 113 that limits the action commanded by queries following the first query 106. In the example shown in Figure 1A, the digital assistant 105 begins by performing a first action: playing music 122 as audio (e.g., track 1) from the speaker 18 of the AED 104. The digital assistant 105 may stream music 122 from a streaming service (not shown), or it may instruct the AED 104 to play music stored in the AED 104. Additionally, the query handler 300 enables round-robin mode, which includes a round-robin queue 350. When round-robin mode is enabled, the digital assistant 105 controls the execution of actions commanded by queries following the first query 106, based on the round-robin queue 350, where the round-robin queue 350 includes a list of current users 316, including multiple users 102a-102c detected by the user detector 310 within the environment of the AED 104. In the example, in response to receiving constraint 113 in the first query 106, the query handler 300 restricts the actions of the digital assistant 105 to playing only music of a specific genre (i.e., pop music).

[0036] Referring to Figure 3, in the example shown in Figure 1A, the query handler 300 enables round-robin mode, including the round-robin queue 350, and adds constraint 113 to the active constraint datastore 330. The query handler 300 maintains a record of the active constraints 332 in the active constraint datastore 330 (for example, stored in the memory hardware 12), and the query handler 300 may restrict the actions that the digital assistant 105 adds to the round-robin queue 350 for it to perform based on the active constraints 332. For example, before the query handler 300 performs the first action associated with command 111 and adds the action "Play Glory" to the round-robin queue (Figure 2A), it may first verify that constraint 113 of command 111 and / or the first query 106 does not conflict with any active constraints 332.

[0037] The AED 104 may inform user 102a (e.g., Barb) who issued the first query 106 that it will enable round-robin mode and use the round-robin queue 350 to control subsequent queries (e.g., queues). For example, the digital assistant 105 may generate a synthesized utterance 123 to the audible output from the speaker 18 of the AED 104 stating, "Barb, I am now enabling round-robin mode to control the queries." In an additional example, the digital assistant 105 provides a notification (e.g., update 352) to the user device 50a associated with user 102a (e.g., Barb) informing user 102a of entries in the round-robin queue 350 and / or any active constraints 232 stored in the active constraint data store 330.

[0038] Referring to Figures 2A to 2C, a graphical user interface (GUI) 200 running on user device 50 may display a round-robin queue 350 containing detected users 102 and queries associated with each detected user 102. As used herein, the GUI 200 may receive user input instructions via one of the following: touch, speech, gesture, gaze, and / or an input device (e.g., mouse or stylus) for interacting with the round-robin queue 350 (via the digital assistant 105). Each listed action may function as a descriptor identifying the respective command 111 for the digital assistant 105 to perform the action. Figure 2A provides an exemplary GUI 200b1 displayed on the screen of user device 50b for informing user 102b of the contents of the round-robin queue 350. Specifically, the GUI 200b1 renders the round-robin queue 350 containing users 102a (Barb), 102b (Jeff), and 102c (Tina) in a repeating order. The round-robin queue 350 further includes placeholders for queries for each of users 102a-102c. As shown in the diagram, the first entry for user 102a (Barb) includes the corresponding action 111 "Play Glory", while subsequent entries for user 102a and the other users 102b and 102c (Jeff and Tina) are empty. In other words, for each of the multiple detected users 102a-102c, the round-robin queue 350 includes the identity of user 102 (e.g., Barb, Jeff, Tina) and a query count of queries containing the corresponding actions received from user 102. As shown in the diagram, user 102a (Barb) has a single query count corresponding to the action "Play Glory", while users 102b (Jeff) and 102c (Tina) have no queries in the round-robin queue 350.

[0039] Additionally, the GUI 200 may render to display an identifier for the current action (e.g., "Playing Glory"), an identifier for the AED 104 (e.g., smart speaker) currently performing the action, an indicator showing that round-robin mode is enabled, and / or the identity of user 102a (e.g., Barb) who issued the first query 106. In embodiments where the first query 106 includes a constraint 113, the GUI 200 renders to display the identifier for the constraint 113 (e.g., pop music). Thus, user 102b may query user device 50b to review the current action of the digital assistant 105 and any constraints 113 that restrict the query issued by user 102b.

[0040] Referring to Figures 4A and 4B, in some embodiments, the AED 104 (or a remote system 130 communicating with the AED 104) also includes an exemplary data store 430 that stores each of the AED 104's registered user data / information for multiple registered users 432a to 432n. Here, each registered user 432 of the AED 104 may undertake a voice registration process to obtain their respective registered speaker vector 154 from audio samples of multiple registered phrases spoken by the registered user 432. For example, the speaker identification model 410 may generate one or more registered speaker vectors 154 from audio samples of registered phrases spoken by each registered user 432 that can be combined, for example, by being averaged or otherwise accumulated, in order to form each registered speaker vector 154. One or more of the registered users 432 may perform the voice registration process using the AED 104, where the microphone 16 captures audio samples of these users speaking registered utterances, and the speaker identification model 410 generates each registered speaker vector 154 therefrom. The model 410 may run on the AED 104, the remote system 130, or a combination thereof. Additionally, one or more of the registered users 432 may register with the AED 104 by providing authorization and authentication credentials to an existing user account on the AED 104, where the existing user account may store the registered speaker vectors 154 obtained from the previous voice registration process, with other devices also linked to the user account.

[0041] In some examples, the registered speaker vector 154 for registered user 432 includes text-dependent registered speaker vectors. For example, text-dependent registered speaker vectors may be extracted from one or more audio samples of each registered user 432 speaking a given term, such as the hotword 110 used to call the AED 104 to wake it up from sleep (e.g., "OK, computer"). In other examples, the registered speaker vector 154 for registered user 432 is obtained from one or more audio samples of each registered user 102 speaking phrases of different terms / words and different lengths, and is text-independent. In these examples, text-independent registered speaker vectors may be obtained over time from audio samples obtained from utterance interactions where user 102 has the AED 104 or other devices linked to the same account.

[0042] Referring to Figure 4A, the speaker identification process 400a identifies user 102a (e.g., Barb) who spoke the first query 106 by first extracting a first speaker identification vector 411 representing the features of the first query 106 issued by user 102a from the initial audio data 402 corresponding to the first query 106. Here, the speaker identification process 400a may run a speaker identification model 410 configured to receive audio data 402 corresponding to the second query 146 as input and to generate the first speaker identification vector 411 as output. The speaker identification model 410 may be a neural network model trained to output the speaker identification vector 411 under machine or human supervision. The speaker identification vector 411 output by the speaker identification model 410 may include an N-dimensional vector having values ​​corresponding to the speech features of the first query 106 associated with user 102a. In some examples, the speaker identification vector 411 is a d-vector. In some examples, the first speaker identification vector 411 includes a set of speaker identification vectors associated with different users who are also authorized to control the AED 104. For example, other authorized users, apart from user 102a who spoke the first query 106, may include other individuals who were present when user 102a spoke the first query 106 that issued the command 111 to perform the first action, and / or individuals that user 102a added / specified as authorized.

[0043] When the first speaker identification vector 411 is output from the model 410, the speaker identification process 400a determines whether the extracted speaker identification vector 411 matches any of the registered speaker vectors 154 stored in the AED 104 (for example, in the memory hardware 12) for the registered users 432a to 432n of the AED 104. As described above, the speaker identification model 410 may generate registered speaker vectors 154 for registered users 200 during the voice registration process. Each registered speaker vector 154 may be used as a reference vector 155 corresponding to a voiceprint or unique identifier representing the voice characteristics of each registered user 432.

[0044] In some embodiments, the speaker identification process 400a uses a comparator 420 that compares a first speaker identification vector 411 with each registered speaker vector 154 associated with each registered user 432a-432n of the AED 104. Here, the comparator 420 may generate a score for each comparison indicating the likelihood that the initial audio data 402 corresponding to the first query 106 corresponds to the identity of each registered user 432, and the identity is accepted when the score meets a threshold. If the score does not meet the threshold, the comparator 420 may reject the identity of the speaker who issued the first query 106. In some embodiments, the comparator 420 calculates the respective cosine distance between the first speaker identification vector 411 and each registered speaker vector 154, and determines that the first speaker identification vector 411 matches one of the registered speaker vectors 154 when the respective cosine distances meet a cosine distance threshold.

[0045] In some examples, the first speaker identification vector 411 is a text-dependent speaker identification vector extracted from one or more word portions corresponding to the first query 106, and each registered speaker vector 154 is also text-dependent to the same one or more words. Using text-dependent speaker vectors can improve the accuracy of determining whether the first speaker identification vector 411 matches any of the registered speaker vectors 154. In other examples, the first speaker identification vector 411 is a text-independent speaker identification vector extracted from the entire initial audio data 402 corresponding to the first query 106.

[0046] When the speaker identification process 400a determines that the first speaker identification vector 411 matches one of the registered speaker vectors 154, the process 400a identifies user 102a, who spoke the first query 106, as the respective registered user 432a associated with one of the registered speaker vectors 154 that matches the extracted speaker identification vector 411. In the illustrated example, the comparator 420 determines the match based on whether the respective cosine distances between the first speaker identification vector 411 and the registered speaker vector 154 associated with registered user 432a satisfy the cosine distance threshold. In some scenarios, the comparator 420 identifies user 102a as the respective registered user 432a associated with the registered speaker vector 154 having the shortest respective cosine distance from the first speaker identification vector 411, if the shortest respective cosine distances also satisfy the cosine distance threshold.

[0047] Conversely, when the speaker identification process 400a determines that the first speaker identification vector 411 does not match any of the registered speaker vectors 154, the process 400a may identify the user 102a who spoke the utterance 106 as a guest user of the AED 104. Thus, the query handler 300 may add the guest user to the round-robin queue 350 and use the first speaker identification vector 411 as a reference speaker vector 155 representing the speech characteristics of the guest user's voice. In some examples, the guest user may be registered with the AED 104, and the AED 104 may store the first speaker identification vector 411 as the registered speaker vector 154 for each new registered user.

[0048] Referring again to Figure 1B, while the digital assistant 105 plays music 122 as playback audio from the speaker 18 of the AED 104, when round-robin mode is enabled, the AED 104 receives a second query 146 containing a command 118 for the digital assistant 105 to take a second action. In the illustrated example, user 102a, who issued the first query 106, also issues a second query 146, “Play Cold Water Next,” containing a command 118 for the digital assistant 105 to play a song (i.e., Cold Water) immediately after playing track #1 (i.e., Glory). Based on having received the second query 146, the query handler 300 resolves the speaker identity of the second query 146 by performing a speaker identification process 400b (Figure 4B) on the audio data 402 corresponding to the second query 146, and determines that the second query 146 was issued by user 102a, who issued the first query 106. As described above with reference to Figure 4A, in an embodiment in which the first query 106 issued by user 102a includes initial audio data 402 (for example, the first query 106 was spoken by the first user 102a), the query handler 300 may first perform a speaker identification process 400a on the initial audio data 402 corresponding to the first query 106 in order to identify the user 102a who issued the first query 106. The speaker identification process 400b may be performed on the data processing hardware 12 of the AED 104. The speaker identification process 400b may also be performed on the remote system 130. If the speaker verification process 400b on the audio data 402 corresponding to the second query 146 indicates that the second query 146 was spoken by a different user 102 than the user 102 who issued the first query 106, the AED 104 may proceed to add the actions associated with the second query 146 to the round-robin queue 350.Conversely, if the speaker verification process 400b for the audio data 402 corresponding to the second query 146 indicates that the second query 146 was spoken by the same user 102 that issued the first query 106, the AED 104 may prevent the performance of the second action (or may require input from at least one or more other users 102 in the environment (e.g., Figure 1B)).

[0049] Referring again to the example in Figure 1B and then to Figure 4B, in response to receiving the second query 146, the AED 104 resolves the identity of user 102 who spoke the second query 146 by performing speaker identification process 400b. Speaker identification process 400b identifies user 102a who spoke the first query 146 by first extracting a second speaker identification vector 412 representing the features of the second query 146 from the audio data 402 corresponding to the first query 146 spoken by user 102a. Here, speaker identification process 400b may run a speaker identification model 410 configured to receive audio data 402 as input and produce a second speaker identification vector 412 as output. As illustrated in Figure 4A, speaker identification model 410 may be a neural network model trained to output the speaker identification vector 412 under machine or human supervision. The second speaker identification vector 412 output by the speaker identification model 410 may include an N-dimensional vector having values ​​corresponding to the speech features of the utterance 146 associated with user 102a. In some examples, the speaker identification vector 412 is a d-vector.

[0050] When the second speaker identification vector 412 is output from the speaker identification model 410, the speaker verification process 400b determines whether the extracted speaker identification vector 412 matches the reference speaker vector 155 associated with the first registered user 432a stored in the AED 104 (for example, in the memory hardware 12). The reference speaker vector 155 associated with the first registered user 432a may include each registered speaker vector 154 associated with the first registered user 432a. As described above, the speaker identification model 410 may generate registered speaker vectors 154 for registered users 432 during the speech registration process. Each registered speaker vector 154 may be used as a reference vector corresponding to a voiceprint or unique identifier representing the voice characteristics of each registered user 432.

[0051] In some embodiments, the speaker verification process 400b uses a comparator 420 that compares a second speaker identification vector 412 with a reference speaker vector 155 associated with a first registered user 432a among the registered users 432. Here, the comparator 420 may generate a comparison score indicating the likelihood that the second query 146 corresponds to the identity of the first registered user 432a, and the identity is accepted when the score meets a threshold. If the score does not meet the threshold, the comparator 420 may reject the identity. In some embodiments, the comparator 420 calculates the respective cosine distances between the second speaker identification vector 412 and the reference speaker vector 155 associated with the first registered user 432a, and determines that the second speaker identification vector matches the reference speaker vector 155 when the respective cosine distances meet a cosine distance threshold.

[0052] When the speaker verification process 400b determines that the second speaker identification vector 412 matches the reference speaker vector 155 associated with the first registered user 432a, the process 400b identifies user 102a, who spoke the second query 146, as the first registered user 432a associated with the reference speaker vector 155. In the illustrated example, the comparator 420 determines the match based on whether the second speaker identification vector 412 and the reference speaker vector 155 associated with the first registered user 432a satisfy the cosine distance threshold. In some scenarios, the comparator 420 identifies user 102a as the respective first registered user 432a associated with the reference speaker vector 155 having the shortest respective cosine distance from the second speaker identification vector 412, if these shortest respective cosine distances also satisfy the cosine distance threshold.

[0053] Referring to Figure 4A above, in some embodiments, the speaker identification process 400a determines that the first speaker identification vector 411 does not match any of the registered speaker vectors 154 and identifies user 102, who spoke the first query 106, as a guest user of the AED 104. Thus, the speaker verification process 400b may first determine whether user 102a, who spoke the first query 106, was identified by the speaker identification process 400a as a registered user 432 or as a guest user. If user 102a is a guest user, the comparator 420 compares the second speaker identification vector 412 with the first speaker identification vector 411 obtained during the speaker identification process 400a. Here, the first speaker identification vector 411 represents the features of the first query 106 spoken by guest user 102a, and is therefore used as a reference vector to verify whether the second query 146 was also spoken by guest user 102a or by another user 102. Here, the comparator 420 may generate a comparison score indicating the likelihood that the second query 146 corresponds to the identity of guest user 102a, and the identity is accepted when the score meets a threshold. If the score does not meet the threshold, the comparator 420 may reject the identity of the guest user who spoke the second query 146. In some embodiments, the comparator 420 calculates the respective cosine distances between the first speaker identification vector 411 and the second speaker vector 412, and determines that the first speaker identification vector 411 matches the second speaker vector 412 when the respective cosine distances meet a cosine distance threshold.

[0054] Referring again to Figure 1B, based on the determination that the second query 146 was spoken by user 102a, who issued the first query 106, the query handler 300 prevents the AED 104 from performing the second action (via the digital assistant 105) and instead prompts at least another user 102 among the multiple users 102a-102c detected in the environment, who is different from user 102a, to have the opportunity to issue the query. In other words, after user 102a is determined to be the issuer of the first query 106 and the speaker of the second query 146, the query handler 300 prevents user 102a from monopolizing the AED 104 by first verifying that the other users 102 in the environment do not wish to issue a query before achieving the second query 146. If the digital assistant 105 confirms that other users 102 in the environment do not wish to issue queries, it may proceed to fulfill the second query 146 issued by user 102a once it has completed fulfilling the first query 106.

[0055] In some embodiments, the query handler 300 (via the digital assistant 105) prompts at least one other user 102 of a plurality of users 102 detected in the environment to provide a third query 148 to perform a third action, which includes a corresponding command 119, before performing a second action, which is contained in a second query 146 issued by user 102a. In these embodiments, prompting at least one other user 102 of the plurality of users 102 includes providing, as output from the AED 104, an option that, if selected, allows the user to issue a third query 148 to perform a third action, which includes a corresponding command 119. For example, the digital assistant 105 may generate a synthesized utterance 123 for an audible output from the speaker 18 of the AED 104 (or a speaker communicating with data processing hardware (e.g., the speaker of user device 50)), which prompts user 102b to issue the query "Jeff, would you like to choose a song next?". In response, user 102b (i.e., Jeff), one of several users 102b, 102c distinct from user 102a, is shown issuing a third query 148, “Yes, play Red next,” near the AED 104. In an additional example, the digital assistant 105 may provide a notification to the user device 50b associated with user 102b (e.g., Jeff) displaying user-selectable options as a graphic element 210 on the screen of the user device 50b, prompting user 102b to have the opportunity to issue the third query 148. Additional or alternative to prompting user 102b audibly, as shown in Figure 2B, the GUI 200b2 allows user 102b to issue (or decline the opportunity to issue) the third query 148 by rendering the graphic element 210 “Jeff, would you like to choose the next song to play?” and “Yes” and “No” to display. Here, the query handler 300 receives a third query 148 when it receives a user input instruction indicating that the user device 50b is selecting one of the graphic elements 210, and the user is selecting one of the available options.

[0056] Referring to Figures 1B and 3, in response to receiving the third query 148, the NLU module 320 running on AED 104 (and / or on remote system 130) may identify the word “Play Red” as command 119 and perform a third action (i.e., play music 122). The query handler 300 may first determine whether the command 119 for performing the third action does not violate any active constraints 232 of the digital assistant 105, and then update the round-robin queue 350 to include the third action. As described above, the active constraints 332 may include constraint 113 included in the first query 106, which restricts the genre of music (e.g., pop music). Here, the query handler 300 verifies that the action of playing “Red” in the third query 148 issued by user 102b does not conflict with constraint 113 which restricts the genre of music to pop music, before adding “Red” to the round-robin queue 350. In some embodiments, the query handler 300 further verifies that the action of playing "Cold Water" included in the second query 146 does not conflict with the constraint 113 that limits the music genre to pop music and / or other active constraints 332 on the digital assistant 105 before updating the round-robin queue 350 to include command 118 in the second query 146 issued by user 102a.

[0057] These examples primarily refer to actions that play music to prevent a single user 102 from monopolizing a playlist, but actions may refer to any category of actions, including but not limited to, search queries, controlling assistant-enabled devices (e.g., smart lights, smart thermostats), and playing other types of media (e.g., podcasts, videos, etc.). For example, a query handler 300 may help users 102 in the environment create a shopping list, thereby ensuring that all users 102 have the opportunity to add items to their shopping list. For example, AED 104 may add multiple items requested by a first user 102 to the shopping list, while prompting a second user 102 to add items in response to the first user 102 issuing a query to add multiple items to the shopping list. Similarly, a query handler 300 may proactively prompt a user 102 to engage in actions such as setting a morning alarm, where AED 104 may suggest that the second user 102 request an alarm in response to receiving a request from the first user 102 to set an alarm. Furthermore, the query handler 300 can ensure that AED 104 does not leave user 102 out of the interaction by engaging / prompting user 102 who has not recently issued a query to participate in the interaction between AED 104 and other users 102.

[0058] Additionally, constraint 332 may include several restrictions on the actions themselves. For example, there may be a time limit on actions that limits the amount of time a user 102 has for each entry in the round-robin queue 350 (e.g., the number of jokes user 102 can request for each turn in the round-robin queue). The time limit in round-robin mode may control how long an event that applies constraint 332 lasts. Similarly, constraint 332 on the threshold number of actions per query may limit the number of actions user 102 can request for each issued query entry in the round-robin queue 350 (e.g., user 102 can request an entire album and / or playlist for each turn in the round-robin queue 350). Likewise, the threshold number of actions may include the total number of user-specific actions in the round-robin queue 350. For example, user 102 can only submit up to 50 additional songs to the round-robin queue 350.

[0059] Continuing with the examples in Figures 1B and 2C, after the query handler 300 verifies that the action to play "Red" in the third query 148 issued by user 102b does not conflict with the constraint 113 that limits the music genre to pop music, the query handler 300 updates the round-robin queue 350 to add "Red" to the identity associated with user 102b in the round-robin queue 350. As shown in Figure 2C, GUI 200b3 renders the round-robin queue 350 containing users 102a (Barb), user 102b (Jeff), and user 102c (Tina) in repeating order. Jeff's action placeholder is updated to include command 119 for the digital assistant 105 to perform the action "Play Red" after AED 105 has completed the execution of the action "Play Glory" in command 111 in the first query 106 (e.g., play). In other words, based on the round-robin queue 350, when the digital assistant 105 completes the execution of the first action 111 of the first query 106 issued by user 102a, it executes the execution of the third action 119 of the third query 148 issued by user 102b. Additionally, the query handler 300 updates the round-robin queue 350 to include the second query 146 issued by user 102a. As shown in the diagram, the action placeholder for the second entry of the barb is updated to include the command 118 for the digital assistant 105 to perform the action "Replay Red" of command 119 of the third query 148, and any additional intervention queries (e.g., replay) submitted by different users of the environment 102 from users 102a and 102b, before the AED 105 completes the execution of the action associated with the third query 148.

[0060] Referring to Figure 1C, the AED 104 may notify user 102a (e.g., Barb), who issued the first query 106, that the second query 146 will be achieved after the third query 148 based on the round-robin queue 350 is achieved. For example, the digital assistant 105 may generate a synthesized utterance 123 for the audible output from the speaker 18 of the AED 104, stating, "After Jeff's selection, Barb, play cold water." In an additional example, the digital assistant 105 provides a notification (e.g., update 352) to the user device 50a associated with user 102a (e.g., Barb) informing user 102a of an updated entry in the round-robin queue 350.

[0061] Figure 5 is a flowchart illustrating the configuration of an exemplary operation of method 500 for processing queries in an environment with multiple users 102. In operation 502, method 500 includes detecting multiple users 102 in the environment of an assistant-enabled device (AED) 104. Method 500 also includes, in operation 504, receiving a first query 106 issued by a first user 102a among the multiple users 102. The first query 106 includes a command 111 for the digital assistant 105 to perform a first action. Method 500 further includes, in operation 506, enabling round-robin mode, which, once enabled, causes the digital assistant 105 to control the execution of actions commanded by queries following the first query 106, based on a round-robin queue 350, where the round-robin queue 350 includes multiple users 102 detected in the environment of the AED 104.

[0062] While the digital assistant 105 is performing the first action and when round-robin mode is enabled, method 500 further includes, in operation 508, receiving audio data 402 corresponding to a second query 146 spoken by one of the multiple users 102 and captured by the AED 104, where the second query 146 includes a command 118 for the digital assistant 105 to perform the second action. In operation 510, method 500 also includes speaker identification in the audio data 402 corresponding to the second query 146 to determine that the second query 146 was spoken by the first user 102a that issued the first query 106.

[0063] Based on the determination that the second query 146 was spoken by the first user 102a who issued the first query 106, method 500 further includes in operation 512 preventing the digital assistant 105 from performing the second action and prompting at least one other user 102 among the multiple users 102 detected in the environment, different from the first user 102a, to have the opportunity to issue a query. In operation 514, method 500 includes receiving a third query 148 issued by the second user 102b among the multiple users 102 detected in the environment, the third query 148 including a command 119 for the digital assistant 105 to perform the third action. Once the digital assistant 105 has completed performing the first action, method 500 further includes in operation 502 performing the third action.

[0064] Figure 6 is a schematic diagram of an exemplary computing device 600 that may be used to implement the systems and methods described herein. The computing device 600 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown herein, their connections and relationships, and their functions are for illustrative purposes only and are not intended to limit the embodiments of the invention described and / or claimed in this specification.

[0065] The computing device 600 includes a processor 610, memory 620, storage device 630, a high-speed interface / controller 640 connected to memory 620 and high-speed expansion port 650, and a low-speed bus 670 and a low-speed interface / controller 660 connected to storage device 630. Each component 610, 620, 630, 640, 650, and 660 is interconnected using various buses and may be mounted on a common motherboard or implemented in other ways as needed. The processor 610 (e.g., data processing hardware 10,132 in Figure 1) processes instructions for execution within the computing device 600, including instructions stored in memory 620 or storage device 630, to display graphical information of a graphical user interface (GUI) on an external input / output device such as a display 680 connected to the high-speed interface 640. In other embodiments, multiple processors and / or multiple buses may be used together with multiple memories and multiple types of memory, as needed. Additionally, multiple computing devices 600 may be connected, with each device performing some of the necessary operations (for example, as a server bank, a group of blade servers, or a multiprocessor system).

[0066] Memory 620 stores information non-temporarily within the computing device 600. Memory 620 (for example, memoryware 12, 134 in Figure 1) may be computer-readable media, volatile memory units, or non-volatile memory units. Non-temporarily memory 620 may be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by the computing device 600. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase-change memory (PCM), and disk or tape.

[0067] The storage device 630 can provide high-capacity storage to the computing device 600. In some embodiments, the storage device 630 is a computer-readable medium. In various different embodiments, the storage device 630 may be a device array including a floppy disk device, a hard disk device, an optical disk device, or a tape device, flash memory or other similar solid-state memory device, or a storage area network or other configuration device. In additional embodiments, a computer program product is tangibly embodied in an information carrier. When executed, the computer program product includes instructions that perform one or more of the methods described above. The information carrier is a computer-readable or machine-readable medium such as memory 620, the storage device 630, or the memory of the processor 610.

[0068] The high-speed controller 640 manages bandwidth-intensive operations for the computing device 600, while the low-speed controller 660 manages low-bandwidth-intensive operations. Such role assignments are merely examples. In some embodiments, the high-speed controller 640 is connected to memory 620, a display 680 (e.g., via a graphics processor or accelerator), and a high-speed expansion port 650 that can accept various expansion cards (not shown). In some embodiments, the low-speed controller 660 is connected to a storage device 630 and a low-speed expansion port 690. The low-speed expansion port 690, which may include various communication ports (e.g., USB, Bluetooth®, Ethernet®, wireless Ethernet), may be connected to one or more input / output devices such as a keyboard, pointing device, scanner, or networking devices such as a switch or router, for example, via a network adapter.

[0069] The computing device 600 can be implemented in many different forms, as shown in the figure. For example, it can be implemented as a standard server 600a, or multiple times in a group of such servers 600a, as a laptop computer 600b, or as part of a rack server system 600c.

[0070] Various embodiments of the systems and technologies described herein can be realized in digital electronic circuits and / or optical circuits, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor. Such a programmable processor may be specialized or general-purpose and may be connected to a storage system, at least one input device, and at least one output device to send and receive data and instructions.

[0071] A software application (i.e., a software resource) can refer to computer software that causes a computing device to perform a task. In some examples, a software application may be called an “application,” “app,” or “program.” Exemplary applications include, but are not limited to, system diagnostic applications, system administration applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and game applications.

[0072] Non-temporary memory can be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by a computing device. Non-temporary memory can be volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase-change memory (PCM), and disk or tape.

[0073] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages ​​and / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” mean any computer program product, non-temporary computer-readable medium, apparatus and / or device (e.g., magnetic disks, optical disks, memory, programmable logic devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, and include machine-readable medium that receives machine instructions as machine-readable signals. The term “machine-readable signal” means any signal used to provide machine instructions and / or data to a programmable processor.

[0074] The processes and logical flows described herein may be performed by one or more programmable processors, also called data processing hardware, executing one or more computer programs to perform functions by acting on input data and producing outputs. Processes and logical flows may also be performed by special-purpose logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). Processors suitable for executing computer programs include, for example, both general-purpose and special-purpose processors, and one or more processors of any type of digital computer. Generally, processors receive instructions and data from read-only memory, random-access memory, or both. The basic elements of a computer are a processor for executing instructions, and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or is operablely connected to receive data from or transmit data to them, or both. However, a computer is not required to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices including EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. Processors and memory may be complemented by or incorporated into dedicated logic circuits.

[0075] To interact with a user, one or more aspects of the present disclosure can be implemented in a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor or touchscreen, and optionally a keyboard and pointing device, such as a mouse or trackball, to which the user can input into the computer. Other types of devices can also be used to provide interaction with a user. For example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback or haptic feedback, and input from the user may be received in any form, such as acoustic input, voice input or haptic input. Furthermore, the computer can interact with the user by sending documents to and receiving documents from the user's device, for example, by sending a web page to a web browser on the user's client device in response to a request received from a web browser.

[0076] Several embodiments have been described. Nevertheless, it is understood that various modifications can be made without departing from the spirit and scope of this disclosure. Accordingly, other embodiments are within the scope of the following claims.

Claims

1. A method (500) executed by a computer, which, when executed by data processing hardware (610), causes the data processing hardware (610) to perform an operation, wherein the operation is Detecting multiple users within the environment of an assistant-enabled device (104), The operation includes receiving a first query (106) issued by a first user among the plurality of users, wherein the first query (106) includes a command (111) for the digital assistant (105) to perform a first action, and the operation further includes This includes enabling round-robin mode, which, once enabled, causes the digital assistant (105) to control the execution of actions commanded by queries following the first query (106), based on a round-robin queue (350), the round-robin queue (350) including the multiple users detected within the environment of the assistant-enabled device (104), and the operation further includes: While the digital assistant (105) is performing the first action, and while the round-robin mode is enabled, The process includes receiving audio data (402) corresponding to a second query (146) spoken by one of the users and captured by the assistant-enabled device (104), wherein the second query (146) includes a command (111) for the digital assistant (105) to perform a second action, and the action further includes: The operation further includes performing speaker identification on the audio data (402) corresponding to the second query (146) in order to determine that the second query (146) was spoken by the first user who issued the first query (106), and the operation further includes Based on the determination that the second query (146) was spoken by the first user who issued the first query (106), To prevent the digital assistant (105) from performing the second action, The operation further includes prompting at least other users among the multiple users detected in the environment, who are different from the first user, to issue a query, and the operation further includes: The operation includes receiving a third query (148) issued by a second user among the multiple users detected in the environment, the third query (148) including a command (111) for the digital assistant (105) to perform a third action, and the operation further includes A method comprising performing the third action once the digital assistant (105) has completed the first action.

2. Prompting at least one of the multiple users detected in the environment includes providing an option that can be selected to perform the third action and to issue the third query (148) from the second user, as output from the user interface (200) of the user device (50) associated with the second user, The method according to claim 1 (500), wherein receiving the third query (148) issued by the second user is based on receiving user input instructions indicating the selection of options available to the user.

3. Providing user-selectable options as output from the user interface (200) includes displaying user-selectable options as graphic elements (210) on the screen of the user device (50) associated with the second user via the user interface (200), wherein the graphic elements (210) prompt the second user to have the opportunity to issue the third query (148), according to claim 2 (500).

4. Providing the user-selectable options as output from the user interface (200) includes providing the user-selectable options as an audible output from a speaker (18) communicating with the data processing hardware (610) via the user interface (200), the audible output prompting the second user to have the opportunity to issue the third query (148), according to claim 2 (500).

5. The first query (106) further includes constraints on subsequent queries, the constraints being: Action category, Time constraints on the action, The aforementioned time limit of the round-robin mode, or The method according to any one of claims 1 to 4 (500), comprising one of the threshold numbers for actions per query.

6. The aforementioned operation is, The determination that the second query (146) does not violate the constraints of the first query (106), The method of claim 5 (500), further comprising updating the round-robin queue (350) to include the second query (146) spoken by the first user.

7. The aforementioned operation is, The first user detects that the assistant-enabled device (104) has left the environment, The method according to claim 6 (500), further comprising updating the round-robin queue (350) to remove the second query (146) spoken by the first user.

8. The method (500) according to any one of claims 1 to 4, wherein detecting the plurality of users within the environment of the assistant-enabled device (104) includes detecting the at least one of the plurality of users based on proximity information (54) of a user device (50) associated with at least one of the plurality of users.

9. Detecting the multiple users within the environment of the assistant-enabled device (104) is Receiving image data (312) corresponding to the scene of the aforementioned environment, The method according to any one of claims 1 to 4 (500), further comprising detecting at least one of the plurality of users based on the image data (312).

10. The method according to any one of claims 1 to 4 (500), wherein detecting the plurality of users within the environment of the assistant-enabled device (104) includes receiving a list indicating each of the plurality of users to be added to the round-robin queue (350).

11. The round-robin queue (350) provides, for each corresponding user among the multiple users detected in the environment, The identity of the corresponding user, The method according to any one of claims 1 to 4 (500), further comprising a query count of the queries received from the corresponding user.

12. The method (500) according to any one of claims 1 to 4, wherein receiving the first query (106) issued by the first user includes receiving a user input instruction from a user device (50) associated with the first user indicating the user's intention to issue the first query (106).

13. The method according to any one of claims 1 to 4 (500), wherein receiving the first query (106) issued by the first user includes receiving initial audio data corresponding to the first query (106) issued by the first user and captured by the assistant-enabled device (104).

14. The operation further includes, after receiving the initial audio data (402) corresponding to the first query (106) issued by the first user, performing speaker identification on the initial audio data (402) in order to identify the first user who issued the first query (106), wherein the speaker identification is performed Extracting a first speaker identification vector (411) representing the features of the first query (106) issued by the first user from the initial audio data (402) corresponding to the first query (106) issued by the first user, This is done by determining whether the extracted speaker identification vector matches any of the registered speaker vectors (154) stored in the assistant-enabled device (104), and each registered speaker vector (154) is associated with a different registered user of the assistant-enabled device (104), and the speaker identification is further performed by The method according to claim 13 (500), which is performed by determining that the first speaker identification vector (411) matches one of the registered speaker vectors (154), and then identifying the first user who issued the first query (106) as the respective registered user associated with one of the registered speaker vectors (154) that matches the extracted speaker identification vector.

15. In order to determine that the second query (146) was spoken by the first user who issued the first query (106), speaker identification is performed on the audio data (402) corresponding to the second query (146). Extracting a second speaker identification vector (412) representing the features of the second query (146) from the audio data (402) corresponding to the second query (146), The method according to any one of claims 1 to 4 (500), comprising determining that the extracted second speaker identification vector matches the first user's reference speaker vector (155).

16. Data processing hardware (610), The system includes memory hardware (620) that communicates with the data processing hardware (610), and the memory hardware (620), when executed by the data processing hardware (610), stores instructions that cause the data processing hardware (610) to perform an action, and the action is Detecting multiple users within the environment of an assistant-enabled device (104), The operation includes receiving a first query (106) issued by a first user among the plurality of users, wherein the first query (106) includes a command (111) for the digital assistant (105) to perform a first action, and the operation further includes This includes enabling round-robin mode, which, once enabled, causes the digital assistant (105) to control the execution of actions commanded by queries following the first query (106), based on a round-robin queue (350), the round-robin queue (350) including the multiple users detected within the environment of the assistant-enabled device (104), and the operation further includes: While the digital assistant (105) is performing the first action, and while the round-robin mode is enabled, The process includes receiving audio data (402) corresponding to a second query (146) spoken by one of the users and captured by the assistant-enabled device (104), wherein the second query (146) includes a command (111) for the digital assistant (105) to perform a second action, and the action further includes: The operation further includes performing speaker identification on the audio data (402) corresponding to the second query (146) in order to determine that the second query (146) was spoken by the first user who issued the first query (106), and the operation further includes Based on the determination that the second query (146) was spoken by the first user who issued the first query (106), To prevent the digital assistant (105) from performing the second action, The operation further includes prompting at least other users among the multiple users detected in the environment, who are different from the first user, to issue a query, and the operation further includes: The operation includes receiving a third query (148) issued by a second user among the multiple users detected in the environment, the third query (148) including a command (111) for the digital assistant (105) to perform a third action, and the operation further includes A system (100) which, once the digital assistant (105) has completed the first action, performs the third action.

17. Prompting at least one of the multiple users detected in the environment includes providing an option that can be selected to perform the third action and to issue the third query (148) from the second user, as output from the user interface (200) of the user device (50) associated with the second user, The system (100) according to claim 16, wherein receiving the third query (148) issued by the second user is based on receiving user input instructions indicating the selection of options available to the user.

18. Providing user-selectable options as output from the user interface (200) includes displaying user-selectable options as graphic elements (210) on the screen of the user device (50) associated with the second user via the user interface (200), the system (100) according to claim 17, wherein the graphic elements (210) prompt the second user to have the opportunity to issue the third query (148).

19. Providing the user-selectable options as output from the user interface (200) includes providing the user-selectable options as an audible output from a speaker (18) communicating with the data processing hardware (610) via the user interface (200), the system (100) according to claim 17, wherein the audible output prompts the second user to have the opportunity to issue the third query (148).

20. The first query (106) further includes constraints on subsequent queries, the constraints being: Action category, Time constraints on the action, The aforementioned time limit of the round-robin mode, or A system (100) according to any one of claims 16 to 19, comprising one of the threshold numbers for actions per query.

21. The aforementioned operation is, The determination that the second query (146) does not violate the constraints of the first query (106), The system (100) according to claim 20, further comprising updating the round-robin queue (350) to include the second query (146) spoken by the first user.

22. The aforementioned operation is, The first user detects that the assistant-enabled device (104) has left the environment, The system (100) according to claim 21, further comprising updating the round-robin queue (350) to remove the second query (146) spoken by the first user.

23. The system (100) according to any one of claims 16 to 19, wherein detecting the plurality of users within the environment of the assistant-enabled device (104) includes detecting the at least one of the plurality of users based on proximity information (54) of a user device (50) associated with at least one of the plurality of users.

24. Detecting the multiple users within the environment of the assistant-enabled device (104) is Receiving image data (312) corresponding to the scene of the aforementioned environment, The system (100) according to any one of claims 16 to 19, further comprising detecting at least one of the plurality of users based on the image data (312).

25. The system (100) according to any one of claims 16 to 19, wherein detecting the plurality of users within the environment of the assistant-enabled device (104) includes receiving a list indicating each of the plurality of users to be added to the round-robin queue (350).

26. The round-robin queue (350) provides, for each corresponding user among the multiple users detected in the environment, The identity of the corresponding user, The system (100) according to any one of claims 16 to 19, further comprising a query count of queries received from the corresponding user.

27. The system (100) according to any one of claims 16 to 19, wherein receiving the first query (106) issued by the first user includes receiving a user input instruction from a user device (50) associated with the first user indicating the user's intention to issue the first query (106).

28. The system (100) according to any one of claims 16 to 19, wherein receiving the first query (106) issued by the first user includes receiving initial audio data corresponding to the first query (106) issued by the first user and captured by the assistant-enabled device (104).

29. The operation further includes, after receiving the initial audio data (402) corresponding to the first query (106) issued by the first user, performing speaker identification on the initial audio data (402) in order to identify the first user who issued the first query (106), wherein the speaker identification is performed Extracting a first speaker identification vector (411) representing the features of the first query (106) issued by the first user from the initial audio data (402) corresponding to the first query (106) issued by the first user, This is done by determining whether the extracted speaker identification vector matches any of the registered speaker vectors (154) stored in the assistant-enabled device (104), and each registered speaker vector (154) is associated with a different registered user of the assistant-enabled device (104), and the speaker identification is further performed by The system (100) according to claim 28, which is performed by determining that the first speaker identification vector (411) matches one of the registered speaker vectors (154), and then identifying the first user who issued the first query (106) as the respective registered user associated with one of the registered speaker vectors (154) that matches the extracted speaker identification vector.

30. In order to determine that the second query (146) was spoken by the first user who issued the first query (106), speaker identification is performed on the audio data (402) corresponding to the second query (146). Extracting a second speaker identification vector (412) representing the features of the second query (146) from the audio data (402) corresponding to the second query (146), The system (100) according to any one of claims 16 to 19, comprising determining that the extracted second speaker identification vector matches the first user's reference speaker vector (155).

31. A program for causing a computer to perform the method described in any one of claims 1 to 4.