Handling voice queries in a multi-user environment
A round-robin scheduling strategy with speaker identification and proactive prompting addresses the challenge of managing conflicting voice queries in multi-user environments, ensuring fair and inclusive device interaction.
Patent Information
- Application Number
- JP2025519858
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-06
- Filing Date
- 2023-09-12
- Publication Date
- 2025-10-09
- Estimated Expiration
- 2043-09-12
AI Technical Summary
In multi-user environments, assistant-enabled devices struggle to manage conflicting voice queries from multiple users effectively, leading to situations where one user monopolizes the device, disrupting the experience for others.
Implementing a round-robin scheduling strategy that manages queries from multiple users by ensuring each user's request is addressed in order, with speaker identification to confirm the user's intent and proactive prompting of other users to submit queries, using a digital assistant to enforce this order.
Ensures fair and inclusive interaction among multiple users by preventing monopolization and allowing all users to participate in device control, enhancing user experience and device efficiency.
Smart Images

Figure 2025533870000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to processing voice queries in a multi-user environment. [Background technology]
[0002] The way users interact with assistant-enabled devices is primarily, but not exclusively, designed through voice input. For example, a user may ask the device to perform an action, including media playback (e.g., music or podcasts), and the device responds by starting to play audio that matches the user's criteria. In situations where a device (e.g., a smart speaker) is widely shared by multiple users in an environment, the device may need to accommodate multiple actions requested by users that may conflict with each other. Summary of the Invention
[0003] One aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations including detecting multiple users in an environment of an assistant-enabled device (AED). The operations also include receiving a first query issued by a first user of the multiple users. The first query includes a command for the digital assistant to perform a first action. The operations further include enabling a round-robin mode, where, when the round-robin mode is enabled, the digital assistant controls performance of actions commanded by queries subsequent to the first query based on a round-robin queue, where the round-robin queue includes multiple users detected in the environment of the AED. While the digital assistant is performing the first action and when the round-robin mode is enabled, the operations further include receiving audio data corresponding to a second query spoken by one of the multiple users and captured by the AED. Where the second query includes a command for the digital assistant to perform a second action. The operations also include performing speaker identification on the audio data corresponding to the second query to determine that the second query was spoken by the first user who issued the first query. Based on determining that the second query was spoken by the first user who issued the first query, the operations also include preventing the digital assistant from performing a second action and prompting at least another user of the plurality of users detected in the environment who is different from the first user to issue a query. The operations further include receiving a third query issued by a second user of the plurality of users detected in the environment, the third query including a command for the digital assistant to perform a third action, and the operations further include performing the third action when the digital assistant completes performance of the first action.
[0004] Embodiments of the present disclosure may include one or more of the following optional features. In some embodiments, prompting at least another user of the plurality of users detected in the environment includes providing, as output from a user interface of a user device associated with the second user, a user-selectable option that, when selected to perform a third action, issues a third query from the second user. Here, receiving the third query issued by the second user is based on receiving a user input indication indicating a selection of the user-selectable option. In these embodiments, providing the user-selectable option as output from the user interface may include displaying, via the user interface, the user-selectable option as a graphical element on a screen of the user device associated with the second user. The graphical element prompts the second user with the opportunity to issue the third query. Additionally or alternatively, providing the user-selectable option as output from the user interface includes providing, via the user interface, the user-selectable option as an audible output from a speaker in communication with the data processing hardware. The audible output prompts the second user with the opportunity to issue the third query.
[0005] In some examples, the first query further includes a constraint for subsequent queries, the constraint including one of a category of actions, a time limit for actions, a time limit for round-robin mode, or a threshold number of actions per query. In these examples, the operations further include determining that the second query does not violate the constraint of the first query and updating the round-robin queue to include the second query spoken by the first user. In these examples, the operations may further include detecting that the first user has left the environment of the assistant-enabled device and updating the round-robin queue to remove the second query spoken by the first user.
[0006] In some embodiments, detecting a plurality of users in the environment of the assistant-enabled device includes detecting at least one of the plurality of users based on proximity information of a user device associated with at least one of the plurality of users. Additionally or alternatively, detecting a plurality of users in the environment of the assistant-enabled device includes receiving image data corresponding to a scene of the environment and detecting at least one of the plurality of users based on the image data. In other embodiments, detecting a plurality of users in the environment of the assistant-enabled device includes receiving a list indicating each user of the plurality of users to add to a round-robin queue. In some examples, the round-robin queue includes, for each corresponding user of the plurality of users detected in the environment, the identity of the corresponding user and a query count of queries received from the corresponding user. In some embodiments, receiving a first query issued by a first user includes receiving a user input indication from a user device associated with the first user indicating a user intent to issue the first query.
[0007] In some examples, receiving a first query issued by a first user includes receiving initial audio data corresponding to the first query issued by the first user and captured by the assistant-enabled device. In these examples, after receiving the initial audio data corresponding to the first query issued by the first user, the operations further include performing speaker identification on the initial audio data to identify the first user who issued the first query. The speaker identification includes extracting a first speaker identification vector representing characteristics of the first query issued by the first user from the initial audio data corresponding to the first query issued by the first user, and determining whether the extracted speaker identification vector matches any enrollment speaker vector stored in the assistant-enabled device. Each enrollment speaker vector is associated with a different respective enrolled user of the assistant-enabled device. When the first speaker identification vector matches one of the enrollment speaker vectors, the operations also include identifying the first user who issued the first query as the respective enrolled user associated with one of the enrollment speaker vectors that matches the extracted speaker identification vector. In some implementations, performing speaker identification on the audio data corresponding to the second query to determine that the second query was spoken by the first user who issued the first query includes extracting a second speaker identification vector from the audio data corresponding to the second query that represents characteristics of the second query, and determining that the extracted second speaker identification vector matches a reference speaker vector of the first user.
[0008] Another aspect of the present disclosure provides a system including data processing hardware and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed by the data processing hardware, cause the data processing hardware to perform operations including detecting multiple users in an environment of an assistant-enabled device (AED). The operations also include receiving a first query issued by a first user of the multiple users. The first query includes a command for the digital assistant to perform a first action. The operations further include enabling a round-robin mode, where, when the round-robin mode is enabled, the digital assistant controls performance of actions commanded by queries subsequent to the first query based on a round-robin queue, where the round-robin queue includes multiple users detected in the environment of the AED. While the digital assistant is performing the first action and when the round-robin mode is enabled, the operations further include receiving audio data corresponding to a second query spoken by one of the multiple users and captured by the AED. Where the second query includes a command for the digital assistant to perform a second action. The operations also include performing speaker identification on the audio data corresponding to the second query to determine that the second query was spoken by the first user who issued the first query. Based on determining that the second query was spoken by the first user who issued the first query, the operations also include preventing the digital assistant from performing a second action and prompting at least another user of the plurality of users detected in the environment, different from the first user, to issue a query. The operations further include receiving a third query issued by a second user of the plurality of users detected in the environment, the third query including a command for the digital assistant to perform a third action, and the operations further include performing the third action when the digital assistant completes performance of the first action.
[0009] This aspect may include one or more of the following optional features. In some implementations, prompting at least another user of the plurality of users detected in the environment includes providing, as output from a user interface of a user device associated with the second user, a user-selectable option that, when selected to perform a third action, issues a third query from the second user. Here, receiving the third query issued by the second user is based on receiving a user input indication indicating a selection of the user-selectable option. In these implementations, providing the user-selectable option as output from the user interface may include displaying, via the user interface, the user-selectable option as a graphical element on a screen of a user device associated with the second user, the graphical element prompting the second user with the opportunity to issue the third query. Additionally or alternatively, providing the user-selectable option as output from the user interface includes providing, via the user interface, the user-selectable option as an audible output from a speaker in communication with the data processing hardware. The audible output prompts the second user with the opportunity to issue the third query.
[0010] In some examples, the first query further includes a constraint for subsequent queries, the constraint including one of a category of actions, a time limit for actions, a time limit for round-robin mode, or a threshold number of actions per query. In these examples, the operations may further include determining that the second query does not violate the constraint of the first query and updating the round-robin queue to include the second query spoken by the first user. In these examples, the operations may further include detecting that the first user has left the environment of the assistant-enabled device and updating the round-robin queue to remove the second query spoken by the first user.
[0011] In some embodiments, detecting a plurality of users in the environment of the assistant-enabled device includes detecting at least one of the plurality of users based on proximity information of a user device associated with at least one of the plurality of users. Additionally or alternatively, detecting a plurality of users in the environment of the assistant-enabled device includes receiving image data corresponding to a scene of the environment and detecting at least one of the plurality of users based on the image data. In other embodiments, detecting a plurality of users in the environment of the assistant-enabled device includes receiving a list indicating each user of the plurality of users to add to a round-robin queue. In some examples, the round-robin queue includes, for each corresponding user of the plurality of users detected in the environment, the identity of the corresponding user and a query count of queries received from the corresponding user. In some embodiments, receiving a first query issued by a first user includes receiving a user input indication from a user device associated with the first user indicating a user intent to issue the first query.
[0012] In some examples, receiving a first query issued by a first user includes receiving initial audio data corresponding to the first query issued by the first user and captured by the assistant-enabled device. In these examples, after receiving the initial audio data corresponding to the first query issued by the first user, the operations further include performing speaker identification on the initial audio data to identify the first user who issued the first query. The speaker identification includes extracting a first speaker identification vector representing characteristics of the first query issued by the first user from the initial audio data corresponding to the first query issued by the first user, and determining whether the extracted speaker identification vector matches any enrollment speaker vector stored in the assistant-enabled device. Each enrollment speaker vector is associated with a different respective enrolled user of the assistant-enabled device. When the first speaker identification vector matches one of the enrollment speaker vectors, the operations also include identifying the first user who issued the first query as the respective enrolled user associated with one of the enrollment speaker vectors that matches the extracted speaker identification vector. In some implementations, performing speaker identification on the audio data corresponding to the second query to determine that the second query was spoken by the first user who issued the first query includes extracting a second speaker identification vector from the audio data corresponding to the second query that represents characteristics of the second query, and determining that the extracted second speaker identification vector matches a reference speaker vector of the first user.
[0013] The details of one or more embodiments of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will become apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]
[0014] [Figure 1A] FIG. 1 is a schematic diagram of an example system including multiple users controlling assistant-enabled devices. [Figure 1B]FIG. 1 is a schematic diagram of an example system including multiple users controlling assistant-enabled devices. [Figure 1C] FIG. 1 is a schematic diagram of an example system including multiple users controlling assistant-enabled devices. [Figure 2A] 1 is an exemplary GUI that may be rendered on a user device screen to display a round robin queue. [Figure 2B] 1 is an exemplary GUI that may be rendered on a user device screen to display a round robin queue. [Figure 2C] 1 is an exemplary GUI that may be rendered on a user device screen to display a round robin queue. [Figure 3] FIG. 1 is a schematic diagram of a query processing process. [Figure 4A] FIG. 1 is a schematic diagram of a speaker identification process. [Figure 4B] FIG. 1 is a schematic diagram of a speaker verification process. [Figure 5] 1 is a flowchart of an exemplary operational arrangement of a method for processing voice queries in a multi-user environment. [Figure 6] FIG. 1 is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION
[0015] Like reference symbols in the various drawings indicate like elements.
[0016] The way in which a user interacts with an assistant-enabled device is designed primarily, but not exclusively, through voice input. For example, a user may ask the device to perform an action, including media playback (e.g., music or podcasts), and the device responds by starting to play audio that matches the user's criteria. In situations where a device (e.g., a smart speaker) is widely shared by multiple users in an environment, the device may need to accommodate multiple actions requested by users that may conflict with each other. If one or more of the multiple users issues multiple individual requests to the device, the device may respond to the individual requests in an ordered manner while actively prompting non-requesting users to submit requests. By ensuring that each user of the environment has an opportunity to submit requests to the device, situations in which individual users monopolize the device are avoided.
[0017] Upon receiving requests from users, the device may implement one or more scheduling strategies to address the individual requests. For example, the device may implement a round-robin scheduling strategy in which each user's first request is addressed individually in order before subsequent requests from each user are addressed. Alternatively, the device may implement a scheduling strategy based on the length of the requested action, the importance of the requested action, or the priority of the requested action. For example, the device may prioritize requests submitted by the event host. Similarly, the device may implement a scheduling strategy that divides users into multiple queues, each queue corresponding to a type of action (e.g., play audio, control lights, etc.).
[0018] In addition to controlling playback actions to prevent a single user from dominating a playlist, the device may also control other types of media, such as podcasts and videos. Similarly, the device may prevent a single user from dominating the device with multiple related questions by prompting other users present to submit questions before answering multiple related questions. Additionally, this may extend to controlling aspects of the home connected to the device. For example, a party host may want to control lighting levels or the type of music played during the party to ensure a relaxing atmosphere. The host may say, "I only allow lights under 60% and jazz music." During the party, the device may prevent other party attendees from adjusting lighting levels or limit the extent to which they can be adjusted, and may prevent or limit bar attendees from requesting music outside of the jazz music genre.
[0019] In addition to restricting individuals, the device may operate to be more inclusive of individuals present in the home. For example, the device may assist individuals in the environment in creating a shopping list, thereby ensuring that all individuals are given the opportunity to add items to the shopping list. For example, the device may prompt individuals to add items to a shopping list in response to the first individual issuing a query to add multiple items to the shopping list. Similarly, the device may proactively prompt individuals to participate in an event, such as setting an alarm in the morning. Here, the device may suggest that the individual request an alarm in response to receiving a request to set an alarm for / from a different individual. Furthermore, the device may ensure that individuals who have not spoken recently are not disengaged from the interaction by engaging / prompting them to participate in an interaction between the device and other individuals who are near them.
[0020] 1A-1C illustrate exemplary systems 100a-100c for processing queries in an environment with multiple users 102, 102a-102n, using a scheduling strategy employing a round-robin mode that proactively engages multiple users 102 detected in the environment based on a round-robin queue 350. Briefly, as described in more detail below, an assistant-enabled device (AED) 104 including a query handler 300 (FIG. 3) detects multiple users 102, 102a-102c in the environment and begins playing music 122 in response to receiving a first query 106 issued by user 102a, "Okay, computer, play Glory, let's have pop music all night long." While the AED 104 is performing the action of playing music 122 as playback audio from speaker 18, the AED 104 receives a second query 146, "Play Cold Water next," spoken by the same user 102a (FIG. 1B). Because the query handler 300 detects / recognizes that other users 102b, 102c are present in the environment, the query handler 300 prompts one or more of the other users 102b, 102c to issue a query to add to the round robin queue 350 before taking any action associated with the second query 146 issued by user 102a.
[0021] The systems 100a-100c include an AED 104 running a digital assistant 105 with which multiple users 102 may interact by issuing queries containing commands to perform actions. In the illustrated example, the AED 104 corresponds to a smart speaker with which multiple users 102 may interact. However, the AED 104 may include other computing devices, such as, but not limited to, a smartphone, tablet, smart display, desktop / laptop, smartwatch, smart glasses / headset, smart appliance, headphones, or vehicle infotainment device. The AED 104 includes data processing hardware 10 and memory hardware 12 that stores instructions that, when executed on the data processing hardware 10, cause the data processing hardware 10 to perform operations. The AED 104 includes an array of one or more microphones 16 configured to capture sounds, such as speech, directed at the AED 104. The AED 104 may also include or communicate with an audio output device (e.g., a speaker) 18 that may output audio, such as music 122 and / or synthesized speech, from the digital assistant 105. Additionally, the AED 104 may include or be in communication with one or more cameras 19 configured to capture images within the environment and output image data 312 (FIG. 3).
[0022] In some configurations, the AED 104 communicates with multiple user devices 50, 50a-50n associated with multiple users 102. In the illustrated example, each user device 50 of the multiple user devices 50a-50c includes a smartphone with which the respective user 102 may interact. However, the user device 50 may include other computing devices, such as, but not limited to, a smartwatch, smart display, smart glasses, smartphone, smart glasses / headset, tablet, smart appliance, headphones, computing device, smart speaker, or other assistant-enabled device. Each user device 50 of the multiple user devices 50a-50n may include at least one microphone 52, 52a-52n present on the user device 50 that communicates with the AED 104. In these configurations, the user device 50 may also communicate with one or more microphones 16 present on the AED 104. Additionally, multiple users 102 may control and / or configure the AED 104 and may interact with the digital assistant 105 using an interface 200, such as a graphical user interface (GUI) 200, 200a-200n, rendered for display on the respective screens of each user device 50.
[0023] As shown in FIGS. 1A-1C and 3, a digital assistant 105 implementing a query handler 300 uses a round robin queue 350 to manage queries issued by multiple users 102. In some implementations, the query handler 300 maintains a log of users 102 detected in the environment and one or more queries received from each of the users 102 in the round robin queue 350 to ensure that the digital assistant 105 performs actions associated with the received queries in an ordered manner according to the round robin queue 350. In this sense, the query handler 300 maintains the order of queries received from multiple users 102 in the round robin queue 350 and prevents a single user 102 from monopolizing the digital assistant 105 with consecutive queries. Additionally, the query handler 300 can facilitate adding additional and / or conflicting queries to the round robin queue 350 from multiple users 102.
[0024] In some examples, the query handler 300 identifies metadata for issued queries that govern the order of actions recorded in the round robin queue 350. Here, the metadata for a query issued to "play" a song may include the length or tempo of the song, and the query handler 300 may execute consecutive short song performances but may sandwich other songs between two long songs to avoid executing consecutive long song performances. Similarly, the query 300 may group similar musical genres or tempos together within the constraints of the round robin queue 350 to provide smoother transitions between actions.
[0025] The GUI 200 of each user device 50 associated with a user 102 may display a round robin queue 350 including the identities of the detected users 102 and queries associated with each user 102. For example, the round robin queue 350 includes, for each user 102 of multiple users 102 detected in the environment, the identity of the corresponding user 102 and a query count of queries received from the corresponding user 102. The query count may refer to the number of songs submitted by the user 102, where each song corresponds to an entry in the round robin queue 350, or may refer to the number of actions submitted by the user 102, where each entry in the round robin queue 350 corresponds to an action, which may include multiple songs. In some configurations, the AED 104 includes a screen and renders the GUI 200 to display the round robin queue 350 on the screen. For example, the AED 104 may include a smart display, tablet, or smart TV in the environment. 2A-2C provide exemplary GUIs 200b1-200b3 displayed on the screen of a user device 50b associated with user 102b to keep user 102b informed of the status of round robin queue 350. Additionally, GUI 200 may render to display an identifier of the current action (e.g., "Playing Glory"), an identifier of the AED 104 (e.g., smart speaker) currently performing the action, and / or the identity of user 102a (e.g., Verb) who initiated the current action being performed by digital assistant 105. As described above, by enabling round robin mode, including query handler 300 managing round robin queue 350, when AED 104 determines (e.g., via query handler 300) that Verb 102a has already issued a first action recorded in round robin queue 350, the AED (e.g., via digital assistant 105) prevents (or at least requires approval from other users 102b, 102c) from performing a second action.
[0026] 1A-1C and 3, while the digital assistant 105 is running, the AED 104 detects multiple users 102a-102c in the environment using the user detector 310 of the query handler 300. For example, the query handler 300 receives proximity information 54 (FIG. 3) about the location of each of the multiple users 102a-102c relative to the AED 104 via the user detector 310. In some implementations, each user device 50a-50c of the multiple users 102 broadcasts proximity information 54 receivable by the user detector 310 that the AED 104 can use to determine the proximity of each user device 50 to the AED 104. The proximity information 54 from each user device 50 may include a wireless communication signal, such as WiFi, Bluetooth, or ultrasound, and the signal strength of the wireless communication signal received by the user detector 310 may correlate with the proximity (e.g., distance) of the user device 50 to the AED 104. In embodiments where a user 102 does not have a user device 50 or has a user device 50 that does not share proximity information 54, the user detector 310 may detect the user 102 based on explicit input (e.g., a guest list) 313 received from the user 102a who issued the first query 106. For example, the user detector 310 may receive a guest list 313 from a seed user 102 (e.g., user 102a) indicating each user 102 of the multiple users 102 and add them to the round-robin queue 350. In other embodiments, the user detector 310 automatically detects multiple users 102 of the environment by receiving image data 312 corresponding to a scene of the environment and acquired by the camera 19. Here, the user detector 310 detects the multiple users 102 based on the received image data 312. Similarly, the user detector 310 may detect multiple users 102 of the environment by performing speaker identification ( FIGS. 4A and 4B ) to resolve the identity of a user 102 in the environment who issues a query by speaking.Here, the user detector 310 detects multiple users 102 based on received audio data 402 associated with the issued queries, and the user detector 310 may continue to detect the user 102 who issued the query (and associated audio data 402) for a threshold time after the user 102 spoke.
[0027] In some embodiments, the user detector 310 maintains a list of previous users 314 present in the environment of the AED 104. Here, the list of previous users 314 may refer to a list of users 102 previously detected by the user detector 310, and the list may include one or more users not present in the most recent (i.e., most recent) detection of the user 102, and / or the most recent detection of the user 102 may include one or more additional users not present in the list of previous users 314. In this example, after receiving the proximity information 54, the image data 312, and / or the guest list, the user detector 310 may determine that the list of previous users 314 does not include the same user 102 associated with the current state of the environment (i.e., the list of current users 316). In other words, the user detector 310 may determine that the users 102 in the list of previous users 314 are different from the users 102 in the list of current users 316. This change between the list of the previous user 314 and the list of the current user 316 triggers the query handler 300 to update the round robin queue 350 to add or remove entries from the round robin queue 350 based on the list of the current user 316 and generate an update 352 for display in the GUI 300 of each user 50 associated with the detected user 102. For example, if user 102a leaves the environment, the user detector 310 may detect that user 102a has left by comparing the list of the current user 314, which does not include user 102a, with the list of the previous user 314, which does include user 102a. In response, the query handler 300 updates the round robin queue 350 to remove any queries issued by user 102a and generates an update to the round robin queue 350. However, in some examples, when the user detector 310 determines that the list of the previous user 314 is the same as the detected list of the current user 316, the user detector 310 may not send the list of the current user 316 to update the round robin queue 350.
[0028] In some examples, when there is a difference between the list of the previous user 314 and the list of the current user 316, the user detector 310 outputs only the list of the current user 316 (thereby triggering the query handler to generate an update 352 to the round robin queue 350). For example, the user detector 310 may be configured with a change threshold, and when the detected difference between the list of the previous user 314 and the list of the current user 316 meets (e.g., exceeds) the change threshold, the user detector 310 outputs the list of the current user 316 to the query handler 300. The threshold may be zero, and a small difference between the list of the previous user 314 and the list of the current user 316 detected by the indication determiner 210 (e.g., as soon as the user 102 enters or exits the environment) may trigger the query handler 300 to update the round robin queue 350. Conversely, the threshold may be higher than zero to prevent unnecessary updates to the round robin queue 350 as a type of queue interruption sensitivity mechanism. For example, the change threshold may be temporary (e.g., a time amount), such that if the user 102 temporarily leaves the environment (e.g., goes to another room) but returns within the threshold time, the AED 104 does not update the round robin queue 350. Here, the user detector 310 updates the round robin queue 350 by outputting only the list of the current user 316 when the difference between the list of the previous user 314 and the list of the current user 316 persists for the threshold time.
[0029] 1A-1C, the user detector 310 may identify the user 102a (e.g., Barb) and the user 102b (e.g., Jeff) by the proximity information 54 received from their respective user devices 50a, 50b. In this example, the user detector 310 may not detect the user 102c by detecting the user device 50c (e.g., the user 102c chooses not to share the proximity information 54 from the user device 50c). However, the user detector 310 still detects the presence of the user 102c (e.g., via speech data, image data, and / or input from other users 102a, 102b) and outputs a list of current users 316. The query handler 300 then adds the list of current users 316, including the detected users 102a-102c, to a round-robin queue 350 that records queries issued by the users 102a-102c.
[0030] Continuing with the example of FIG. 1A , user 102a of multiple users 102a-102c is shown issuing a first query 106, “Okay, computer, play Glory, let's have pop music all night,” in the vicinity of AED 104. Here, the first query 106 issued by user 102a is spoken by user 102a and includes initial audio data 402 ( FIG. 3 ) corresponding to the first query 106. The first query 106 may further include a user input indication indicating the user's intent to issue the first query via any one of touch, speech, gesture, gaze, and / or an input device (e.g., a mouse or stylus) to interact with the AED 104. Optionally, based on receiving the initial audio data 402 corresponding to the first query 106, the query handler 300 performs a speaker identification process 400a (FIG. 4A) on the audio data 402 to resolve the identity of the speaker of the first query 106 by determining that the first query 106 was issued by the user 102a. In other implementations, the user 102a issues the first query 106 without speaking. In these implementations, the user 102a issues the first query 106 via a user device 50a associated with the user 102a (e.g., by entering text corresponding to the first query 106 into a GUI 200a displayed on the screen of the user device 50a, selecting the first query 106 displayed on the screen of the user device 50a, etc.). Here, the AED 104 may resolve the identity of the user 102 who issued the first query 106 by recognizing the user device 50a associated with the user 102a.
[0031] The microphone 16 of the AED 104 receives the first query 106 and processes initial audio data 402 corresponding to the first query 106. The initial processing of the audio data 402 may include filtering the audio data 402 and converting the audio data 402 from an analog signal to a digital signal. Once the AED 104 processes the audio data 402, the AED may store the audio data 402 in a buffer in the memory hardware 12 for further processing. With the audio data 402 in the buffer, the AED 104 may use the hotword detector 108 to detect whether the audio data 402 includes a hotword. The hotword detector 108 is configured to identify hotwords included in the audio data 402 without performing speech recognition on the audio data 402.
[0032] In some implementations, the hot word detector 108 is configured to identify hot words in an initial portion of the first query 106. In this example, the hot word detector 108 may determine that the first query 106, "Okay, computer, play Glory, let's have pop music all night long," contains the hot word 110, "Okay, computer," if the hot word detector 108 detects acoustic features in the audio data 402 that are characteristic of the hot word 110. The acoustic features may be Mel-Frequency Cepstral Coefficients (MFCCs), which are a representation of the short-term power spectrum of the first query 106, or may be Mel-scale filter bank energies of the first utterance 106. For example, the hot word detector 108 may detect that the first query 106, "Okay, computer, play Glory, let's play pop music all night long," includes the hot word 110, "Okay, computer," based on generating MFCCs from the audio data 402 and classifying the MFCCs as including MFCCs similar to MFCCs characteristic of the hot word "Okay, computer" stored in a hot word model of the hot word detector 108. As another example, the hot word detector 108 may detect that the first query 106, "Okay, computer, play Glory, let's play pop music all night long," includes the hot word 110, "Okay, computer," based on generating MFCCs from the audio data 402 and classifying the MFCCs as including MFCCs similar to MFCCs characteristic of the hot word "Okay, computer" stored in a hot word model of the hot word detector 108.
[0033] When the hot word detector 108 determines that the initial audio data 402 corresponding to the first query 106 includes the hot word 110, the AED 104 may trigger a wake-up process to begin speech recognition on the audio data 402 corresponding to the first query 106. For example, FIG. 3 shows an AED 104 including a speech recognizer 170 employing an automatic speech recognition model 172 that may perform speech recognition or semantic interpretation on the audio data 402 corresponding to the first query 106. The speech recognizer 170 may perform speech recognition on the portion of the audio data 402 that follows the hot word 110. In this example, the speech recognizer 170 may identify the words "Play Glory, let's play pop music all night long" in the first query 106.
[0034] In some examples, the AED 104 is configured to communicate with a remote system 130 over the network 120. The remote system 130 may include remote resources, such as remote data processing hardware 132 (e.g., a remote server or CPU) and / or remote memory hardware 134 (e.g., a remote database or other storage hardware). The query handler 300 may execute on the remote system 130 in addition to or instead of the AED 104. The AED 104 may utilize the remote resources to perform various functions related to speech processing and / or synthesis and playback communication. In some implementations, the speech recognizer 170 is located on the remote system 130 in addition to or instead of the AED 104. When the hot word detector 108 triggers the AED 104 to wake up in response to detecting the hot word 110 in the first query 106, the AED 104 may transmit initial audio data 402 corresponding to the first query 106 to the remote system 130 over the network 120. Here, the AED 104 may transmit the portion of the initial audio data 402 that includes the hotword 110 for the remote system 130 to verify the presence of the hotword 110. Alternatively, the AED 104 may transmit only the portion of the initial audio data 402 that corresponds to the portion of the utterance 106 that follows the hotword 110 to the remote system 130, and the remote system 130 executes the speech recognizer 170 to perform speech recognition and returns a transcription of the initial audio data 402 to the AED 104.
[0035] 1A-1C and 3, the query handler 300 may further include a natural language understanding (NLU) module 320 that performs semantic interpretation on the first query 106 to identify queries / commands directed to the AED 104. Specifically, the NLU module 320 identifies words of the first query 106 identified by the speech recognizer 170 and performs semantic interpretation to identify any spoken commands in the first query 106. The NLU module 320 of the AED 104 (and / or remote system 130) identifies the words “play glory” as a command 111 to perform a first action (i.e., play music 122) and identifies the words “let's play pop music all tonight” as a constraint 113 that restricts the action commanded by queries following the first query 106. In the example shown in FIG. 1A, the digital assistant 105 begins performing a first action of playing music 122 as playback audio (e.g., track 1) from the speaker 18 of the AED 104. The digital assistant 105 may stream music 122 from a streaming service (not shown), or the digital assistant 105 may instruct the AED 104 to play music stored on the AED 104. Additionally, the query handler 300 enables a round-robin mode including a round-robin queue 350. When the round-robin mode is enabled, the digital assistant 105 controls the performance of actions commanded by queries subsequent to the first query 106 based on the round-robin queue 350. Here, the round-robin queue 350 includes a list of current users 316, including multiple users 102a-102c detected by the user detector 310 in the environment of the AED 104. In the example, in response to receiving the constraint 113 in the first query 106, the query handler 300 limits the actions of the digital assistant 105 to only playing music of a particular genre (i.e., pop music).
[0036] 1A, with reference to FIG. 3, the query handler 300 enables a round-robin mode including a round-robin queue 350 and adds constraints 113 to the active constraints data store 330. The query handler 300 maintains a record of active constraints 332 in the active constraints data store 330 (e.g., stored in memory hardware 12), and the query handler 300 may limit the actions added to the round-robin queue 350 for the digital assistant 105 to perform based on the active constraints 332. For example, before the query handler 300 performs the first action associated with the command 111 and adds the action "play glory" to the round-robin queue (FIG. 2A), it may first verify that the constraints 113 of the command 111 and / or the first query 106 do not conflict with any of the active constraints 332.
[0037] The AED 104 may inform the user 102a (e.g., Verb) who issued the first query 106 that it has enabled round-robin mode and will use the round-robin queue 350 to control subsequent queries (e.g., queues). For example, the digital assistant 105 may generate synthesized speech 123 for audible output from the speaker 18 of the AED 104 stating, "Verb, now enable round-robin mode to control queries." In a further example, the digital assistant 105 provides a notification (e.g., update 352) to the user device 50a associated with the user 102a (e.g., Verb) informing the user 102a of the entries in the round-robin queue 350 and / or any active constraints 232 stored in the active constraints data store 330.
[0038] 2A-2C, a graphical user interface (GUI) 200 executing on a user device 50 may display a round robin queue 350 including detected users 102 and queries associated with each detected user 102. As used herein, the GUI 200 may receive user input instructions via any one of touch, speech, gesture, gaze, and / or an input device (e.g., a mouse or stylus) for interacting with the round robin queue 350 (via the digital assistant 105). Each listed action may serve as a descriptor identifying a respective command 111 for the digital assistant 105 to perform the action. FIG. 2A provides an exemplary GUI 200b1 displayed on the screen of a user device 50b for informing a user 102b of the contents of the round robin queue 350. Specifically, the GUI 200b1 renders a round robin queue 350 including user 102a (Barb), user 102b (Jeff), and user 102c (Tina) in repeating order. Round robin queue 350 further includes a query placeholder for each of users 102a-102c. As shown, the first entry for user 102a (Barb) includes the corresponding action 111, "Play Glory," while subsequent entries for user 102a and other users 102b-102c (Jeff, Tina) are empty. In other words, round robin queue 350 includes, for each of the multiple detected users 102a-102c, a query count for queries that include the identity of the user 102 (e.g., Barb, Jeff, Tina) and the action received from the corresponding user 102. As shown, user 102a (Barb) has a single query count corresponding to the action, "Play Glory," while user 102b (Jeff) and user 102c (Tina) do not have any queries in round robin queue 350.
[0039] Additionally, GUI 200 may render to display an identifier of the current action (e.g., "Playing Glory"), an identifier of the AED 104 (e.g., smart speaker) currently performing the action, an indicator that round-robin mode is enabled, and / or the identity of user 102a (e.g., Verb) who issued first query 106. In an embodiment in which first query 106 includes a constraint 113, GUI 200 renders an identifier of constraint 113 (e.g., pop music) for display. Thus, user 102b may query user device 50b to review the current action of digital assistant 105 and any constraints 113 that restrict the query issued by user 102b.
[0040] 4A and 4B , in some implementations, the AED 104 (or a remote system 130 in communication with the AED 104) also includes an exemplary data store 430 that stores enrolled user data / information for each of multiple enrolled users 432a-432n of the AED 104. Here, each enrolled user 432 of the AED 104 may undertake a speech enrollment process to obtain a respective enrollment speaker vector 154 from audio samples of multiple enrollment phrases spoken by the enrolled user 432. For example, the speaker identification model 410 may generate one or more enrollment speaker vectors 154 from audio samples of enrollment phrases spoken by each enrolled user 432, which may be combined, e.g., averaged or otherwise accumulated, to form the respective enrollment speaker vector 154. One or more of the enrollment users 432 may perform a voice enrollment process using the AED 104, with the microphone 16 capturing audio samples of these users speaking enrollment utterances, from which the speaker identification model 410 generates respective enrollment speaker vectors 154. The model 410 may run on the AED 104, the remote system 130, or a combination thereof. Additionally, one or more of the enrollment users 432 may enroll with the AED 104 by providing authorization and authentication credentials to an existing user account on the AED 104. Here, the existing user account may store the enrollment speaker vectors 154 obtained from a previous voice enrollment process, with other devices also linked to the user account.
[0041] In some examples, the enrollment speaker vector 154 of an enrolled user 432 includes a text-dependent enrollment speaker vector. For example, the text-dependent enrollment speaker vector may be extracted from one or more audio samples of each enrolled user 432 speaking a predetermined term, such as a hot word 110 (e.g., "Okay, computer") used to invoke the AED 104 to wake up from a sleep state. In other examples, the enrollment speaker vector 154 of an enrolled user 432 is obtained from one or more audio samples of each enrolled user 102 speaking phrases of different lengths with different terms / words, and is text-independent. In these examples, the text-independent enrollment speaker vector may be obtained over time from audio samples obtained from speech interactions the user 102 has with the AED 104 or other devices linked to the same account.
[0042] 4A , a speaker identification process 400a identifies the user 102a (e.g., Verb) who spoke the first query 106 by first extracting a first speaker identification vector 411 representing characteristics of the first query 106 issued by the user 102a from initial audio data 402 corresponding to the first query 106. Here, the speaker identification process 400a may execute a speaker identification model 410 configured to receive audio data 402 corresponding to the second query 146 as input and generate the first speaker identification vector 411 as output. The speaker identification model 410 may be a neural network model trained under machine or human supervision to output the speaker identification vector 411. The speaker identification vector 411 output by the speaker identification model 410 may include an N-dimensional vector having values corresponding to speech features of the first query 106 associated with the user 102a. In some examples, the speaker identification vector 411 is a d-vector. In some examples, the first speaker identification vector 411 includes a set of speaker identification vectors each associated with a different user who is also authorized to control the AED 104. For example, aside from the user 102a who spoke the first query 106, the other authorized users may include other individuals who were present when the user 102a spoke the first query 106 to issue the command 111 to perform the first action, and / or individuals that the user 102a added / designated as authorized.
[0043] Once the first speaker identification vector 411 is output from the model 410, the speaker identification process 400a determines whether the extracted speaker identification vector 411 matches any of the enrollment speaker vectors 154 stored in the AED 104 (e.g., in the memory hardware 12) for the enrolled users 432a-432n of the AED 104. As described above, the speaker identification model 410 may generate the enrollment speaker vectors 154 for the enrolled users 200 during the voice enrollment process. Each enrollment speaker vector 154 may be used as a reference vector 155 corresponding to a voiceprint or unique identifier representative of the voice characteristics of the respective enrolled user 432.
[0044] In some implementations, the speaker identification process 400a uses a comparator 420 that compares the first speaker identification vector 411 with each enrollment speaker vector 154 associated with each enrolled user 432a-432n of the AED 104. Here, the comparator 420 may generate a score for each comparison indicating the likelihood that the initial audio data 402 corresponding to the first query 106 corresponds to the identity of the respective enrolled user 432, and when the score meets a threshold, the identity is accepted. When the score does not meet the threshold, the comparator 420 may reject the identity of the speaker who issued the first query 106. In some implementations, the comparator 420 calculates a respective cosine distance between the first speaker identification vector 411 and each enrollment speaker vector 154 and determines that the first speaker identification vector 411 matches one of the enrollment speaker vectors 154 when the respective cosine distance meets a cosine distance threshold.
[0045] In some examples, the first speaker identification vector 411 is a text-dependent speaker identification vector extracted from a portion of one or more words corresponding to the first query 106, and each enrollment speaker vector 154 is also text-dependent on the same one or more words. Using a text-dependent speaker vector can improve accuracy in determining whether the first speaker identification vector 411 matches any of the enrollment speaker vectors 154. In other examples, the first speaker identification vector 411 is a text-independent speaker identification vector extracted from the entire initial audio data 402 corresponding to the first query 106.
[0046] When the speaker identification process 400a determines that the first speaker identification vector 411 matches one of the enrollment speaker vectors 154, the process 400a identifies the user 102a who spoke the first query 106 as the respective enrolled user 432a associated with one of the enrollment speaker vectors 154 that matches the extracted speaker identification vector 411. In the illustrated example, the comparator 420 determines a match based on the respective cosine distances between the first speaker identification vector 411 and the enrollment speaker vector 154 associated with the enrollment user 432a satisfying a cosine distance threshold. In some scenarios, the comparator 420 identifies the user 102a as the respective enrolled user 432a associated with the enrollment speaker vector 154 having the shortest respective cosine distance from the first speaker identification vector 411 if the shortest respective cosine distance also meets the cosine distance threshold.
[0047] Conversely, when the speaker identification process 400a determines that the first speaker identification vector 411 does not match any of the enrollment speaker vectors 154, the process 400a may identify the user 102a who spoke the utterance 106 as a guest user of the AED 104. Accordingly, the query handler 300 may add the guest user to the round robin queue 350 and use the first speaker identification vector 411 as a reference speaker vector 155 that represents the speech characteristics of the guest user's voice. In some examples, the guest user may register with the AED 104, and the AED 104 may store the first speaker identification vector 411 as the enrollment speaker vector 154 for each of the newly enrolled users.
[0048] Referring again to FIG. 1B, while the digital assistant 105 plays music 122 as playback audio from the speaker 18 of the AED 104, when round-robin mode is enabled, the AED 104 receives a second query 146 including a command 118 for the digital assistant 105 to perform a second action. In the illustrated example, the user 102a who issued the first query 106 also issues the second query 146 "Play Cold Water Next," including a command 118 for the digital assistant 105 to play a song (i.e., Cold Water) immediately after playing track #1 (i.e., Glory). Based on receiving the second query 146, the query handler 300 resolves the identity of the speaker of the second query 146 by performing a speaker identification process 400b (FIG. 4B) on the audio data 402 corresponding to the second query 146 and determines that the second query 146 was issued by the user 102a who issued the first query 106. 4A , in an implementation in which the first query 106 issued by the user 102a includes initial audio data 402 (e.g., the first query 106 was spoken by the first user 102a), the query handler 300 may first perform a speaker identification process 400a on the initial audio data 402 corresponding to the first query 106 to identify the user 102a who issued the first query 106. The speaker identification process 400b may execute on the data processing hardware 12 of the AED 104. The speaker identification process 400b may also execute on the remote system 130. If the speaker verification process 400b on the audio data 402 corresponding to the second query 146 indicates that the second query 146 was spoken by a different user 102 than the user 102 who issued the first query 106, the AED 104 may proceed with adding an action associated with the second query 146 to the round robin queue 350.Conversely, if the speaker verification process 400b for the audio data 402 corresponding to the second query 146 indicates that the second query 146 was spoken by the same user 102 that issued the first query 106, the AED 104 may prevent the second action from being performed (or may at least require input from one or more other users 102 in the environment (e.g., of FIG. 1B )).
[0049] Referring again to FIG. 4B with reference to the example of FIG. 1B, in response to receiving the second query 146, the AED 104 resolves the identity of the user 102 who spoke the second query 146 by performing a speaker identification process 400b. The speaker identification process 400b identifies the user 102a who spoke the first query 146 by first extracting a second speaker identification vector 412 representing characteristics of the second query 146 from audio data 402 corresponding to the first query 146 spoken by the user 102a. Here, the speaker verification process 400b may execute a speaker identification model 410 configured to receive the audio data 402 as input and generate the second speaker identification vector 412 as output. As illustrated in FIG. 4A, the speaker identification model 410 may be a neural network model trained to output the speaker identification vector 412 under machine or human supervision. The second speaker identification vector 412 output by the speaker identification model 410 may include an N-dimensional vector having values corresponding to speech features of the utterance 146 associated with the user 102a. In some examples, the speaker identification vector 412 is a d-vector.
[0050] Once the second speaker identification vector 412 is output from the speaker identification model 410, the speaker verification process 400b determines whether the extracted speaker identification vector 412 matches a reference speaker vector 155 associated with the first enrolled user 432a stored in the AED 104 (e.g., in the memory hardware 12). The reference speaker vector 155 associated with the first enrolled user 432a may include each enrollment speaker vector 154 associated with the first enrolled user 432a. As described above, the speaker identification model 410 may generate the enrollment speaker vector 154 for the enrolled user 432 during a voice enrollment process. Each enrollment speaker vector 154 may be used as a reference vector corresponding to a voiceprint or unique identifier representing the voice characteristics of the respective enrolled user 432.
[0051] In some implementations, the speaker verification process 400b uses a comparator 420 to compare the second speaker identification vector 412 with a reference speaker vector 155 associated with a first enrolled user 432a of the enrolled users 432. The comparator 420 may generate a score for the comparison indicating the likelihood that the second query 146 corresponds to the identity of the first enrolled user 432a, and when the score meets a threshold, the identity is accepted. When the score does not meet the threshold, the comparator 420 may reject the identity. In some implementations, the comparator 420 calculates the respective cosine distances between the second speaker identification vector 412 and the reference speaker vector 155 associated with the first enrolled user 432a, and determines that the second speaker identification vector matches the reference speaker vector 155 when the respective cosine distances meet a cosine distance threshold.
[0052] When the speaker verification process 400b determines that the second speaker identification vector 412 matches the reference speaker vector 155 associated with the first enrolled user 432a, the process 400b identifies the user 102a who spoke the second query 146 as the first enrolled user 432a associated with the reference speaker vector 155. In the illustrated example, the comparator 420 determines a match based on the second speaker identification vector 412 and the reference speaker vector 155 associated with the first enrolled user 432a satisfying a cosine distance threshold. In some scenarios, the comparator 420 identifies the user 102a as the respective first enrolled user 432a associated with the reference speaker vector 155 having the shortest respective cosine distance from the second speaker identification vector 412 if this shortest respective cosine distance also satisfies the cosine distance threshold.
[0053] 4A above, in some implementations, the speaker identification process 400a determines that the first speaker identification vector 411 does not match any of the enrollment speaker vectors 154 and identifies the user 102 who spoke the first query 106 as a guest user of the AED 104. Accordingly, the speaker verification process 400b may first determine whether the user 102a who spoke the first query 106 was identified by the speaker identification process 400a as an enrollment user 432 or as a guest user. When the user 102a is a guest user, the comparator 420 compares the second speaker identification vector 412 with the first speaker identification vector 411 obtained during the speaker identification process 400a. Here, the first speaker identification vector 411 represents characteristics of the first query 106 spoken by the guest user 102a, and is therefore used as a reference vector for verifying whether the second query 146 was also spoken by the guest user 102a or by another user 102. Here, the comparator 420 may generate a comparison score indicating the likelihood that the second query 146 corresponds to the identity of the guest user 102a, and if the score meets a threshold, the identity is accepted. If the score does not meet the threshold, the comparator 420 may reject the identity of the guest user who spoke the second query 146. In some implementations, the comparator 420 calculates the respective cosine distances between the first speaker identification vector 411 and the second speaker vector 412, and determines that the first speaker identification vector 411 matches the second speaker vector 412 when the respective cosine distances meet a cosine distance threshold.
[0054] 1B , based on determining that the second query 146 was spoken by the user 102a who issued the first query 106, the query handler 300 prevents the AED 104 (via the digital assistant 105) from taking the second action and instead prompts at least another user 102 of the multiple users 102a-102c detected in the environment who is different from the user 102a to issue a query. In other words, after the user 102a is determined as the issuer of the first query 106 and the speaker of the second query 146, the query handler 300 prevents the user 102a from occupying the AED 104 by first verifying that the other users 102 in the environment do not wish to issue a query before completing the second query 146. If other users 102 of the environment confirm that they do not want to issue a query, the digital assistant 105 may continue to complete the second query 146 issued by the user 102a when the completion of the first query 106 is completed.
[0055] In some embodiments, the query handler 300 (via the digital assistant 105) prompts at least another user 102 of the plurality of users 102 detected in the environment to provide a third query 148 to perform a third action including a corresponding command 119 before performing the second action included in the second query 146 issued by the user 102a. In these embodiments, prompting the at least another user 102 of the plurality of users 102 includes providing, as output from the AED 104, a user-selectable option that, when selected, issues the third query 148 to perform the third action including the corresponding command 119. For example, the digital assistant 105 may generate synthesized speech 123 for audible output from the speaker 18 of the AED 104 (or a speaker in communication with the data processing hardware (e.g., the speaker of the user device 50)), which prompts the user 102b for the opportunity to issue the query, "Jeff, would you like to select a song next?" In response, user 102b (i.e., Jeff), one of multiple users 102b, 102c different from user 102a, is shown issuing a third query 148, “Yes, play Red next,” near AED 104. In a further example, digital assistant 105 may provide a notification to user device 50b associated with user 102b (e.g., Jeff), displaying user-selectable options as a graphical element 210 on the screen of user device 50b, where the graphical element 210 prompts user 102b with the opportunity to issue third query 148. Additionally or alternatively to audibly prompting user 102b, as shown in FIG. 2B , GUI 200b2 may enable user 102b to issue (or decline the opportunity to issue) third query 148 by rendering graphical element 210, “Jeff, would you like to select a song to play next?”, along with “Yes” and “No” to display. Here, the query handler 300 receives the third query 148 when the user device 50b receives a user input indication indicating a selection of a user-selectable option that selects one of the graphic elements 210.
[0056] 1B and 3, in response to receiving the third query 148, the NLU module 320 executing on the AED 104 (and / or executing on the remote system 130) may identify the word "play red" as a command 119 and perform a third action (i.e., play music 122). The query handler 300 may first determine whether the command 119 for performing the third action does not violate any active constraints 232 of the digital assistant 105, and then update the round robin queue 350 to include the third action. As described above, the active constraints 332 may include the constraint 113 included in the first query 106, where the constraint 113 restricts the music genre (e.g., pop music). Here, before adding "red" to the round robin queue 350, the query handler 300 verifies that the action of playing "red" in the third query 148 issued by the user 102b does not contradict the constraint 113 that restricts the music genre to pop music. In some implementations, before updating the round robin queue 350 to include the command 118 in the second query 146 issued by the user 102a, the query handler 300 further verifies that the action of playing "Cold Water" included in the second query 146 does not conflict with the constraint 113 limiting the genre of music to pop music and / or other active constraints 332 on the digital assistant 105.
[0057] While these examples primarily refer to actions such as playing music to avoid dominating a playlist for a single user 102, actions may refer to any category of actions, including, but not limited to, search queries, controlling assistant-enabled devices (e.g., smart lights, smart thermostats), and playing other types of media (e.g., podcasts, videos, etc.). For example, the query handler 300 may assist users 102 of an environment in creating a shopping list, thereby ensuring that all users 102 are given the opportunity to add items to the shopping list. For example, the AED 104 may add multiple items requested by a first user 102 to the shopping list, while prompting a second user 102 to add items in response to the first user 102 issuing a query to add multiple items to the shopping list. Similarly, the query handler 300 may proactively prompt users 102 to be included in an action, such as setting a morning alarm. Here, the AED 104 may suggest that the second user 102 request an alarm in response to receiving a request to set an alarm from the first user 102. Additionally, the query handler 300 can ensure that the AED 104 does not leave users 102 out of the interaction by engaging / prompting users 102 who have not recently issued a query to participate in an interaction between the AED 104 and other users 102.
[0058] Additionally, the constraints 332 may include some limits on the actions themselves. For example, a time limit on actions may be placed that limits the amount of time a user 102 has to make a turn in each entry of the round robin queue 350 (e.g., the number of jokes a user 102 can request each turn in the round robin queue). A time limit for round robin mode may control how long an event applying the constraint 332 lasts. Similarly, a threshold number of actions per query constraint 332 may limit the number of actions a user 102 may request per submitted query entry in the round robin queue 350 (e.g., a user 102 may request an entire album and / or playlist per turn in the round robin queue 350). Similarly, the threshold number of actions may include the total number of actions per user in the round robin queue 350. For example, a user 102 may only submit up to 50 additional songs to the round robin queue 350.
[0059] Continuing with the examples of FIGS. 1B and 2C, after query handler 300 verifies that the action to play "Red" in third query 148 issued by user 102b is consistent with constraint 113 limiting the music genre to pop music, query handler 300 updates round robin queue 350 to add "Red" to the identity associated with user 102b in round robin queue 350. As shown in FIG. 2C, GUI 200b3 renders round robin queue 350 with user 102a (Barb), user 102b (Jeff), and user 102c (Tina) in repeating order. Jeff's action placeholder has been updated to include command 119 for digital assistant 105 to perform the action "Play Red" once AED 105 completes (e.g., plays) the action "Play Glory" of command 111 in first query 106. In other words, based on the round robin queue 350, when the digital assistant 105 completes the execution of the first action 111 of the first query 106 issued by the user 102a, it executes the execution of the third action 119 of the third query 148 issued by the user 102b. Additionally, the query handler 300 updates the round robin queue 350 to include the second query 146 issued by the user 102a. As shown, the action placeholder of the second entry of the verb is updated to include a command 118 for the digital assistant 105 to perform the action "play red" of the command 119 of the third query 148 and the execution (e.g., playing) of any additional intervening queries submitted by users 102a, 102b, or different users of the environment 102, before the AED 105 completes the execution of the action associated with the third query 148.
[0060] 1C, the AED 104 may notify the user 102a (e.g., Barb) who issued the first query 106 that the second query 146 will be fulfilled after the fulfillment of the third query 148 based on the round robin queue 350. For example, the digital assistant 105 may generate synthesized speech 123 for audible output from the speaker 18 of the AED 104 stating, "After Jeff's selection, Barb, play Cold Water." In an additional example, the digital assistant 105 provides a notification (e.g., update 352) to the user device 50a associated with the user 102a (e.g., Barb) informing the user 102a of the updated entries in the round robin queue 350.
[0061] 5 is a flowchart of an exemplary operational arrangement of a method 500 for processing queries in an environment with multiple users 102. At operation 502, the method 500 includes detecting multiple users 102 in the environment of an assistant-enabled device (AED) 104. The method 500 also includes receiving a first query 106 issued by a first user 102a of the multiple users 102 at operation 504. The first query 106 includes a command 111 for the digital assistant 105 to perform a first action. At operation 506, the method 500 further includes enabling a round-robin mode, whereby the digital assistant 105 controls the performance of an action commanded by a query subsequent to the first query 106 based on a round-robin queue 350. Here, the round-robin queue 350 includes multiple users 102 detected in the environment of the AED 104.
[0062] While the digital assistant 105 is performing the first action and when round-robin mode is enabled, the method 500 further includes, at operation 508, receiving audio data 402 corresponding to a second query 146 spoken by one of the multiple users 102 and captured by the AED 104. Here, the second query 146 includes a command 118 for the digital assistant 105 to perform the second action. At operation 510, the method 500 also includes performing speaker identification on the audio data 402 corresponding to the second query 146 to determine that the second query 146 was spoken by the first user 102a who issued the first query 106.
[0063] Based on determining that the second query 146 was spoken by the first user 102a who issued the first query 106, the method 500 further includes, at operation 512, preventing the digital assistant 105 from performing the second action and prompting at least another user 102 of the plurality of users 102 detected in the environment, different from the first user 102a, to issue a query. At operation 514, the method 500 includes receiving a third query 148 issued by a second user 102b of the plurality of users 102 detected in the environment, the third query 148 including a command 119 for the digital assistant 105 to perform a third action. When the digital assistant 105 completes the performance of the first action, the method 500 further includes, at operation 502, performing the third action.
[0064] 6 is a schematic diagram of an exemplary computing device 600 that may be used to implement the systems and methods described herein. Computing device 600 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown here, their connections and relationships, and their functions are for illustrative purposes only and are not intended to limit the scope of the invention(s) described and / or claimed herein.
[0065] Computing device 600 includes a processor 610, memory 620, a storage device 630, a high-speed interface / controller 640 that connects to memory 620 and a high-speed expansion port 650, and a low-speed interface / controller 660 that connects to a low-speed bus 670 and storage device 630. Each of the components 610, 620, 630, 640, 650, and 660 is interconnected using various buses and may be mounted on a common motherboard or implemented in other ways as desired. Processor 610 (e.g., data processing hardware 10, 132 of FIG. 1 ) processes instructions for execution within computing device 600, including instructions stored in memory 620 or storage device 630, and can display graphical information for a graphical user interface (GUI) on an external input / output device, such as a display 680 connected to high-speed interface 640. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and multiple types of memory, as desired. Also, multiple computing devices 600 may be connected (eg, as a server bank, a group of blade servers, or a multi-processor system) with each device performing some of the required operations.
[0066] The memory 620 stores information non-temporarily within the computing device 600. The memory 620 (e.g., memory hardware 12, 134 in FIG. 1) may be a computer-readable medium, a volatile memory unit(s), or a non-volatile memory unit(s). The non-transitory memory 620 may be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by the computing device 600. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase-change memory (PCM), and disk or tape.
[0067] Storage device 630 can provide mass storage for computing device 600. In some embodiments, storage device 630 is a computer-readable medium. In various different implementations, storage device 630 may be a device array including a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or a device in a storage area network or other configuration. In additional embodiments, a computer program product is tangibly embodied on an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as memory 620, storage device 630, or the memory of processor 610.
[0068] High-speed controller 640 manages bandwidth-intensive operations for computing device 600, while low-speed controller 660 manages low-bandwidth-intensive operations. This role assignment is merely exemplary. In some implementations, high-speed controller 640 is connected to memory 620, display 680 (e.g., via a graphics processor or accelerator), and high-speed expansion port 650, which can accept various expansion cards (not shown). In some implementations, low-speed controller 660 is connected to storage device 630 and low-speed expansion port 690. Low-speed expansion port 690, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), may connect to one or more input / output devices, such as a keyboard, pointing device, scanner, etc., or to a networking device, such as a switch or router, e.g., via a network adapter.
[0069] The computing device 600, as shown, can be implemented in many different forms. For example, it can be implemented as a standard server 600a, or multiple times in a group of such servers 600a, as a laptop computer 600b, or as part of a rack server system 600c.
[0070] Various implementations of the systems and techniques described herein may be realized in digital electronic and / or optical circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementation in one or more computer programs executable and / or interpretable by a programmable system including at least one programmable processor. Such programmable processor may be special-purpose or general-purpose and may be coupled to receive and transmit data and instructions from a storage system, at least one input device, and at least one output device.
[0071] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform a task. In some examples, a software application may be referred to as an "application," "app," or "program." Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0072] Non-transitory memory may be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by a computing device. Non-transitory memory may be volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.
[0073] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language, and / or assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0074] The processes and logic flows described herein may be implemented by one or more programmable processors, also referred to as data processing hardware, executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be implemented by special-purpose logic circuitry, such as an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit). Processors suitable for executing computer programs include, for example, both general-purpose and special-purpose processors, as well as any one or more processors of any kind of digital computer. Generally, a processor receives instructions and data from a read-only memory, a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or is operably connected to receive data from or transmit data to them, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0075] To interact with a user, one or more aspects of the present disclosure can be implemented in a computer that has a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, or a touch screen for displaying information to the user, and optionally a keyboard and pointing device, such as a mouse or trackball, through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user. For example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user can be received in any form, such as acoustic input, speech input, or tactile input. Furthermore, a computer can interact with a user by sending documents to and receiving documents from a device used by the user, for example, by sending a web page to a web browser on the user's client device in response to a request received from the web browser.
[0076] Although several embodiments have been described, it will be understood that various modifications can be made without departing from the spirit and scope of the present disclosure. Accordingly, other embodiments are within the scope of the following claims.
Claims
1. A computer-implemented method (500) that, when executed by data processing hardware (610), causes the data processing hardware (610) to perform an operation, the operation comprising: Detecting a plurality of users within an environment of the assistant-enabled device (104); and receiving a first query (106) issued by a first user of the plurality of users, the first query (106) including a command (111) for a digital assistant (105) to perform a first action, the operation further comprising: and enabling a round robin mode, wherein when the round robin mode is enabled, the digital assistant controls the performance of actions commanded by queries subsequent to the first query based on a round robin queue, the round robin queue including the plurality of users detected within the environment of the assistant-enabled device, and the operation further includes: While the digital assistant (105) is performing the first action and when the round robin mode is enabled, receiving audio data (402) corresponding to a second query (146) spoken by one of the plurality of users and captured by the assistant-enabled device (104), the second query (146) including a command (111) for the digital assistant (105) to perform a second action, the operation further comprising: performing speaker identification on the audio data (402) corresponding to the second query (146) to determine that the second query (146) was spoken by the first user who issued the first query (106), the operations further comprising: based on determining that the second query (146) was spoken by the first user who issued the first query (106); Preventing the digital assistant (105) from performing the second action; and prompting at least another user of the plurality of users detected in the environment that is different from the first user for an opportunity to issue a query, the operations further comprising: receiving a third query (148) issued by a second user of the plurality of users detected in the environment, the third query (148) including a command (111) for the digital assistant (105) to perform a third action, the operation further comprising: When the digital assistant (105) completes performance of the first action, the computer-implemented method includes performing the third action.
2. Prompting at least the other user of the plurality of users detected in the environment includes providing, as output from a user interface (200) of a user device (50) associated with the second user, a user-selectable option that, when selected to perform the third action, issues the third query (148) from the second user; 10. The method of claim 1, wherein receiving the third query issued by the second user is based on receiving a user input indication indicating a selection of a user-selectable option.
3. 3. The method (500) of claim 2, wherein providing the user-selectable options as output from the user interface (200) includes displaying the user-selectable options as a graphical element (210) on a screen of the user device (50) associated with the second user via the user interface (200), the graphical element (210) prompting the second user for the opportunity to issue the third query (148).
4. 3. The method of claim 2, wherein providing the user-selectable options as output from the user interface includes providing the user-selectable options as an audible output from a speaker in communication with the data processing hardware via the user interface, the audible output prompting the second user of the opportunity to issue the third query.
5. The first query (106) further includes a constraint on a subsequent query, the constraint comprising: Action category, time limits for actions, a time limit for said round robin mode, or The method (500) of any one of claims 1 to 4, comprising one of a threshold number of actions per query.
6. The operation is determining that the second query (146) does not violate the constraints of the first query (106); and updating the round robin queue to include the second query spoken by the first user.
7. The operation is Detecting that the first user has left the environment of the assistant-enabled device (104); 7. The method of claim 6, further comprising: updating the round robin queue to remove the second query spoken by the first user.
8. 8. The method of claim 1, wherein detecting the plurality of users in the environment of the assistant-enabled device includes detecting at least one of the plurality of users based on proximity information of a user device associated with the at least one of the plurality of users.
9. Detecting the plurality of users within the environment of the assistant-enabled device (104) includes: receiving image data (312) corresponding to a scene of the environment; Detecting at least one of the plurality of users based on the image data.
10. 10. The method of claim 1, wherein detecting the plurality of users in the environment of the assistant-enabled device includes receiving a list indicating each user of the plurality of users to add to the round-robin queue.
11. The round robin queue (350) for each corresponding one of the plurality of users detected in the environment: the identity of the corresponding user; A method (500) according to any one of claims 1 to 10, comprising: a query count of queries received from said corresponding users.
12. 12. The method (500) of claim 1, wherein receiving the first query (106) issued by the first user comprises receiving, from a user device (50) associated with the first user, a user input indication indicating a user intent to issue the first query (106).
13. The method (500) of any one of claims 1 to 12, wherein receiving the first query (106) issued by the first user includes receiving initial audio data corresponding to the first query (106) issued by the first user and captured by the assistant-enabled device (104).
14. The operations further include, after receiving the initial audio data (402) corresponding to the first query (106) issued by the first user, performing speaker identification on the initial audio data (402) to identify the first user who issued the first query (106), wherein performing speaker identification includes: extracting a first speaker identification vector (411) representing characteristics of the first query (106) issued by the first user from the initial audio data (402) corresponding to the first query (106) issued by the first user; and determining that the extracted speaker identification vector matches any enrollment speaker vector (154) stored in the assistant-enabled device (104), each enrollment speaker vector (154) being associated with a different respective enrolled user of the assistant-enabled device (104), and the speaker identification is further performed by 14. The method (500) of claim 13, wherein, based on determining that the first speaker identification vector (411) matches one of the enrollment speaker vectors (154), the method (500) is performed by identifying the first user who issued the first query (106) as the respective enrolled user associated with one of the enrollment speaker vectors (154) that matches the extracted speaker identification vector.
15. performing speaker identification on the audio data (402) corresponding to the second query (146) to determine that the second query (146) was spoken by the first user who issued the first query (106); extracting a second speaker identification vector (412) from the audio data (402) corresponding to the second query (146), the second speaker identification vector (412) representing characteristics of the second query (146); and determining that the extracted second speaker identification vector matches a reference speaker vector (155) of the first user.
16. data processing hardware (610); and memory hardware (620) in communication with the data processing hardware (610), the memory hardware (620) storing instructions that, when executed by the data processing hardware (610), cause the data processing hardware (610) to perform operations, the operations including: Detecting a plurality of users within an environment of the assistant-enabled device (104); and receiving a first query (106) issued by a first user of the plurality of users, the first query (106) including a command (111) for a digital assistant (105) to perform a first action, the operation further comprising: and enabling a round robin mode, wherein when the round robin mode is enabled, the digital assistant controls the performance of actions commanded by queries subsequent to the first query based on a round robin queue, the round robin queue including the plurality of users detected within the environment of the assistant-enabled device, and the operation further includes: While the digital assistant (105) is performing the first action and when the round robin mode is enabled, receiving audio data (402) corresponding to a second query (146) spoken by one of the plurality of users and captured by the assistant-enabled device (104), the second query (146) including a command (111) for the digital assistant (105) to perform a second action, the operation further comprising: performing speaker identification on the audio data (402) corresponding to the second query (146) to determine that the second query (146) was spoken by the first user who issued the first query (106), the operations further comprising: based on determining that the second query (146) was spoken by the first user who issued the first query (106); Preventing the digital assistant (105) from performing the second action; and prompting at least another user of the plurality of users detected in the environment that is different from the first user for an opportunity to issue a query, the operations further comprising: receiving a third query (148) issued by a second user of the plurality of users detected in the environment, the third query (148) including a command (111) for the digital assistant (105) to perform a third action, the operation further comprising: The system (100) includes performing the third action when the digital assistant (105) completes the performance of the first action.
17. Prompting at least the other user of the plurality of users detected in the environment includes providing, as output from a user interface (200) of a user device (50) associated with the second user, a user-selectable option that, when selected to perform the third action, issues the third query (148) from the second user; 17. The system of claim 16, wherein receiving the third query issued by the second user is based on receiving a user input indication indicating a selection of an option selectable by the user.
18. 18. The system of claim 17, wherein providing the user-selectable options as output from the user interface includes displaying the user-selectable options as a graphical element on a screen of the user device associated with the second user via the user interface, the graphical element prompting the second user for the opportunity to issue the third query.
19. 18. The system of claim 17, wherein providing the user-selectable options as output from the user interface includes providing the user-selectable options as an audible output from a speaker in communication with the data processing hardware via the user interface, the audible output prompting the second user of the opportunity to issue the third query.
20. The first query (106) further includes a constraint on a subsequent query, the constraint comprising: Action category, time limits for actions, a time limit for said round robin mode, or The system (100) of any one of claims 16 to 19, comprising one of a threshold number of actions per query.
21. The operation is determining that the second query (146) does not violate the constraints of the first query (106); and updating the round robin queue to include the second query spoken by the first user.
22. The operation is Detecting that the first user has left the environment of the assistant-enabled device (104); and updating the round robin queue to remove the second query spoken by the first user.
23. Detecting the plurality of users in the environment of the assistant-enabled device includes detecting at least one of the plurality of users based on proximity information of a user device associated with the at least one of the plurality of users.
24. Detecting the plurality of users within the environment of the assistant-enabled device (104) includes: receiving image data (312) corresponding to a scene of the environment; Detecting at least one of the plurality of users based on the image data.
25. Detecting the plurality of users in the environment of the assistant-enabled device includes receiving a list indicating each user of the plurality of users to add to the round-robin queue.
26. The round robin queue (350) for each corresponding one of the plurality of users detected in the environment: the identity of the corresponding user; A system (100) according to any one of claims 16 to 25, comprising: a query count of queries received from said corresponding user.
27. 27. The system (100) of claim 16, wherein receiving the first query (106) issued by the first user comprises receiving, from a user device (50) associated with the first user, a user input indication indicating a user intent to issue the first query (106).
28. Receiving the first query (106) issued by the first user includes receiving initial audio data corresponding to the first query (106) issued by the first user and captured by the assistant-enabled device (104). The system (100) of any one of claims 16 to 27.
29. The operations further include, after receiving the initial audio data (402) corresponding to the first query (106) issued by the first user, performing speaker identification on the initial audio data (402) to identify the first user who issued the first query (106), wherein performing speaker identification includes: extracting a first speaker identification vector (411) representing characteristics of the first query (106) issued by the first user from the initial audio data (402) corresponding to the first query (106) issued by the first user; and determining that the extracted speaker identification vector matches any enrollment speaker vector (154) stored in the assistant-enabled device (104), each enrollment speaker vector (154) being associated with a different respective enrolled user of the assistant-enabled device (104), and the speaker identification is further performed by The system (100) of claim 28 is performed by identifying the first user who issued the first query (106) as the respective enrolled user associated with one of the enrollment speaker vectors (154) that matches the extracted speaker identification vector, based on determining that the first speaker identification vector (411) matches one of the enrollment speaker vectors (154).
30. performing speaker identification on the audio data (402) corresponding to the second query (146) to determine that the second query (146) was spoken by the first user who issued the first query (106); extracting a second speaker identification vector (412) from the audio data (402) corresponding to the second query (146), the second speaker identification vector (412) representing characteristics of the second query (146); and determining that the extracted second speaker identification vector matches a reference speaker vector of the first user.
Citation Information
Patent Citations
Operation support device, operation support system, and operation support method
JP2019207640A
Voice recognition device, terminal, voice recognition method, and voice recognition program
JP2020064267A
Information processor, information processing method, program and information processing system
JP2021015202A
Resolving conflicting commands received by an electronic device
US20200177410A1