Handling conflicting queries on shared devices
The method addresses conflicting queries on shared devices by identifying speaker identities and generating compromise actions, effectively managing query conflicts and conserving resources.
Patent Information
- Application Number
- JP2025519841
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-06
- Filing Date
- 2023-10-03
- Publication Date
- 2025-11-05
AI Technical Summary
Existing technologies fail to effectively handle conflicting queries on a shared device, particularly in environments where multiple users interact with voice-enabled devices.
A computer-implemented method that determines a digital assistant to receive conflicting queries from multiple users, identify speaker identities, and generate compromise actions using query embeddings and proximity information to manage conflicts.
The method effectively manages conflicting queries by identifying speaker identities and generating compromise actions, reducing the frequency of query overrides and conserving computational resources.
Smart Images

Figure 2025536238000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to handling conflicting queries on a shared device. [Background technology]
[0002] The way in which a user interacts with an assistant-enabled device is designed to be primarily, but not exclusively, through voice input. For example, a user may ask the device to perform an action, including media playback (e.g., music or podcasts), and the device responds by starting to play audio that matches the user's criteria. In an example where a device (e.g., a smart speaker) is commonly shared by multiple users in an environment, the device may need to process multiple actions requested by users that may conflict with each other. Summary of the Invention
[0003] One aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform an operation, the operation including receiving a first query issued by a first user, the first query specifying a first long-term operation to be performed by a digital assistant. While the digital assistant is performing the first long-term operation, the operation also includes receiving a second query, the second query specifying a second long-term operation to be performed by the digital assistant, and further including determining that the second query was issued by another user different from the first user. Based on determining that the second query was received from the other user, the operation also includes using a query resolver to determine that performing the second long-term operation would conflict with the first long-term operation, and based on determining that performing the second long-term operation would conflict with the first long-term operation, identifying one or more compromise operations to be performed by the digital assistant. The operation further includes instructing the digital assistant to perform a selected compromise action among the identified one or more compromise actions.
[0004] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, receiving the second query includes receiving audio data corresponding to the second query, where the second query is spoken by another user and captured by an assistant-enabled device running a digital assistant, and determining that the second query was issued by a user other than the first user includes performing speaker identification on the audio data to determine that the second query was spoken by a user other than the first user who issued the query. In these implementations, performing speaker identification on the audio data to determine that the second query was spoken by a user other than the first user who issued the query includes extracting a speaker identification vector representing characteristics of the second query from the audio data corresponding to the second query, and determining that the speaker identification vector extracted from the audio data corresponding to the second query does not match at least one of a reference speaker vector of the first user or matches an enrollment speaker vector associated with the other user. In some examples, receiving the first query issued by the first user includes receiving, from a user device associated with the first user, a user input indication indicating the user's intent to issue the first query. Additionally or alternatively, receiving the first query issued by the first user includes receiving audio data corresponding to the first query spoken by the first user and captured by an assistant-enabled device running a digital assistant.
[0005] In some embodiments, identifying one or more compromise actions to be performed by the digital assistant includes identifying criteria associated with a first query, identifying criteria associated with a second query, and using a query embedding model to generate a first query embedding based on the criteria associated with the first query and a second query embedding based on the criteria associated with the second query. The operations also include determining a combined embedding based on the first query embedding and the second query embedding, and identifying at least one compromise action that maps to the combined embedding in the embedding space. In these embodiments, the identified criteria associated with the first query may include a first preference for a type of media content for playback from an assistant-enabled device running the digital assistant, the identified criteria associated with the second query may include a second preference for a type of media content for playback from the assistant-enabled device, and the identified at least one compromise action may include a third preference for a type of media content for playback from the assistant-enabled device. Additionally, the type of media content may include music, the first preference for the type of media content may include a first genre of music, and the second preference for the type of media content may include a second genre of music. Alternatively, the identified criteria associated with the first query may include a first value for a setting of a home automation device, the identified criteria associated with the second query may include a second value for the setting of the home automation device, and the identified at least one compromise action may include adjusting the first value of the setting of the home automation device to a new value. Here, the home automation device may include a smart thermostat, a smart light, a smart speaker, or a smart display.
[0006] In some examples, the operations further include obtaining a home graph indicating at least two assistant-enabled devices in the same environment as the first user and the other user that are capable of performing the first long-term operation and the second long-term operation. Here, identifying one or more compromise operations for the digital assistant to perform includes identifying a first assistant-enabled device from the home graph as a candidate digital assistant for performing the first long-term operation, and identifying a second assistant-enabled device from the home graph as a candidate digital assistant for simultaneously performing the second long-term operation while the digital assistant performs the long-term operation on the first assistant-enabled device. In these examples, the operations may further include obtaining, from the home graph, proximity information for each of at least two AEDs in the same environment as the first user and the other user, and proximity information for each of the first user who issued the first query and the other user who issued the second query. In these examples, identifying a first assistant-enabled device from the home graph as a candidate digital assistant to perform a first long-term operation, and identifying a second assistant-enabled device from the home graph as a candidate digital assistant to simultaneously perform a second long-term operation, is based on proximity information for each of at least two AEDs and proximity information for each of the first user and the other user.
[0007] In some embodiments, the digital assistant performs a first long-term operation on a first assistant-enabled device, and instructing the digital assistant to perform the selected compromise operation includes instructing the digital assistant to simultaneously perform a second long-term operation on a second assistant-enabled device while the digital assistant is performing the first long-term operation on the first assistant-enabled device. In these examples, after instructing the digital assistant to perform the second long-term operation on the second assistant-enabled device, the operation may further include instructing the digital assistant to adjust performance of the first long-term operation on the first assistant-enabled device. In some embodiments, once multiple compromise operations for the digital assistant to perform are identified, the operation further includes determining a respective score associated with each compromise operation among the multiple compromise operations and selecting the compromise operation among the multiple compromise operations as the compromise operation having the highest respective score. In these embodiments, the operation may further include determining that the respective scores associated with the selected compromise operation meet a threshold. Here, instructing the digital assistant to perform the compromise operation is based on the respective scores associated with the selected compromise operation meeting a threshold. In some examples, the operation further includes prompting the first user and / or other users to provide confirmation that the digital assistant will perform the selected compromise operation, and receiving a positive confirmation from the first user and / or other users that the digital assistant will perform the selected compromise operation, and instructing the digital assistant to perform the selected compromise operation is based on the received positive confirmation.
[0008] Another aspect of the present disclosure provides a system including data processing hardware and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed by the data processing hardware, cause the data processing hardware to perform operations, the operations including receiving a first query issued by a first user, the first query specifying a first long-term operation to be performed by the digital assistant. While the digital assistant is performing the first long-term operation, the operations also include receiving a second query, the second query specifying a second long-term operation to be performed by the digital assistant, and determining that the second query was issued by another user different from the first user. Based on determining that the second query was received from the other user, the operations also include using a query resolver to determine that performing the second long-term operation would conflict with the first long-term operation, and based on determining that performing the second long-term operation would conflict with the first long-term operation, identifying one or more compromise operations to be performed by the digital assistant. The operation further includes instructing the digital assistant to perform a selected compromise action among the identified one or more compromise actions.
[0009] This aspect may include one or more of the following optional features. In some implementations, receiving the second query includes receiving audio data corresponding to the second query, wherein the second query is spoken by another user and captured by an assistant-enabled device running a digital assistant, and determining that the second query was issued by a user other than the first user includes performing speaker identification on the audio data to determine that the second query was spoken by a user other than the first user who issued the query. In these implementations, performing speaker identification on the audio data to determine that the second query was spoken by a user other than the first user who issued the query includes extracting a speaker identification vector representing characteristics of the second query from the audio data corresponding to the second query, and determining that the speaker identification vector extracted from the audio data corresponding to the second query does not match at least one of a reference speaker vector of the first user or a match with an enrollment speaker vector associated with the other user. In some examples, receiving the first query issued by the first user includes receiving, from a user device associated with the first user, a user input indication indicating the user's intent to issue the first query. Additionally or alternatively, receiving the first query issued by the first user includes receiving audio data corresponding to the first query spoken by the first user and captured by an assistant-enabled device running a digital assistant.
[0010] In some embodiments, identifying one or more compromise actions to be performed by the digital assistant includes identifying criteria associated with a first query, identifying criteria associated with a second query, and using a query embedding model to generate a first query embedding based on the criteria associated with the first query and a second query embedding based on the criteria associated with the second query. The operations also include determining a combined embedding based on the first query embedding and the second query embedding, and identifying at least one compromise action that maps to the combined embedding in the embedding space. In these embodiments, the identified criteria associated with the first query may include a first preference for a type of media content for playback from an assistant-enabled device running the digital assistant, the identified criteria associated with the second query may include a second preference for a type of media content for playback from the assistant-enabled device, and the identified at least one compromise action may include a third preference for a type of media content for playback from the assistant-enabled device. Additionally, the type of media content may include music, the first preference for the type of media content may include a first genre of music, and the second preference for the type of media content may include a second genre of music. Alternatively, the identified criteria associated with the first query may include a first value for a setting of a home automation device, the identified criteria associated with the second query may include a second value for the setting of the home automation device, and the identified at least one compromise action may include adjusting the first value of the setting of the home automation device to a new value. Here, the home automation device may include a smart thermostat, a smart light, a smart speaker, or a smart display.
[0011] In some examples, the operations further include obtaining a home graph indicating at least two assistant-enabled devices in the same environment as the first user and the other user that are capable of performing the first long-term operation and the second long-term operation. Here, identifying one or more compromise operations for the digital assistant to perform includes identifying a first assistant-enabled device from the home graph as a candidate digital assistant for performing the first long-term operation, and identifying a second assistant-enabled device from the home graph as a candidate digital assistant for simultaneously performing the second long-term operation while the digital assistant performs the long-term operation on the first assistant-enabled device. In these examples, the operations may further include obtaining, from the home graph, proximity information for each of at least two AEDs in the same environment as the first user and the other user, and proximity information for each of the first user who issued the first query and the other user who issued the second query. In these examples, identifying a first assistant-enabled device from the home graph as a candidate digital assistant to perform a first long-term operation, and identifying a second assistant-enabled device from the home graph as a candidate digital assistant to simultaneously perform a second long-term operation, is based on proximity information for each of at least two AEDs and proximity information for each of the first user and the other user.
[0012] In some embodiments, the digital assistant performs a first long-term operation on a first assistant-enabled device, and instructing the digital assistant to perform the selected compromise operation includes instructing the digital assistant to simultaneously perform a second long-term operation on a second assistant-enabled device while the digital assistant is performing the first long-term operation on the first assistant-enabled device. In these examples, after instructing the digital assistant to perform the second long-term operation on the second assistant-enabled device, the operation may further include instructing the digital assistant to adjust performance of the first long-term operation on the first assistant-enabled device. In some embodiments, once multiple compromise operations for the digital assistant to perform are identified, the operation further includes determining a respective score associated with each compromise operation among the multiple compromise operations and selecting the compromise operation among the multiple compromise operations as the compromise operation having the highest respective score. In these embodiments, the operation may further include determining that the respective scores associated with the selected compromise operation meet a threshold. Here, instructing the digital assistant to perform the compromise operation is based on the respective scores associated with the selected compromise operation meeting a threshold. In some examples, the operation further includes prompting the first user and / or other users to provide confirmation that the digital assistant will perform the selected compromise operation, and receiving a positive confirmation from the first user and / or other users that the digital assistant will perform the selected compromise operation, and instructing the digital assistant to perform the selected compromise operation is based on the received positive confirmation.
[0013] The details of one or more embodiments of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will become apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]
[0014] [Figure 1A] FIG. 1 is a schematic diagram of an example system including multiple users controlling assistant-enabled devices. [Figure 1B] FIG. 1 is a schematic diagram of an example system including multiple users controlling assistant-enabled devices. [Figure 1C] FIG. 1 is a schematic diagram of an example system including multiple users controlling assistant-enabled devices. [Figure 2A] 1 is an exemplary graphical user interface rendered on a screen of a user device for displaying long term actions. [Figure 2B] 1 is an exemplary graphical user interface rendered on a screen of a user device for displaying long term actions. [Figure 3] FIG. 1 is a schematic diagram of a query handling process. [Figure 4A] FIG. 1 is a schematic diagram of a speaker identification process. [Figure 4B] FIG. 1 is a schematic diagram of a speaker verification process. [Figure 5] An embedding space for the query handling process. [Figure 6] 1 is a flowchart of an exemplary operational arrangement of a method for handling voice queries in a multi-user environment. [Figure 7] FIG. 1 is a schematic diagram of an example computing device that may be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION
[0015] Like reference symbols in the various drawings indicate like elements.
[0016] The way in which a user interacts with an assistant-enabled device is designed to be primarily, but not exclusively, through voice input. For example, a user may ask the device to perform an action, including media playback (e.g., music or podcasts), and the device responds by starting to play audio that matches the user's criteria. In an example in which a device (e.g., a smart speaker) is commonly shared by multiple users in an environment, the device may need to process multiple actions requested by users that may conflict with each other. If one or more of the multiple users issue multiple individual requests for the device, subsequent requests may override existing actions being performed by the device. Rather than overriding previous requests, the device may attempt to provide multiple users with a compromise that encompasses the respective preferences of the multiple users. Ensuring that each user in the environment is considered reduces the frequency with which requests are unnecessarily overridden without the consent of the initiating user.
[0017] In addition to controlling actions such as playing music to accommodate conflicting requests, the device may control other types of media, such as podcasts and videos, as well as home automation, such as adjusting lighting levels and controlling air conditioning levels. Similarly, the device may prevent other users from overriding the first user's request by identifying compromises and prompting the user to accept the compromise before overriding the first user's request. This saves computational resources for processing conflicting requests, as well as the time it would take the first user to reinstate their original request if the original request was overridden without consent. Additionally, this may extend to controlling aspects of the home connected to the device. For example, a party host may set lighting levels during the party to ensure a calming atmosphere. The host might utter, "Set the lights to 60%." During the party, the device may prevent other party attendees from adjusting lighting levels, or limit the extent to which they can adjust them, by incorporating the attendees' lighting requests into compromises that the host can agree to or reject.
[0018] The device may additionally operate to resolve conflicts between individuals present within the home. For example, the device may assist individuals in the environment in creating a shopping list, thereby ensuring that any conflicting items are resolved by providing the individuals with compromise suggestions for adding items to the shopping list. For example, the device may recommend items to add to the shopping list in response to two individual requests that conflict or are similar enough to be combined. Similarly, the device may proactively mediate disagreements between individuals. For example, the device may engage / prompt individuals with conflicting views with a compromise that suits both individuals, thereby resolving the disagreement.
[0019] 1A-1C illustrate exemplary systems 100a-100c for handling queries in an environment with multiple users 102, 102a-102n, using a query handler that balances queries from multiple users 102 detected in the environment by providing compromises. Briefly, as described in more detail below, a digital assistant 105 including a query handler 300 (FIG. 3) detects multiple users 102, 102a-102b in the environment and begins playing music 122 in response to receiving a first query 106 issued by user 102a: "Okay, computer, play Red from the pop music playlist." While the digital assistant 105 is performing a long-running operation playing music 122 as playback audio from speaker 18, the digital assistant 105 receives a second query 146: "Play Canon in D" spoken by another user 102b (FIG. 1B) different from user 102a. Because the query handler 300 detects / recognizes that the other user 102b is different from the user 102a and that the second query 146 issued by the user 102b conflicts with the first query 106 issued by the user 102a, the query handler 300 identifies one or more compromise actions 354, 354a-354n (FIG. 3) for the digital assistant 105 to perform.
[0020] The systems 100a-100c include an assistant-enabled electronic device (AED) 104 (i.e., also referred to as the "primary AED 104") and multiple secondary assistant-enabled devices (AEDs) 103, 103a-103n located throughout the environment. In the illustrated example, the environment may correspond to a home having one and two floors, where a first smart speaker 104 (i.e., AED 104) is located on the first floor and a second smart speaker 103a, a smart light 103bc, and a smart thermostat 103c are located on the second floor. However, the AED 104 and / or secondary AEDs 103 may include other computing devices, such as, but not limited to, smartphones, tablets, smart displays, desktops / laptops, smartwatches, smart glasses / headsets, smart appliances, headphones, or vehicle infotainment devices. As illustrated, a digital assistant 105 runs on the AED 104 with which multiple users 102 may interact by issuing queries containing commands to perform long-term actions. The AED 104 includes data processing hardware 10 and memory hardware 12 that stores instructions that, when executed on the data processing hardware 10, cause the data processing hardware 10 to perform operations. The AED 104 includes an array of one or more microphones 16 configured to capture sound, such as speech, directed at the AED 104. The AED 104 may also include or communicate with an audio output device (e.g., speaker) 18 that may output audio, such as music 122 and / or synthesized speech, from the digital assistant 105. Additionally, the AED 104 may include or communicate with one or more cameras 19 configured to capture images of the environment and output image data 312 ( FIG. 3 ).
[0021] In some implementations, each secondary AED 103 broadcasts proximity information 107, 107a-107c that can be received by environmental detector 310 (FIG. 3), and a digital assistant 105 running on an AED 104 can use the proximity information 107, 107a-107c to determine the presence of each secondary AED 103. Additionally, the digital assistant 105 can use the proximity information 107 of each secondary AED 103 to infer a home graph and understand the spatial proximity of each secondary AED 103 to each other and to the AED 104 running the digital assistant 105 (e.g., to determine which secondary AEDs 103 can be partitioned / compartmentalized to perform conflicting long-term operations). The proximity information 107 from each secondary AED 103 may include a wireless communication signal, such as WiFi, Bluetooth, or ultrasound, where the signal strength of the wireless communication signal received by the environment detector 310 may correlate with the proximity (e.g., distance) of the secondary AED 103 to the AED 104. Additionally or alternatively, the proximity information 107 from each secondary AED 103 may be determined by playing a stationary sound at the AED 104 and determining which of the secondary AEDs 103 in the environment detect the stationary sound to establish an approximate distance between the AED 104 and the secondary AED 103.
[0022] In some configurations, the digital assistant 105 communicates with multiple user devices 50, 50a-50n associated with multiple users 102. In the illustrated example, each user device 50 of the multiple user devices 50a-50c includes a smartphone with which the respective user 102 may interact. However, the user device 50 may include other computing devices, such as, but not limited to, a smartwatch, a smart display, smart glasses, a smartphone, smart glasses / headset, a tablet, a smart appliance, headphones, a computing device, a smart speaker, or other assistant-enabled devices. Each user device 50 of the multiple user devices 50a-50n may include at least one microphone 52, 52a-52n present on the user device 50 that communicates with the digital assistant 105. In these configurations, the user device 50 may also communicate with one or more microphones 16 present on the AED 104. Additionally, multiple users 102 may control and / or configure the AED 104 and secondary AED 103 and may interact with the digital assistant 105 using an interface 200, such as a graphical user interface (GUI) 200, rendered for display on the respective screens of each user device 50.
[0023] 1A-1C and 3, a digital assistant 105 executing a query handler 300 uses a compromise generator 350 to manage queries issued by multiple users 102. In some implementations, the query handler 300 includes a query resolver 340 that identifies conflicts between one or more queries received from each of the users 102 in the environment, and the compromise generator 350 that executes a compromise model 352 that identifies a compromise action 354 for the digital assistant 105 to perform. In this sense, the query handler 300 balances the competing interests of multiple users 102 while minimizing the frequency with which a user's 102 query is overridden or interrupted by subsequent queries issued by other users 102 in the environment.
[0024] 2A and 2B, GUIs 200a, 200b of user devices 50 associated with users 102 may display a current long-term action 111 (e.g., playing music 122) to keep each user 102 informed of active long-term actions being performed by the digital assistant 105. In some configurations, the AED 104 includes a screen and renders the GUI 200 to display the active long-term action on the screen. For example, the AED 104 may include a smart display, tablet, or smart TV in the environment. 2A provides an exemplary GUI 200 displayed on the screen of a user device 50 associated with a user 102, which may additionally display for display: an identifier of the current long-running action 111 (e.g., "Playing Red"), an identifier of the AED 104 (e.g., smart speaker 1) currently performing the long-running action, an indication of the next song 202 to be played during the long-running action (e.g., Play Next), and / or the identity of the user 102a (e.g., Barb) who initiated the current long-running action being performed by the digital assistant 105. As described above, the query handler 300 manages long-running actions such that when the query handler 300 determines that another user 102 has issued a second long-running action that conflicts with a first long-running action issued by the user 102a (e.g., Barb), the digital assistant 105 prevents the second action from being performed (or at least requests an attempt to implement a compromise).
[0025] 1A-1C , the environmental detector 310 may identify the user 102a (e.g., Barb) and the user 102b (e.g., Jeff) by the proximity information 54 received from their respective user devices 50a, 50b. However, in other examples, the environmental detector 310 may detect that the users 102a, 102b have detected their respective user devices 50a, 50b (e.g., one or more of the users 102a, 102b have chosen not to share the proximity information 54 from their respective user devices 50a, 50b). However, the environmental detector 310 still detects the presence of the users 102a, 102b (e.g., via speech data, image data, and / or input from the users 102a, 102b). Additionally, the environmental detector may identify the secondary AEDs 103a-103c by the proximity information 107 received from their respective secondary AEDs 103a-103c. The environment detector 310 then generates a home graph that includes the identities and locations of the users 102a, 102b and the secondary AED 103 within the current environment 316. When a conflict occurs, the query handler 300 may generate a compromise action 354 using the home graph that represents the current environment 316.
[0026] Continuing with the example of FIG. 1A , user 102a of multiple users 102a-102c is shown issuing a first query 106, "OK, computer, play Red from the pop music playlist," in proximity to AED 104. Here, the first query 106 issued by user 102a is spoken by user 102a and includes initial audio data 402 ( FIG. 3 ) for the first query 106. The first query 106 may further include a user input indication indicating the user's intent to issue the first query via any one of touch, speech, gesture, gaze, and / or an input device (e.g., a mouse or stylus) to interact with AED 104. Optionally, based on receiving the initial audio data 402 corresponding to the first query 106, the query handler 300 performs a speaker identification process 400a (FIG. 4A) on the audio data 402 and resolves the identity of the speaker of the first query 106 by determining that the first query 106 was issued by the user 102a. In other implementations, the user 102a issues the first query 106 without speaking. In these implementations, the user 102a may issue the first query 106 via a user device 50a associated with the user 102a (e.g., by entering text corresponding to the first query 106 into a GUI 200 displayed on a screen of the user device 50a associated with the user 102a, by selecting the first query 106 displayed on the screen of the user device 50a, etc.). Here, the AED 104 may resolve the identity of the user 102 who issued the first query 106 by recognizing the user device 50a associated with the user 102a.
[0027] The microphone 16 of the AED 104 receives the first query 106 and processes initial audio data 402 corresponding to the first query 106. The initial processing of the audio data 402 may include filtering the audio data 402 and converting the audio data 402 from an analog signal to a digital signal. Once the AED 104 processes the audio data 402, the AED may store the audio data 402 in a buffer in the memory hardware 12 for further processing. With the audio data 402 in the buffer, the AED 104 may use the hot word detector 108 to detect whether the audio data 402 includes a hot word. The hot word detector 108 is configured to identify hot words included in the audio data 402 without performing voice recognition on the audio data 402.
[0028] In some implementations, the hot word detector 108 is configured to identify hot words that are in an initial portion of the first query 106. In this example, the hot word detector 108 may determine that the first query 106, "Okay, computer, play Red from the pop music playlist," contains the hot word 110, "Okay, computer," if the hot word detector 108 detects acoustic features in the audio data 402 that are characteristic of the hot word 110. The acoustic features may be Mel-Frequency Cepstral Coefficients (MFCCs), which are a representation of the short-term power spectrum of the first query 106, or may be Mel-scale filter bank energies of the first utterance 106. For example, the hot word detector 108 may detect that the first query 106, "Okay, computer, play Red from a pop music playlist," includes the hot word 110, "Okay, computer," based on generating MFCCs from the audio data 402 and classifying the MFCCs as including MFCCs similar to MFCCs characteristic of the hot word "Okay, computer" stored in a hot word model of the hot word detector 108. As another example, the hot word detector 108 may detect that the first query 106, "Okay, computer, play Glory and let's have pop music all night long," includes the hot word 110, "Okay, computer," based on generating mel-scale filter bank energies from the audio data 402 and classifying the mel-scale filter bank energies as including mel-scale filter bank energies similar to mel-scale filter bank energies characteristic of the hot word "Okay, computer" stored in a hot word model of the hot word detector 108.
[0029] When the hot word detector 108 determines that the initial audio data 402 corresponding to the first query 106 includes the hot word 110, the AED 104 may trigger a wake-up process to begin speech recognition on the audio data 402 corresponding to the first query 106. For example, FIG. 3 shows the AED 104 including a speech recognizer 170 employing an automatic speech recognition model 172 that may perform speech recognition or semantic interpretation on the audio data 402 corresponding to the first query 106. The speech recognizer 170 may perform speech recognition on the portion of the audio data 402 that follows the hot word 110. In this example, the speech recognizer 170 may identify the words "Play Glory and let's keep it pop music all night" in the first query 106.
[0030] In some examples, the AED 104 is configured to communicate with a remote system 130 over the network 120. The remote system 130 may include remote resources, such as remote data processing hardware 132 (e.g., a remote server or CPU) and / or remote memory hardware 134 (e.g., a remote database or other storage hardware). The query handler 300 may execute on the remote system 130 in addition to or instead of the AED 104. The AED 104 may utilize the remote resources to perform various functions related to speech processing and / or synthesis and playback communications. In some implementations, the speech recognizer 170 is located on the remote system 130 in addition to or instead of the AED 104. When the hot word detector 108 triggers the AED 104 to wake up in response to detecting the hot word 110 in the first query 106, the AED 104 may transmit initial audio data 402 corresponding to the first query 106 to the remote system 130 over the network 120. Here, the AED 104 may transmit the portion of the initial audio data 402 that includes the hotword 110 for the remote system 130 to verify the presence of the hotword 110. Alternatively, the AED 104 may transmit only the portion of the initial audio data 402 that corresponds to the portion of the utterance 106 after the hotword 110 to the remote system 130, and the remote system 130 executes the speech recognizer 170 to perform speech recognition and returns a transcription of the initial audio data 402 to the AED 104.
[0031] 1A-1C and 3, the query handler 300 further includes a natural language understanding (NLU) module 320 that performs semantic interpretation on the first query 106 to identify a query / command directed to the AED 104. Specifically, the NLU module 320 identifies words in the first query 106 identified by the speech recognizer 170 and performs semantic interpretation to identify any voice commands in the first query 106. The NLU module 320 of the AED 104 (and / or remote system 130) may identify the words "Play Red" as a command specifying a first long-term action 111 of the digital assistant 105 (i.e., play music 122) and may identify the words "from a pop music playlist" as criteria 113 for the digital assistant 105 to play music of a particular genre (e.g., pop music) while performing the first long-term action 111. 1A , the digital assistant 105 begins executing a first long-term action 111, playing music 122 as playback audio (e.g., track 1) from the speaker 18 of the AED 104. The digital assistant 105 may stream the music 122 from a streaming service (not shown), or the digital assistant 105 may instruct the AED 104 to play music stored on the AED 104. While the exemplary long-term action 111 includes music playback, the long-term action may also include other types of media playback, such as videos, podcasts, and / or audiobooks. The long-term action 111 may also include home automation (e.g., adjusting light levels, controlling a thermostat, etc.).
[0032] 1A, with reference to FIG. 3, the query handler 300 adds long-term actions 111 to the data store of active actions 330. The query handler 300 maintains records of active actions 332 in the environment (e.g., light intensity of a smart light bulb, temperature setpoint of a smart thermostat, etc.) in the data store of active actions 330 (e.g., stored in memory hardware 12, 134), and the query handler 300 may limit the long-term actions performed by the digital assistant 105 based on the active actions 332. For example, before performing the first long-term action 111, the query handler 300 may first use the query resolver 340 to verify that the first long-term action 111 in the first query 106 does not conflict with any active actions 332 in the environment.
[0033] The AED 104 may notify the user 102a (e.g., Barb) who issued the first query 106 that the first long-term action 111 is running. For example, the digital assistant 105 may generate synthesized speech 123 for audible output from the speaker 18 of the AED 104 stating, "Barb, I'm playing Red from Pop Music now." In a further example, the digital assistant 105 provides a notification to the user device 50a associated with the user 102a (e.g., Barb) informing the user 102a of the approved first long-term action 111 and / or any active actions 332 stored in the data store 330 of active actions.
[0034] 2A , a graphic user interface (GUI) 200a executing on the user device 50 may display a first long-term action 111 and / or active actions 332 executing on secondary devices 103 in the environment. As used herein, the GUI 200a may receive user input indications via any one of touch, speech, gesture, gaze, and / or an input device (e.g., a mouse or stylus) for interacting with the digital assistant 105. For example, the GUI 200a of FIG. 2A may render for display an identifier of the current long-term action 111 (e.g., “Play Red”), an identifier of the AED 104 (e.g., Smart Speaker 1) currently performing the long-term action 111, an identifier of the criteria 113 (e.g., pop music) associated with the first query 106, an indication of the next song 202 to be played during the long-term action (e.g., Play Next), and / or the identity of the user 102a (e.g., Barb) who issued the first query 106. In an implementation in which the first query 106 includes a criterion 113 (e.g., pop music), the GUI 200a renders for display an identifier of the criterion 113. Additionally, the GUI 200a renders an identifier of the smart thermostat 103c and an indication of an active long-term operation 332 that the setpoint is 70 degrees, as well as an identifier of the smart light bulb 103b and an indication of an active long-term operation 332 that the light is on. Thus, the user 102 can consult the user device 50 to review the criteria 113, which may limit the active long-term operation of the digital assistant 105 within the environment and queries issued by the user 102.
[0035] 4A and 4B , in some implementations, the AED 104 (or a remote system 130 in communication with the AED 104) also includes an exemplary data store 430 that stores data / information for each of a plurality of enrolled users 432a-432n of the AED 104. Here, each enrolled user 432 of the AED 104 may undergo a speech enrollment process to obtain a respective enrollment speaker vector 154 from audio samples of a plurality of enrollment phrases spoken by the enrolled user 432. For example, the speaker identification model 410 may generate one or more enrollment speaker vectors 154 from audio samples of enrollment phrases spoken by each enrolled user 432, which may be combined, e.g., averaged or otherwise accumulated, to form a respective enrollment speaker vector 154. One or more of the enrollment users 432 may perform a voice enrollment process using the AED 104, with the microphone 16 capturing audio samples of these users uttering enrollment utterances, from which the speaker identification model 410 generates respective enrollment speaker vectors 154. The model 410 may run on the AED 104, the remote system 130, or a combination thereof. Additionally, one or more of the enrollment users 432 may register with the AED 104 by providing authorization and authentication credentials to an existing user account with the AED 104. Here, the existing user account may store the enrollment speaker vectors 154 obtained from a previous voice enrollment process, with other devices also linked to the user account.
[0036] In some examples, the enrollment speaker vector 154 of an enrolled user 432 includes a text-dependent enrollment speaker vector. For example, the text-dependent enrollment speaker vector may be extracted from one or more audio samples of each enrolled user 432 speaking a predetermined term, such as a hot word 110 (e.g., "OK, computer") used to summon the AED 104 to wake it from a sleep state. In other examples, the enrollment speaker vector 154 of an enrolled user 432 is obtained from one or more audio samples of each enrolled user 102 speaking phrases of different lengths with different terms / words, and is text-independent. In these examples, the text-independent enrollment speaker vector may be obtained over time from audio samples obtained from speech interactions the user 102 has with the AED 104 or other devices linked to the same account.
[0037] 4A , a speaker identification process 400a identifies a user 102a (e.g., Barb) who uttered a first query 106 by first extracting a first speaker identification vector 411 representing characteristics of the first query 106 issued by the user 102a from initial audio data 402 corresponding to the first query 106. Here, the speaker identification process 400a may execute a speaker identification model 410 configured to receive audio data 402 corresponding to the second query 146 as input and generate the first speaker identification vector 411 as output. The speaker identification model 410 may be a neural network model trained under machine or human supervision to output the speaker identification vector 411. The speaker identification vector 411 output by the speaker identification model 410 may include an N-dimensional vector having values corresponding to speech features of the first query 106 associated with the user 102a. In some examples, the speaker identification vector 411 is a d-vector. In some examples, the first speaker identification vector 411 includes a set of speaker identification vectors each associated with a different user who is also authorized to control the AED 104. For example, aside from the user 102a who uttered the first query 106, the other authorized users would include other individuals who were present when the user 102a uttered the first query 106 to issue the command 111 to perform the first action, and / or individuals that the user 102a added / designated as authorized.
[0038] Once the first speaker identification vector 411 is output from the model 410, the speaker identification process 400a determines whether the extracted speaker identification vector 411 matches any of the enrollment speaker vectors 154 stored in the AED 104 (e.g., in the memory hardware 12) for the enrolled users 432a-432n of the AED 104. As described above, the speaker identification model 410 may generate the enrollment speaker vectors 154 for the enrolled users 432 during the voice enrollment process. Each enrollment speaker vector 154 may be used as a reference vector 155 corresponding to a voiceprint or unique identifier representing the vocal characteristics of the respective enrolled user 432.
[0039] In some implementations, the speaker identification process 400a uses a comparator 420 that compares the first speaker identification vector 411 to a respective enrollment speaker vector 154 associated with each enrolled user 432a-432n of the AED 104. Here, the comparator 420 may generate a score for each comparison indicating the likelihood that the initial audio data 402 corresponding to the first query 106 matches the identity of the respective enrolled user 432, and if the score meets a threshold, the identity is accepted. If the score does not meet the threshold, the comparator 420 may reject the identity of the speaker who issued the first query 106. In some implementations, the comparator 420 calculates a respective cosine distance between the first speaker identification vector 411 and each enrollment speaker vector 154 and determines that the first speaker identification vector 411 matches one of the enrollment speaker vectors 154 if the respective cosine distance meets a cosine distance threshold.
[0040] In some examples, the first speaker identification vector 411 is a text-dependent speaker identification vector extracted from a portion of one or more words corresponding to the first query 106, and each enrollment speaker vector 154 is also text-dependent on the same one or more words. Using a text-dependent speaker vector can improve accuracy in determining whether the first speaker identification vector 411 matches any of the enrollment speaker vectors 154. In other examples, the first speaker identification vector 411 is a text-independent speaker identification vector extracted from the entire initial audio data 402 corresponding to the first query 106.
[0041] If the speaker identification process 400a determines that the first speaker identification vector 411 matches one of the enrollment speaker vectors 154, the process 400a identifies the user 102a who spoke the first query 106 as the respective enrolled user 432a associated with one of the enrollment speaker vectors 154 that matches the extracted speaker identification vector 411. In the illustrated example, the comparator 420 determines a match based on the respective cosine distances between the first speaker identification vector 411 and the enrollment speaker vector 154 associated with the enrollment user 432a satisfying a cosine distance threshold. In some scenarios, the comparator 420 identifies the user 102a as the respective enrolled user 432a associated with the enrollment speaker vector 154 having the shortest respective cosine distance from the first speaker identification vector 411 if this shortest respective cosine distance also satisfies the cosine distance threshold.
[0042] Conversely, if the speaker identification process 400a determines that the first speaker identification vector 411 does not match any of the enrollment speaker vectors 154, the process 400a may identify the user 102a who spoke the utterance 106 as a guest user of the AED 104. Accordingly, the query handler 300 may add the guest user and use the first speaker identification vector 411 as the reference speaker vector 155 that represents the speech characteristics of the guest user's voice. In some examples, the guest user may register with the AED 104, and the AED 104 could store the first speaker identification vector 411 as the enrollment speaker vector 154 for each of the newly enrolled users.
[0043] 1B, while the digital assistant 105 is performing the first long-term action 111 of playing music 122 as playback audio from the speaker 18 of the AED 104, the digital assistant 105 receives a second query 146 that specifies a second long-term action 112 for the digital assistant 105 to perform. In the illustrated example, another user 102b, different from the user 102a who issued the first query 106, issues the second query 146 "Play Canon in D," which includes a command for the digital assistant 105 to perform the second long-term action 112 of playing a song (i.e., Canon in D), with associated criteria 115 for the digital assistant 105 to play a particular genre of music (e.g., classical music). Based on receiving the second query 146, the query handler 300 resolves the identity of the speaker of the second query 146 by performing a speaker identification process 400b (FIG. 4B) on audio data 402 corresponding to the second query 146 and determines that the second query 146 was issued by another user 102b different from the user 102a who issued the first query 106. As described above with reference to FIG. 4A, in an implementation in which the first query 106 issued by the user 102a includes initial audio data 402 (e.g., the first query 106 was spoken by the first user 102a), the query handler 300 may first perform a speaker identification process 400a on the initial audio data 402 corresponding to the first query 106 to identify the user 102a who issued the first query 106.
[0044] The speaker identification process 400b may execute on the data processing hardware 12 of the AED 104. The speaker identification process 400b may also execute on the remote system 130. If the speaker verification process 400b for the audio data 402 corresponding to the second query 146 indicates that the second query 146 was spoken by the same user 102a who issued the first query 106, the digital assistant 105 may proceed to execute the second long-term action 105 without first determining whether the first long-term action 111 and the second long-term action 112 conflict. In other words, if the same user 102a issued both queries 106, 146, the query handler 300 may not be needed to resolve conflicts between the users 102. Conversely, if the speaker verification process 400b for the audio data 402 corresponding to the second query 146 indicates that the second query 146 was spoken by another user 102b than the user 102a who issued the first query 106, the query handler 300 may prevent the second long-term operation 112 from being executed (or at least require input from one or more other users 102 in the environment (e.g., in Figures 1B and 2B)).
[0045] Referring again to FIG. 4B with reference to the example of FIG. 1B , in response to receiving the second query 146, the AED 104 resolves the identity of the user 102 who uttered the second query 146 by executing a speaker identification process 400b. The speaker identification process 400b identifies the user 102a who uttered the first query 146 by first extracting a second speaker identification vector 412 representing characteristics of the second query 146 from audio data 402 corresponding to the first query 146 uttered by the user 102a. Here, the speaker verification process 400b may execute a speaker identification model 410 configured to receive the audio data 402 as input and generate the second speaker identification vector 412 as output. As illustrated in FIG. 4A , the speaker identification model 410 may be a neural network model trained to output the speaker identification vector 412 under machine or human supervision. The second speaker identification vector 412 output by the speaker identification model 410 may include an N-dimensional vector having values corresponding to speech features of the utterance 146 associated with the user 102a. In some examples, the speaker identification vector 412 is a d-vector.
[0046] Once the second speaker identification vector 412 is output from the speaker identification model 410, the speaker verification process 400b determines whether the extracted speaker identification vector 412 matches a reference speaker vector 155 associated with the first enrolled user 432a stored on the AED 104 (e.g., in the memory hardware 12). The reference speaker vector 155 associated with the first enrolled user 432a may include each enrollment speaker vector 154 associated with the first enrolled user 432a. As described above, the speaker identification model 410 may generate the enrollment speaker vector 154 for the enrolled user 432 during the voice enrollment process. Each enrollment speaker vector 154 may be used as a reference vector corresponding to a voiceprint or unique identifier representing the voice characteristics of the respective enrolled user 432.
[0047] In some implementations, the speaker verification process 400b uses a comparator 420 that compares the second speaker identification vector 412 with a reference speaker vector 155 associated with a first enrolled user 432a of the enrolled users 432. The comparator 420 may generate a score for the comparison that indicates the likelihood that the second query 146 corresponds to the identity of the first enrolled user 432a, and if the score meets a threshold, the identity is accepted. If the score does not meet the threshold, the comparator 420 may reject the identity. In some implementations, the comparator 420 calculates the respective cosine distances between the second speaker identification vector 412 and the reference speaker vector 155 associated with the first enrolled user 432a, and determines that the second speaker identification vector matches the reference speaker vector 155 if the respective cosine distances meet a cosine distance threshold.
[0048] If the speaker verification process 400b determines that the second speaker identification vector 412 matches the reference speaker vector 155 associated with the first enrolled user 432a, the process 400b identifies the user 102a who uttered the second query 146 as the first enrolled user 432a associated with the reference speaker vector 155. In the illustrated example, the comparator 420 determines a match based on the second speaker identification vector 412 and the reference speaker vector 155 associated with the first enrolled user 432a satisfying a cosine distance threshold. In some scenarios, the comparator 420 identifies the user 102a as the respective first enrolled user 432a associated with the reference speaker vector 155 having the shortest respective cosine distance from the second speaker identification vector 412 if the shortest respective cosine distance also meets the cosine distance threshold.
[0049] 4A above, in some implementations, the speaker identification process 400a determines that the first speaker identification vector 411 does not match any of the enrollment speaker vectors 154 and identifies the user 102 who uttered the first query 106 as a guest user of the AED 104. Accordingly, the speaker verification process 400b may first determine whether the user 102a who uttered the first query 106 was identified by the speaker identification process 400a as an enrollment user 432 or as a guest user. When the user 102a is a guest user, the comparator 420 compares the second speaker identification vector 412 with the first speaker identification vector 411 obtained during the speaker identification process 400a. Here, the first speaker identification vector 411 represents characteristics of the first query 106 uttered by the guest user 102a, and is therefore used as a reference vector for verifying whether the second query 146 was also uttered by the guest user 102a or another user 102. Here, the comparator 420 may generate a comparison score indicating the likelihood that the second query 146 corresponds to the identity of the guest user 102a, and if the score meets a threshold, the identity is accepted. If the score does not meet the threshold, the comparator 420 may reject the identity of the guest user who issued the second query 146. In some implementations, the comparator 420 calculates the respective cosine distances between the first speaker identification vector 411 and the second speaker vector 412, and determines that the first speaker identification vector 411 matches the second speaker vector 412 if the respective cosine distances meet a cosine distance threshold.
[0050] 1B , based on determining that the second query 146 was uttered by another user 102b different from the user 102a who issued the first query 106, the NLU module 320 running on the AED 104 (and / or running on the remote system 130) may identify the words "Play Canon in D" as a command specifying the second long-term action 112 (i.e., play music 122). Here, the query handler 300 may first determine, using the query resolver 340, whether performing the second long-term action 112 conflicts with the first long-term action 111 before allowing the digital assistant 105 to perform the second long-term action 112. For example, the query handler 300 may determine whether the first long-term action 111 and the second long-term action 112 invoke the same function of the digital assistant 105 (e.g., playing music) or different functions (e.g., playing music and changing the brightness setting of a smart light bulb). In an example where a first long-running action 111 and a second long-running action 112 invoke the same function, the query resolver 340 determines that the long-running actions conflict, and the query handler 300 outputs the conflicting long-running actions 111, 112 to determine one or more compromise actions for the users 102a, 102b.
[0051] In some examples, when the query resolver 340 determines that the second long-term action 112 conflicts with the first long-term action 111, it only outputs the first long-term action 111 and the second long-term action 112 (thereby triggering the query handler 300 to identify one or more compromise actions 354). Conversely, if the first long-term action 111 and the second long-term action 112 invoke different functions, the query resolver 340 determines that the second long-term action 112 does not conflict with the long-term action 111 of leaf 1. Here, when a conflict exists that represents competing interests between user 102a and user 102b, the query resolver 340 only outputs the long-term actions 111, 112, thereby prompting the query handler 300 to identify one or more compromise actions 354. Additionally, as described above, the query resolver 340 may verify that the second long-running operation 112 in the first query 106 does not conflict with the active operations 332 stored in the data store 300 of active operations before executing the second long-running operation 112.
[0052] In the example, the query resolver 340 determines that the second long action 112, playing Canon in D, conflicts with the first long action 111, playing Red, because executing the second long action 112 via the speaker 18 of the AED 104 would necessarily interrupt the execution of the first long action 111, which is currently playing on the speaker 18 of the AED 104. Based on determining that the second user 102b issued the second query 146 and determining that the second query 146 conflicts with the first query 106 issued by the user 102a, the query handler 300 (via the digital assistant 105) prevents the AED 104 from expiring the second long action 112 and instead generates one or more compromise actions 354, 354a-354n to which the users 102a, 102b may agree. In other words, after user 102a is determined as the issuer of the first query 106 and user 102b is determined as the issuer of the second query 146, the query handler 300 attempts to respect the first long-running action 111 and the second long-running action 112 by determining a compromise solution.
[0053] 3 , in some implementations, the query handler 300 includes a compromise generator 350 that uses one or more approaches to generate one or more compromise actions 354, 354a-354n. For example, the compromise generator 350 may execute a compromise model 352 configured to receive the first query 106 and the second query 146 as input and output one or more compromise actions 354 that combine the first query 106 and the second query 146. For example, if the first query requests "party music" and the second query requests "classical music," the compromise generator 350 may output a compromise action to play "upbeat classical music." Similarly, if the first query requests turning up the lights and the second query requests turning down the lights, the compromise generator 350 may output a compromise of a medium level of lighting to accommodate the criteria of both queries.
[0054] Compromise model 352 may be a neural network model trained under machine or human supervision to output compromise actions 354. In other implementations, compromise generator 350 includes multiple compromise models (e.g., some that include neural networks and some that do not). In these implementations, compromise generator 350 may select which of the multiple compromise models to use as compromise model 352 based on the category of action with which the query is associated.
[0055] Continuing with the example, compromise generator 350 identifies criteria (e.g., Red) 113 associated with first query 106 and criteria (e.g., Canon in D) 115 associated with second query 146. Compromise model 352 receives as input the criteria 113 associated with first query 106 and the criteria 115 associated with second query 146 and generates as output a first query embedding 502 ( FIG. 5 ) based on the criteria 113 associated with first query 106 and a second query embedding 504 ( FIG. 5 ) based on the criteria 115 associated with second query 146. 3 with reference to FIG. 5, the compromise model 352 may determine a combined embedding 506 based on the first query embedding 502 and the second query embedding 504 and identify at least one compromise operation 354 that maps (e.g., pre-mapped or pre-defined) to the combined embedding 506 within the embedding space 500. For example, the compromise model 352 may average the first query embedding 502 and the second query embedding 504 to generate the combined embedding 506. In other examples, if the space between the first query embedding 502 and the second query embedding 504 is too large, the compromise model 352 generates the compromise operation 354 without pre-mapping the compromise operation 354 to the combined embedding 506. In some implementations, the combined embedding 506 includes a conflict score that predicts the degree of conflict between the first query embedding 502 and the second query embedding 504.
[0056] In other examples, the identified criteria associated with the first query include a first value for a setting of a home automation device (e.g., a smart thermostat, a smart light, a smart speaker, or a smart display), while the second query includes a second value for the setting of the home automation device. Here, the identified at least one compromise action 354 includes adjusting the first value for the setting of the home automation device to a new value. For example, the compromise model 352 may include a heuristic model that analyzes the first value and the second value and determines an average value between the first value and the second value to set as the new value. In some examples, the home automation device corresponds to a secondary AED 103 in the environment.
[0057] Continuing with the music playback example, as shown in FIG. 5 , an average embedding 506 is determined based on query preferences. Here, the identified criteria 113 associated with a first query embedding 502 includes a first preference 512 in the embedding space 500 for a type of media content (i.e., pop music) for playback from an AED 104 running the digital assistant 105. Additionally, the identified criteria 115 associated with a second query embedding 504 includes a second preference 514 in the embedding space 500 for a type of media content (i.e., classical music) for playback from an AED 104 running the digital assistant 105. Here, although the query 106, 146 does not explicitly state the preferences 512, 514, the compromise generator 350 infers the preferences 512, 514 based on the genres of music associated with the songs "Red" and "Canon in D." In this example, the type of media content includes music, the first preference 512 includes a first genre of music (i.e., pop), and the second preference 514 includes a second genre of music (i.e., classical). As shown, the first preference 512 and the second preference 514 overlap in a region that includes a third preference 516 in the embedding space 500 for the type of media content for a third genre of music (i.e., violin pop covers). Based on the combined embedding 506, the compromise model 352 generates as output a compromise action 354 that continues playing the media but changes the long-term action to the third preference 516 of playing violin pop covers. In an example where the first query preference and the second query preference do not overlap (e.g., classical music and hair metal), the compromise model 350 may infer that no type of music exists that combines classical music and hair metal, and therefore may not output the compromise action 354 that combines the preferences.
[0058] 3, in some implementations, the compromise model 352 of the compromise generator 350 determines whether the conflicting long-term operation can be offloaded to a second AED 103 in the environment (i.e., splitting / separating the AEDs) rather than aborting the current long-term operation. As described above with reference to FIGS. 1A-1C during the execution of the digital assistant 105, the AED 104 uses the environment detector 310 of the query handler 300 to detect multiple users 102a, 102b and secondary devices 103a-103c in the environment. For example, the query handler 300 receives proximity information 54 about the locations of each of the multiple users 102a, 102b and proximity information 107 about the locations of each of the secondary AEDs 103 relative to the AED 104. By monitoring the respective presence and relative locations of the user 102 and the secondary AED 103 within the environment, the environmental detector 310 can provide the compromise generator 350 with a home graph representing the current environment 316 when the query resolver 340 identifies a conflict between the first long-term operation 111 and the second long-term operation 112. In the home graph representing the current environment 316, the environmental detector 310 can identify which secondary AEDs 103 are available, far enough away from the AED 104 so as not to interfere with an active long-term operation, and / or close enough to the user 102 requesting the conflicting long-term operation that the requesting user 102 may prefer the secondary AED 103 to perform the conflicting long-term operation.
[0059] In some implementations, user devices 50a-50c of multiple users 102 broadcast proximity information 54, and each secondary AED 103a-103c broadcasts proximity information 107 receivable by the environment detector 310 that the AED 104 may use to determine the proximity of each user device 50 and secondary user device 103 to the AED 104. The proximity information 54 from each user device 50 and the proximity information 107 from each secondary AED 103 and AED 104 may include wireless communication signals, such as WiFi, Bluetooth, or ultrasound, where the signal strength of the wireless communication signal received by the environment detector 310 may correlate with the proximity (e.g., distance) of the user device 50 and / or secondary AED 103 to the AED 104.
[0060] In embodiments where the user 102 does not have a user device 50 or has a user device 50 that does not share proximity information 54, the environment detector 310 may detect the user 102 based on explicit input (e.g., a guest list) 313 received from the user 102a who issued the first query 106. For example, the environment detector 310 receives the guest list 313 from a seed user 102 (e.g., user 102a) indicating the identity of each user 102 of the multiple users 102. Alternatively, the environment detector 310 detects one or more of the users 102 by performing speaker identification ( FIGS. 4A and 4B ) on speech corresponding to audio data 402 detected in the environment. In other embodiments, the environment detector 310 automatically detects the multiple users 102 and / or secondary AEDs 103 in the environment by receiving image data 312 corresponding to a scene of the environment and acquired by the camera 19. Here, the environment detector 310 detects multiple users 102 and / or secondary AEDs 103 based on the received image data 312 .
[0061] In some implementations, the environment detector 310 maintains a home graph representing the current environment 316 of the user 102 and the secondary AED 103 as the user 102 moves throughout the environment. Here, the home graph shows the user 102 and the secondary AED 103 relative to each other, to rooms / floors within the environment, and to the AED 104. For example, if the user 102b leaves the first floor of the environment, the environment detector 310 may detect that the user 102b is closer to the secondary AED 103b (e.g., smart speaker 2) and may prefer to have the secondary AED 103a execute the conflicting long-term action issued by the user 102b. In response to receiving the home graph representing the current environment 316 from the environment detector 310, the compromise generator 350 may generate one or more additional compromise actions 354, including fulfilling the conflicting queries 106, 146 on separate AEDs.
[0062] In some examples, the compromise generator 350 is configured with a modification threshold, such that when the confidence scores of each of the one or more compromise solutions 354 meet (e.g., exceed) the threshold, the compromise generator 354 outputs the one or more compromise solutions 354 to the user 102. Here, the compromise generator 350 determines a respective confidence score associated with each compromise action 354 among the plurality of compromise actions 354 and selects the compromise action 354 among the plurality of compromise actions 354 as the compromise action 354 having the highest respective confidence score. The threshold may be zero, such that all compromise solutions 354 (e.g., even undesirable compromises) are output to the user 102. Conversely, the threshold may be higher than zero to avoid unnecessary compromise solutions 354 that are likely to be rejected by the user 102. Furthermore, in some implementations, compromises are not possible. For example, in an environment with only a single AED 104, executing a second query on a second AED would not be included in the compromise solution 354. Similarly, the compromise generator 350 may determine that “heavy metal” and “soul” music cannot be combined, and therefore no compromise exists. In some implementations, prompting the digital assistant 105 to perform the compromise action 354 is based on the respective confidence scores associated with the selected compromise actions 354 meeting a threshold. In other words, when the compromise generator 350 identifies multiple compromise solutions 354, each with a respective confidence score, the compromise generator 350 may select the compromise action 354 with the highest respective confidence score and / or may provide a list of the best n compromise actions 354 for the user to select. In some examples, the query handler 500 automatically performs the compromise actions 354 when their respective confidence scores exceed a threshold, rather than prompting the user 102 to select.
[0063] 1B , after determining that the second query 146 was issued by another user 102b than the first user 102a who issued the first query 102a and determining that performing the second long-term action 112 conflicts with the first long-term action 111, the query handler 300 identifies / generates one or more compromise actions 354 for the digital assistant 105. As described above, the query handler (via the compromise generator 350) determines that the first query 106 (e.g., pop music) and the second query 146 (e.g., classical music) are similar enough to be combined, and generates a compromise action 354a that plays a violin pop cover. Additionally, the query handler 300 identifies the AED 104 as a candidate for the digital assistant 105 to continue performing the first long-term operation 111, and identifies the AED 103 (e.g., smart speaker 103a) to simultaneously perform the second long-term operation 112 while the digital assistant 105 continues performing the first long-term operation 111 on the AED 104 as a compromise operation 354b. For example, the query handler 300 may obtain proximity information 107 of at least two secondary AEDs 103 within the environments of the users 102a and 102b, and respective proximity information 54 of the first user 102a and the other user 102b, from the home graph representing the current environment 316. Here, identifying the AED 104 from the home graph as a candidate for the digital assistant 105 to perform the first long-term operation 111, and identifying the assistant-enabled device 103a from the home graph as a candidate for the digital assistant 105 to simultaneously perform the second long-term operation 112, is based on the proximity information 107 of each of the AEDs 103a, 104, and the proximity information 54 of each of the first user 102a and the other user 102b.
[0064] In some embodiments, the query handler 300 presents the identified compromise actions 354a, 354b to the users 102a, 102b (via the digital assistant 105) and prompts one or more of the users 102a, 102b to provide confirmation that the digital assistant 105 will perform the selected compromise action 354. In these embodiments, prompting the user 102 includes providing, as an output from the AED 104, a user-selectable option that, when selected, provides positive confirmation that the digital assistant 105 will perform the selected compromise action 354. For example, the digital assistant 105 may generate synthesized speech 123 for audible output from the speaker 18 of the AED 104 (or a speaker in communication with the data processing hardware (e.g., the speaker of the user device 50)), which prompts the seed user 102a to instruct the digital assistant 105 to perform the identified compromise solution 354, "Barb, do you want to switch to a violin pop cover, or play Canon in D on smart speaker 2?" In response, the user 102a (i.e., Barb) is shown providing confirmation that the digital assistant 105 will perform the compromise action 354b by issuing a third query 148, "Play Canon in D on smart speaker 2," in the vicinity of the AED 104. In response to receiving a positive confirmation from the user 102a, the digital assistant 105 performs a selected compromise action 354b that plays the second long-term action 112 on the secondary AED 103a (i.e., smart speaker 2) while simultaneously performing the first long-term action 111 on the AED 104.
[0065] In addition to or instead of audibly prompting the user 102a, 102, the digital assistant 105 may additionally provide a notification to the user device 50 associated with the user 102 to display one or more user-selectable options of compromise solutions 354 as a graphic element 210 on the screen of the user device 50, wherein the graphic element 210 prompts the user 102 to provide a response to the digital assistant 105 performing the compromise operation 354. As shown in FIG. 2B, GUI 200b displays the graphic element 210 "Conflict Detected. Want a Compromise?" (i.e., compromise action 354a), "Play Violin Pop Cover" (i.e., compromise action 354b), and "No" to allow each user 102 of the device to issue (or decline the opportunity to issue) a third query 148 that instructs digital assistant 105 to perform compromise action 354. Now, when user device 50 receives a user input indication indicating selection of the user-selectable option that selects graphic element 210 "Play Canon in D on Smart Speaker 2," query handler 300 receives a positive confirmation from user 102 that digital assistant 105 will perform compromise action 354b.
[0066] Referring to FIG. 1C , while the digital assistant 105 is executing a first long-term action 111 of music 122 as playback audio (e.g., track 1) on the AED 104, the query handler 300 instructs the digital assistant 105 to simultaneously execute a second long-term action 112 of playing music 122 as playback audio (e.g., track 2) from the secondary AED 103a (i.e., smart speaker 2). As shown, the user 102b has moved to the second floor of the environment to be closer to the requested long-term action 112 of playing Canon in D. In some examples, after instructing the digital assistant 105 to execute the second long-term action 112 on the secondary AED 103a, the query handler 300 instructs the digital assistant 105 to coordinate the execution of the first long-term action 111 on the AED 104. Here, the query handler 300 may instruct the digital assistant to lower the volume of the music 122 output by the AED 104 so that the volume of the music 122 output by the AED 104 does not interfere with the music 122 output by the secondary AED 103a. In another example, if the long-term operation involves two or more home automation devices (e.g., smart light bulb 103b), the query handler 300 may instruct the first device to lower its brightness when a second device is turned on in an adjacent room to maintain a calm environment.
[0067] While the examples primarily refer to avoiding interruptions in long-running activities playing music, long-running activities may refer to any category of actions, including, but not limited to, search queries, controlling assistant-enabled devices (e.g., smart lights, smart thermostats), and playing / adjusting other types of media (e.g., podcasts, videos, etc.). For example, query handler 300 may help users 102 in an environment create shopping lists by resolving conflicts between items on the shopping list by recommending items that all users 102 agree on. Additionally, query handler 300 may enable digital assistant 105 to mediate disagreements between users 102 by engaging / prompting users 102 with compromises they may not have considered themselves.
[0068] 6 includes a flowchart of an exemplary arrangement of operations of a method 600 for handling conflicting queries on a shared device. At operation 602, the method 600 includes receiving a first query 106 issued by a first user 102a, the first query 106 specifying a first long-term operation 111 to be performed by the digital assistant 105. While the digital assistant 105 is performing the first long-term operation 111, the method 600 also includes receiving a second query 146 at operation 604, the second query 146 specifying a second long-term operation 112 to be performed by the digital assistant 105.
[0069] At operation 606, method 600 further includes determining whether second query 146 was issued by another user 102b different from first user 102. Based on determining that second query 146 was received from another user 102b, method 600 also includes, at operation 608, using query resolver 340, determining that performing second long-term operation 112 would conflict with first long-term operation 111. Based on determining that performing second long-term operation 112 would conflict with first long-term operation 111, method 600 also includes, at operation 610, identifying one or more compromise operations 352 for digital assistant 105 to perform. At operation 312, method 600 also includes instructing digital assistant 105 to perform a selected compromise operation 352 from the identified one or more compromise operations 352.
[0070] 7 is a schematic diagram of an exemplary computing device 700 that may be used to implement the systems and methods described herein. Computing device 700 is intended to represent various types of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown, their connections and relationships, and their functions are for illustrative purposes only and are not intended to limit the implementation of the invention as described and / or claimed herein.
[0071] Computing device 700 includes a processor 710, memory 720, a storage device 730, a high-speed interface / controller 740 connecting to memory 720 and a high-speed expansion port 750, and a low-speed interface / controller 760 connecting to a low-speed bus 770 and storage device 730. Each of the components 710, 720, 730, 740, 750, and 760 are interconnected using various buses and may be mounted on a common motherboard or otherwise mounted as desired. Processor 710 (e.g., data processing hardware 10, 132 of FIG. 1 ) processes instructions for execution in computing device 700, including instructions stored in memory 720 or storage device 730, and can display graphical information for a graphical user interface (GUI) on an external input / output device, such as a display 780 connected to high-speed interface 740. In other embodiments, multiple processors and / or multiple buses may be used, along with multiple memories and multiple types of memories, as desired. Additionally, multiple computing devices 700 may be connected together (eg, as a bank of servers, a group of blade servers, or a multi-processor system) with each device providing a portion of the required operations.
[0072] The memory 720 stores information non-temporarily within the computing device 700. The memory 720 (e.g., memory hardware 12, 134 in FIG. 1) may be a computer-readable medium, a volatile memory unit(s), or a non-volatile memory unit(s). The non-volatile memory 720 may be a physical device used to temporarily or permanently store programs (e.g., instruction sequences) or data (e.g., program state information) for use by the computing device 700. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase-change memory (PCM), and disk or tape.
[0073] Storage device 730 can provide mass storage for computing device 700. In some embodiments, storage device 730 is a computer-readable medium. In various different implementations, storage device 730 may be an array of devices, including a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or a storage area network or other configuration of devices. In further embodiments, a computer program product is tangibly embodied on an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as memory 720, storage device 730, or memory on processor 710.
[0074] The high-speed controller 740 manages the bandwidth-intensive operations of the computing device 700, while the low-speed controller 760 manages the bandwidth-intensive operations. This allocation of roles is merely exemplary. In some implementations, the high-speed controller 740 is coupled to the memory 720, the display 780 (e.g., via a graphics processor or accelerator), and the high-speed expansion port 750, which can accept various expansion cards (not shown). In some implementations, the low-speed controller 760 is coupled to the storage device 730 and the low-speed expansion port 790. The low-speed expansion port 790 may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet) and may be coupled to one or more input / output devices such as a keyboard, pointing device, scanner, etc., or to a network device such as a switch or router, e.g., via a network adapter.
[0075] Computing device 700 may be implemented in a number of different forms, as shown in the figure: for example, it may be implemented as a standard server 700a or as a repetition within a group of such servers 700a, as a laptop computer 700b, or as part of a rack server system 700c.
[0076] Various implementations of the systems and techniques described herein may be realized in digital electronic and / or optical circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may be specialized or general-purpose, and may include implementation in one or more computer programs executable and / or interpretable by a programmable system including at least one programmable processor, at least one input device, and at least one output device, coupled to receive data and instructions from, and transmit data and instructions to, the storage system.
[0077] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform tasks. In some examples, a software application may be referred to as an "application," "app," or "program." Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0078] Non-transitory memory may be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by a computing device. Non-transitory memory may be volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.
[0079] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language, and / or assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0080] The processes and logic flows described herein may be implemented by one or more programmable processors, also referred to as data processing hardware, executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be implemented by special-purpose logic circuitry, such as an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit). Processors suitable for executing computer programs include, for example, both general-purpose and special-purpose processors, as well as any one or more processors of any kind of digital computer. Generally, a processor receives instructions and data from a read-only memory, a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or is operably connected to receive data from or transmit data to them, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0081] To interact with a user, aspects of the present invention can be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, or a touch screen, for displaying information to the user, and optionally a keyboard and pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user. For example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user can be received in any form, such as acoustic input, speech input, or tactile input. Additionally, a computer can interact with a user by sending and receiving documents to devices used by the user, for example, by sending a web page to a web browser on the user's client device in response to a request received from the web browser.
[0082] Although several embodiments have been described, it will be understood that various modifications can be made without departing from the spirit and scope of the present disclosure. Accordingly, other embodiments are within the scope of the following claims.
Claims
1. A computer-implemented method (600) that, when executed by data processing hardware (10, 132), causes the data processing hardware (10, 132) to perform an operation, the operation comprising: Receiving a first query (106) issued by a first user (102a), the first query (106) specifying a first long-term action (111) to be performed by the digital assistant (105), the action further comprising: While the digital assistant (105) is performing the first long-term operation (111), receiving a second query (146), the second query (146) specifying a second long-term action (112) to be performed by the digital assistant (105), the action further comprising: determining that the second query (146) was issued by a user (102b) different from the first user (102a); determining, based on determining that the second query (146) was received from the other user (102b), that performing the second long-term operation (112) using a query resolver (340) would conflict with the first long-term operation (111); Identifying one or more compromise actions (354) for the digital assistant (105) to perform based on determining that performing the second long-term action (112) would conflict with the first long-term action (111); Instructing the digital assistant (105) to perform a selected compromise action (354) among the identified one or more compromise actions (354); A computer-implemented method (600) comprising:
2. Receiving the second query (146) includes receiving audio data (402) corresponding to the second query (146), the second query (146) being spoken by the other user (102b) and captured by an assistant-enabled device (104) running the digital assistant (105); 2. The method of claim 1, wherein determining that the second query was issued by a different user than the first user includes performing speaker identification on the audio data to determine that the second query was spoken by the different user than the first user who issued the query.
3. performing speaker identification on the audio data to determine that the second query was spoken by the other user, comprising: extracting a speaker identification vector (412) from the audio data (402) corresponding to the second query (146), the speaker identification vector (412) being characteristic of the second query (146); The speaker identification vector (412) extracted from the audio data (402) corresponding to the second query (146) is does not match the reference speaker vector (155) of the first user (102a), or Does it match the enrollment speaker vector (154) associated with the other user (102b)? and determining that the at least one of:
4. 4. The method of claim 1, wherein receiving the first query issued by the first user comprises receiving, from a user device associated with the first user, a user input indication indicating a user intent to issue the first query.
5. Receiving the first query (106) issued by the first user (102a) includes receiving audio data (402) corresponding to the first query (106) spoken by the first user (102a) and captured by an assistant-enabled device (104) running the digital assistant (105). A method (600) according to any one of claims 1 to 4.
6. Identifying the one or more compromise actions (354) for the digital assistant (105) to perform includes: identifying criteria (113) associated with the first query (106); identifying criteria (115) associated with the second query (146); generating a first query embedding (502) based on the criteria (113) associated with the first query (106) using a query embedding model (352); generating a second query embedding (504) based on the criteria (115) associated with the second query (146) using the query embedding model (352); determining a combined embedding (506) based on the first query embedding (502) and the second query embedding (504); identifying at least one compromise operation (354) in an embedding space (500) that maps to the combined embedding (506); The method (600) of any of claims 1 to 5, comprising:
7. The identified criteria (113) associated with the first query (106) include a first preference (512) for a type of media content for playback from an assistant-enabled device (104) running the digital assistant (105); the identified criteria (115) associated with the second query (146) include a second preference (514) for the type of media content for playback from the assistant-enabled device (104); The identified at least one compromise action (354) includes a third preference (516) for the type of media content for playback from the assistant-enabled device (104). The method (600) of claim 6.
8. the type of media content includes music; the first preference (512) for the type of media content comprises a first genre of music; 8. The method of claim 7, wherein the second preference of the type of media content comprises a second genre of music.
9. the identified criteria (113) associated with the first query (106) include a first value for a setting of a home automation device (103); the identified criteria (115) associated with the second query (146) includes a second value for the setting of the home automation device (103); The method (600) according to any of claims 6 to 8, wherein the identified at least one compromise action (354) comprises adjusting the first value for the setting of the home automation device (103) to a new value.
10. 10. The method (600) of claim 9, wherein the home automation device (103) comprises a smart thermostat, a smart light, a smart speaker, or a smart display.
11. The operation is Obtaining a home graph showing at least two assistant-enabled devices (103) that are in the same environment as the first user (102a) and the other user (102b) and that are capable of performing the first long-term action (111) and the second long-term action (112); further comprising Identifying the one or more compromise actions (354) to be performed by the digital assistant (105) includes identifying a first assistant-enabled device (103) from the home graph as a candidate for the digital assistant (105) to perform the first long-term action (111), and identifying a second assistant-enabled device (103) from the home graph as a candidate for the digital assistant (105) to simultaneously perform the second long-term action (112) while the digital assistant (105) is performing the long-term action (111) on the first assistant-enabled device (103). A method (600) according to any one of claims 1 to 10.
12. The operation is Obtaining proximity information (107) of each of the at least two AEDs (103) in the same environment as the first user (102a) and the other user (102b) from the home graph; Obtaining proximity information (107) of the first user (102a) who issued the first query (106) and the other user (102b) who issued the second query (146); further comprising The method (600) of claim 11, wherein identifying the first assistant-enabled device (103) from the home graph as a candidate for the digital assistant (105) to perform the first long-term operation (111) and identifying the second assistant-enabled device (103) from the home graph as a candidate for the digital assistant (105) to simultaneously perform the second long-term operation (112) are based on the proximity information (107) of each of the at least two AEDs (103) and the proximity information (107) of each of the first user (102a) and the other user (102b).
13. The digital assistant (105) executes the first long-term action (111) on a first assistant-enabled device (103); Instructing the digital assistant (105) to perform the selected compromise operation (354) includes instructing the digital assistant (105) to simultaneously perform the second long-term operation (112) on the second assistant-enabled device (103) while the digital assistant (105) is performing the first long-term operation (111) on the first assistant-enabled device (103), The method (600) according to any of claims 1 to 12.
14. 14. The method of claim 13, wherein the operation further includes, after instructing the digital assistant to perform the second long-term operation on the second assistant-enabled device, instructing the digital assistant to adjust performance of the first long-term operation on the first assistant-enabled device.
15. Once multiple compromise actions (354) for the digital assistant (105) to perform are identified, the actions include: determining a respective score associated with each compromise action (354) in the plurality of compromise actions (354); selecting the compromise action (354) among the plurality of compromise actions (354) as the compromise action (354) having the highest respective score; The method (600) of any of claims 1 to 14, further comprising:
16. The operation is determining that the respective scores associated with the selected compromise actions (354) meet a threshold; further comprising 16. The method (600) of claim 15, wherein instructing the digital assistant (105) to perform the compromise action (354) is based on the respective scores associated with the selected compromise action (354) satisfying the threshold.
17. The operation is Prompting the first user (102a) and / or the other user (102b) to provide confirmation that the digital assistant (105) will perform the selected compromise action (354); Receiving a positive confirmation from the first user (102a) and / or the other user (102b) that the digital assistant (105) will perform the selected compromise action (354); further comprising Instructing the digital assistant (105) to perform the selected compromise action (354) is based on the received positive confirmation. The method (600) according to any of the preceding claims.
18. A system (100), comprising: data processing hardware (10, 132); and memory hardware (12, 134) in communication with the data processing hardware (10, 132), the memory hardware (12, 134) storing instructions that, when executed by the data processing hardware (10, 132), cause the data processing hardware (10, 132) to perform operations, the operations comprising: Receiving a first query (106) issued by a first user (102a), the first query (106) specifying a first long-term action (111) to be performed by the digital assistant (105), the action further comprising: While the digital assistant (105) is performing the first long-term operation (111), receiving a second query (146), the second query (146) specifying a second long-term action (112) to be performed by the digital assistant (105), the action further comprising: determining that the second query (146) was issued by a user (102b) different from the first user (102a); determining, based on determining that the second query (146) was received from the other user (102b), that performing the second long-term operation (112) using a query resolver (340) would conflict with the first long-term operation (111); Identifying one or more compromise actions (354) for the digital assistant (105) to perform based on determining that performing the second long-term action (112) would conflict with the first long-term action (111); Instructing the digital assistant (105) to perform a selected compromise action (354) among the identified one or more compromise actions (354); A system (100) comprising:
19. Receiving the second query (146) includes receiving audio data (402) corresponding to the second query (146), the second query (146) being spoken by the other user (102b) and captured by an assistant-enabled device (104) running the digital assistant (105); 20. The system of claim 18, wherein determining that the second query was issued by a different user than the first user includes performing speaker identification on the audio data to determine that the second query was spoken by the different user than the first user who issued the query.
20. performing speaker identification on the audio data to determine that the second query was spoken by the other user, comprising: extracting a speaker identification vector (412) from the audio data (402) corresponding to the second query (146), the speaker identification vector (412) being characteristic of the second query (146); The speaker identification vector (412) extracted from the audio data (402) corresponding to the second query (146) is does not match the reference speaker vector (155) of the first user (102a), or Does it match the enrollment speaker vector (154) associated with the other user (102b)? and determining that the at least one of:
21. 21. The system (100) of claim 18, wherein receiving the first query (106) issued by the first user (102a) comprises receiving, from a user device (50) associated with the first user (102a), a user input indication indicating the user's intent to issue the first query (106).
22. Receiving the first query (106) issued by the first user (102a) includes receiving audio data (402) corresponding to the first query (106) spoken by the first user (102a) and captured by an assistant-enabled device (104) running the digital assistant (105). The system (100) described in any one of claims 18 to 21.
23. Identifying the one or more compromise actions (354) for the digital assistant (105) to perform includes: identifying criteria (113) associated with the first query (106); identifying criteria (115) associated with the second query (146); generating a first query embedding (502) based on the criteria (113) associated with the first query (106) using a query embedding model (352); generating a second query embedding (504) based on the criteria (115) associated with the second query (146) using the query embedding model (352); determining a combined embedding (506) based on the first query embedding (502) and the second query embedding (504); identifying at least one compromise operation (354) in an embedding space (500) that maps to the combined embedding (506); A system (100) according to any of claims 18 to 22, comprising:
24. The identified criteria (113) associated with the first query (106) include a first preference (512) for a type of media content for playback from an assistant-enabled device (104) running the digital assistant (105); the identified criteria (115) associated with the second query (146) include a second preference (514) for the type of media content for playback from the assistant-enabled device (104); The identified at least one compromise action (354) includes a third preference (516) for the type of media content for playback from the assistant-enabled device (104).
24. The system (100) of claim 23.
25. the type of media content includes music; the first preference (512) for the type of media content comprises a first genre of music; 25. The system of claim 24, wherein the second preferences of the type of media content include a second genre of music.
26. the identified criteria (113) associated with the first query (106) include a first value for a setting of a home automation device (103); the identified criteria (115) associated with the second query (146) includes a second value for the setting of the home automation device (103); The system (100) of any of claims 23 to 25, wherein the identified at least one compromise action (354) comprises adjusting the first value for the setting of the home automation device (103) to a new value.
27. 27. The system (100) of claim 26, wherein the home automation device (103) comprises a smart thermostat, a smart light, a smart speaker, or a smart display.
28. The operation is Obtaining a home graph showing at least two assistant-enabled devices (103) that are in the same environment as the first user (102a) and the other user (102b) and that are capable of performing the first long-term action (111) and the second long-term action (112); further comprising Identifying the one or more compromise operations (354) to be performed by the digital assistant (105) includes identifying a first assistant-enabled device (103) from the home graph as a candidate for the digital assistant (105) to perform the first long-term operation (111), and identifying a second assistant-enabled device (103) from the home graph as a candidate for the digital assistant (105) to simultaneously perform the second long-term operation (112) while the digital assistant (105) is performing the long-term operation (111) on the first assistant-enabled device (103). The system (100) of any one of claims 18 to 27.
29. The operation is Obtaining proximity information (107) of each of the at least two AEDs (103) in the same environment as the first user (102a) and the other user (102b) from the home graph; Obtaining proximity information (107) of the first user (102a) who issued the first query (106) and the other user (102b) who issued the second query (146); further comprising 29. The system (100) of claim 28, wherein identifying the first assistant-enabled device (103) from the home graph as a candidate for the digital assistant (105) to perform the first long-term operation (111) and identifying the second assistant-enabled device (103) from the home graph as a candidate for the digital assistant (105) to simultaneously perform the second long-term operation (112) are based on the proximity information (107) of each of the at least two AEDs (103) and the proximity information (107) of each of the first user (102a) and the other user (102b).
30. The digital assistant (105) executes the first long-term action (111) on a first assistant-enabled device (103); Instructing the digital assistant (105) to perform the selected compromise operation (354) includes instructing the digital assistant (105) to simultaneously perform the second long-term operation (112) on the second assistant-enabled device (103) while the digital assistant (105) is performing the first long-term operation (111) on the first assistant-enabled device (103). A system (100) as described in any one of claims 18 to 29.
31. The system (100) of claim 30 further includes, after instructing the digital assistant (105) to perform the second long-term operation (112) on the second assistant-enabled device (103), instructing the digital assistant (105) to adjust performance of the first long-term operation (111) on the first assistant-enabled device (103).
32. Once multiple compromise actions (354) for the digital assistant (105) to perform are identified, the actions include: determining a respective score associated with each compromise action (354) in the plurality of compromise actions (354); selecting the compromise action (354) among the plurality of compromise actions (354) as the compromise action (354) having the highest respective score; The system (100) of any of claims 18 to 31, further comprising:
33. The operation is determining that the respective scores associated with the selected compromise actions (354) meet a threshold; further comprising 33. The system (100) of claim 32, wherein instructing the digital assistant (105) to perform the compromise action (354) is based on the respective scores associated with the selected compromise action (354) satisfying the threshold.
34. The operation is Prompting the first user (102a) and / or the other user (102b) to provide confirmation that the digital assistant (105) will perform the selected compromise action (354); Receiving a positive confirmation from the first user (102a) and / or the other user (102b) that the digital assistant (105) will perform the selected compromise action (354); further comprising A system (100) according to any one of claims 18 to 33, wherein instructing the digital assistant (105) to perform the selected compromise action (354) is based on the received positive confirmation.
Citation Information
Patent Citations
Voice recognition robot system, voice recognition robot, controller for voice recognition robot, communication terminal for controlling voice recognition robot, and program
JP2016090655A
Apparatus and method for processing control command based on voice agent, and agent device
JP2017076393A
Information providing apparatus, information providing method, and vehicle
JP2019003244A
Resolving conflicting commands received by an electronic device
US20200177410A1