Handling contradictory queries on shared device
By using query digester and trade-off operation identification and execution methods in a multi-user shared environment, the conflict problem of digital assistants when handling multiple user queries is solved, and system stability and user experience are improved.
Patent Information
- Application Number
- CN202380071340.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-06
- Filing Date
- 2023-10-03
- Publication Date
- 2025-05-06
AI Technical Summary
In a multi-user shared environment, the digital assistant needs to handle possible competing queries issued by multiple users, resulting in operational conflicts and resource competition.
By using the query digester to determine that performing the second long-term operation will conflict with the first long-term operation, and identifying the tradeoff operation to be performed by the digital assistant, instructing the digital assistant to perform the identified tradeoff operation to resolve the conflict.
It effectively solves the problem of query conflicts in multi-user environments, reduces unnecessary resource overwriting and interrupts of user requests, and improves system stability and user experience.
Smart Images

Figure CN119948477A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to handling contradictory queries on a shared device. Background Art
[0002] The manner in which users interact with assistant-enabled devices is designed primarily, if not exclusively, with the aid of voice input. For example, a user may ask a device to perform an action involving media playback (e.g., music or a podcast), where the device responds by initiating playback of audio matching the user's criteria. In instances where a device (e.g., a smart speaker) is shared by multiple users in an environment, the device may need to cope with multiple actions requested by users that may compete with each other. Summary of the invention
[0003] One aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations, the operations including: receiving a first query issued by a first user, the first query specifying a first long-term operation to be performed by a digital assistant. While the digital assistant is performing the first long-term operation, the operations further include: receiving a second query, the second query specifying a second long-term operation to be performed by the digital assistant; and determining that the second query is issued by another user different from the first user. Based on determining that the second query is received from another user, the operations further include: using a query resolver to determine that performing the second long-term operation will conflict with the first long-term operation; and based on determining that performing the second long-term operation will conflict with the first long-term operation, identifying one or more compromise operations to be performed by the digital assistant. The operations further include: instructing the digital assistant to perform a selected compromise operation from among the one or more identified compromise operations.
[0004] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, receiving a second query includes receiving audio data corresponding to the second query, the second query being spoken by another user and captured by a device that supports the digital assistant, and determining that the second query is issued by another user different from the first user includes performing speaker recognition on the audio data to determine that the second query is spoken by another user different from the first user who issued the query. In these implementations, performing speaker recognition on the audio data to determine that the second query is spoken by another user includes extracting a speaker identification vector representing the characteristics of the second query from the audio data corresponding to the second query, and determining that the speaker identification vector extracted from the audio data corresponding to the second query is at least one of the following: not matching a reference speaker vector of the first user, or matching a registered speaker vector associated with another user. In some examples, receiving a first query issued by a first user includes receiving a user input indication from a user device associated with the first user indicating the user's intention to issue the first query. Additionally or alternatively, receiving a first query issued by a first user includes receiving audio data corresponding to the first query spoken by the first user and captured by a device that supports the digital assistant.
[0005] In some implementations, identifying one or more tradeoff operations for the digital assistant to perform includes: identifying criteria associated with a first query; identifying criteria associated with a second query; generating a first query embedding based on the criteria associated with the first query using a query embedding model; and generating a second query embedding based on the criteria associated with the second query using a query embedding model. The operations also include: determining a combined embedding based on the first query embedding and the second query embedding; and identifying at least one tradeoff operation mapped to the combined embedding in the embedding space. In these implementations, the identified criteria associated with the first query may include a first preference for the type of media content played back from an assistant-supported device that executes the digital assistant, the identified criteria associated with the second query may include a second preference for the type of media content played back from the assistant-supported device, and the identified at least one tradeoff operation may include a third preference for the type of media content played back from the assistant-supported device. Additionally, the type of media content may include music, wherein the first preference for the type of media content includes a first music genre, and the second preference for the type of media content includes a second music genre. Alternatively, the identified criteria associated with the first query includes a first value for a setting of a home automation device, the identified criteria associated with the second query includes a second value for the setting of the home automation device, and the identified at least one tradeoff operation includes adjusting the first value for the setting of the home automation device to a new value. Here, the home automation device may include a smart thermostat, a smart light, a smart speaker, or a smart display.
[0006] In some examples, the operations further include obtaining a family map indicating at least two assistant-supporting devices that are within the same environment as the first user and the other user and that are capable of performing a first long-term operation and a second long-term operation. Here, identifying one or more compromise operations to be performed by the digital assistant includes: identifying the first assistant-supporting device from the family map as a candidate for causing the digital assistant to perform the first long-term operation, and identifying the second assistant-supporting device from the family map as a candidate for causing the digital assistant to perform the second long-term operation while the digital assistant performs the long-term operation on the first assistant-supporting device. In these examples, the operations may further include: obtaining proximity information of each of at least two AEDs (assistant-supporting devices) within the same environment as the first user and the other user from the family map, and obtaining proximity information of each of the first user who issued the first query and the other user who issued the second query. In these examples, identifying a first assistant-supporting device from the household graph as a candidate for causing the digital assistant to perform a first long-term operation and identifying a second assistant-supporting device from the household graph as a candidate for causing the digital assistant to simultaneously perform a second long-term operation is based on proximity information for each of at least two AEDs and proximity information for each of the first user and the other user.
[0007] In some implementations, the digital assistant performs a first long-term operation on a first support assistant's device; and instructing the digital assistant to perform the selected compromise operation includes instructing the digital assistant to perform a second long-term operation on a second support assistant's device while the digital assistant is performing the first long-term operation on the first support assistant's device. In these examples, after instructing the digital assistant to perform the second long-term operation on the second support assistant's device, the operations may further include instructing the digital assistant to adjust the execution of the first long-term operation on the first support assistant's device. In some implementations, when multiple compromise operations are identified for the digital assistant to perform, the operations further include: determining a corresponding score associated with each of the multiple compromise operations, and selecting the compromise operation with the highest corresponding score among the multiple compromise operations as the compromise operation. In these implementations, the operations may further include determining that the corresponding score associated with the selected compromise operation satisfies a threshold. Here, instructing the digital assistant to perform the compromise operation is based on the corresponding score associated with the selected compromise operation satisfying the threshold. In some examples, the operations further include prompting the first user and / or other users to provide confirmation that the digital assistant is to perform the selected compromise operation, and receiving affirmative confirmation from the first user and / or other users that the digital assistant is to perform the selected compromise operation, and wherein instructing the digital assistant to perform the selected compromise operation is based on the received affirmative confirmation.
[0008] Another aspect of the present disclosure provides a system that includes data processing hardware and memory hardware that communicates with the data processing hardware. The memory hardware stores instructions that, when executed by the data processing hardware, cause the data processing hardware to perform operations, including receiving a first query issued by a first user, the first query specifying a first long-term operation to be performed by the digital assistant. When the digital assistant is performing the first long-term operation, the operations also include: receiving a second query that specifies a second long-term operation to be performed by the digital assistant; and determining that the second query is issued by another user different from the first user. Based on determining that the second query is received from another user, the operations also include: using a query resolver to determine that performing the second long-term operation will conflict with the first long-term operation; and based on determining that performing the second long-term operation will conflict with the first long-term operation, identifying one or more compromise operations to be performed by the digital assistant. The operations further include: instructing the digital assistant to perform a selected compromise operation from the one or more identified compromise operations.
[0009] This aspect may include one or more of the following optional features. In some implementations, receiving a second query includes receiving audio data corresponding to the second query, the second query being spoken by another user and captured by a device that supports the digital assistant, and determining that the second query is issued by another user different from the first user includes performing speaker recognition on the audio data to determine that the second query is spoken by another user different from the first user who issued the query. In these implementations, performing speaker recognition on the audio data to determine that the second query is spoken by another user includes extracting a speaker identification vector representing characteristics of the second query from the audio data corresponding to the second query, and determining that the speaker identification vector extracted from the audio data corresponding to the second query is at least one of the following: not matching a reference speaker vector of the first user, or matching a registered speaker vector associated with another user. In some examples, receiving a first query issued by the first user includes receiving a user input indication from a user device associated with the first user indicating the user's intention to issue the first query. Additionally or alternatively, receiving a first query issued by the first user includes receiving audio data corresponding to the first query spoken by the first user and captured by a device that supports the digital assistant.
[0010] In some implementations, identifying one or more tradeoff operations for the digital assistant to perform includes: identifying criteria associated with a first query; identifying criteria associated with a second query; generating a first query embedding based on the criteria associated with the first query using a query embedding model; and generating a second query embedding based on the criteria associated with the second query using a query embedding model. The operations also include: determining a combined embedding based on the first query embedding and the second query embedding; and identifying at least one tradeoff operation mapped to the combined embedding in the embedding space. In these implementations, the identified criteria associated with the first query may include a first preference for the type of media content played back from an assistant-supported device that executes the digital assistant, the identified criteria associated with the second query may include a second preference for the type of media content played back from the assistant-supported device, and the identified at least one tradeoff operation may include a third preference for the type of media content played back from the assistant-supported device. Additionally, the type of media content may include music, wherein the first preference for the type of media content includes a first music genre, and the second preference for the type of media content includes a second music genre. Alternatively, the identified criteria associated with the first query includes a first value for a setting of a home automation device, the identified criteria associated with the second query includes a second value for the setting of the home automation device, and the identified at least one tradeoff operation includes adjusting the first value for the setting of the home automation device to a new value. Here, the home automation device may include a smart thermostat, a smart light, a smart speaker, or a smart display.
[0011] In some examples, the operations further include obtaining a family map indicating at least two assistant-supporting devices that are within the same environment as the first user and the other user and that are capable of performing a first long-term operation and a second long-term operation. Here, identifying one or more compromise operations to be performed by the digital assistant includes: identifying the first assistant-supporting device from the family map as a candidate for causing the digital assistant to perform the first long-term operation, and identifying the second assistant-supporting device from the family map as a candidate for causing the digital assistant to perform the second long-term operation while the digital assistant performs the long-term operation on the first assistant-supporting device. In these examples, the operations may further include: obtaining proximity information of each of at least two AEDs (assistant-supporting devices) within the same environment as the first user and the other user from the family map, and obtaining proximity information of each of the first user who issued the first query and the other user who issued the second query. In these examples, identifying a first assistant-supporting device from the household graph as a candidate for causing the digital assistant to perform a first long-term operation and identifying a second assistant-supporting device from the household graph as a candidate for causing the digital assistant to simultaneously perform a second long-term operation is based on proximity information for each of at least two AEDs and proximity information for each of the first user and the other user.
[0012] In some implementations, the digital assistant performs a first long-term operation on a first support assistant's device; and instructing the digital assistant to perform the selected compromise operation includes instructing the digital assistant to perform a second long-term operation on a second support assistant's device while the digital assistant is performing the first long-term operation on the first support assistant's device. In these examples, after instructing the digital assistant to perform the second long-term operation on the second support assistant's device, the operations may further include instructing the digital assistant to adjust the execution of the first long-term operation on the first support assistant's device. In some implementations, when multiple compromise operations are identified for the digital assistant to perform, the operations further include: determining a corresponding score associated with each of the multiple compromise operations, and selecting the compromise operation with the highest corresponding score among the multiple compromise operations as the compromise operation. In these implementations, the operations may further include determining that the corresponding score associated with the selected compromise operation satisfies a threshold. Here, instructing the digital assistant to perform the compromise operation is based on the corresponding score associated with the selected compromise operation satisfying the threshold. In some examples, the operations further include prompting the first user and / or other users to provide confirmation that the digital assistant is to perform the selected compromise operation, and receiving affirmative confirmation from the first user and / or other users that the digital assistant is to perform the selected compromise operation, and wherein instructing the digital assistant to perform the selected compromise operation is based on the received affirmative confirmation.
[0013] The details of one or more implementations of the present disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figures 1A to 1C is a schematic diagram of an example system including multiple users controlling an assistant-enabled device.
[0015] Figure 2A and Figure 2B is an example graphical user interface rendered on a screen of a user's device to display long-term operations.
[0016] Figure 3 is a schematic diagram of the query handling process.
[0017] Figure 4A is a schematic diagram of the speaker identification process.
[0018] Figure 4B is a schematic diagram of the speaker verification process.
[0019] Figure 5 is the embedding space used for query processing.
[0020] Figure 6 is a flow diagram of an example operational arrangement of a method for handling voice queries in an environment with multiple users.
[0021] Figure 7 is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein.
[0022] Like reference numerals in the various drawings indicate like elements. DETAILED DESCRIPTION
[0023] The manner in which users interact with assistant-enabled devices is designed primarily, if not exclusively, with the aid of voice input. For example, a user may ask a device to perform an action that includes media playback (e.g., music or a podcast), where the device responds by initiating playback of audio that matches the user's criteria. In instances where a device (e.g., a smart speaker) is shared by multiple users in an environment, the device may need to deal with multiple actions requested by users that may compete with each other. In the event that one or more of the multiple users issues multiple separate requests to the device, subsequent requests may override existing operations being performed by the device. Rather than overwriting previous requests, a device may attempt to offer multiple users a compromise that includes the preferences of each of the multiple users. By ensuring that each user in the environment has been considered, the frequency of unnecessarily overriding requests without the consent of the initiating user is reduced.
[0024] In addition to controlling the action of playing music to accommodate conflicting requests, the device can also control other types of media such as podcasts and videos, as well as home automation such as adjusting light levels, controlling air conditioning levels, etc. Similarly, the device can prevent another user from overriding the initial user's request by identifying a compromise before overriding the initial user's request and prompting the user to agree to the compromise. This saves computing resources used to handle conflicting requests, as well as the time required for the initial user to restore the original request when the original request is overridden without consent. This can also be extended to control various aspects of the home connected to the device. For example, the host of a party can set the lighting level during the party to ensure a soothing atmosphere. The host can say "set the lights to 60%". For the duration of the party, the device can prevent or limit the extent to which other participants at the party can adjust the lighting level by incorporating lighting requests from participants into a compromise that the host can agree or reject.
[0025] The device may additionally operate to resolve conflicts between individuals present in the household. For example, the device may assist individuals in the environment in creating shopping lists, thereby ensuring that any conflicting items are resolved by offering the individuals a compromise to add items to the shopping list. For example, the device may recommend items to be added to the shopping list in response to two individuals requesting that they conflict or be similar enough to be combined. Similarly, the device may actively mediate disagreements between individuals. For example, the device may communicate with / prompt individuals with conflicting opinions about a compromise that is appropriate for both individuals, thereby resolving the disagreement.
[0026] Figures 1A to 1C An example system 100a-c is shown for handling queries in an environment having multiple users 102 (102a-n) using a query handler that balances queries detected in the environment from the multiple users 102 by providing a tradeoff. Briefly, and as described in more detail below, a query handler 300 ( Figure 3 ) detects multiple users 102 (102a-b) within the environment and, in response to receiving a first query 106 issued by user 102a, "Ok computer, play Red from my Pop Music playlist", starts playing music 122. While the digital assistant 105 is performing a long-term operation of playing music 122 from the speaker 18 as playback audio, the digital assistant 105 receives a query issued by another user 102b ( Figure 1B) uttered the second query 146 “Play Canon in D”. Because the query handler 300 detects / recognizes that the other user 102b is different from the user 102a, and the second query 146 issued by the user 102b conflicts with the first query 106 issued by the user 102a, the query handler 300 identifies one or more compromise operations 354 (354a–n) for the digital assistant 105 to perform. Figure 3 ).
[0027] The system 100a-100c includes an assistant-enabled device (AED) 104 (i.e., also referred to as a "primary AED 104") and a plurality of auxiliary assistant-enabled devices (AEDs) 103 (103a-n) located throughout the environment. In the example shown, the environment may correspond to a home having a first floor and a second floor, wherein a first smart speaker 104 (i.e., AED 104) is located on the first floor, and a second smart speaker 103a, a smart light 103bc, and a smart thermostat 103c are located on the second floor. However, the AED 104 and / or the auxiliary AED 103 may include other computing devices, such as, but not limited to, a smart phone, a tablet computer, a smart display, a desktop / laptop computer, a smart watch, smart glasses / headsets, a smart appliance, a headset, or a vehicle infotainment device. As shown, a digital assistant 105 is executed on the AED 104, and a plurality of users 102 may interact with the AED by issuing queries including commands to perform long-term actions. The AED 104 includes data processing hardware 10 and memory hardware 12 that stores instructions that, when executed on the data processing hardware 10, cause the data processing hardware 10 to perform operations. The AED 104 includes an array of one or more microphones 16 that are configured to capture acoustic sounds, such as speech directed toward the AED 104. The AED 104 may also include or communicate with an audio output device (e.g., a speaker) 18 that may output audio, such as music 122 and / or synthesized speech from a digital assistant 105. Additionally, the AED 104 may include or communicate with one or more cameras 19 that are configured to capture images within an environment and output image data 312 ( Figure 3 ).
[0028] In some implementations, each auxiliary AED 103 broadcast may be received by an environmental detector 310 ( Figure 3) received by the digital assistant 105 executing on the AED 104, which can be used by the digital assistant 105 to determine the presence of each auxiliary AED 103. The digital assistant 105 can additionally use the proximity information 107 of each auxiliary AED 103 to infer a home graph to understand the spatial proximity of each auxiliary AED 103 relative to each other and relative to the AED 104 executing the digital assistant 105 (e.g., for determining which auxiliary AEDs 103 can be split / partitioned to perform conflicting long-term operations). The proximity information 107 from each auxiliary AED 103 can include wireless communication signals, such as WiFi, Bluetooth, or ultrasound, wherein the signal strength of the wireless communication signals received by the environment detector 310 can be related to the proximity (e.g., distance) of the auxiliary AED 103 relative to the AED 104. Additionally or alternatively, proximity information 107 from each secondary AED 103 may be determined by playing a stationary sound at the AED 104 to determine which of the secondary AEDs 103 in the environment detected the stationary sound in order to establish an approximate distance between the AED 104 and the secondary AEDs 103 .
[0029] In some configurations, the digital assistant 105 communicates with multiple user devices 50 (50a-n) associated with multiple users 102. In the example shown, each user device 50 in the multiple user devices 50a-c includes a smart phone with which the corresponding user 102 can interact. However, the user device 50 may include other computing devices, such as but not limited to a smart watch, a smart display, smart glasses, a smart phone, smart glasses / headsets, a tablet computer, a smart appliance, a headset, a computing device, a smart speaker, or another assistant-enabled device. Each user device 50 in the multiple user devices 50a-n may include at least one microphone 52 (52a-n) resident on the user device 50 that communicates with the digital assistant 105. In these configurations, the user device 50 may also communicate with one or more microphones 16 resident on the AED 104. Additionally, multiple users 102 may control and / or configure the AED 104 and auxiliary AED 103 , as well as interact with the digital assistant 105 using an interface 200 , such as a graphical user interface (GUI) 200 rendered for display on a respective screen of each user device 50 .
[0030] like Figures 1A to 1C and Figure 3As shown, the digital assistant 105 implementing the query handler 300 uses a tradeoff generator 350 to manage queries issued by multiple users 102. In some implementations, the query handler 300 includes a query resolver 340 that identifies conflicts between one or more queries received from each of the users 102 in the environment and the tradeoff generator 350 that executes a tradeoff model 352 that identifies tradeoff operations 354 to be performed by the digital assistant 105. In this sense, the query handler 300 balances the competing interests of the multiple users 102 while minimizing the frequency with which a user 102's query is overwritten / interrupted by subsequent queries issued by other users 102 in the environment.
[0031] refer to Figure 2A and Figure 2B , the GUI 200a, 200b of the user device 50 associated with the user 102 can display the current long-term operation 111 (e.g., playing music 122) to keep each user 102 aware of the active long-term operation being performed by the digital assistant 105. In some configurations, the AED 104 includes a screen, and the GUI 200 is rendered to display the active long-term operation on the screen. For example, the AED 104 may include a smart display, a tablet computer, or a smart TV within the environment. Figure 2A An example GUI 200a is provided for display on a screen of a user device 50 associated with a user 102, which may additionally render for display: an identifier of a current long-term operation 111 (e.g., "Playing Red"), an identifier of an AED 104 that is currently performing the long-term operation (e.g., Smart Speaker 1), an indication of the next song 202 to be played during the long-term operation (e.g., Playing Next), and / or the identity of the user 102a (e.g., Barb) who initiated the current long-term operation being performed by the digital assistant 105. As described above, the query handler 300 manages the long-term operations such that when the query handler 300 determines that another user 102 has issued a second long-term operation that conflicts with a first long-term operation issued by user 102a (e.g., Barb), the digital assistant 105 will prevent the execution of the second action (or at least require a compromise that attempts to achieve the second action).
[0032] Reference again Figures 1A to 1C, the environment detector 310 can identify the user 102a (e.g., Barb) and the user 102b (e.g., Jeff) via the proximity information 54 received from their respective user devices 50a, 50b. However, in other examples, the environment detector 310 can detect the users 102a, 102b who detect the respective user devices 50a, 50b (e.g., one or more of the users 102a, 102b choose not to share the proximity information 54 from their respective user devices 50a, 50b). However, the environment detector 310 still detects the presence of the users 102a, 102b (e.g., via voice data, image data, and / or input from the users 102a, 102b). In addition, the environment detector can identify the auxiliary AEDs 103a-c via the proximity information 107 received from the respective auxiliary AEDs 103. Thereafter, the environment detector 310 generates a household graph that includes the identities and locations of the users 102a, 102b and the auxiliary AEDs 103 within the current environment 316. The query handler 300 may use the household graph representing the current environment 316 to generate a compromise operation 354 when a conflict arises.
[0033] continue Figure 1A , a user 102a of a plurality of users 102a-c issues a first query 106 "Ok computer, play Red from my Pop Music playlist" near an AED 104. Here, the first query 106 issued by the user 102a is spoken by the user 102a and includes initial audio data 402 ( Figure 3 ). The first query 106 may also include a user input indication indicating a user's intent to issue the first query via any of touch, voice, gesture, gaze, and / or an input device (e.g., a mouse or stylus) for interacting with the AED 104. Optionally, based on receiving initial audio data 402 corresponding to the first query 106, the query handler 300 performs a speaker identification process 400a ( Figure 4A) and determine that the first query 106 was issued by the user 102a to resolve the identity of the speaker of the first query 106. In other implementations, the user 102a issues the first query 106 without speaking. In these implementations, the user 102a may issue the first query 106 via a user device 50a associated with the user 102a (e.g., enter text corresponding to the first query 106 into the GUI 200 displayed on the screen of the user device 50a associated with the user 102a, select the first query 106 displayed on the screen of the user device 50a, etc.). Here, the AED 104 can resolve the identity of the user 102 who issued the first query 106 by recognizing the user device 50a associated with the user 102a.
[0034] The microphone 16 of the AED 104 receives the first query 106 and processes initial audio data 402 corresponding to the first query 106. The initial processing of the audio data 402 may involve filtering the audio data 402 and converting the audio data 402 from an analog signal to a digital signal. As the AED 104 processes the audio data 402, the AED may store the audio data 402 in a buffer of the memory hardware 12 for additional processing. In the case where there is audio data 402 in the buffer, the AED 104 may use the hot word detector 108 to detect whether the audio data 402 includes a hot word. The hot word detector 108 is configured to identify hot words included in the audio data 402 without performing speech recognition on the audio data 402.
[0035] In some implementations, the hotword detector 108 is configured to identify hotwords in the initial portion of the first query 106. In this example, if the hotword detector 108 detects an acoustic feature that is characteristic of the hotword 110 in the audio data 402, the hotword detector 108 may determine that the first query 106 "Ok computer, play Red from my PopMusic playlist" includes the hotword 110 "okcomputer". The acoustic feature may be a Mel-frequency cepstral coefficient (MFCC), which is a representation of a short-term power spectrum of the first query 106, or may be a Mel-scale filter bank energy of the first query 106. For example, the hot word detector 108 may detect that the first query 106 “Ok computer, play Red from my Pop Music playlist” includes the hot word 110 “ok computer” based on generating MFCCs from the audio data 402 and classifying that the MFCCs include MFCCs similar to MFCCs that are characteristics of the hot word “ok computer” stored in the hot word model of the hot word detector 108. As another example, the hot word detector 108 may detect that the first query 106 “Ok computer, play Glory, and let's stick to pop music tonight” includes the hot word 110 “ok computer” based on generating Mel-scale filter bank energies from the audio data 402 and classifying that the Mel-scale filter bank energies include Mel-scale filter bank energies similar to Mel-scale filter bank energies that are characteristics of the hot word “ok computer” stored in the hot word model of the hot word detector 108.
[0036] When the hotword detector 108 determines that the initial audio data 402 corresponding to the first query 106 includes the hotword 110, the AED 104 can trigger a wake-up process to initiate speech recognition of the audio data 402 corresponding to the first query 106. For example, Figure 3The AED 104 is shown to include a speech recognizer 170 employing an automatic speech recognition model 172 that can perform speech recognition or semantic interpretation on the audio data 402 corresponding to the first query 106. The speech recognizer 170 can perform speech recognition on the portion of the audio data 402 that follows the hot word 110. In this example, the speech recognizer 170 can recognize the words "play Glory, and let's stick to pop music tonight" in the first query 106.
[0037] In some examples, the AED 104 is configured to communicate with a remote system 130 via a network 120. The remote system 130 may include remote resources, such as remote data processing hardware 132 (e.g., a remote server or CPU) and / or remote memory hardware 134 (e.g., a remote database or other storage hardware). The query handler 300 may be executed on the remote system 130 in addition to or in lieu of the AED 104. The AED 104 may utilize the remote resources to perform various functionalities related to speech processing and / or synthesized playback communications. In some implementations, the speech recognizer 170 is located on the remote system 130 in addition to or in lieu of the AED 104. After the hotword detector 108 triggers the AED 104 to wake up in response to detecting the hotword 110 in the first query 106, the AED 104 may transmit initial audio data 402 corresponding to the first query 106 to the remote system 130 via the network 120. Here, the AED 104 may transmit the portion of the initial audio data 402 including the hotword 110 for the remote system 130 to confirm the presence of the hotword 110. Alternatively, the AED 104 may transmit only the portion of the initial audio data 402 corresponding to the portion of the utterance 106 following the hotword 110 to the remote system 130, wherein the remote system 130 executes the speech recognizer 170 to perform speech recognition and returns a transcription of the initial audio data 402 to the AED 104.
[0038] Continue to refer Figures 1A to 1C and Figure 3, the query handler 300 may also include a natural language understanding (NLU) module 320 that performs semantic interpretation on the first query 106 to identify queries / commands directed to the AED 104. Specifically, the NLU module 320 identifies words in the first query 106 that are identified by the speech recognizer 170, and performs semantic interpretation to identify any voice commands in the first query 106. The NLU module 320 of the AED 104 (and / or the remote system 130) may identify the words "play Red" as a command to the digital assistant 105 for a specified first long-term operation 111 (i.e., play music 122), and identify the words "from my Pop Music playlist" as a criterion 113 for the digital assistant 105 to play music of a specific genre (e.g., pop music) when performing the first long-term operation 111. Figure 1A In the example shown, the digital assistant 105 begins to perform a first long-term operation 111 of playing music 122 as playback audio (e.g., track #1) from the speaker 18 of the AED 104. The digital assistant 105 can stream the music 122 from a streaming service (not shown), or the digital assistant 105 can instruct the AED 104 to play music stored on the AED 104. Although the example long-term operation 111 includes music playback, the long-term operation can include other types of media playback, such as videos, podcasts, and / or audio books. The long-term operation 111 can also include home automation (e.g., adjusting light levels, controlling thermostats, etc.).
[0039] exist Figure 1A In the example shown, refer to Figure 3 , the query handler 300 adds the long-term operation 111 to the active operation data store 330. The query handler 300 maintains a record of the active operations 332 in the environment (e.g., light intensity in a smart light bulb, temperature set point in a smart thermostat, etc.) in the active operation data store 330 (e.g., stored on the memory hardware 12, 134), and the query handler 300 can limit the long-term operations performed by the digital assistant 105 based on the active operations 332. For example, before executing the first long-term operation 111, the query handler 300 can first use the query resolver 340 to verify that the first long-term operation 111 in the first query 106 does not conflict with any active operations 332 in the environment.
[0040] The AED 104 can notify the user 102a (e.g., Barb) who issued the first query 106 that the first long-term operation 111 is being performed. For example, the digital assistant 105 can generate synthesized speech 123 for audible output from the speaker 18 of the AED 104, the synthesized speech stating "Barb, now playing Red from Pop Music". In an additional example, the digital assistant 105 provides a notification to the user device 50a associated with the user 102a (e.g., Barb) to inform the user 102a of the approved first long-term operation 111 and / or any active operations 332 stored in the active operations data repository 330.
[0041] refer to Figure 2A , a graphical user interface (GUI) 200a executed on the user device 50 can display the first long-term operation 111 and / or any active operations 332 being performed on the auxiliary device 103 within the environment. As used herein, the GUI 200a can receive user input indications via any of touch, voice, gesture, gaze, and / or an input device (e.g., a mouse or stylus) for interacting with the digital assistant 105. For example, Figure 2A The GUI 200a in the embodiment may render the following for display: an identifier of the current long-term operation 111 (e.g., "Playing Red"), an identifier of the AED 104 currently performing the long-term operation 111 (e.g., smart speaker 1), an identifier of the criterion 113 associated with the first query 106 (e.g., Pop Music), an indication of the next song 202 to be played during the long-term operation (e.g., Playing Next), and / or an identity of the user 102a who issued the first query 106 (e.g., Barb). In an implementation where the first query 106 includes the criterion 113 (e.g., pop music), the GUI 200a renders the identifier of the criterion 113 for display. Additionally, the GUI 200a renders an identifier of the smart thermostat 103c and an indication of an active long-term operation 332 with a set point at 70 degrees, and an identifier of the smart light bulb 103b and an indication of an active long-term operation 332 with the light in an on state. Thus, user 102 can consult user device 50 to view active long-term operations within the environment of digital assistant 105 as well as any criteria 113 that may limit queries issued by user 102.
[0042] refer to Figure 4A and Figure 4BIn some implementations, the AED 104 (or a remote system 130 in communication with the AED 104) also includes an example data repository 430 that stores enrollment user data / information for each of a plurality of enrollment users 432a-n of the AED 104. Here, each of the enrollment users 432 of the AED 104 may undergo a voice enrollment process to obtain a corresponding enrollment speaker vector 154 from audio samples of a plurality of enrollment phrases spoken by the enrollment user 432. For example, the speaker recognition model 410 may generate one or more enrollment speaker vectors 154 from audio samples of the enrollment phrases spoken by each of the enrollment users 432, which may be combined, such as averaged or otherwise accumulated, to form a corresponding enrollment speaker vector 154. One or more of the enrollment users 432 may use the AED 104 to undergo a voice enrollment process, wherein the microphone 16 captures audio samples of the users speaking the enrollment utterances, and the speaker recognition model 410 generates corresponding enrollment speaker vectors 154 from the audio samples. Model 410 may be executed on AED 104, remote system 130, or a combination thereof. Additionally, one or more of registered users 432 may register with AED 104 by providing authorization and authentication credentials to an existing user account of AED 104. Here, the existing user account may store an enrollment speaker vector 154 obtained from a previous voice enrollment process of another device that is also linked to the user account.
[0043] In some examples, the registered speaker vectors 154 for the registered users 432 include text-dependent registered speaker vectors. For example, the text-dependent registered speaker vectors may be extracted from one or more audio samples of the corresponding registered user 432, the corresponding registered user speaking a predetermined term, such as a hot word 110 (e.g., "Ok computer") for invoking the AED 104 to wake up from a sleep state. In other examples, the registered speaker vectors 154 for the registered users 432 are text-independent registered speaker vectors obtained from one or more audio samples of the corresponding registered users 102 speaking phrases having different terms / words and different lengths. In these examples, the text-independent registered speaker vectors may be obtained over time from audio samples obtained from voice interactions of the user 102 with the AED 104 or other devices linked to the same account.
[0044] refer to Figure 4A, the speaker identification process 400a identifies the user 102a (e.g., Barb) who spoke the first query 106 by first extracting a first speaker identification vector 411 representing the characteristics of the first query 106 issued by the user 102a from the initial audio data 402 corresponding to the first query 106. Here, the speaker identification process 400a can execute a speaker identification model 410, which is configured to receive the audio data 402 corresponding to the second query 146 as input and generate a first speaker identification vector 411 as output. The speaker identification model 410 can be a neural network model trained under machine or human supervision to output the speaker identification vector 411. The speaker identification vector 411 output by the speaker identification model 410 can include an N-dimensional vector having values corresponding to the speech characteristics of the first query 106 associated with the user 102a. In some examples, the speaker identification vector 411 is a d-vector. In some examples, the first speaker identification vector 411 includes a set of speaker identification vectors, each speaker identification vector in the set of speaker identification vectors being associated with a different user who is also authorized to control the AED 104. For example, in addition to the user 102a who uttered the first query 106, other authorized users may include other individuals who were present when the user 102a uttered the first query 106 (issued the command 111 to perform the first action) and / or individuals that the user 102a added / designated as authorized individuals.
[0045] Once the first speaker identification vector 411 is output from the model 410, the speaker identification process 400a determines whether the extracted speaker identification vector 411 matches any of the enrollment speaker vectors 154 stored on the AED 104 (e.g., stored in the memory hardware 12) for the enrollment users 432a-n of the AED 104. As described above, the speaker identification model 410 can generate the enrollment speaker vectors 154 for the enrollment users 432 during the voice enrollment process. Each enrollment speaker vector 154 can be used as a reference vector 155 corresponding to a voiceprint or unique identifier representing characteristics of the voice of the corresponding enrollment user 432.
[0046] In some implementations, the speaker identification process 400a uses a comparator 420 that compares the first speaker identification vector 411 with the corresponding enrollment speaker vector 154 associated with each enrolled user 432a-n of the AED 104. Here, the comparator 420 can generate a score for each comparison that indicates the likelihood that the initial audio data 402 corresponding to the first query 106 corresponds to the identity of the corresponding enrolled user 432, and when the score meets a threshold, the identity is accepted. When the score does not meet the threshold, the comparator 420 can reject the identity of the speaker who issued the first query 106. In some implementations, the comparator 420 calculates a corresponding cosine distance between the first speaker identification vector 411 and each enrollment speaker vector 154, and determines that the first speaker identification vector 411 matches one of the enrollment speaker vectors 154 when the corresponding cosine distance meets the cosine distance threshold.
[0047] In some examples, the first speaker discriminant vector 411 is a text-dependent speaker discriminant vector extracted from a portion of one or more words corresponding to the first query 106, and each enrollment speaker vector 154 is also a text-dependent enrollment speaker vector on the same one or more words. The use of text-dependent speaker vectors can improve the accuracy of determining whether the first speaker discriminant vector 411 matches any enrollment speaker vector in the enrollment speaker vectors 154. In other examples, the first speaker discriminant vector 411 is a text-independent speaker discriminant vector extracted from the entire initial audio data 402 corresponding to the first query 106.
[0048] When the speaker identification process 400a determines that the first speaker discriminant vector 411 matches one of the enrollment speaker vectors 154, the process 400a identifies the user 102a who spoke the first query 106 as the corresponding registered user 432a associated with the enrollment speaker vector 154 that matches the extracted speaker discriminant vector 411. In the example shown, the comparator 420 determines the match based on the corresponding cosine distance between the first speaker discriminant vector 411 and the enrollment speaker vector 154 associated with the enrollment user 432a satisfying a cosine distance threshold. In some scenarios, the comparator 420 identifies the user 102a as the corresponding registered user 432a associated with the enrollment speaker vector 154 having the shortest corresponding cosine distance from the first speaker discriminant vector 411, provided that the shortest corresponding cosine distance also satisfies the cosine distance threshold.
[0049] Conversely, when speaker identification process 400a determines that first speaker identification vector 411 does not match any of the enrollment speaker vectors 154, process 400a may identify user 102a who uttered utterance 106 as a guest user of AED 104. Accordingly, query handler 300 may add the guest user and use first speaker identification vector 411 as a reference speaker vector 155 representing the speech characteristics of the guest user's voice. In some examples, the guest user may register with AED 104, and AED 104 may store first speaker identification vector 411 as the corresponding enrollment speaker vector 154 for the newly registered user.
[0050] Return to reference Figure 1B , when the digital assistant 105 performs the first long-term operation 111 of playing music 122 from the speaker 18 of the AED 104 as playback audio, the digital assistant 105 receives a second query 146 that specifies a second long-term operation 112 for the digital assistant 105 to perform. In the example shown, another user 102b, different from the user 102a that issued the first query 106, issues the second query 146 "Play Canon in D," which includes a command for the digital assistant 105 to perform the second long-term operation 112 of playing a song (i.e., Canon in D), with associated criteria 115 for the digital assistant 105 to play music of a particular genre (e.g., classical music). Based on receiving the second query 146, the query handler 300 performs a speaker identification process 400b ( Figure 4B ) to resolve the identity of the speaker of the second query 146, and determine that the second query 146 is issued by another user 102b different from the user 102a that issued the first query 106. Figure 4A As described, in an implementation where a first query 106 issued by a user 102a includes initial audio data 402 (e.g., the first query 106 is spoken by the first user 102a), the query handler 300 may first perform a speaker identification process 400a on the initial audio data 402 corresponding to the first query 106 to identify the user 102a who issued the first query 106.
[0051] The speaker identification process 400b may be executed on the data processing hardware 12 of the AED 104. The speaker identification process 400b may also be executed on the remote system 130. If the speaker verification process 400b performed on the audio data 402 corresponding to the second query 146 indicates that the second query 146 was spoken by the same user 102a who issued the first query 106, the digital assistant 105 may proceed to perform the second long-term operation 105 without first determining whether the first long-term operation 111 and the second long-term operation 112 conflict. In other words, when the same user 102a issues both queries 106, 146, the query handler 300 may not be required to resolve the conflict between the users 102. Conversely, if the speaker verification process 400 b performed on the audio data 402 corresponding to the second query 146 indicates that the second query 146 was spoken by another user 102 b than the user 102 a that issued the first query 106, the query handler 300 may prevent the execution of the second long-term operation 112 (or at least require input from one or more other users 102 in the environment (e.g., in Figure 1B and Figure 2B middle)).
[0052] Reference again Figure 4B ,refer to Figure 1B In the example of , in response to receiving the second query 146, the AED 104 resolves the identity of the user 102 who spoke the second query 146 by executing a speaker identification process 400b. The speaker identification process 400b identifies the user 102a who spoke the first query 146 by first extracting a second speaker identification vector 412 representing characteristics of the second query 146 from the audio data 402 corresponding to the first query 146 spoken by the user 102a. Here, the speaker verification process 400b can execute a speaker identification model 410 that is configured to receive the audio data 402 as input and generate a second speaker identification vector 412 as output. As described above in Figure 4A As discussed in , the speaker identification model 410 can be a neural network model trained under machine or human supervision to output a speaker identification vector 412. The second speaker identification vector 412 output by the speaker identification model 410 can include an N-dimensional vector having values corresponding to speech features of the utterance 146 associated with the user 102a. In some examples, the speaker identification vector 412 is a d-vector.
[0053] Once the second speaker identification vector 412 is output from the speaker identification model 410, the speaker verification process 400b determines whether the extracted speaker identification vector 412 matches a reference speaker vector 155 associated with the first enrolled user 432a stored on the AED 104 (e.g., stored in the memory hardware 12). The reference speaker vector 155 associated with the first enrolled user 432a may include a corresponding enrolled speaker vector 154 associated with the first enrolled user 432a. As discussed above, the speaker identification model 410 may generate the enrollment speaker vectors 154 of the enrollment users 432 during the voice enrollment process. Each enrollment speaker vector 154 may be used as a reference vector corresponding to a voiceprint or unique identifier representing characteristics of the voice of the corresponding enrolled user 432.
[0054] In some implementations, the speaker verification process 400b uses a comparator 420 that compares the second speaker identification vector 412 with a reference speaker vector 155 associated with a first registered user 432a in the registered users 432. Here, the comparator 420 can generate a score for the comparison that indicates the likelihood that the second query 146 corresponds to the identity of the first registered user 432a, and when the score meets a threshold, the identity is accepted. When the score does not meet the threshold, the comparator 420 can reject the identity. In some implementations, the comparator 420 calculates a corresponding cosine distance between the second speaker identification vector 412 and the reference speaker vector 155 associated with the first registered user 432a, and when the corresponding cosine distance meets the cosine distance threshold, determines that the second speaker identification vector matches the reference speaker vector 155.
[0055] When the speaker verification process 400b determines that the second speaker identification vector 412 matches the reference speaker vector 155 associated with the first registered user 432a, the process 400b identifies the user 102a who spoke the second query 146 as the first registered user 432a associated with the reference speaker vector 155. In the example shown, the comparator 420 determines the match based on the corresponding cosine distances between the second speaker identification vector 412 and the reference speaker vector 155 associated with the first registered user 432a satisfying the cosine distance threshold. In some scenarios, the comparator 420 identifies the user 102a as the corresponding first registered user 432a associated with the reference speaker vector 155 having the shortest corresponding cosine distance from the second speaker identification vector 412, provided that the shortest corresponding cosine distance also satisfies the cosine distance threshold.
[0056] Refer to the above Figure 4AIn some implementations, the speaker identification process 400a determines that the first speaker discriminant vector 411 does not match any of the enrollment speaker vectors 154, and identifies the user 102a who spoke the first query 106 as a guest user of the AED 104. Therefore, the speaker verification process 400b can first determine whether the user 102a who spoke the first query 106 is identified by the speaker identification process 400a as a registered user 432 or a guest user. When the user 102a is a guest user, the comparator 420 compares the second speaker discriminant vector 412 with the first speaker discriminant vector 411 obtained during the speaker identification process 400a. Here, the first speaker discriminant vector 411 represents the characteristics of the first query 106 spoken by the guest user 102a, and is therefore used as a reference vector to verify whether the second query 146 is also spoken by the guest user 102a or another user 102. Here, the comparator 420 can generate a compared score indicating the likelihood that the second query 146 corresponds to the identity of the guest user 102a, and when the score meets a threshold, the identity is accepted. When the score does not meet the threshold, the comparator 420 can reject the identity of the guest user who spoke the second query 146. In some implementations, the comparator 420 calculates the corresponding cosine distance between the first speaker discriminant vector 411 and the second speaker discriminant vector 412, and determines that the first speaker discriminant vector 411 matches the second speaker discriminant vector 412 when the corresponding cosine distance meets the cosine distance threshold.
[0057] Return to reference Figure 1B , based on determining that the second query 146 was spoken by another user 102b than the user 102a that issued the first query 106, the NLU module 320 executing on the AED 104 (and / or executing on the remote system 130) can recognize the words "play Canon in D" as a command specifying the second long-term operation 112 (i.e., play music 122). Here, before allowing the digital assistant 105 to execute the second long-term operation 112, the query handler 300 can first use the query resolver 340 to determine whether executing the second long-term operation 112 conflicts with the first long-term operation 111. For example, the query handler 300 can determine whether the first long-term operation 111 and the second long-term operation 112 invoke the same function of the digital assistant 105 (e.g., play music) or different functions (e.g., play music and change the brightness setting on the smart light bulb). In an example where the first long-term operation 111 and the second long-term operation 112 invoke the same function, the query resolver 340 determines that the long-term operations conflict and outputs the conflicting long-term operations 111 , 112 for the query handler 300 to determine one or more compromise operations for the users 102a , 102b .
[0058] In some examples, the query resolver 340 outputs only the first long-term operation 111 and the second long-term operation 112 when it determines that the second long-term operation 112 conflicts with the first long-term operation 111 (thereby triggering the query handler 300 to identify one or more compromise operations 354). Conversely, in the case where the first long-term operation 111 and the second long-term operation 112 call different functions, the query resolver 340 determines that the second long-term operation 112 does not conflict with the first long-term operation 111. Here, the query resolver 340 outputs only the long-term operations 111, 112, thereby prompting the query handler 300 to identify one or more compromise operations 354 when there is a conflict representing the competing interests between the user 102a and the user 102b. In addition, as discussed above, before executing the second long-term operation 112, the query resolver 340 can verify that the second long-term operation 112 in the first query 106 does not conflict with the active operation 332 stored in the active operation data repository 300.
[0059] In this example, the query resolver 340 determines that the second long-term operation 112 of playing Canon in D conflicts with the first long-term operation 111 of playing Red because the execution of the second long-term operation 112 via the speaker 18 of the AED 104 must interrupt the execution of the first long-term operation 111 currently playing on the speaker 18 of the AED 104. Based on the determination that the second user 102b issued the second query 146 and the determination that the second query 146 conflicts with the first query 106 issued by the user 102a, the query handler 300 prevents the AED 104 (via the digital assistant 105) from executing the second long-term operation 112 and instead generates one or more compromise operations 354 (354a-n) that the users 102a, 102b can agree on. In other words, after the user 102a is determined to be the issuer of the first query 106 and the user 102b is determined to be the issuer of the second query 146, the query handler 300 attempts to respect the first long-term operation 111 and the second long-term operation 112 by determining a compromise solution.
[0060] Return to reference Figure 3In some implementations, the query handler 300 includes a tradeoff generator 350 that generates one or more tradeoff operations 354 (354a-n) using one or more methods. For example, the tradeoff generator 350 may execute a tradeoff model 352 that is configured to receive a first query 106 and a second query 146 as inputs and generate one or more tradeoff operations 354 that combine the first query 106 and the second query 146 as outputs. For example, if the first query requests "party music" and the second query requests "classical music", the tradeoff generator 350 may output a tradeoff operation to play "upbeat classical music". Similarly, if the first query requests to turn on the lights and the second query requests to turn off the lights, the tradeoff generator 350 may output a tradeoff for a medium lighting level to accommodate the criteria in both queries.
[0061] The tradeoff model 352 may be a neural network model trained under machine or human supervision to output the tradeoff operation 354. In other implementations, the tradeoff generator 350 includes multiple tradeoff models (e.g., some tradeoff models include neural networks, some tradeoff models do not include neural networks). In these implementations, the tradeoff generator 350 may select which of the multiple tradeoff models to use as the tradeoff model 352 based on the category of the action associated with the query.
[0062] Continuing with this example, the tradeoff generator 350 identifies the criteria (e.g., Red) 113 associated with the first query 106 and the criteria (e.g., Canon in D) 115 associated with the second query 146. The tradeoff model 352 receives the criteria 113 associated with the first query 106 and the criteria 115 associated with the second query 146 as inputs and generates a first query embedding 502 ( Figure 5 ) and generating a second query embedding 504 based on the criteria 115 associated with the second query 146 ( Figure 5 ) as output. Figure 3 In the example, refer to Figure 5, the tradeoff model 352 can determine a combined embedding 506 based on the first query embedding 502 and the second query embedding 504, and identify at least one tradeoff operation 354 that is mapped (e.g., pre-mapped or pre-defined) to the combined embedding 506 in the embedding space 500. For example, the tradeoff model 352 averages the first query embedding 502 and the second query embedding 504 to generate the combined embedding 506. In other examples, where the space between the first query embedding 502 and the second query embedding 504 is too large, the tradeoff model 352 generates the tradeoff operation 354 without pre-mapping the tradeoff operation 354 to the combined embedding 506. In some implementations, the combined embedding 506 includes a conflict score that predicts a degree of conflict between the first query embedding 502 and the second query embedding 504.
[0063] In other examples, the identified criteria associated with the first query include a first value for a setting of a home automation device (e.g., a smart thermostat, a smart light, a smart speaker, or a smart display), and the second query includes a second value for the setting of the home automation device. Here, the identified at least one compromise operation 354 includes adjusting the first value of the setting for the home automation device to a new value. For example, the compromise model 352 may include a heuristic model that parses the first value and the second value and determines an average between the first value and the second value to set as the new value. In some examples, the home automation device corresponds to an auxiliary AED 103 in the environment.
[0064] Continue the music playback example, such as Figure 5As shown, an average embedding 506 is determined based on preferences for the query. Here, the identified criteria 113 associated with the first query embedding 502 includes a first preference 512 in the embedding space 500 for the type of media content played back from the AED 104 executing the digital assistant 105 (i.e., pop music). In addition, the identified criteria 115 associated with the second query embedding 504 includes a second preference 514 in the embedding space 500 for the type of media content played back from the AED 104 executing the digital assistant 105 (i.e., classical music). Here, although the queries 106, 146 do not explicitly state the preferences 512, 514, the tradeoff generator 350 infers the preferences based on the music genres associated with the songs "Red" and "Canon in D". In this example, the type of media content includes music, wherein the first preference 512 includes the first music genre (i.e., pop), and the second preference 514 includes the second music genre (i.e., classical). As shown, the first preference 512 and the second preference 514 overlap in an area that includes a third preference 516 for a third music genre (i.e., violin pop cover) for the type of media content in the embedding space 500. Based on the combined embedding 506, the tradeoff model 352 generates a tradeoff operation 354 as an output to continue playing back the media, but to change the long-term operation to the third preference 516 for playing violin pop cover. In an example where the preference for the first query and the preference for the second query do not overlap (e.g., classical music and hair metal music), the tradeoff model 350 may infer that a music genre combining classical music and hair metal music does not exist, and therefore may not output a tradeoff operation 354 combining the preferences.
[0065] Return to reference Figure 3 In some implementations, the tradeoff model 352 of the tradeoff generator 350 determines whether the conflicting long-term operation can be offloaded (ie, split / partitioned AED) to a second AED 103 in the environment, rather than interrupting the current long-term operation. Figures 1A to 1CAs described, during execution of the digital assistant 105, the AED 104 detects a plurality of users 102a, 102b and auxiliary devices 103a-c in the environment using the environment detector 310 of the query handler 300. For example, the query handler 300 receives proximity information 54 of the location of each of the plurality of users 102a, 102b, and proximity information 107 of each of the auxiliary AEDs 103 relative to the location of the AED 104. By monitoring the presence and relative location of each of the users 102 and the auxiliary AEDs 103 within the environment, when the query resolver 340 identifies a conflict between the first long-term operation 111 and the second long-term operation 112, the environment detector 310 can provide a home map representing the current environment 316 to the compromise generator 350. In a home map representing the current environment 316, the environment detector 310 can identify which auxiliary AEDs 103 are available, far enough away from the AED 104 to not interfere with the active long-term operation, and / or close enough to the user 102 requesting the conflicting long-term operation so that the requesting user 102 can prefer the auxiliary AED 103 to perform the conflicting long-term operation.
[0066] In some implementations, each user device 50a-c of the plurality of users 102 broadcasts proximity information 54, and each auxiliary AED 103a-c broadcasts proximity information 107 that can be received by the environmental detector 310, which the AED 104 can use to determine the proximity of each user device 50 and auxiliary user device 103 relative to the AED 104. The proximity information 54 from each user device 50 and the proximity information 107 from each auxiliary AED 103 and AED 104 can include wireless communication signals, such as WiFi, Bluetooth, or ultrasound, wherein the signal strength of the wireless communication signals received by the environmental detector 310 can be related to the proximity (e.g., distance) of the user device 50 and / or the auxiliary AED 103 relative to the AED 104.
[0067] In an implementation where the user 102 does not have a user device 50 or has a user device 50 that does not share proximity information 54, the environment detector 310 can detect the user 102 based on explicit input (e.g., a guest list) 313 received from the user 102a that issued the first query 106. For example, the environment detector 310 receives the guest list 313 from the seed user 102 (e.g., user 102a), the guest list indicating the identity of each user 102 in the plurality of users 102. Alternatively, the environment detector 310 detects the user 102 by performing speaker recognition (SIR) on the utterance corresponding to the audio data 402 detected within the environment. Figure 4A and Figure 4B) to detect one or more of the users 102. In other implementations, the environment detector 310 automatically detects multiple users 102 and / or auxiliary AEDs 103 in the environment by receiving image data 312 corresponding to a scene of the environment and obtained by the camera 19. Here, the environment detector 310 detects multiple users 102 and / or auxiliary AEDs 103 based on the received image data 312.
[0068] In some implementations, as the user 102 moves throughout the environment, the environment detector 310 maintains a home map representing the current environment 316 of the user 102 and the auxiliary AEDs 103. Here, the home map indicates the user 102 and the auxiliary AEDs 103 in relation to each other, to rooms / floors within the environment, and to the AEDs 104. For example, if the user 102b leaves the first floor of the environment, the environment detector 310 can detect that the user 102b is closer to the auxiliary AED 103a (e.g., SmartSpeaker2), and can prefer to cause the auxiliary AED 103a to perform the conflicting long-term operation issued by the user 102b. In response to receiving the home map representing the current environment 316 from the environment detector 310, the compromise generator 350 can generate one or more additional compromise operations 354 including fulfilling the conflicting queries 106, 146 for the individual AEDs.
[0069] In some examples, the compromise generator 350 is configured with a change threshold, and when the corresponding confidence scores of the one or more compromise solutions 354 meet the threshold (e.g., exceed the threshold), the compromise generator 354 outputs the one or more compromise solutions 354 to the user 102. Here, the compromise generator 350 determines the corresponding confidence score associated with each of the multiple compromise operations 354, and selects the compromise operation 354 with the highest corresponding confidence score among the multiple compromise operations 354 as the compromise operation 354. The threshold can be zero, where all the compromise solutions 354 (e.g., even undesirable compromises) are output to the user 102. Conversely, the threshold can be higher than zero to avoid unnecessary compromise solutions 354 that may be rejected by the user 102. In addition, in some implementations, a compromise is not possible. For example, in an environment with only a single AED 104, executing the second query on the second AED will not be included in the compromise solution 354. Similarly, the compromise generator 350 may determine that "heavy metal" and "soul" music cannot be combined, and therefore no compromise exists. In some implementations, instructing the digital assistant 105 to perform the compromise operation 354 is based on the corresponding confidence score associated with the selected compromise operation 354 satisfying a threshold. In other words, when the compromise generator 350 identifies multiple compromise solutions / operations 354, each having a corresponding confidence score, the compromise generator 350 may select the compromise operation 354 with the highest corresponding confidence score and / or provide an n-best list of compromise operations 354 for the user to select from. In some examples, when the corresponding confidence score exceeds the threshold, the query handler 500 automatically performs the compromise operation 354, rather than prompting the user 102 to make a selection.
[0070] Reference again Figure 1B, after determining that the second query 146 was issued by another user 102b other than the first user 102a that issued the first query 102a, and determining that the execution of the second long-term operation 112 conflicts with the first long-term operation 111, the query handler 300 identifies / generates one or more compromise operations 354 for the digital assistant 105. As discussed above, the query handler (via the compromise generator 350) determines that the first query 106 (e.g., pop music) and the second query 146 (e.g., classical music) are similar enough to be combined, and generates a compromise operation 354a to play a violin pop music cover. In addition, the query handler 300 identifies the AED 104 as a candidate for causing the digital assistant 105 to continue to execute the first long-term operation 111 and identifies the second AED 103 (e.g., smart speaker 103a) as a candidate for executing the second long-term operation 112 as a compromise operation 354b while the digital assistant 105 continues to execute the first long-term operation 111 on the AED 104. For example, the query handler 300 may obtain proximity information 107 of at least two auxiliary AEDs 103 within the environment of the users 102a and 102b, and proximity information 54 of each of the first user 102a and the other user 102b from the family graph representing the current environment 316. Here, identifying the AED 104 as a candidate for causing the digital assistant 105 to perform the first long-term operation 111 from the family graph and identifying the assistant-enabled device 103a as a candidate for causing the digital assistant 105 to simultaneously perform the second long-term operation 112 from the family graph is based on the proximity information 107 of each of the AEDs 103a, 104 and the proximity information 54 of each of the first user 102a and the other user 102b.
[0071] In some implementations, the query handler 300 presents the identified trade-off operations 354a, 354b to the users 102a, 102b (via the digital assistant 105), and prompts one or more of the users 102a, 102b to provide confirmation that the digital assistant 105 is to perform the selected trade-off operation 354. In these implementations, prompting the user 102 includes providing a user-selectable option as an output from the AED 104 that, when selected, provides an affirmative confirmation that the digital assistant 105 is to perform the selected trade-off operation 354. For example, the digital assistant 105 may generate synthesized speech 123 for audible output from the speaker 18 of the AED 104 (or a speaker in communication with the data processing hardware (e.g., a speaker of the user device 50)) that prompts the seed user 102a to instruct the digital assistant 105 to perform the identified compromise solution 354 “Barb, would you like to switch to violin pop covers, or play Canon in D on Smart Speaker 2?” In response, the user 102a (i.e., Barb) is shown providing confirmation to the digital assistant to perform the compromise operation 354b by issuing a third query 148 “Play Canon in D on Smart Speaker 2” near the AED 104. In response to receiving a positive confirmation from the user 102a, the digital assistant 105 performs the selected compromise operation 354b of playing the second long-term operation 112 on the auxiliary AED 103a (i.e., smart speaker 2) while executing the first long-term operation 111 on the AED 104.
[0072] In addition to or in lieu of audibly prompting the user 102a, 102, the digital assistant 105 may additionally provide a notification to the user device 50 associated with the user 102 to display user-selectable options for one or more trade-off solutions 354 as graphical elements 210 on the screen of the user device 50, the graphical elements 210 prompting the user 102 to provide confirmation that the digital assistant 105 is to perform the trade-off operation 354. Figure 2BAs shown, the GUI 200b renders the following graphical elements 210 for display: “A conflict was detected, would you like a compromise?”, “Play Violin Pop Covers” (i.e., compromise operation 354a), “Play Canon in D on Smart Speaker2” (i.e., compromise operation 354b), and “No” for the third query 148 that allows the corresponding user 102 of the device to issue (or choose to exit from the opportunity to issue) an instruction for the digital assistant 105 to perform the compromise operation 354. Here, when the user device 50 receives a user input indication indicating a selection of the user-selectable option of selecting the graphical element 210 “Play Canon in D on Smart Speaker2”, the query handler 300 receives an affirmative confirmation from the user 102 that the digital assistant 105 is to perform the compromise operation 354b.
[0073] refer to Figure 1C , while the digital assistant 105 is performing a first long-term operation 111 of playing music 122 as playback audio (e.g., track #1) on the AED 104, the query handler 300 simultaneously instructs the digital assistant 105 to perform a second long-term operation 112 of playing music 122 as playback audio (e.g., track #2) from the auxiliary AED 103a (i.e., smart speaker 2). As shown, the user 102b has moved to the second floor of the environment to be closer to the requested long-term operation 112 of playing Canonin D. In some examples, after instructing the digital assistant 105 to perform the second long-term operation 112 on the auxiliary AED 103a, the query handler 300 instructs the digital assistant 105 to adjust the execution of the first long-term operation 111 on the AED 104. Here, the query handler 300 may instruct the digital assistant to lower the volume of the music 122 output by the AED 104 so that it does not interfere with the music 122 output by the auxiliary AED 103a. In other examples, where long-term operation involves two or more home automation devices (eg, smart light bulbs 103b), when a second device is turned on in an adjacent room, the query handler 300 may instruct a first device to reduce its brightness to maintain a calm environment.
[0074] While the examples primarily refer to avoiding interruptions to a long-term operation of playing music, long-term operations may refer to any category of actions, including but not limited to search queries, controls for assistant-enabled devices (e.g., smart lights, smart thermostats), and playing / adjusting other types of media (e.g., podcasts, videos, etc.), etc. For example, the query handler 300 may help users 102 in the environment create shopping lists by resolving conflicts between items on the shopping list by recommending items that all users 102 agree on. In addition, the query handler 300 may enable the digital assistant 105 to mediate disagreements between users 102 by communicating with / prompting the user to trade-offs that the user 102 himself may not have considered.
[0075] Figure 6 A flow chart of an example operational arrangement of a method 600 for handling contradictory queries on a shared device. At operation 602, the method 600 includes receiving a first query 106 issued by a first user 102a, the first query 106 specifying a first long-term operation 111 to be performed by a digital assistant 105. While the digital assistant 105 is performing the first long-term operation 111, the method 600 also includes receiving a second query 146 at operation 604, the second query specifying a second long-term operation 112 to be performed by the digital assistant 105.
[0076] At operation 606, the method 600 further includes determining whether the second query 146 is issued by another user 102b that is different from the first user 102a. Based on the determination that the second query 146 is received from another user 102b, the method 600 also includes determining, using the query resolver 340 at operation 608, that performing the second long-term operation 112 will conflict with the first long-term operation 111. Based on the determination that performing the second long-term operation 112 will conflict with the first long-term operation 111, the method 600 also includes identifying, at operation 610, one or more compromise operations 352 for the digital assistant 105 to perform. At operation 312, the method 600 also includes instructing the digital assistant 105 to perform a selected compromise operation 352 from among the identified one or more compromise operations 352.
[0077] Figure 7 700 is a schematic diagram of an example computing device 700 that can be used to implement the systems and methods described in this document. Computing device 700 is intended to represent various forms of digital computers, such as laptops, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown here, their connections and relationships, and their functions are intended to be exemplary only and are not intended to limit implementations of the inventions described and / or claimed in this document.
[0078] The computing device 700 includes a processor 710, a memory 720, a storage device 730, a high-speed interface / controller 740 connected to the memory 720 and a high-speed expansion port 750, and a low-speed interface / controller 760 connected to a low-speed bus 770 and the storage device 730. Each of the components 710, 720, 730, 740, 750, and 760 is interconnected using various buses and can be installed on a common motherboard or installed in other ways as appropriate. The processor 710 (e.g., the data processing hardware 10, 132 of Figure 1) can process instructions for execution within the computing device 700, including instructions stored in the memory 720 or on the storage device 730, to display graphical information of a graphical user interface (GUI) on an external input / output device such as a display 780 coupled to the high-speed interface 740. In other implementations, multiple processors and / or multiple buses and multiple memories and multiple types of memories can be used as appropriate. Furthermore, multiple computing devices 700 may be connected (eg, as a server bank, a blade server group, or a multi-processor system), with each device providing portions of the necessary operations.
[0079] The memory 720 stores information non-temporarily within the computing device 700. The memory 720 (e.g., memory hardware 12, 134 of FIG. 1) may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. The non-temporary memory 720 may be a physical device for temporarily or permanently storing programs (e.g., instruction sequences) or data (e.g., program state information) for use by the computing device 700. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware, such as bootloaders). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.
[0080] The storage device 730 is capable of providing mass storage for the computing device 700. In some implementations, the storage device 730 is a computer-readable medium. In various implementations, the storage device 730 may be a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or an array of devices, including devices in a storage area network or other configuration. In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable medium or a machine-readable medium, such as the memory 720, the storage device 730, or a memory on the processor 710.
[0081] The high-speed controller 740 manages bandwidth-intensive operations of the computing device 700, while the low-speed controller 760 manages less bandwidth-intensive operations. This division of responsibilities is exemplary only. In some implementations, the high-speed controller 740 is coupled to the memory 720, the display 780 (e.g., through a graphics processor or accelerator), and the high-speed expansion port 750 that can accept various expansion cards (not shown). In some implementations, the low-speed controller 760 is coupled to the storage device 730 and the low-speed expansion port 790. The low-speed expansion port 790, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), may be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner; or to a networking device, such as a switch or a router, for example, through a network adapter.
[0082] As shown, the computing device 700 may be implemented in a variety of different forms. For example, the computing device may be implemented as a standard server 700a or multiple times in a group of such servers 700a, as a laptop computer 700b, or as part of a rack server system 700c.
[0083] Various implementations of the systems and techniques described herein can be implemented in digital electronic and / or optical circuit systems, integrated circuit systems, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system that includes at least one programmable processor that can be either special purpose or general purpose and can be coupled to receive data and instructions from and transmit data and instructions to a storage system, at least one input device, and at least one output device.
[0084] A software application (i.e., software resource) may refer to computer software that causes a computing device to perform tasks. In some examples, a software application may be referred to as an "application," "app," or "program." Example applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0085] Non-transitory memory can be a physical device used to temporarily or permanently store programs (e.g., instruction sequences) or data (e.g., program state information) for use by a computing device. Non-transitory memory can be volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., commonly used for firmware, such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.
[0086] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0087] The process and logic flow described in this specification can be performed by one or more programmable processors, also referred to as data processing hardware, which execute one or more computer programs to perform functions by operating on input data and generating outputs. The process and logic flow can also be performed by a dedicated logic circuit system, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). For example, a processor suitable for executing a computer program includes both a general-purpose microprocessor and a special-purpose microprocessor, and any one or more processors of any type of digital computer. Typically, the processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as a disk, a magneto-optical disk, or an optical disk, or be operably coupled to receive data from one or more mass storage devices or transfer data to one or more mass storage devices or both. However, a computer does not necessarily have such a device. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and memory devices, including, for example: semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0088] To provide interaction with a user, one or more aspects of the present disclosure may be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen, and possibly a keyboard and pointing device, such as a mouse or trackball, through which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including acoustic, voice, or tactile input. In addition, the computer may interact with the user by sending documents to and receiving documents from a device used by the user; for example, by sending a web page to a web browser on a user's client device in response to a request received from the web browser.
[0089] A variety of implementations have been described. However, it should be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Therefore, other implementations are within the scope of the appended claims.
Claims
1. A computer-implemented method (600) which, when executed by data processing hardware (10, 132), causes the data processing hardware (10, 132) to perform operations comprising: Receiving a first query (106) issued by a first user (102a), the first query (106) specifying a first long-term operation (111) to be performed by a digital assistant (105); When the digital assistant (105) is performing the first long-term operation (111): receiving a second query (146) specifying a second long-term operation (112) to be performed by the digital assistant (105); determining that the second query (146) was issued by another user (102b) different from the first user (102a); Based on determining that the second query (146) was received from the other user (102b), determining using a query resolver (340) that performing the second long-term operation (112) will conflict with the first long-term operation (111); as well as Based on determining that performing the second long-term operation (112) will conflict with the first long-term operation (111), identifying one or more trade-off operations (354) to be performed by the digital assistant (105); and The digital assistant (105) is instructed to perform a selected trade-off operation (354) from among the identified one or more trade-off operations (354).
2. The method (600) of claim 1, wherein: Receiving the second query (146) includes receiving audio data (402) corresponding to the second query (146), the second query (146) being spoken by the other user (102b) and captured by an assistant-supporting device (104) executing the digital assistant (105); and Determining that the second query (146) was issued by another user (102b) different from the first user (102a) includes: performing speaker recognition on the audio data (402) to determine that the second query (146) was spoken by the other user (102b) different from the first user (102a) who issued the query (106).
3. The method (600) of claim 2, wherein: Performing speaker identification on the audio data (402) to determine that the second query (146) was spoken by the other user (102b) includes: extracting a speaker identification vector (412) characteristic of the second query (146) from the audio data (402) corresponding to the second query (146); and Determining that the speaker identification vector (412) extracted from the audio data (402) corresponding to the second query (146) is at least one of: does not match a reference speaker vector (155) for the first user (102a); or A matching is performed with an enrolled speaker vector (154) associated with the other user (102b).
4. The method (600) of any one of claims 1 to 3, wherein: Receiving the first query (106) issued by the first user (102a) includes receiving a user input indication from a user device (50) associated with the first user (102a) indicating a user intent to issue the first query (106).
5. The method (600) of any one of claims 1 to 4, wherein: Receiving the first query (106) issued by the first user (102a) includes: receiving audio data (402) corresponding to the first query (106) spoken by the first user (102a) and captured by a device (104) supporting the digital assistant (105).
6. The method (600) according to any one of claims 1 to 5, wherein: Identifying the one or more tradeoff operations (354) to be performed by the digital assistant (105) includes: identifying criteria (113) associated with the first query (106); identifying criteria (115) associated with the second query (146); generating a first query embedding (502) based on the criteria (113) associated with the first query (106) using a query embedding model (352); generating a second query embedding (504) based on the criteria (115) associated with the second query (146) using the query embedding model (352); determining a combined embedding (506) based on the first query embedding (502) and the second query embedding (504); and At least one trade-off operation (354) is identified that maps the combined embedding (506) into the embedding space (500).
7. The method (600) of claim 6, wherein The identified criteria (113) associated with the first query (106) includes a first preference (512) for the type of media content played back from an assistant-enabled device (104) executing the digital assistant (105); the identified criteria (115) associated with the second query (146) includes a second preference (514) for the type of media content played back from the assistant-enabled device (104); and The identified at least one tradeoff operation (354) includes a third preference (516) for the type of media content played back from the assistant-enabled device (104).
8. The method (600) of claim 7, wherein: The types of media content include music; The first preference (512) for the type of the media content includes a first music genre; and The second preference (514) for the type of the media content includes a second music genre.
9. The method (600) of any one of claims 6 to 8, wherein: The identified criteria (113) associated with the first query (106) includes a first value for a setting of a home automation device (103); the identified criteria (115) associated with the second query (146) includes a second value for the setting of the home automation device (103); and The identified at least one compromise operation (354) includes adjusting the first value of the setting for the home automation device (103) to a new value.
10. The method (600) of claim 9, wherein: The home automation device (103) includes a smart thermostat, a smart light, a smart speaker or a smart display.
11. The method (600) of any one of claims 1 to 10, wherein: The operations further include: obtaining a family map indicating at least two assistant-supported devices (103) within the same environment as the first user (102a) and the other user (102b) and capable of performing the first long-term operation (111) and the second long-term operation (112), Wherein, identifying the one or more compromise operations (354) to be performed by the digital assistant (105) includes: identifying a first assistant-supporting device (103) from the family map as a candidate for causing the digital assistant (105) to perform the first long-term operation (111), and identifying a second assistant-supporting device (103) from the family map as a candidate for causing the digital assistant (105) to perform the second long-term operation (112) simultaneously when the digital assistant (105) performs the first long-term operation (111) on the first assistant-supporting device (103).
12. The method (600) of claim 11, wherein: The operations further include: Obtaining proximity information (107) of each of the at least two AEDs (103) within the same environment as the first user (102a) and the other user (102b) from the family map; and obtaining proximity information (107) of each of the first user (102a) that issued the first query (106) and the other user (102b) that issued the second query (146), Wherein, identifying the first support assistant's device (103) from the family map as a candidate for causing the digital assistant (105) to perform the first long-term operation (111) and identifying the second support assistant's device (103) from the family map as a candidate for causing the digital assistant (105) to simultaneously perform the second long-term operation (112) is based on the proximity information (107) of each of the at least two AEDs (103) and the proximity information (107) of each of the first user (102a) and the other user (102b).
13. The method (600) of any one of claims 1 to 12, wherein: The digital assistant (105) performs the first long-term operation (111) on a first assistant-supported device (103); and Instructing the digital assistant (105) to perform the selected compromise operation (354) includes: while the digital assistant (105) is performing the first long-term operation (111) on the first assistant-supporting device (103), simultaneously instructing the digital assistant (105) to perform the second long-term operation (112) on the second assistant-supporting device (103).
14. The method (600) of claim 13, wherein: The operation further includes: after instructing the digital assistant (105) to perform the second long-term operation (112) on the device (103) of the second support assistant, instructing the digital assistant (105) to adjust the execution of the first long-term operation (111) on the device (103) of the first support assistant.
15. The method (600) of any one of claims 1 to 14, wherein: When a plurality of trade-off operations (354) are identified for the digital assistant (105) to perform, the operations further include: determining a respective score associated with each trade-off operation (354) among the plurality of trade-off operations (354); and A trade-off operation (354) having the highest corresponding score among the plurality of trade-off operations (354) is selected as the trade-off operation (354).
16. The method (600) of claim 15, wherein: The operations further include: determining that the corresponding score associated with the selected tradeoff operation (354) satisfies a threshold, Wherein, instructing the digital assistant (105) to perform the compromise operation (354) is based on the corresponding score associated with the selected compromise operation (354) satisfying the threshold.
17. The method (600) of any one of claims 1 to 16, wherein: The operations further include: prompting the first user (102a) and / or the other user (102b) to provide confirmation that the digital assistant (105) is to perform the selected trade-off operation (354); and receiving a positive confirmation from the first user (102a) and / or the other user (102b) that the digital assistant (105) is to perform the selected trade-off operation (354), Wherein, instructing the digital assistant (105) to perform the selected compromise operation (354) is based on the received positive confirmation.
18. A system (100), comprising: Data processing hardware (10, 132); as well as Memory hardware (12, 134) in communication with the data processing hardware (10, 132), the memory hardware (12, 134) storing instructions that, when executed on the data processing hardware (10, 132), cause the data processing hardware (10, 132) to perform operations, the operations comprising: Receiving a first query (106) issued by a first user (102a), the first query (106) specifying a first long-term operation (111) to be performed by a digital assistant (105); When the digital assistant (105) is performing the first long-term operation (111): receiving a second query (146) specifying a second long-term operation (112) to be performed by the digital assistant (105); determining that the second query (146) was issued by another user (102b) different from the first user (102a); Based on determining that the second query (146) was received from the other user (102b), determining using a query resolver (340) that performing the second long-term operation (112) will conflict with the first long-term operation (111); as well as Based on determining that performing the second long-term operation (112) will conflict with the first long-term operation (111), identifying one or more trade-off operations (354) to be performed by the digital assistant (105); and The digital assistant (105) is instructed to perform a selected trade-off operation (354) from among the identified one or more trade-off operations (354).
19. The system (100) of claim 18, wherein: Receiving the second query (146) includes receiving audio data (402) corresponding to the second query (146), the second query (146) being spoken by the other user (102b) and captured by an assistant-supporting device (104) executing the digital assistant (105); and Determining that the second query (146) was issued by another user (102b) different from the first user (102a) includes: performing speaker recognition on the audio data (402) to determine that the second query (146) was spoken by the other user (102b) different from the first user (102a) who issued the query (106).
20. The system (100) of claim 19, wherein: Performing speaker identification on the audio data (402) to determine that the second query (146) was spoken by the other user (102b) includes: extracting a speaker identification vector (412) characteristic of the second query (146) from the audio data (402) corresponding to the second query (146); and Determining that the speaker identification vector (412) extracted from the audio data (402) corresponding to the second query (146) is at least one of: does not match a reference speaker vector (155) for the first user (102a); or A matching is performed with an enrolled speaker vector (154) associated with the other user (102b).
21. The system (100) of any one of claims 18 to 20, wherein: Receiving the first query (106) issued by the first user (102a) includes receiving a user input indication from a user device (50) associated with the first user (102a) indicating a user intent to issue the first query (106).
22. The system (100) of any one of claims 18 to 21, wherein: Receiving the first query (106) issued by the first user (102a) includes: receiving audio data (402) corresponding to the first query (106) spoken by the first user (102a) and captured by a device (104) supporting the digital assistant (105).
23. The system (100) of any one of claims 18 to 22, wherein: Identifying the one or more tradeoff operations (354) to be performed by the digital assistant (105) includes: identifying criteria (113) associated with the first query (106); identifying criteria (115) associated with the second query (146); generating a first query embedding (502) based on the criteria (113) associated with the first query (106) using a query embedding model (352); generating a second query embedding (504) based on the criteria (115) associated with the second query (146) using the query embedding model (352); determining a combined embedding (506) based on the first query embedding (502) and the second query embedding (504); and At least one trade-off operation (354) is identified that maps the combined embedding (506) into the embedding space (500).
24. The system (100) of claim 23, wherein The identified criteria (113) associated with the first query (106) includes a first preference (512) for the type of media content played back from an assistant-enabled device (104) executing the digital assistant (105); the identified criteria (115) associated with the second query (146) includes a second preference (514) for the type of media content played back from the assistant-enabled device (104); and The identified at least one tradeoff operation (354) includes a third preference (516) for the type of media content played back from the assistant-enabled device (104).
25. The system (100) of claim 24, wherein: The types of media content include music; The first preference (512) for the type of the media content includes a first music genre; and The second preference (514) for the type of the media content includes a second music genre.
26. The system (100) of any one of claims 23 to 25, wherein: The identified criteria (113) associated with the first query (106) includes a first value for a setting of a home automation device (103); the identified criteria (115) associated with the second query (146) includes a second value for the setting of the home automation device (103); and The identified at least one compromise operation (354) includes adjusting the first value of the setting for the home automation device (103) to a new value.
27. The system (100) of claim 26, wherein: The home automation device (103) includes a smart thermostat, a smart light, a smart speaker or a smart display.
28. The system (100) of any one of claims 18 to 27, wherein: The operations further include: obtaining a family map indicating at least two assistant-supported devices (103) within the same environment as the first user (102a) and the other user (102b) and capable of performing the first long-term operation (111) and the second long-term operation (112), Wherein, identifying the one or more compromise operations (354) to be performed by the digital assistant (105) includes: identifying a first assistant-supporting device (103) from the family map as a candidate for causing the digital assistant (105) to perform the first long-term operation (111), and identifying a second assistant-supporting device (103) from the family map as a candidate for causing the digital assistant (105) to perform the second long-term operation (112) simultaneously when the digital assistant (105) performs the first long-term operation (111) on the first assistant-supporting device (103).
29. The system (100) of claim 28, wherein: The operations further include: Obtaining proximity information (107) of each of the at least two AEDs (103) within the same environment as the first user (102a) and the other user (102b) from the family map; and obtaining proximity information (107) of each of the first user (102a) that issued the first query (106) and the other user (102b) that issued the second query (146), Wherein, identifying the first support assistant's device (103) from the family map as a candidate for causing the digital assistant (105) to perform the first long-term operation (111) and identifying the second support assistant's device (103) from the family map as a candidate for causing the digital assistant (105) to simultaneously perform the second long-term operation (112) is based on the proximity information (107) of each of the at least two AEDs (103) and the proximity information (107) of each of the first user (102a) and the other user (102b).
30. The system (100) of any one of claims 18 to 29, wherein: The digital assistant (105) performs the first long-term operation (111) on a first assistant-supported device (103); and Instructing the digital assistant (105) to perform the selected compromise operation (354) includes: while the digital assistant (105) is performing the first long-term operation (111) on the first assistant-supporting device (103), simultaneously instructing the digital assistant (105) to perform the second long-term operation (112) on the second assistant-supporting device (103).
31. The system (100) of claim 30, wherein: The operation further includes: after instructing the digital assistant (105) to perform the second long-term operation (112) on the device (103) of the second support assistant, instructing the digital assistant (105) to adjust the execution of the first long-term operation (111) on the device (103) of the first support assistant.
32. The system (100) of any one of claims 18 to 31, wherein: When a plurality of trade-off operations (354) are identified for the digital assistant (105) to perform, the operations further include: determining a respective score associated with each trade-off operation (354) among the plurality of trade-off operations (354); and A trade-off operation (354) having the highest corresponding score among the plurality of trade-off operations (354) is selected as the trade-off operation (354).
33. The system (100) of claim 32, wherein: The operations further include: determining that the corresponding score associated with the selected tradeoff operation (354) satisfies a threshold, Wherein, instructing the digital assistant (105) to perform the compromise operation (354) is based on the corresponding score associated with the selected compromise operation (354) satisfying the threshold.
34. The system (100) of any one of claims 18 to 33, wherein: The operations further include: prompting the first user (102a) and / or the other user (102b) to provide confirmation that the digital assistant (105) is to perform the selected trade-off operation (354); and receiving a positive confirmation from the first user (102a) and / or the other user (102b) that the digital assistant (105) is to perform the selected trade-off operation (354), Wherein, instructing the digital assistant (105) to perform the selected compromise operation (354) is based on the received positive confirmation.