Handling continued sessions across multiple devices
By using the first device to detect subsequent requests and notify other devices of user locations in the automation assistant device ecosystem, the problems of resource waste and information exposure are solved, and more efficient audio data processing and user experience improvement are achieved.
Patent Information
- Application Number
- CN202380071633.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-17
- Filing Date
- 2023-10-16
- Publication Date
- 2025-07-04
AI Technical Summary
In the multiple automation assistant device ecosystem, it is difficult to effectively determine which device is best suited to process subsequent audio data requests from users, resulting in waste of resources and potentially sensitive information exposure.
By processing the user's initial request on the first device, determining whether there is a subsequent request, and sending a notification to other devices to detect the user's location, allowing the device closest to the user to process the subsequent audio data, stopping unnecessary processing, and providing context information to ensure the correct resolution of the request.
Improve resource utilization efficiency, reduce unnecessary processing and information exposure, ensure that subsequent requests are processed on the most suitable device, and improve user experience.
Smart Images

Figure CN120266203A_ABST
Abstract
Description
Background Art
[0001] Humans can have a human - machine conversation with an interactive software application referred to herein as an "automation assistant" (also known as a "chatbot", "interactive personal assistant", "intelligent personal assistant", "personal voice assistant", "conversational agent", etc.). For example, a person (who can be referred to as a "user" when interacting with the automation assistant) can provide an explicit input (e.g., a command, query, and / or request) to the automation assistant, which can cause the automation assistant to generate and provide a responsive output to control one or more Internet of Things (IoT) devices, and / or to perform one or more other functions (e.g., assistant actions). This explicit input provided by the user can be, for example, a typed natural - language input and / or an oral natural - language input (i.e., an oral utterance), which in some cases can be converted into text (or other semantic representation) and then further processed.
[0002] In some cases, the automation assistant can include an automation - assistant client that is executed locally by an assistant device and directly participated in by the user, and a cloud - based counterpart that utilizes the almost unlimited resources of the cloud to help the automation - assistant client respond to user input. For example, the automation - assistant client can provide the audio data (or its text conversion) of the user's oral utterance, and optionally data indicating the user identity (e.g., credentials), to the cloud - based counterpart. The cloud - based counterpart can perform various processes on the explicit input to return the result to the automation - assistant client, which can then provide the corresponding output to the user. In other cases, the automation assistant can be executed specifically locally by the assistant device, and these assistant devices are directly operated by the user to reduce latency.
[0003] Many users can use the automation assistant to perform daily routine tasks via assistant actions. For example, users can routinely provide one or more explicit user inputs that cause the automation assistant to check the weather, check the traffic on the way to work, start the vehicle, and / or other explicit user inputs that cause the automation assistant to perform other assistant actions while the user is having breakfast. As another example, users can routinely provide one or more explicit user inputs that cause the automation assistant to play a specific playlist, track an exercise, and / or other explicit user inputs that cause the automation assistant to perform other assistant actions to prepare for the user to go for a run. However, in some situations, multiple devices can be near the user and can be configured to process the user's requests. Therefore, determining which device is most suitable for processing audio data can improve the user experience when interacting with multiple devices configured in a connected - device ecosystem. Summary of the Invention
[0004] Some implementations disclosed herein involve selecting a second device to process audio data to continue a session with a user, the session being initiated by an automated assistant executing on a first device. A user may invoke the automated assistant on the first device, and the first device may process a request provided with the invocation (e.g., before and / or after the invocation). In addition to processing the request to determine one or more actions to perform, the automated assistant may also determine that there may be a follow-up request. In response, the automated assistant may provide a notification that there may be a follow-up request to one or more other devices linked to the first device via a linked device ecosystem. Each notified device (and / or the automated assistant executing at least in part on each notified device) may process sensor data from sensors of the corresponding device and determine whether the user is present near one of the other devices. If the user is detected to be present near one of the other devices, the device may send an indication to the first device that the user has changed location and that subsequent audio data may be processed by that device. In response, the first device may stop processing subsequent audio data, and the device closest to the user may begin processing the audio data to await a follow-up request.
[0005] As an example, a user may invoke a first intelligent device (e.g., "OK SmartSpeaker" for invoking a smart speaker) and issue a query such as "What is the weather like here today?". The automated assistant executing on the smart speaker can determine and provide a response (e.g., "It is going to be 75 and sunny today"), and further determine that the user may use another related query (e.g., "What about in Miami") to follow up on that query. In response, the automated assistant executing on the smart speaker (or another application executing on the smart speaker) can provide a notification to one or more other linked devices that are part of a linked device ecosystem that includes the smart speaker (e.g., a second smart speaker in another room, a smart TV, a smartphone at another location). The notification can indicate that the user may be issuing a follow-up query that can be captured by the microphone of one of the other devices. Each of the other devices can then process sensor data from the sensors of the corresponding device (i.e., each automated assistant processes sensor data from the sensors of the device on which the automated assistant is executing) to determine whether the user has moved and is now present near the device. For example, a smartphone can use accelerometer data to determine whether the user has picked up and / or moved the smartphone, a device with a camera can determine whether the user has entered the field of view of the camera, and so on. Once it is determined that the user is present near one of the other devices, that device can provide an indication to the first device (i.e., the device that processed the first request) to indicate that it will handle the subsequent follow-up request. In response, the first device can stop processing the subsequent audio data and the device closest to the user can start processing the audio data to capture any follow-up requests provided by the user.
[0006] In some implementations, the processing of audio data can be delayed such that the automated assistant does not start processing the subsequent audio data immediately after performing an action, but can be delayed until a later time. For example, a user may request to play a song via one or more speakers of a device, and the audio data can be processed when the song ends and / or near the end of the song. In a manner similar to that described previously, the presence of the user can be determined based on sensor data from sensors of devices in the ecosystem to determine whether the user has moved while the song is playing. If the user is detected to be near a device other than the device that first received the invocation, the automated assistant of the first device can stop processing the audio data, and the device closer to the user near the end of the song can start processing the audio data.
[0007] In some implementations, the same automated assistant can execute on multiple devices. For example, the automated assistant can execute partially on a first device, a second device, and partially on a cloud-based device such that the cloud-based device executes some or all of the processing of the audio data. In these cases, all devices can access the context of a session initiated by the first device and can thus utilize it when processing subsequent requests. However, in some cases, the device processing subsequent audio data may require the context of the previous session to process the follow-up request. For example, a first automated assistant running on a first device can process an initial request, determine that a follow-up request may occur, and send a notification for processing sensor data to other linked devices to determine whether the user has changed location. If one of the devices provides a notification indicating that the user is present near the device, the first device can provide the context of the previous request such that the second device can resolve any context ambiguities in the follow-up request. For example, the user can invoke the automated assistant on the first device by "OK, Speaker" and then continue to query "what’s the weather today". This query can be followed by "How about tomorrow", which requires the context from the previous query to resolve the intent of the request. Thus, the first device executing the first automated assistant can provide the context of the previous session for a second automated assistant executing on a second device such that the second automated assistant can resolve the intent of "How about tomorrow".
[0008] In some implementations, the first automated assistant can determine a user profile associated with the user who invoked the first automated assistant. For example, the first automated assistant can utilize text-dependent (TD) and / or text-independent (TI) speaker verification to identify the user profile associated with the speaker. In some implementations, the first automated assistant can provide an indication of the speaker of the first request (e.g., an indication of the user profile associated with the speaker, a vector embedding of the speaker) along with the notification such that one or more other devices can determine whether the same speaker is co-present with one or more of the other devices.
[0009] For example, one of the other devices can use sensor data (e.g., accelerometer data) from one or more sensors of the device to determine that a person is co-located with the device, and can process some audio data and / or visual data, and determine whether the co-located person is a user of, e.g., TI, by comparing the captured audio data with a vector representing a user who speaks one or more phrases. Thus, in some implementations, the other device can determine whether the same speaker as the user who uttered the initial request (and invocation) is co-located, and only further process the audio data to identify the follow-up request when it determines that the user is co-located with the device.
[0010] In some implementations, a user can utter an invocation phrase that invokes multiple automated assistants, which either execute on the same device or on separate devices. For example, the user can utter an invocation phrase such as “OK Assistant A and Assistant B, how tall is Jack Smith?” (note: Jack Smith is a fictional man with wide name recognition), which can simultaneously invoke a first assistant (i.e., “OK Assistant A”) and a second automated assistant (i.e., “OK Assistant B”). In some implementations, both automated assistants can process the audio data and determine that there may be a follow-up query. Additionally, each automated assistant can determine the likelihood that the follow-up query will likely be directed to it or to the other automated assistant. For example, the first automated assistant can be configured to handle queries in the audio data, while the second automated assistant can be configured to handle media playback. If the invoked automated assistant determines that the follow-up query is unlikely to be directed to it (or is more likely to be directed to another automated assistant), it can stop processing the subsequent audio data and instead allow the more likely automated assistant to handle the follow-up query.
[0011] As another example, a user may invoke an automated assistant executing on a different device by indicating, in the invocation, the device that is executing the separate automated assistant. For example, the user may say “OK Kitchen Devices, show me a weather map of Miami”. In response, an automated assistant executing on multiple devices identified as “Kitchen Devices” (e.g., a smart speaker and a smart screen with a graphical interface) may be invoked and begin processing the request. The smart speaker may determine that, if there is a follow-up query, the follow-up query may require a graphical interface. In response, the smart speaker may stop processing subsequent audio data and instead allow the smart screen, which includes a graphical interface, to handle the subsequent follow-up query (e.g., “Now show me Los Angeles”).
[0012] The foregoing description is provided as an overview of only some implementations disclosed herein. Those implementations and other implementations are described in more detail herein.
[0013] It should be understood that all combinations of the foregoing concepts and additional concepts described in more detail herein are considered to be part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are contemplated as part of the subject matter disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a block diagram of an example environment in which implementations disclosed herein may be implemented.
[0015] Figure 2 is an illustration of an example ecosystem of connected devices.
[0016] Figure 3 is a timing diagram showing an example method for selectively processing audio data in a continued conversation mode.
[0017] Figure 4A and Figure 4B are flowcharts showing example methods in accordance with various implementations disclosed herein.
[0018] Figure 5 is a flowchart showing an example method in accordance with various implementations disclosed herein.
[0019] Figure 6A and Figure 6B are flowcharts showing example methods in accordance with various implementations disclosed herein.
[0020] Figure 7 illustrates an example architecture of a computing device. Detailed implementation manners
[0021] Now turning to Figure 1 FIG. 6 illustrates an example environment in which the techniques disclosed herein may be implemented. The example environment includes a plurality of assistant input devices 106 and one or more cloud-based automated assistant components 119. One or more (e.g., all) of the assistant input devices 106 may execute respective instances of a corresponding automated assistant client 118. However, in some implementations, one or more of the assistant input devices 106 may optionally lack an instance of the corresponding automated assistant client 118 and still include engines and hardware components (e.g., microphone 109, speaker 108, speech recognition engine, natural language processing engine, speech synthesis engine, etc.) for receiving and processing user input directed to the automated assistant. Instances of the automated assistant client 118 may be applications separate from (e.g., installed “on top of”) the operating system of the corresponding assistant input device 106 or may alternatively be implemented directly by the operating system of the corresponding assistant input device 106. As further described below, each instance of the automated assistant client 118 may optionally interact with one or more cloud-based automated assistant components 119 in response to various requests provided by a corresponding user interface component 107 of any of the corresponding assistant input devices 106. Additionally, and also as described below, other engines of the assistant input device 106 may optionally interact with one or more cloud-based automated assistant components 119.
[0022] One or more cloud-based automated assistant components 119 may be implemented on one or more computing systems (e.g., servers collectively referred to as “the cloud” or “remote” computing systems) that are communicatively coupled to the corresponding assistant input devices 106 via one or more local area networks (“LANs”, including Wi-Fi LANs, Bluetooth networks, near field communication networks, mesh networks, etc.), wide area networks (“WANs”, including the Internet, etc.), and / or other networks. The communicative coupling of the cloud-based automated assistant components 119 to the assistant input devices 106 is generally represented by 110 in FIG. 6. Moreover, in some implementations, these assistant input devices 106 may be communicatively coupled to each other via one or more networks (e.g., LANs and / or WANs). Figure 1 in FIG. 6
[0023] An instance of the automated assistant client 118, through its interaction with one or more of the cloud-based automated assistant components 119, can form something that appears to the user to be a logical instance of an automated assistant, and using this logical instance, the user can participate in a human-machine conversation. For example, a first automated assistant can be covered by the first automated assistant client 118 of the first assistant input device 106 and one or more cloud-based automated assistant components 119. A second automated assistant can be covered by the second automated assistant client 118 of the second assistant input device 106 and one or more cloud-based automated assistant components 119. The first automated assistant and the second automated assistant can also be simply referred to as "automated assistant" in this document. Therefore, it should be understood that each user interacting with the automated assistant client 118 executed on one or more assistant input devices 106 can actually interact with his or her own logical instance of an automated assistant (or a logical instance of an automated assistant shared among a family or other user groups and / or shared among multiple automated assistant clients 118). Although Figure 1 only multiple assistant input devices 106 are shown in, it is understood that the cloud-based automated assistant components 119 can additionally serve many additional groups of assistant input devices. Additionally, although the various engines of the cloud-based automated assistant components 119 are described in this document as being implemented separately from the automated assistant client 118 (e.g., at a server), it should be understood that this is for example only and does not mean a limitation. For example, one or more (e.g., all) of the engines described with respect to the cloud-based automated assistant components 119 can be implemented locally by one or more assistant input devices 106.
[0024] The assistant input device 106 can include, for example, one or more of the following: a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a computing device of the user's vehicle (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), an interactive stand-alone speaker (e.g., with or without a display), a smart appliance (such as a smart TV or a smart washer / dryer), a wearable device of the user including a computing device (e.g., the user's watch with a computing device, the user's glasses with a computing device, a virtual or augmented reality computing device), and / or any IoT device capable of receiving user input directed to the automated assistant. Additional and / or alternative assistant input devices can be provided. In some implementations, multiple assistant input devices 106 can be associated with each other in various ways to facilitate the execution of the techniques described in this document. For example, in some implementations, multiple assistant input devices 106 can be connected via one or more networks (e.g., via Figure 1are communicatively coupled to and associated with each other. For example, it may be the case where multiple assistant input devices 106 are deployed in a specific area or environment (such as a home, a building, etc.). Additionally or alternatively, in some implementations, multiple assistant input devices 106 may be associated with each other by virtue of being members of a coordinated ecosystem that is selectively accessible by at least one or more users (e.g., an individual, a family, an employee of an organization, other predefined groups, etc.). In some of those implementations, the ecosystem of multiple assistant input devices 106 may be manually and / or automatically associated with each other in a device topology representation of the ecosystem.
[0025] In various implementations, one or more assistant input devices 106 may include one or more corresponding sensors 105 that are configured to provide sensor data indicative of one or more environmental conditions present in the environment of the corresponding device upon approval of the corresponding user. In some of those implementations, the automated assistant may identify one or more assistant input devices 106 to satisfy an oral utterance of a user associated with the ecosystem. The oral utterance may be satisfied by rendering responsive content at one or more assistant input devices 106 (e.g., auditorily and / or visually), by causing one or more assistant input devices 106 to be controlled based on the oral utterance, and / or by causing one or more assistant input devices 106 to perform any other action to satisfy the oral utterance.
[0026] The corresponding sensors 105 may be in various forms. Some assistant input devices 106 may be equipped with one or more digital cameras that are configured to capture and provide a signal indicative of detected motion within their field of view. Additionally or alternatively, some assistant input devices 106 may be equipped with other types of light-based sensors 105, such as passive infrared (“PIR”) sensors that measure infrared (“IR”) light radiated from objects within their field of view. Additionally or alternatively, some assistant input devices 106 may be equipped with sensors 105 that detect acoustic (or pressure) waves, such as one or more microphones.
[0027] Additionally or alternatively, in some implementations, sensor 105 can be configured to detect other phenomena associated with an environment that includes at least a portion of the ecosystem. For example, in some embodiments, a given one of assistant devices 106 can be equipped with sensor 105 that detects various types of wireless signals (e.g., waves such as radio waves, ultrasonic waves, electromagnetic waves, etc.) transmitted by other assistant devices (e.g., mobile devices, wearable computing devices, etc.) carried / operated by a specific user and / or other assistant devices in the ecosystem. For example, some of assistant devices 106 can be configured to transmit waves that are imperceptible to humans, such as ultrasonic or infrared waves, which can be detected by one or more assistant input devices 106 (e.g., via an ultrasonic / infrared receiver such as a microphone with ultrasonic capabilities). Also, for example, in some embodiments, a given one of assistant devices 106 can be equipped with sensor 105 to detect the movement of the device (e.g., an accelerometer), the temperature near the device, and / or other environmental conditions that can be detected near the device (e.g., a heart monitor that can detect the user's current heart rate).
[0028] Additionally or alternatively, various assistant devices can transmit other types of waves that are imperceptible to humans, such as radio waves (e.g., Wi-Fi, Bluetooth, cellular, etc.), which can be detected by other assistant devices (e.g., mobile devices, wearable computing devices, etc.) carried / operated by the user and used to determine the specific location of the operating user. In some implementations, GPS and / or Wi-Fi triangulation can be used, for example, based on GPS and / or Wi-Fi signals to / from the assistant devices to detect the location of a person. In other implementations, other wireless signal characteristics (such as time of flight, signal strength, etc.) can be used alone or in combination by various assistant devices to determine the location of a specific person based on signals transmitted by other assistant devices carried / operated by a specific user.
[0029] Additionally or alternatively, in some implementations, one or more assistant input devices 106 may perform speaker recognition to identify a user based on their voice. For example, some instances of an automated assistant may be configured to match a voice to a user profile, e.g., for the purpose of providing / limiting access to various resources. A variety of techniques have been utilized for user identification and / or authorization of automated assistants. For example, when identifying a user, some automated assistants utilize text-dependent (TD) techniques, which are constrained to invocation phrases for the assistant (e.g., "OK Assistant" and / or "Hey Assistant"). Using such techniques, an enrollment process is performed, in which the user is explicitly prompted to provide one or more instances of verbal utterances of the invocation phrase to which the TD features are constrained. A speaker feature (e.g., a speaker embedding) of the user can then be generated by processing an instance of audio data, where each instance in the instance captures a corresponding one of the verbal utterances in the verbal utterance. For example, a speaker feature can be generated by using a TD machine learning model to process each instance in the instance of audio data to generate a corresponding speaker embedding for each utterance in the utterance. The speaker feature can then be generated as a function of the speaker embeddings and stored (e.g., on the device) for TD techniques. For example, the speaker feature can be a cumulative speaker embedding, which is a function (e.g., an average) of the speaker embeddings. It has also been proposed to utilize text-independent (TI) techniques as a supplement or alternative to TD techniques. TI features are not constrained to a subset of phrases as in TD. Similar to TD, TI can also utilize the speaker features of the user and can generate those features based on user utterances obtained through an enrollment process and / or other verbal interactions, although more instances of user utterances may be required to generate useful TI speaker features.
[0030] After generating the speaker feature, the speaker feature can be used to identify the user who uttered the verbal utterance. For example, when a user utters another verbal utterance, the audio data capturing the verbal utterance can be processed to generate utterance features, those utterance features can be compared to the speaker feature, and based on the comparison, the profile associated with the speaker feature can be identified. As a specific example, a speaker recognition model can be used to process the audio data to generate an utterance embedding, and the utterance embedding can be compared to the previously generated speaker embedding of the user to identify the user's profile. For example, if the distance metric between the generated utterance embedding and the speaker embedding of the user meets a threshold, then the user can be identified as the user who uttered the verbal utterance.
[0031] Each assistant input device 106 further includes a corresponding user interface component 107, and each user interface component may include one or more user interface input devices (e.g., a microphone, a touch screen, a keyboard, and / or other input devices) and / or one or more user interface output devices (e.g., a display, a speaker, a projector, and / or other output devices). As an example, the user interface component 107 of the assistant input device 106 may include only a speaker 108 and a microphone 109, while the user interface component 107 of another assistant input device 106 may include a speaker 108, a touch screen, and a microphone 109.
[0032] Each of the assistant input devices 106 and / or any other computing device operating one or more of the cloud-based automated assistant components 119 may include one or more memories for storing data and software applications, one or more processors for accessing the data and executing the applications, and other components that facilitate communication over a network. Operations performed by one or more of the assistant input devices 106 and / or by the automated assistant may be distributed across multiple computer systems. The automated assistant may be implemented, for example, as a computer program running on one or more computers located in one or more locations and coupled to each other via a network (e.g., Figure 1 the network 110).
[0033] As described above, in various implementations, each assistant input device 106 may operate a corresponding automated assistant client 118. In various embodiments, each automated assistant client 118 may include a corresponding voice capture / text-to-speech (TTS) / speech-to-text (STT) module 114 (also simply referred to herein as "voice capture / TTS / STT module 114"). In other implementations, one or more aspects of the corresponding voice capture / TTS / STT module 114 may be implemented separately from the corresponding automated assistant client 118 (e.g., by one or more cloud-based automated assistant components 119).
[0034] Each corresponding voice capture / TTS / STT module 114 may be configured to perform one or more functions, including, for example: capturing the user's voice (voice capture, e.g., via the corresponding microphone 109); converting the captured audio to text and / or to other representations or embeddings (STT) using a voice recognition model stored in a database; and / or converting text to speech (TTS) using a voice synthesis model stored in a database. Instances of these models may be stored locally at each corresponding assistant input device 106 and / or accessible by the assistant input device (e.g., via Figure 1Network 110). In some implementations, since one or more assistant input devices 106 may be relatively limited in terms of computing resources (e.g., processor cycles, memory, battery, etc.), the corresponding voice capture / TTS / STT module 114 local to each assistant input device 106 may be configured to use a speech recognition model to convert a limited number of different spoken phrases into text (or into other forms, such as low-dimensional embeddings). Other voice inputs may be sent to one or more cloud-based automated assistant components 119, which may include a cloud-based TTS module 116 and / or a cloud-based STT module 117.
[0035] The cloud-based STT module 117 may be configured to utilize the almost infinite resources of the cloud and use a speech recognition model to convert the audio data captured by the voice capture / TTS / STT module 114 into text (which can then be provided to the natural language processing (NLP) module 122). The cloud-based TTS module 116 may be configured to utilize the almost infinite resources of the cloud and use a speech synthesis model to convert text data (e.g., text formulated by the automated assistant) into computer-generated voice output. In some implementations, the cloud-based TTS module 116 may provide the computer-generated voice output to one or more assistant devices 106 for direct output, for example, using the corresponding speakers 108 of the respective assistant devices. In other implementations, the text data generated by the automated assistant using the cloud-based TTS module 116 (e.g., client device notifications included in a command) may be provided to the voice capture / TTS / STT module 114 of the corresponding assistant device, which can then locally convert the text data into computer-generated voice using a speech synthesis model and cause the computer-generated voice to be rendered via the local speaker 108 of the corresponding assistant device.
[0036] The NLP module 122 processes natural language inputs generated by the user via the assistant input device 106 and may generate annotated outputs for use by one or more other components of the automated assistant, the assistant input device 106. For example, the NLP module 122 may process natural language free-form inputs generated by the user via one or more corresponding user interface input devices of the assistant input device 106. The annotated output generated based on processing the natural language free-form input may include one or more annotations of the natural language input and optionally include one or more (e.g., all) terms of the natural language input.
[0037] In some implementations, the NLP module 122 is configured to identify and annotate various types of syntactic information in natural language input. For example, the NLP module 122 may include a speech tagger that is configured to annotate terms by their syntactic roles. In some implementations, the NLP module 122 may additionally and / or alternatively include an entity tagger (not depicted) that is configured to annotate entity references in one or more segments (such as references to people (including, for example, literary characters, celebrities, public figures, etc.), organizations, locations (real and fictional), etc.). In some implementations, data about entities may be stored in one or more databases, such as stored in a knowledge graph (not depicted). In some implementations, the knowledge graph may include nodes representing known entities (and in some cases, entity attributes), and edges connecting the nodes and representing relationships between the entities.
[0038] The entity tagger of the NLP module 122 may annotate references to entities at a high granularity level (e.g., such that all references to an entity category (such as people) can be identified) and / or at a lower granularity level (e.g., such that all references to a specific entity (such as a specific person) can be identified). The entity tagger may rely on the content of the natural language input to resolve specific entities and / or may optionally communicate with a knowledge graph or other entity database to resolve specific entities.
[0039] In some implementations, the NLP module 122 may additionally and / or alternatively include a coreference resolver (not depicted) that is configured to group or "cluster" references to the same entity based on one or more contextual cues. For example, the coreference resolver may be utilized to resolve the term "it" in the natural language input "lockit" to "front door lock" based on the "front door lock" mentioned in a client device notification rendered immediately prior to receiving the natural language input "lockit".
[0040] In some implementations, one or more components of the NLP module 122 can depend on annotations from one or more other components of the NLP module 122. For example, in some implementations, a named entity tagger can depend on annotations from a coreference resolver and / or a dependency parser to annotate all mentions of a specific entity. Also, for example, in some implementations, a coreference resolver can depend on annotations from a dependency parser to cluster referents of the same entity. In some implementations, when processing a specific natural language input, one or more components of the NLP module 122 can use related data other than the specific natural language input to determine one or more annotations—such as an assistant input device notification rendered immediately prior to receiving the natural language input on which the assistant input device notification is based.
[0041] In some implementations, one or more assistant input devices 106 can include a second automation assistant 120 that includes one or more components sharing features with the components described herein with respect to the automation assistant client 118 and / or the cloud-based components 119. For example, in addition to or instead of including the automation assistant client 118, one or more assistant input devices 106 can include a second automation assistant 120 that can include a voice capture component, TTS, STT, NLP, and / or one or more fulfillment engines for processing requests received from a user. In some implementations, the second automation assistant 120 can include one or more cloud-based components that are different from the cloud-based components 119. For example, the second automation assistant 120 can be a stand-alone automation assistant having the ability to process audio data, recognize one or more wake words and / or invocation phrases, process additional audio data to identify a request included in the audio data, and cause one or more actions to be performed in response to the request.
[0042] Reference Figure 2 , for illustrative purposes, a diagram of a connected device ecosystem is provided. In some implementations, the ecosystem can include more components, fewer components, and / or different components than those shown in Figure 2 . However, for example, for the purposes of the example described further in relation to Figure 3 - FIG. 6, the ecosystem will be used as an example that includes a kitchen speaker 205, a bedroom speaker 210, and a living room speaker 215, each speaker including one or more components of the assistant input device 106. For example, the kitchen speaker 205 can be executing the automation assistant client 118, which further includes Figure 1 one or more components of the automation assistant client 118 shown in Figure 2The ecosystem can include devices that are connected to each other and are further located in the user environment but at different locations. For example, a user can place the kitchen speaker 205 on the kitchen countertop, the bedroom speaker 210 in the bedroom, and the living room speaker in the living room. Thus, for explanatory purposes only, the speakers of devices 205, 210, and 215 can be heard by the user (and thus can respond to requests) and / or the microphones of devices 205, 210, and 215 can capture audio data of the user speaking a call phrase and / or a request when the user is near the device (i.e., in the same vicinity).
[0043] In addition, for exemplary purposes only, the examples described herein will assume that the kitchen speaker 205 and the bedroom speaker share one or more cloud-based assistant components 219 that perform one or more actions. However, it should be understood that one or more assistant input devices 106 can include Figure 1 one or more of the components shown in as components of the cloud-based assistant and can be executed by devices of the ecosystem that include one or more components. For example, the living room speaker 215 can include one or more of the cloud-based assistant components 219 and, in some configurations, can receive and process requests in the same manner as described herein for the kitchen speaker 205 and the bedroom speaker 210 that share the cloud-based automation assistant component 219 (i.e., the living room speaker 215 can be a device that does not include the automation assistant client 118 but only includes the second automation assistant 120). In addition, one or more other components (such as information related to the context of the session between the automation assistant and the user) can be accessible to one or more assistant input devices 106. Thus, for a session that starts with the automation assistant client of the kitchen speaker 205, the context can be available to the bedroom speaker 210 but not to the living room speaker 215.
[0044] In some implementations, an automation assistant executed on one device in the device ecosystem can be invoked by a user performing an action (e.g., touching the device, performing a gesture captured by the device's camera) and / or speaking a call phrase that indicates the user's interest in performing one or more actions on the automation assistant. For example, the user can say "OK Kitchen Assistant", and the automation assistant of the kitchen speaker 205 can process the audio before and / or after the call to determine whether the audio data includes a request. The audio data captured by the microphone of the kitchen speaker 205 can be processed by the automation assistant client running on the device and / or at least partially by one or more of the cloud-based components 219 using STT, NLP, and / or ASR.
[0045] Once it is determined that the audio data includes a request, the action processing engine 180 (shown in Figure 1 as a component in the cloud-based automated assistant component 119, but can additionally or alternatively be a component of the automated assistant client 118) can determine one or more actions to perform and cause the action to be executed. For example, a user may say "OK KitchenAssistant", followed by "how tall is Jack Smith". In response, the action processing engine 180 can generate a response to the request (or "query") and provide the response to the query via the microphone 109 of the kitchen speaker 205.
[0046] In some implementations, the user can activate a continue session mode before starting a session with one or more automated assistants, which allows the user to continue subsequent requests without explicitly re-invoking the automated assistant. For example, in some implementations, after the start action and / or after the execution of an action, the invoked automated assistant can continue to process the audio data to determine whether the audio data includes a follow-up request and / or query. For example, a user can use the invocation phrase "OK Assistant, what is the weather in Miami" to invoke the automated assistant, and the automated assistant can generate and provide a response. Subsequently, the automated assistant can continue to process the audio data to determine whether the user has provided a follow-up request. This can include, for example, the automated assistant client locally processing the audio data, not providing full processing until it is determined that the audio data includes a follow-up request, only processing a limited amount of audio data (e.g., 5 seconds after the response to the first query has been provided), and / or other limited processing that can determine whether the user has an additional request after the first request. In some implementations, the automated assistant client 118 and / or one or more cloud-based automated assistant components 119 can determine whether a query is likely to be followed by a follow-up query. For example, for the initial query "What is the weather in Miami", the automated assistant 118 can determine that the user may issue an additional query such as "How about in Orlando" (e.g., a related query) and process the audio data to determine whether the user has submitted an additional query.
[0047] In some implementations, the automated assistant can process audio data when and / or after an action is performed in response to a user's request. For example, the user can provide a request such as "Play the song we will rock you", and in response, the automated assistant can cause the song "We Will Rock You" to be played via the speaker of the device on which the automated assistant is being executed. Once the song is completed and / or as the song is about to be completed, the automated assistant can start processing the audio data to wait for the user to request a new song after the current song is completed. Thus, in some implementations, computational resources are saved by only processing the audio data when the user is likely to make a new request, rather than continuously processing and wasting resources when the user is less likely to make a new request.
[0048] Once the invoked automated assistant 118 determines that there may be a follow-up request (e.g., the "continue conversation" mode is active and the request is of a type that can be followed by another request without an intermediate invocation phrase), the automated assistant 118 can send a notification to one or more other devices that the user can provide a follow-up request, each device executing the automated assistant and / or an automated assistant client. Referring again to Figure 1 , the assistant input device 106 includes a co-location monitor 130 that can determine whether the user is co-located with the device based on sensor data generated by the sensors 105, camera 111, and / or microphone 109. As shown, the co-location monitor 130 is shown as a component of the assistant input device 106. However, in some implementations, the co-location monitor 130 can be a component of the automated assistant client 118, the second automated assistant 120, the cloud-based automated assistant component 119, and / or shared among one or more automated assistant components.
[0049] In some implementations, a notification can be provided to one or more other devices in the device ecosystem when it is detected that the user is moving away from the invoked device and / or the user is currently changing location. For example, the invoked automated assistant 118 can process sensor data from one or more sensors of the assistant device 106 to determine whether the user is located near the device 106. If the automated assistant 118 determines based on the sensor data that the user is changing location to a position farther from the device 106, the automated assistant 118 can provide a notification to other devices in the device ecosystem so that those devices can start processing the sensor data to determine whether the user is changing location to a position closer to one of those other devices.
[0050] In some implementations, the invoked automated assistant 118 can utilize the notification to provide an indication of the user who invoked the automated assistant (and subsequently provided the request). For example, a user can invoke the kitchen speaker 205 by saying "OK Kitchen Speaker", and one or more components of the automated assistant and / or one or more shared components 219 executing on the kitchen speaker 205 can utilize one or more TD speaker recognition models to determine a vector based on speech. The vector can be utilized to identify a user profile associated with the speaker by comparing the vector with one or more stored vectors of the user, such as those generated during user registration (e.g., the user can be prompted to say "OK Kitchen Speaker" one or more times, and the profile of the user saying the phrase can be stored along with the user profile). Also, for example, the kitchen speaker 205 can utilize one or more TI models to process a request provided by the user and can generate a TI vector that can be provided to one or more other devices of the ecosystem along with the notification. Thus, in some implementations, a user profile indicator and / or a vector representing at least a portion of the user's utterance can be provided to, for example, the bedroom speaker 210 and / or the living room speaker 215 such that other devices can determine whether the co-present user is the same user as the user who uttered the initial request (and / or the invocation phrase) based on the processing of audio data captured by the device.
[0051] The assistant input device 106 further includes a speaker recognition engine 140 that can determine whether the user associated with the user profile is the same as the user who uttered the invocation phrase and / or the request based on processing audio and / or camera data.
[0052] Reference Figure 3 , shows a timing diagram that illustrates one or more implementations described herein. As shown, the kitchen speaker 205 processes the invocation 305 and further processes the request 310. In response, the kitchen speaker 205 performs one or more actions 312. In addition to processing the request 310, the kitchen speaker 205 (i.e., the automated assistant executing on the kitchen speaker 205) can further determine that there may be a follow-up request, which can be determined by the content of the request and that the "continue conversation" mode is active. Once the action is performed 312 (or while the action is being performed), the kitchen speaker 205 can process 315 subsequent audio data to identify a follow-up request that may be included in the subsequent audio data.
[0053] The kitchen speaker 205 provides a notification 320 to other devices in the linked device ecosystem. In some implementations, the notification 320 can include an indication that the user may be co-located with the device and determine whether / when this occurs. In some implementations, call phrases and / or requests can be processed to determine the user and / or user profile, and an indication of the user or the user speaker profile (e.g., a vector representing the user's speaking voice) can be provided along with the notification 320.
[0054] In response to receiving the notification 320, the bedroom speaker 210 and the living room speaker 215 can process sensor data 325 using components that share one or more characteristics with the co-location monitor 130 to determine whether the user is co-located with each device. The co-location monitor 130 can use sensor data from, for example, one or more sensors 105, microphones 109, and / or cameras 11 to determine whether the user is present near the respective devices 210 and 215. If the user is detected near a device 335, such as the bedroom speaker 210, additional processing can be performed to determine whether the co-located user is the same user who uttered the call phrase and / or provided the request. Thus, as an additional optional step, the user profile provided with the notification at 320 can be verified at 340. As shown, the bedroom speaker 210 processes its sensor data 325, while the living room speaker 215 processes its sensor data 330. Additionally, as shown, the bedroom speaker 210 determines user co-location 335 and optionally verifies that the co-located user is the same user who uttered the call and / or initial request indicated by the kitchen speaker 205.
[0055] In response to determining user co-location, the bedroom speaker 210 can provide one or more indications 345 to the kitchen speaker 205 (and optionally, to other devices in the ecosystem) to indicate that the user is now co-located with the bedroom speaker 210. As shown, once the indication is received, the living room speaker 215 stops processing the sensor data 330. However, in some implementations, the living room speaker 215 can continue to process the sensor data to determine whether the user has moved and subsequently determine whether the user is co-located with the living room speaker 215. For example, the user may be quickly walking from the kitchen to the bedroom and then to the living room. Thus, in some implementations, the user's co-location with the bedroom speaker 210 can be only temporary, and after determining temporary co-location with the bedroom speaker 210 via the bedroom speaker 210, the user can subsequently be co-located with the living room speaker 215.
[0056] When the kitchen speaker 205 receives an indication that the user (or a user) is co-located with the bedroom speaker 210, the kitchen speaker 205 may stop processing subsequent audio data 350. Additionally, the bedroom speaker 210 may start processing subsequent audio data 360. In some implementations, the bedroom speaker 210 may start processing subsequent audio data 360 before it sends the indication 345 to ensure that there is always at least one device processing the subsequent audio data. In some implementations, the method may continue where the bedroom speaker 210 sends an indication to other devices that it is processing the subsequent audio data and a further notification that there may be a follow-up request. Other devices (including the kitchen device 205) may process sensor data to determine if the user has moved again and may be co-located with one of the other devices.
[0057] In some implementations, the kitchen device 205 may provide the previously requested context 355 to the co-located device (e.g., 210) before the co-located device starts processing the subsequent audio data. In some cases, the co-located device may already have the context, such as when the invoked automated assistant and the co-located automated assistant are both components of the same automated assistant sharing one or more components (e.g., two clients of the same automated assistant where the context is stored on a cloud-based component shared by the two clients). However, in the case where the invoked automated assistant and the co-located automated assistant are different automated assistants, the co-located automated assistant may need the previously requested context before it can process the subsequent request.
[0058] For example, referring to Figure 2 , the user may invoke the automated assistant executing on the kitchen speaker 205 with "OK Kitchen Assistant" and further provide a request of "how old is Jack Smith?". The request may be processed and a notification indicating that there may be a follow-up request may be further provided to the bedroom speaker 210 and the living room speaker 215. Subsequently, the living room speaker 215, which does not share components with the kitchen speaker 205, may determine that the user is co-located. The living room speaker 215 may send an indication to the kitchen speaker 205, and the kitchen speaker 205 may provide the previously requested context. The living room speaker 215 may need the context to process the follow-up request, such as determining the meaning of "he" in the follow-up request of "how old is he?". However, in the case where the bedroom speaker 210 instead provides the indication of the user being co-located, the context may not be necessary because the kitchen speaker 205 and the bedroom speaker 210 both share the cloud-based component 219 and thus may already share the previously requested context information as well as other resources.
[0059] In some implementations, a user may use a single invocation phrase to invoke multiple automated assistants, and both automated assistants may initially process the request. For example, the user may say the invocation phrase "OK Assistant 1 and Assistant 2, show me a weather map of Miami", and in response, both the automated assistant client 118 and the second automated assistant 120 may process the request. One or both of the automated assistants may determine that there may be a follow-up request based on, for example, the content of the request and / or the activation of a continuing session mode. In some implementations, one of the invoked automated assistants may determine that a follow-up request is unlikely to be directed to it. For example, the second automated assistant 120 may not be configured to process the type of request that may follow the initial request. In response, the second automated assistant 120 may stop processing subsequent audio data, thus saving processing resources for assistant input and further limiting the potential for sensitive information to be unnecessarily provided to components that do not need the information.
[0060] Similarly, in some implementations, a user may invoke multiple automated assistants running on different devices, and each automated assistant may begin processing subsequent audio data to determine if the user has provided a follow-up request. For example, the user may invoke all of the automated assistants in Figure 3 by calling "OK Speakers", followed by the request "Show me a weather map of Miami". As an example, the living room speaker 215 may not include a graphical interface and thus may not be able to display a weather map. In response, if the automated assistant executing on the living room speaker 215 determines that a follow-up request may require a graphical interface that it does not have, the living room speaker 215 may stop processing subsequent audio data and thus may not be able to process any follow-up request "Show me a weather map of Miami".
[0061] Reference Figure 4A and Figure 4B provide flowcharts showing a method for determining whether to continue processing subsequent audio data when a user is co-located with another device in a linked device ecosystem. For convenience, the operations of the method are described with reference to a system that performs these operations - such as the system shown in Figure 1 . This system of the method includes one or more processors and / or other components of a client device. Additionally, while the operations of the method are shown in a specific order, this is not intended to be limiting. One or more operations may be reordered, omitted, or added.
[0062] At step 405, audio data after and / or before the processing of the automated assistant call is processed. The call phrase can call an automated assistant that shares one or more features with the automated assistant client 118 and / or the second automated assistant 120. For example, in some implementations, the called automated assistant 118 can be a client that is executed partially on the assistant input device 106 and partially on another device, and has a cloud-based automated assistant component 119. In some implementations, the called automated assistant 120 can be a stand-alone automated assistant that is executed only on the assistant input device 106. In some implementations, the called automated assistant can be executed on a device that is part of a linked device ecosystem located in the user environment. For example, the called automated assistant can share features with Figure 2 one of the devices shown in
[0063] At step 410, the called automated assistant determines that the audio data includes a request for the automated assistant to perform one or more actions. The called automated assistant 118 or 120 can determine that the audio data before and / or after the call to the automated assistant includes a request for the automated assistant to perform one or more actions by processing the spoken words. For example, one or more components of the automated assistant can utilize ASR, NLP, and / or STT to process the audio data to determine the user's intent and provide instructions to the action processing engine 180 to perform one or more actions. For example, the user can say "how tall is Jack Smith", and the query can be processed and provided to the action processing engine 180, and the action processing engine 180 can provide the query to one or more search engines to determine a response.
[0064] At step 415, the automated assistant causes one or more actions to be performed. For example, the action processing engine 180 can provide instructions to one or more other components to perform actions (such as changing the state of a smart home appliance), respond to a query, and / or cause one or more other actions to be performed in response to a user's request. In some implementations, the context of the request and / or response can be locally stored by the automated assistant fulfilling the request, such that for subsequent responses, the content of the previous request and / or response can be utilized to determine the intent and / or meaning of one or more terms in the subsequent request. For example, for the request "how tall is Jack Smith", the context information can include "Jack Smith", such that for a subsequent request "how old is he", the term "he" can be understood to mean "Jack Smith".
[0065] At step 420, the invoked automated assistant determines that the subsequent audio data may include a further assistant request. In some implementations, the user can activate a "continue session" mode such that the automated assistant 118 can continue to process a limited portion of the subsequent audio data to determine if a request has been spoken, without the user having to speak the invocation phrase again. Also, for example, in some implementations, the automated assistant 118 can determine that the initially processed request is of a type that may be followed by a related request. For example, the user can provide a request of "What’s the weather like today", and the automated assistant 118 can determine that the user may (or might) follow that request with a related request such as "how about tomorrow".
[0066] At step 425, the automated assistant begins to process the subsequent audio data to determine if the subsequent audio data includes a further assistant request. For example, the user can speak a request of "play the song we will rock you", and in response, the automated assistant can cause the requested song to be played. Near the end of the song playback (or after the song has started and / or ended), the automated assistant can begin to process the audio data to determine if the user has requested a new song. In some implementations, after processing the subsequent audio data for a period of time without determining that a request is included in the audio data, the automated assistant can stop processing the audio data. For example, the automated assistant client 118 can process 5 seconds of audio and stop processing the subsequent audio data if no subsequent request is included in the audio data, until the automated assistant is invoked again.
[0067] At step 430, the automated assistant transmits a notification to other devices in the linked device ecosystem that includes the device on which the automated assistant is executing. The notification can indicate that there may be further assistant requests and can cause the notified devices (e.g., the automated assistants executing on each device) to process sensor data generated by the sensors of the corresponding device to determine whether the user (or a user) is co-located with the device. For example, subsequent audio data can first be processed by the kitchen speaker automated assistant. Notifications can be provided to devices in the bedroom and living room indicating that the user may change location but may still be interested in continuing the conversation. Thus, by providing the notification, other devices can utilize one or more techniques to determine whether the user is co-located (e.g., process audio data, camera data). In some implementations, an indication of the user who made the call can be provided with the notification. For example, the automated assistant that was called can use text-dependent and / or text-independent speaker recognition to process the call (and / or request) to identify the associated user profile. The other devices can then perform limited processing on the audio data (or visual data) to determine whether the user determined to be co-located with the device is the same user whose profile was provided with the notification. At step 435, the automated assistant receives an indication that the user (or a user) is co-located with one of the devices in the linked device ecosystem. In some implementations, the co-located device can begin processing subsequent audio data.
[0068] At step 440, the automated assistant stops processing subsequent audio data. Processing of the subsequent audio data can be stopped in response to receiving the indication. In some implementations, the automated assistant that made the initial call can further provide context to the automated assistant of the co-located device. For example, for an initial request of "how tall is Jack Smith", context can be provided to the co-located automated assistant (e.g., if the co-located automated assistant does not already have context) such that for a subsequent request of "how old is he", the intent of the term "he" can be parsed (e.g., to refer to "Jack Smith").
[0069] Reference Figure 5 , a flowchart is provided that illustrates a method for determining whether to process subsequent audio data when a user is co-located with a device in a linked device ecosystem. For convenience, the operations of the method are described with reference to a system that performs these operations - such as Figure 1 the system shown in
[0070] At step 505, receive a notification indicating that subsequent audio data may include further assistant requests from an uninvoked automated assistant and from an invoked automated assistant. This notification may share one or more characteristics with the notification previously described with respect to step 430. For example, the notification may include a reference to the user profile of the user who uttered the invocation and / or initial request and / or information related to the user profile.
[0071] At step 510, process sensor data of the device executing the automated assistant to determine if a user is co-located. The sensor data may include data generated by a microphone, a camera, an accelerometer, and / or other sensors previously described with respect to sensor 105. In some implementations, co-location may be determined by a component that shares one or more characteristics with the co-location monitor 130. Additionally, in some implementations, a component that shares one or more characteristics with the speaker recognition engine 140 may process limited audio data (and additionally or alternatively, visual data) to determine the profile of the user co-located with the device. For example, the speaker recognition engine 140 may utilize a text-independent speaker recognition model to process a limited portion of the audio data, compare the output to the embeddings of users who uttered one or more other phrases, and determine if the user who uttered the one or more other phrases is the same as the co-located user (i.e., the similarity between the embeddings of the voice profiles is within a threshold). If it is determined that a user (or the user) is co-located, then at step 515, provide an indication to the invoked automated assistant to indicate that a user (or the user) is co-located. In response, the initially invoked automated assistant may stop processing subsequent audio data, as previously described. At step 520, the co-located automated assistant begins processing subsequent audio data, as previously described.
[0072] Reference Figure 6A - Figure 6B , a flowchart is provided showing a method for determining whether to process subsequent audio data when a user is co-located with a device in a linked device ecosystem. For convenience, the operations of the method are described with reference to a system that performs these operations - such as Figure 1 the system shown in
[0073] At step 605, a call to invoke multiple automated assistants is received. For example, the phrase "OK Assistant 1" can be used to invoke the first automated assistant, and the call phrase "OK Assistant 2" can be used to invoke the second automated assistant, and the user can say calls such as "OK Assistants", "OK Assistant 1 and Assistant 2", and / or another call phrase that can be recognized by multiple automated assistants by processing the audio data. In some implementations, the user can say a call phrase that can invoke multiple assistants running on the same device (e.g., "Assistant 1" can be, for example, automated assistant client 118, and "Assistant 2" can be automated assistant 120). In some implementations, the invoked automated assistants can execute on separate devices. For example, the user can say call phrases such as "OK Speaker 1 and Speaker 2" and / or "OK Speakers", and the automated assistant executing on both "Speaker 1" and "Speaker 2" can be invoked by this phrase.
[0074] At step 610, the request after and / or before the call is processed. The request can be processed by one or more of the invoked automated assistants. For example, automated assistant client 118 can process the audio data before and / or after the call as described above, and the second automated assistant 120 can process the audio data, each using its separate components (e.g., ASR, NLP, STT). In some implementations, one or more automated assistants can perform one or more actions in response to a request included in the audio data.
[0075] At step 615, one of the automated assistants processing the request determines that there may be a further assistant request. This step can share one or more features with Figure 4A step 420. For example, the user can activate the "continue session" mode, which allows the user to say subsequent requests without the user invoking the automated assistant a second time. Also, for example, the request can include one or more terms and / or be of a type that may be followed by an additional request. For example, the user can say the request "what is the weather today", and one or more of the invoked automated assistants can determine that the user may follow this request with an additional request (e.g., "what about tomorrow").
[0076] At step 620, the invoked automated assistant processes subsequent audio data. This processing can be performed by any one or all of the initially invoked automated assistants. In some implementations, the automated assistant can process limited audio data to determine whether a request is included, and continue processing only if the automated assistant determines that the audio data may contain a request.
[0077] At step 625, a subsequent request is received. The subsequent request can be included in the subsequent audio data being processed by the automated assistant. For example, a user can say "OK Assistants, what is the weather today", and "Assistant 1" can respond with "it is going to be 75 and sunny today". Both "Assistant 1" and "Assistant 2" can continue to process the subsequent audio data to determine whether the user has provided an additional request. When the user says "show me a weather map", "Assistant 1" and "Assistant 2" can process this phrase to determine that the user has spoken a follow-up request.
[0078] At step 630, one of the automated assistants determines that the subsequent request is not directed to it but to another automated assistant. For example, the user can provide a request of "show me a weather map" as a follow-up request, and "Assistant 2" can execute it on a smart speaker not equipped with a visual display. Thus, for example, "Assistant 2" can determine that the follow-up request of "show me a weather map" is not directed to it. In response, at step 635, the automated assistant that is not the target of the subsequent request stops processing the subsequent audio data.
[0079] Figure 7 Is a block diagram of an example computing device 710 that can optionally be used to perform one or more aspects of the techniques described herein. Computing device 710 generally includes at least one processor 714 that communicates with a plurality of peripheral devices via a bus subsystem 712. These peripheral devices can include a storage subsystem 724 (including, for example, a memory subsystem 725 and a file storage subsystem 726), a user interface output device 720, a user interface input device 722, and a network interface subsystem 716. The input and output devices allow a user to interact with computing device 710. The network interface subsystem 716 provides an interface to an external network and is coupled to a corresponding interface device in other computing devices.
[0080] The user interface input device 722 may include a keyboard, a pointing device (such as a mouse, trackball, touchpad, or graphics tablet), a scanner, a touch screen incorporated into a display, an audio input device (such as a voice recognition system, microphone), and / or other types of input devices. Generally speaking, the use of the term "input device" is intended to include all possible types of devices and ways of inputting information into the computing device 710 or into a communication network.
[0081] The user interface output device 720 may include a display subsystem, a printer, a fax machine, or a non-visual display (such as an audio output device). The display subsystem may include a cathode ray tube (CRT), a flat panel device (such as a liquid crystal display (LCD)), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide a non-visual display, such as via an audio output device. Generally speaking, the use of the term "output device" is intended to include all possible types of devices and ways of outputting information from the computing device 710 to a user or to another machine or computing device.
[0082] The storage subsystem 724 stores programming and data constructs that provide some or all of the functionality described in the modules herein. For example, the storage subsystem 724 may include selected aspects for performing the methods of FIGS. 4 through 6 and / or the logic for implementing Figure 1 the various components depicted
[0083] These software modules are typically executed by the processor 714, either alone or in combination with other processors. The memory 725 used in the storage subsystem 724 may include multiple memories, including a main random access memory (RAM) 730 for storing instructions and data during program execution and a read-only memory (ROM) 732 in which fixed instructions are stored. The file storage subsystem 726 may provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical disk drive, or a removable media cartridge. Modules implementing the functionality of certain implementations may be stored in the storage subsystem 724 by the file storage subsystem 726 or in other machines accessible by the processor 714.
[0084] The bus subsystem 712 provides an mechanism for enabling the various components and subsystems of the computing device 710 to communicate with each other as intended. Although the bus subsystem 712 is schematically shown as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0085] The computing device 710 may be of different types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, forFigure 7 The description of the computing device 710 depicted is only intended as a specific example for illustrative purposes of some implementations. Many other configurations of the computing device 710 are possible, which have more or fewer components compared to Figure 7 the depicted computing device.
[0086] In some implementations, a method is disclosed, and the method includes the following steps: in response to determining that a user's invocation input is intended for a first assistant device in a linked assistant device ecosystem: processing audio data captured via one or more first microphones of the first assistant device and after and / or before the invocation input; based on processing the audio data, determining that an assistant request is included in the user's spoken words captured by the audio data; causing one or more actions corresponding to the assistant request to be executed; based on the assistant request, determining that subsequent audio data may include a further assistant request. In response to a continue assistant session mode being active and in response to determining that subsequent audio data may include a further assistant request, the method further includes the following steps: processing subsequent audio data captured via the first microphone of the first assistant device to determine whether the subsequent audio data includes a further assistant request; and transmitting a notification to a plurality of additional assistant devices in the linked assistant device ecosystem such that each additional assistant device temporarily processes corresponding device-specific sensor data to determine whether the user is co-located with the additional assistant device. The method further includes the following steps: before any determination that the subsequent audio data includes a further assistant request, receiving an indication from a specific additional assistant device among the additional devices that the additional assistant device has determined that the user is co-located with the additional assistant device; and in response to receiving the indication: ceasing to process the subsequent audio data captured via the first microphone of the first assistant device.
[0087] These implementations and other implementations of the techniques disclosed herein may include one or more of the following features.
[0088] In some implementations, the method further includes the following steps: determining a user identifier of the user; and providing the user identifier along with the notification. In some of those implementations, determining the user identifier includes: processing the invocation input using one or more text-dependent speaker verification models to generate a speaker profile; and identifying the user identifier based on an association between the user and the speaker profile. In other of those implementations, determining the user identifier includes: processing the audio data using one or more text-independent speaker verification models to generate a speaker profile; and identifying the user identifier based on an association between the user and the speaker profile.
[0089] In some implementations, the method further includes transmitting context data generated based on one or more actions to a specific attachment device, where the specific attachment device uses the context data to determine whether the attached subsequent audio data includes a further assistant request and / or to process a further assistant request included in the attached subsequent audio data.
[0090] In some implementations, the device-specific sensor data of at least one additional assistant device includes sensor data generated by one or more of a microphone, an accelerometer, and / or a camera.
[0091] In other implementations, another method is disclosed, and the method includes the steps of: receiving a notification from an invoked automated assistant executing on a first device in a linked assistant device ecosystem that further includes a client device, the notification indicating that the invoked automated assistant is processing additional audio input received via one or more microphones of the first device to determine whether the additional audio data includes a further assistant request; and processing sensor data by an additional automated assistant executing on the client device to determine whether the user is co-located with the client device, where the sensor data is generated by one or more sensors of the client device. In response to determining that the user is co-located with the client device, the method further includes: providing an indication to the invoked automated assistant that the additional automated assistant has determined that the user is co-located with the client device, where providing the indication causes the invoked automated assistant to stop processing the audio input; and processing additional subsequent audio data captured via one or more microphones of the client device to determine whether the additional subsequent audio data includes a further assistant request.
[0092] These implementations and other implementations of the techniques disclosed herein may include one or more of the following features.
[0093] In some implementations, providing the indication causes the invoked automated assistant to stop processing the subsequent audio data.
[0094] In some implementations, the method further includes: receiving an identifier of a user who invoked the automated assistant; determining, based on limited processing of the subsequent audio data and / or visual data, whether the co-located user is the user who invoked the automated assistant; and transmitting a notification only if the co-located user is identified as the user associated with the identifier. In some of those implementations, determining whether the co-located user is the user who invoked the automated assistant includes performing text-dependent and / or text-independent analysis on at least a portion of the subsequent audio data to generate an embedding, and the method further includes: comparing the embedding with a previously generated embedding associated with the user's profile, where the profile is identified based on the identifier.
[0095] In some implementations, the method further includes: receiving, from a called automated assistant, context information related to one or more previous requests of a user and / or one or more responses to the user. In some of those implementations, the method further includes: using the context information to process subsequent audio data to determine one or more actions to be performed in response to a subsequent request included in the subsequent audio data.
[0096] In some implementations, a system is disclosed and the system includes: linking a first assistant device and a second assistant device in an assistant device ecosystem, wherein the first assistant device is configured to: in response to determining that a call input of a user is intended for the first assistant device: process audio data captured by one or more first microphones of the first assistant device and after and / or before the call input; based on processing the audio data, determine that an assistant request is included in the user's spoken words captured by the audio data; cause one or more actions corresponding to the assistant request to be performed; based on the assistant request, determine that subsequent audio data may include a further assistant request; in response to determining that the subsequent audio data may include a further assistant request, the first device is further configured to: process the subsequent audio data captured by the first microphone of the first assistant device to determine whether the subsequent audio data includes a further assistant request; and transmit a notification to a second assistant device in the linked assistant device ecosystem, the notification indicating that the subsequent audio data may include a further assistant request. The second assistant device is configured to: process sensor data to determine whether a user is co-located with the second assistant device, wherein the sensor data is generated by one or more sensors of the second assistant device; in response to determining that the co-located user is co-located with the second assistant device: provide an indication to the first assistant device that the co-located user is co-located with the second assistant device, wherein providing the indication causes the first assistant device to stop processing audio input; and process additional subsequent audio data captured by one or more microphones of a client device to determine whether a further assistant request is included in the additional subsequent audio data.
[0097] These implementations and other implementations of the techniques disclosed herein may include one or more of the following features.
[0098] In some implementations, the first assistant device is further configured to: determine a user identifier of the user; and provide the user identifier to the second device, and the second assistant device is further configured to: determine whether the co-present user is the user based on at least a portion of the user identifier and the sensor data; and provide an indication only if the co-present user is the user. In some of those implementations, when determining the user identifier, the first assistant device is further configured to: process the call input using one or more text-dependent speech models to generate a speaker profile; and identify the user identifier based on the association between the user and the speaker profile.
[0099] In some implementations, when determining the user identifier, the second assistant device is further configured to: process the audio data using one or more text-independent speech models to generate a speaker profile; and identify the user identifier based on the association between the user and the speaker profile.
[0100] In some implementations, the sensor data includes visual data captured by one or more cameras of the second assistant device.
[0101] In some implementations, the first assistant device is further configured to: transmit context data generated based on one or more actions to the second assistant device. In some of those implementations, the second assistant device is further configured to: use the context data to determine a further assistant request included in the additional subsequent audio data.
[0102] In some implementations, both the first assistant device and the second assistant device are execution instances of the same automated assistant in combination with one or more shared cloud-based components.
[0103] In certain implementations discussed herein that may collect or use personal information about a user (e.g., user data extracted from other electronic communications, information about the user's social network, the user's location, the user's time, the user's biometric information, and the user's activities and demographic information, relationships between users, etc.), the user is provided with one or more opportunities to control whether information is collected, whether personal information is stored, whether personal information is used, and how information about the user is collected, stored, and used. That is, the systems and methods discussed herein collect, store, and / or use user personal information, particularly when explicit authorization to do so is received from the relevant user.
[0104] For example, users are provided with control over whether a program or feature collects user information about that specific user or other users associated with the program or feature. One or more options are presented to each user whose personal information is to be collected to allow control over the information collection associated with that user, providing permission or authorization regarding whether information is collected and regarding which portions of the information are collected. For example, one or more such control options can be provided to the user via a communication network. Additionally, specific data can be processed in one or more ways before it is stored or used such that personally identifiable information is removed. As an example, a user's identity can be processed such that no personally identifiable information can be determined. As another example, a user's geographic location can be generalized to a larger area such that the user's specific location cannot be determined.
[0105] Although several implementations have been described and illustrated herein, various other means and / or structures can be used for performing the functions and / or obtaining the results and / or one or more of the advantages described herein, and each such variation and / or modification is considered to be within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and the actual parameters, dimensions, materials, and / or configurations will depend on one or more specific applications for which this teaching is used. Those skilled in the art will recognize or be able to ascertain using no more than routine experimentation many equivalents to the implementations described herein. Accordingly, it is to be understood that the foregoing implementations are presented by way of example only, and it is to be understood that implementations can be practiced otherwise than as specifically described and claimed within the scope of the appended claims and their equivalents. Implementations of the present disclosure relate to each and every individual feature, system, article of manufacture, material, kit, and / or method described herein. Additionally, any combination of two or more such features, systems, articles of manufacture, materials, kits, and / or methods is included within the scope of the present disclosure if such features, systems, articles of manufacture, materials, kits, and / or methods are not mutually inconsistent.
Claims
1. A method implemented by one or more processors, comprising: In response to determining that a user's invocation input is directed to a first assistant device in a linked assistant device ecosystem: Processing audio data captured by one or more first microphones of the first assistant device and after and / or before the invocation input; Based on processing the audio data, determining that an assistant request is included in the user's spoken words captured by the audio data; Causing one or more actions corresponding to the assistant request to be executed; Based on the assistant request, determining that subsequent audio data may include further assistant requests; In response to a continued assistant session mode being active and in response to determining that subsequent audio data may include further assistant requests: Processing subsequent audio data captured by the first microphone of the first assistant device to determine whether the subsequent audio data includes further assistant requests; And Transmitting a notification to a plurality of additional assistant devices in the linked assistant device ecosystem such that each of the additional assistant devices temporarily processes corresponding device-specific sensor data to determine whether the user is co-located with the additional assistant device; Before any determination that the subsequent audio data includes further assistant requests, receiving an indication from a specific additional assistant device among the additional devices that the user is co-located with the additional assistant device; and In response to receiving the indication: Stopping processing the subsequent audio data captured by the first microphone of the first assistant device.
2. The method of claim 1, further comprising: Determining a user identifier of the user; And Providing the user identifier along with the notification.
3. The method according to claim 2, wherein, Determining the user identifier includes: Processing the invocation input using one or more text-dependent speaker verification models to generate a speaker profile; and Identifying the user identifier based on an association between the user and the speaker profile.
4. The method according to claim 2, wherein, Determining the user identifier includes: Processing the audio data using one or more text-independent speaker verification models to generate a speaker profile; and Identifying the user identifier based on an association between the user and the speaker profile.
5. The method of claim 1, further comprising: Transmitting context data generated based on the one or more actions to the specific additional device, wherein the specific additional device uses the context data to determine whether the additional subsequent audio data includes further assistant requests and / or to process further assistant requests included in the additional subsequent audio data.
6. The method according to claim 1, wherein, The device-specific sensor data of at least one of the additional assistant devices includes sensor data generated by one or more of a microphone, an accelerometer, and / or a camera.
7. A method implemented by one or more processors of a client device, comprising: Receive a notification from an invoked automated assistant executing on a first device in the link assistant device ecosystem that also includes the client device, the notification indicating that the invoked automated assistant is processing additional audio input received via one or more microphones of the first device to determine whether further assistant requests are included in the additional audio data; Process sensor data by an additional automated assistant executing on the client device to determine whether the user is co-located with the client device, wherein the sensor data is generated by one or more sensors of the client device; In response to determining that the user is co-located with the client device: Provide an indication to the invoked automated assistant that the additional automated assistant has determined that the user is co-located with the client device, wherein providing the indication causes the invoked automated assistant to stop processing the audio input; and Process additional subsequent audio data captured via one or more microphones of the client device to determine whether further assistant requests are included in the additional subsequent audio data.
8. The method according to claim 7, wherein, Providing the indication causes the invoked automated assistant to stop processing the subsequent audio data.
9. The method of claim 7, further comprising: Receive an identifier of the user invoking the automated assistant; Based on limited processing of the subsequent audio data and / or visual data, determine whether the co-located user is the user who invoked the automated assistant; And Transmit the notification only if the co-located user is identified as the user associated with the identifier.
10. The method according to claim 9, wherein, Determining whether the co-located user is the user who invoked the automated assistant includes performing text-dependent and / or text-independent analysis on at least a portion of the subsequent audio data to generate an embedding, and further comprising: Compare the embedding with a previously generated embedding associated with the user's profile, wherein the profile is identified based on the identifier.
11. The method according to claim 7, further comprising: Receive context information from the invoked automated assistant related to one or more previous requests of the user and / or one or more responses to the user.
12. The method according to claim 11, further comprising: Utilize the context information to process the subsequent audio data to determine one or more actions to be performed in response to subsequent requests included in the subsequent audio data.
13. A system, comprising: A first assistant device and a second assistant device in a link assistant device ecosystem, wherein the first assistant device is configured to: In response to determining that a user's invocation input is intended for the first assistant device: Process audio data captured via one or more first microphones of the first assistant device and after and / or before the invocation input; Based on processing the audio data, determine that an assistant request is included in the user's spoken words captured by the audio data; Cause one or more actions corresponding to the assistant request to be performed; Based on the assistant request, determine that subsequent audio data may include further assistant requests; In response to determining that subsequent audio data may include further assistant requests: Process subsequent audio data captured by the first microphone of the first assistant device to determine whether the subsequent audio data includes a further assistant request; and Transmit a notification to the second assistant device in the linked assistant device ecosystem, the notification indicating that the subsequent audio data may include a further assistant request; and wherein the second assistant device is configured to: Process sensor data to determine whether a user is co-located with the second assistant device, wherein the sensor data is generated by one or more sensors of the second assistant device; In response to determining that the co-located user is co-located with the second assistant device: Provide an indication to the first assistant device that the co-located user is co-located with the second assistant device, wherein providing the indication causes the first assistant device to stop processing audio input; and Process additional subsequent audio data captured by one or more microphones of the client device to determine whether the additional subsequent audio data includes a further assistant request.
14. The system according to claim 13, wherein, The first assistant device is further configured to: Determine a user identifier of the user; and Provide the user identifier to the second device, wherein the second assistant device is further configured to: Determine whether the co-located user is the user based on at least a portion of the user identifier and the sensor data; and Provide the indication only if the co-located user is the user.
15. The system according to claim 14, wherein, When determining the user identifier, the first assistant device is further configured to: Utilize one or more text-dependent speech models to process the call input to generate a speaker profile; And Identify the user identifier based on the association between the user and the speaker profile.
16. The system according to claim 14, wherein, When determining the user identifier, the second assistant device is further configured to: Utilize one or more text-independent speech models to process the audio data to generate a speaker profile; and Identify the user identifier based on the association between the user and the speaker profile.
17. The system according to claim 13, wherein, The sensor data includes visual data captured by one or more cameras of the second assistant device.
18. The system according to claim 13, wherein, The first assistant device is further configured to: Transmit context data generated based on the one or more actions to the second assistant device.
19. The system according to claim 18, wherein, The second assistant device is further configured to: Utilize the context data to determine a further assistant request included in the additional subsequent audio data.
20. The system according to claim 13, wherein, The first assistant device and the second assistant device are both execution instances of the same automated assistant in combination with one or more shared cloud-based components.