Inferring semantic tags for an assistant device based on device-specific signals

By analyzing the device-specific signals of the assistant device automatically infer and assigning semantic tags, the problems of inaccurate device tags and low user interaction efficiency are solved, and the response efficiency and resource utilization of automated assistant devices are improved.

CN115605859BActive Publication Date: 2025-07-08GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080100906.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-10-29
Filing Date
2020-12-14
Publication Date
2025-07-08
Estimated Expiration
2040-12-14

AI Technical Summary

Technical Problem

In the prior art, users need to manually assign or automate the Assistant Device guessing tags, resulting in inaccurate device tags or need to be updated manually, especially when the device is moved or added, resulting in inefficient user interaction.

Method used

By analyzing device-specific signals such as query, command, ambient noise and user preferences, automatically infer and assign semantic tags, update device topology representations, and reduce user input.

Benefits of technology

Improves the accuracy of device tags and the response efficiency of automation assistant devices, reduces user input and network traffic, and saves computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115605859B_ABST
    Figure CN115605859B_ABST
Patent Text Reader

Abstract

Embodiments can identify a given assistant device among multiple assistant devices in an ecosystem, obtain a device-specific signal generated by the given assistant device, process the device-specific signal to generate candidate semantic labels for the given assistant device, select a given semantic label for the given assistant device from among the candidate semantic labels, and assign the given semantic label to the given assistant device in a device topology representation of the ecosystem. Embodiments can optionally receive an oral utterance including a query or command at the assistant device, determine that the semantic attributes of the query or command match the given semantic label for the given assistant device, and cause the given assistant device to satisfy the query or command.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Humans are able to interact with an interactive software application, referred to herein as an "automation assistant" (also known as a "chatbot", "interactive personal assistant", "intelligent personal assistant", "personal voice assistant", "dialogue agent", etc.), in a human-computer dialogue. For example, a human (who may be referred to as a "user" when interacting with the automation assistant) can provide an input (e.g., a command, query, and / or request) to the automation assistant, which can cause the automation assistant to generate and provide a response output to control one or more Internet of Things (IoT) devices and / or perform one or more other functions. The input provided by the user can be, for example, an oral natural language input (i.e., an oral utterance), which can be converted into text (or other semantic representation) in some cases and then further processed, and / or a typed natural language input.

[0002] In some cases, the automation assistant can include an automation assistant client that is executed locally by an assistant device and directly interacted with by the user, and a cloud-based counterpart that utilizes the nearly infinite resources of the cloud to assist the automation assistant client in responding to the user's input. For example, the automation assistant client can provide an audio recording (or its text conversion) of the user's oral utterance to the cloud-based counterpart, and optionally provide data indicating the user's identity (e.g., credentials). The cloud-based counterpart can perform various processes on the query to return the results to the automation assistant client, which can then provide the corresponding output to the user.

[0003] Many users can interact with the automation assistant using multiple assistant devices. For example, some users can have a coordinated "ecosystem" of assistant devices that can receive user input directed to the automation assistant and / or be controlled by the automation assistant, such as one or more smart phones, one or more tablet computers, one or more vehicle computing systems, one or more wearable computing devices, one or more smart TVs, one or more interactive standalone speakers, and / or one or more IoT devices and other assistant devices. The user can use any of these assistant devices (assuming the automation assistant client is installed and the assistant device can receive input) to interact with the automation assistant in a human-computer dialogue. In some cases, these assistant devices can be scattered around the user's primary residence, secondary residence, workplace, and / or other structures. For example, mobile assistant devices such as smart phones, tablets, smart watches, etc. can be on the user's person and / or anywhere the user last placed them. Other assistant devices, such as traditional desktop computers, smart TVs, interactive standalone speakers, and IoT devices, can be more stationary, but can also be located in various places (e.g., rooms) within the user's home or workplace.

[0004] There are technologies that enable a user (e.g., a single user, multiple users in a household, colleagues, cohabitants, etc.) to manually assign tags to assistant devices in an ecosystem of assistant devices and then interact with or control any of the assistant devices using an automated assistant client of any of the assistant devices. For example, a user can issue an oral command “Show me some chili recipes on the kitchen device” to the automated assistant client of an assistant device, causing the assistant device (or another assistant device in the ecosystem) to retrieve search results for chili recipes and present the search results to the user via the kitchen device. However, such technologies require the user to specify a particular assistant device via a previously assigned tag (e.g., “kitchen device”) that the user may have forgotten, or require the automated assistant to guess the “best” device to provide the search results (e.g., the device closest to the user). Additionally, if a particular assistant device is newly introduced to the ecosystem or moved within the ecosystem, the tag assigned to the particular assistant device by the user may not represent the particular assistant device. SUMMARY OF THE INVENTION

[0005] Embodiments described herein relate to assigning semantic tags to corresponding assistant devices in a device topology representation of an ecosystem including multiple assistant devices. The semantic tags assigned to the corresponding assistant devices can be inferred based on one or more device-specific signals associated with the corresponding assistant devices. The one or more device-specific signals can include, for example, one or more queries (if any) previously received at the corresponding assistant device, one or more commands (if any) previously executed at the corresponding assistant device, instances of ambient noise previously detected at the corresponding assistant device (and optionally only if voice reception is active at the corresponding assistant device), unique identifiers (or tags) for any other assistant devices located in proximity to the corresponding assistant device, and / or user preferences of a user associated with the ecosystem, the user preferences being determined based on user interactions with multiple assistant devices in the ecosystem. Each of the one or more device-specific signals associated with the corresponding assistant device can be processed to classify each of them into one or more semantic categories from a plurality of different semantic categories. One or more candidate semantic tags can be generated for the corresponding assistant device based on the semantic categories into which one or more of the device-specific signals are classified. Additionally, a given semantic tag for a given assistant device among the one or more candidate semantic tags can be selected and assigned to the given assistant device in the device topology representation of the ecosystem.

[0006] For example, assume that a given assistant device is an interactive standalone speaker device with a display located in the primary residence of a user associated with an ecosystem. Further assume that multiple queries related to retrieving food recipes have been received and executed at the given assistant device and / or multiple commands related to setting a timer have been received and executed at the given assistant device, assume that an instance of ambient noise has been detected at a client device, assume that a unique identifier (or tag) associated with another assistant device corresponding to a "smart oven" in the ecosystem has been detected at the given assistant device, and assume that the user preferences of the user associated with the ecosystem indicate that the user likes a virtual chef named Johnny Flay. In this example, further assume that the queries related to retrieving food recipes are classified into "recipe", "kitchen", and / or "cooking" categories, and the commands related to setting a timer are classified into "timer" and / or "cooking" categories, further assume that the instance of ambient noise is classified into "kitchen" and / or "cooking" categories based on the ambient noise capturing cooking sounds (e.g., food sizzling in a skillet, knife cutting food, microwave in use, etc.), further assume that the unique identifier (or tag) of the "smart oven" associated with the other assistant device is classified into "kitchen" and / or "cooking" categories, and further assume that the virtual chef is classified into the "cooking" category (or a more specific category of "Johnny Flay"). Thus, candidate semantic tags "recipe display device", "kitchen display device", "cooking display device", "timer display device", and "Johnny Flay device" can be generated for the interactive standalone speaker device with a display. Additionally, a given semantic tag from among the candidate semantic tags can be assigned to the interactive standalone speaker device with a display in a device topology representation of the ecosystem of the user's primary residence.

[0007] In some embodiments, a given semantic tag can be automatically assigned to a given assistant device in the device topology representation of an ecosystem. For example, if the confidence level associated with the given semantic tag meets a threshold confidence level, the given semantic tag can be automatically assigned to the given assistant device in the device topology representation of the ecosystem. When processing one or more device-specific signals associated with the given assistant device, the confidence level associated with the given semantic tag can be determined. For example, the confidence level associated with the given assistant device can be based on the number of one or more device-specific signals classified into one or more semantic categories in the semantic category. For example, if nine queries related to retrieving food recipes have been received at the given assistant device and only one query related to retrieving weather information has been received at the given assistant device, the semantic tag "cooking display device" or "recipe display device" can be automatically assigned to the given assistant device in the device topology representation of the ecosystem (even if the given assistant device is not located in the user's kitchen). For example, the confidence level associated with the given assistant device can be a measure determined based on the output generated by processing one or more device-specific signals using a semantic classifier and / or an environmental noise detection model. For example, a semantic classifier can be used to process previously received queries or commands (or their corresponding texts) to classify each of the queries or commands into one or more semantic categories in the semantic category, an environmental noise detection model can be used to process instances of environmental noise to classify each of the instances of environmental noise into one or more semantic categories in the semantic category, and a semantic classifier can be used to process unique identifiers (or tags) to classify each of the unique identifiers (or tags) along with the corresponding measures into one or more semantic categories in the semantic category. As another example, if the given semantic tag is unique (relative to other assistant devices in the ecosystem that are in close proximity to the given assistant device in terms of location), the given semantic tag can be automatically assigned to the given assistant device in the device topology representation of the ecosystem.

[0008] In some additional or alternative embodiments, a given semantic tag can be assigned to a given assistant device in a device topology representation of the ecosystem in response to receiving user input that assigns the given semantic tag to the given assistant device. For example, a prompt can be generated to request a selection of the given semantic tag from among one or more candidate semantic tags from a user associated with the ecosystem. The prompt can be rendered at the user's client device (e.g., the given assistant device or another client device of the user, such as a mobile phone), and the given semantic tag can be assigned to the given assistant device in response to receiving a selection of the given semantic tag. For example, assume that one or more candidate semantic tags include "cooking display device", "recipe display device", and "weather display device". In this case, the prompt can include each of the candidate semantic tags and request the user to select the given semantic tag from among these candidate semantic tags that should be assigned to the given assistant device (and optionally replace an existing semantic tag). Although the above example is described with respect to a single semantic tag being assigned to a given assistant device, it should be understood that this is for illustrative purposes and is not meant to be limiting. For example, the assistant devices described herein can be assigned multiple semantic tags such that each of the assistant devices is stored in association with a list of semantic tags.

[0009] In various embodiments, and after one or more semantic tags among the semantic tags are assigned to corresponding assistant devices in a device topology representation of the ecosystem, the semantic tags assigned to the assistant devices according to the techniques described herein can also be used to process spoken utterances received at one or more assistant devices in the ecosystem. For example, audio data corresponding to the spoken utterance can be processed to identify semantic attributes included in the spoken utterance. Additionally, an embedding of the identified semantic attributes (e.g., a word2vec representation) can be generated and compared with multiple embeddings (e.g., corresponding word2vec representations) of the corresponding semantic tags assigned to assistant devices in the ecosystem. Further, based on the comparison, it can be determined that the semantic attribute matches a given embedding among the multiple embeddings of the corresponding semantic tags. For example, assume that the embedding is a word2vec representation. In this example, the cosine distance between the word2vec representation of the semantic attribute and each of the word2vec representations of the corresponding semantic tags can be determined, and the given semantic tag associated with the corresponding cosine distance that satisfies a distance threshold can be used to determine that the semantic attribute of the spoken utterance matches the given semantic tag (e.g., an exact match or a soft match). Thus, a given assistant device associated with the given semantic tag can be selected to satisfy the spoken utterance. Additionally or alternatively, when selecting a given assistant device to satisfy the spoken utterance, the proximity of the user to the given assistant device and / or the device capabilities of the given assistant device can be considered.

[0010] By using the techniques described herein to infer semantic tags and assign semantic tags to assistant devices in an ecosystem, the device topology representation of the ecosystem can be kept up-to-date without the need for multiple (or even any) user interface inputs to do so. Additionally, the semantic tags assigned to the assistant devices are semantically meaningful to the user because the semantic tags assigned to the corresponding assistant devices are selected based on the usage of the corresponding assistant devices and / or the corresponding portions of the ecosystem in which the corresponding assistant devices are located. Thus, when an oral utterance is received at one or more of the assistant devices in the ecosystem, the automated assistant devices can more robustly and / or accurately select one or more of the assistant devices that are most suitable to satisfy the oral utterance. Accordingly, the amount and / or duration of user input received by one or more of the assistant devices in the ecosystem can be reduced because the user associated with the ecosystem does not need to specify a particular device to satisfy the oral utterance or repeat the oral utterance in the case where an incorrect device has been selected to satisfy the oral utterance, thereby saving computing resources and / or network resources at the assistant devices by reducing network traffic. Additionally, the amount of user input received by one or more of the assistant devices in the ecosystem can be reduced because the user does not need to manually update the device topology representation via a software application associated with the ecosystem when an assistant device is newly added to the ecosystem, moved within the ecosystem, or located within a portion of the ecosystem that has been repurposed (e.g., a room in the user's primary residence has been changed from a study to a bedroom).

[0011] The foregoing description is provided as an overview of only some embodiments of the present disclosure. Further descriptions of these and other embodiments are described in more detail herein. As a non-limiting example, the various embodiments are described in more detail in the claims included herein.

[0012] Additionally, some embodiments include one or more processors of one or more computing devices, wherein the one or more processors are operable to execute instructions stored in an associated memory, and wherein the instructions are configured to cause the execution of any of the methods described herein. Some embodiments also include one or more non-transitory computer-readable storage media storing computer instructions that are executable by the one or more processors to execute any of the methods described herein.

[0013] It should be understood that all combinations of the foregoing concepts and additional concepts described in more detail herein are considered to be part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are considered to be part of the subject matter disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1is a block diagram of an example environment in which embodiments disclosed herein may be implemented.

[0015] Figure 2A and Figure 2B depicts some examples related to assigning a given semantic tag to a given assistant device newly added to and / or moved within the ecosystem of an assistant device, according to various embodiments.

[0016] Figure 3 is a flowchart illustrating an example method of assigning a given semantic tag to a given assistant device in an ecosystem, according to various embodiments.

[0017] Figure 4 is a flowchart illustrating an example method of using the assigned semantic tag when fulfilling a query or command received at an assistant device in an ecosystem, according to various embodiments.

[0018] Figure 5 depicts an example architecture of a computing device, according to various embodiments. Detailed Description

[0019] There are a large number of intelligent, multi-sensing network-connected devices (also referred to herein as assistant devices), such as smart phones, tablet computers, vehicle computing systems, wearable computing devices, smart TVs, interactive standalone speakers (e.g., with or without a display), sound speakers, home alarms, door locks, cameras, lighting systems, treadmills, thermostats, weighing scales, smart beds, irrigation systems, garage door openers, appliances, baby monitors, fire alarms, humidity detectors, etc. Generally, multiple assistant devices are within the scope of a structure such as a home, or within multiple related structures such as a user's primary residence and the user's secondary residence, the user's vehicle, and / or the user's workplace.

[0020] In addition, there are a large number of assistant devices each including a logical instance of an automated assistant (also referred to herein as an assistant input device) that can form an automated assistant client. These assistant input devices can be dedicated solely to assistant functions (e.g., an interactive standalone speaker and / or standalone audio / video device that includes only an assistant client and an associated interface and is dedicated solely to assistant functions), or can perform assistant functions in addition to other functions (e.g., a mobile phone or tablet that includes an assistant client as one of multiple applications). In addition, some IoT devices can also be assistant input devices. For example, some IoT devices can include an automated assistant client and at least a speaker and / or a microphone, which (at least partially) serve as a user interface output and / or input device for the assistant interface of the automated assistant client. Although some assistant devices may not implement an automated assistant client or have means for interacting with a user interface (e.g., a speaker and / or a microphone), they can still be controlled by an automated assistant (also referred to herein as an assistant non-input device). For example, a smart bulb may not include an automated assistant client, a speaker, and / or a microphone, but can receive commands and / or requests via an automated assistant to control the functions of the smart light (e.g., turn on / off, dim, change color, etc.).

[0021] Various techniques have been proposed for tagging and / or grouping assistant devices (including both assistant input devices and assistant non-input devices) within an ecosystem of assistant devices. For example, when adding a new assistant device to the ecosystem, a user associated with the ecosystem can manually assign a tag (or unique identifier) to the new assistant device in a device topology representation of the ecosystem and / or manually add the new assistant device to a group of assistant devices in the ecosystem via a software application (e.g., via an automated assistant application, a software application associated with the ecosystem, a software application associated with the new assistant device, etc.). As described herein, the tags initially assigned to assistant devices may be forgotten by the user, or may not be semantically meaningful in terms of how to utilize the assistant devices or where the assistant devices are located within the ecosystem. In addition, if an assistant device moves within the ecosystem, the user may need to manually change the tag assigned to the assistant device and / or manually change the group to which the assistant device is assigned via a software application. Otherwise, the tag assigned to the assistant device and / or the group to which the assistant device is assigned may not accurately reflect the location or use of the assistant device, and / or may not be semantically meaningful for the assistant device. For example, if a smart speaker labeled "living room speaker" is located in the living room of the user's primary house, but the smart speaker is moved to the kitchen of the user's primary house, the smart speaker may still be labeled "living room speaker," even though the tag does not indicate the location of the assistant device, unless the user manually changes the tag in the device topology representation of the ecosystem of the user's primary house.

[0022] The device topology representation can include tags (or unique identifiers) associated with respective assistant devices. Additionally, the device topology representation can specify tags (or unique identifiers) associated with respective assistant devices. The device attributes of a given assistant device can indicate, for example, one or more input and / or output modalities supported by the respective assistant device. For example, the device attributes of a stand-alone speaker-only assistant client device can indicate that it can provide audible output but not visual output. The device attributes of a given assistant device can additionally or alternatively, for example, identify one or more states of the given assistant device that can be controlled; identify the party (e.g., a first party (1P) or a third party (3P)) that manufactures, distributes, and / or creates the firmware for the assistant device; and / or identify a unique identifier of the given assistant device, such as a fixed identifier provided by the 1P or 3P or a label assigned by a user to the given assistant device. According to various embodiments disclosed herein, the device topology representation can optionally further specify: which smart devices can be locally controlled by which assistant devices; the local addresses of the assistant devices that can be locally controlled (or the local addresses of the hubs that can directly locally control these assistant devices); the local signal strength and / or other preference indicators among the respective assistant devices. Additionally, according to various embodiments disclosed herein, the device topology representation (or a variant thereof) can be locally stored at each of a plurality of assistant devices for locally controlling tags and / or locally assigning tags to assistant devices. Additionally, the device topology representation can specify groups associated with respective assistant devices that can be defined at various granularity levels. For example, multiple smart lights in the living room of a user's primary house can be considered to belong to a "living room lights" group. Additionally, if the living room of the primary house also includes a smart speaker, all assistant devices located in the living room can be considered to belong to a "living room assistant devices" group.

[0023] The automated assistant can detect various events occurring in the ecosystem based on one or more signals generated by one or more assistant devices in the assistant device. For example, the automated assistant can use an event detection model or rule to process one or more of the signals to detect these events. In addition, the automated assistant can cause one or more actions to be performed based on the output generated according to one or more of the signals of the events occurring in the ecosystem. In some embodiments, the detected event can be a device-related event associated with one or more assistant devices in the assistant device (e.g., an assistant input device and / or an assistant non-input device). For example, a given assistant device in the assistant device can detect when it is newly added to the ecosystem based on one or more wireless signals generated by the given assistant device in the assistant device (and optionally, a unique identifier associated with the given assistant device included in one or more wireless signals). As another example, a given assistant device in the assistant device can detect when it has moved within the ecosystem based on being surrounded by one or more different assistant devices that were previously surrounding the given assistant device in the assistant device (and optionally determined based on the respective unique identifiers of one or more different assistant devices). In these embodiments, one or more actions performed by the automated assistant can include, for example, in response to determining that a given assistant device in the assistant device has been newly introduced to the ecosystem or has moved locations within the ecosystem, determining a semantic label for the given assistant device in the assistant device, and causing the semantic label to be assigned to the given assistant device in the device topology representation of the ecosystem.

[0024] In some additional or alternative embodiments, the detected event can be an acoustic event captured via a respective microphone of one or more assistant devices. The automated assistant can cause audio data capturing the acoustic event to be processed using an acoustic event model. The acoustic events detected by the acoustic event model can include, for example, detecting a hot word invoking the automated assistant included in an uttered speech using a hot word detection model, detecting ambient noise in an ecosystem (and optionally when voice reception is active at a given assistant device among the assistant devices) using an ambient noise detection model. Detecting a specific sound in the ecosystem using a sound detection model (e.g., glass breaking, dog barking, cat meowing, doorbell ringing, smoke alarm going off, or carbon monoxide detector going off), and / or other acoustically related events that can be detected using a respective sound event detection model. For example, assume that audio data is detected via a respective microphone of at least one of the assistant devices. In this example, the automated assistant can cause the audio data to be processed by a hot word detection model of at least one of the assistant devices to determine whether the audio data captures a hot word to invoke the automated assistant. Additionally, the automated assistant can additionally or alternatively cause the audio data to be processed by an ambient noise detection model of at least one of the assistant devices to classify any ambient (or background) noise captured in the audio data into one or more different semantic classes of ambient noise (e.g., movie or TV sounds, cooking sounds, and / or other different classes of sounds). Additionally, the automated assistant can additionally or alternatively cause the audio data to be processed by a sound detection model of at least one of the assistant devices to determine whether any specific sound is captured in the audio data.

[0025] The embodiments described herein relate to inferring and determining semantic labels for an assistant device based on one or more signals generated by each of the respective devices. These embodiments also relate to assigning semantic labels to the assistant device in a device topology representation of the ecosystem. The semantic labels can be automatically assigned to the assistant device or can be presented to a user associated with the ecosystem to request a selection of one or more semantic labels to be assigned to the assistant device. Additionally, these embodiments relate to subsequently using the semantic labels to determine whether an uttered speech includes a term or phrase that matches any of the semantic labels when processing the uttered speech, and when determining that the uttered speech includes a term or phrase that matches one of the semantic labels, using the assistant device associated with the matching semantic label to fulfill the uttered speech.

[0026] Now turning to Figure 1 , an example environment is illustrated in which the techniques disclosed herein can be implemented. The example environment includes: a plurality of assistant input devices 106 1-N(This document is also simply referred to as "Assistant Input Device 106"), one or more cloud-based automated assistant components 119, one or more assistant non-input systems 180, and one or more assistant non-input devices 185 1-N (This document is also simply referred to as "Assistant Non-Input Device 185"), a device activity database 191, a machine learning ("ML") model database, and a device topology database 193. Figure 1 The assistant input device 106 and the assistant non-input device 185 can also be collectively referred to as "assistant devices" herein.

[0027] One or more (e.g., all) assistant input devices 106 are capable of executing corresponding instances of the corresponding automated assistant client 118. 1-N However, in some embodiments, one or more assistant input devices 106 can optionally lack instances of the corresponding automated assistant client 118 and still include engines and hardware components (e.g., microphones, speakers, speech recognition engines, natural language processing engines, speech synthesis engines, etc.) for receiving and processing user input directed to the automated assistant. Instances of the automated assistant client 118 1-N can be applications separate from the operating system of the corresponding assistant input device 106 (e.g., installed "on top" of the operating system), or can alternatively be directly implemented by the operating system of the corresponding assistant input device 106. As further described below, each instance of the automated assistant client 118 1-N can optionally interact with one or more cloud-based automated assistant components 119 in response to various requests provided by the corresponding user interface component 107 of any one of the corresponding assistant input devices 106. Additionally, and as also described below, other engines of the assistant input device 106 can optionally interact with one or more cloud-based automated assistant components 119. 1-N 1-N

[0028] One or more cloud-based automated assistant components 119 can be implemented on one or more computing systems (e.g., servers collectively referred to as "cloud" or "remote" computing systems) that are communicatively coupled to the corresponding assistant input device 106 via one or more local area networks ("LAN", including Wi-Fi LAN, Bluetooth networks, near field communication networks, mesh networks, etc.) and / or wide area networks ("WAN", including the Internet, etc.). The communicative coupling of the cloud-based automated assistant components 119 to the assistant input device 106 is typically indicated by Figure 1 1101. Additionally, in some embodiments, the assistant input devices 106 can be communicatively coupled to each other via one or more networks (e.g., LAN and / or WAN), typically indicated by Figure 1 1102. ​​

[0029] One or more cloud-based automated assistant components 119 are also communicatively coupled via one or more networks (e.g., LAN and / or WAN) to one or more assistant non-input systems 180. The communicative coupling of the cloud-based automated assistant components 119 to the assistant non-input systems 180 is typically indicated by Figure 1 1103. Additionally, the assistant non-input systems 180 are each communicatively coupled via one or more networks (e.g., LAN and / or WAN) to one or more assistant non-input devices 185 (e.g., groups). For example, a first assistant non-input system 180 can be communicatively coupled to a first group of one or more assistant non-input devices 185 and receive data from the first group of one or more assistant non-input devices 185, a second assistant non-input system 180 can be communicatively coupled to a second group of one or more assistant non-input devices 185 and receive data from the second group of one or more assistant non-input devices 185, and so on. The communicative coupling of the assistant non-input systems 180 to the assistant non-input devices 185 is typically indicated by Figure 1 1104.

[0030] Instances of the automated assistant client 118, through its interaction with one or more cloud-based automated assistant components 119, can form a logical instance of an automated assistant 120 that, from the user's perspective, appears to be an automated assistant with which the user can interact in a human-machine conversation. Two instances of such an automated assistant 120 are depicted in Figure 1 The first automated assistant 120A enclosed by the dashed line includes the automated assistant client 1181 of the assistant input device 1061 and one or more cloud-based automated assistant components 119. The second automated assistant 120B enclosed by the dotted line includes the automated assistant client 118 N of the assistant input device 106 N and one or more cloud-based automated assistant components 119. Thus, it should be understood that each user who interacts with the automated assistant client 118 executing on one or more assistant input devices 106 can actually interact with a logical instance of his or her own automated assistant 120 (or a logical instance of the automated assistant 120 shared among a household or other user group). For simplicity and brevity, the term "automated assistant" as used herein will refer to the combination of the automated assistant client 118 executing on a corresponding assistant input device in the assistant input devices 106 and one or more cloud-based automated assistant components 119 (which can be shared among multiple automated assistant clients 118). Although only multiple assistant input devices 106 are illustrated in Figure 1 it should be understood that the cloud-based automated assistant components 119 can additionally serve many other groups of assistant input devices.

[0031] The assistant input device 106 may include, for example, one or more of the following: a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a computing device of a user's vehicle (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), an interactive standalone speaker (e.g., with or without a display), a smart appliance such as a smart TV, a wearable device of the user including a computing device (e.g., a watch of the user with a computing device, glasses of the user with a computing device, a virtual or augmented reality computing device), and / or any IoT device capable of receiving user input directed to the automated assistant 120. Additional and / or alternative assistant input devices may be provided. The assistant non-input device 185 may include many of the same devices as the assistant input device 106 but is not capable of receiving user input directed to the automated assistant 120 (e.g., does not include a user interface input component). Although the assistant non-input device 185 does not receive user input directed to the automated assistant 120, the assistant non-input device 185 may still be controlled by the automated assistant 120.

[0032] In some embodiments, the plurality of assistant input devices 106 and assistant non-input devices 185 can be related to each other in various ways to facilitate the performance of the techniques described herein. For example, in some embodiments, the plurality of assistant input devices 106 and assistant non-input devices 185 may be communicatively coupled to each other by communicating via one or more networks (e.g., via Figure 1 the network 110). For example, this may be the case where the plurality of assistant input devices 106 and assistant non-input devices 185 are deployed across a particular area or environment such as a home, a building, etc. Additionally or alternatively, in some embodiments, the plurality of assistant input devices 106 and assistant non-input devices 185 may be related to each other by being members of a coordinated ecosystem that is at least selectively accessible by one or more users (e.g., an individual, a family, an employee of an organization, other predefined groups, etc.). In some of these embodiments, the ecosystem of the plurality of assistant input devices 106 and assistant non-input devices 185 can be manually and / or automatically related to each other in a device topology representation of the ecosystem stored in the device topology database 193.

[0033] The assistant non-input system 180 can include one or more first-party (1P) systems and / or one or more third-party (3P) systems. A 1P system refers to a system controlled by the same party that controls the automated assistant 120 referenced herein. As used herein, a 3P system refers to a system controlled by a party different from the party that controls the automated assistant 120 referenced herein.

[0034] The assistant non-input system 180 can receive from the assistant non-input device 185 and / or (e.g., viaFigure 1 One or more cloud-based automated assistant components 119 communicatively coupled to the network 110 of receive data and selectively transmit data (e.g., status, status changes, and / or other data) to the assistant non-input device 185 and / or one or more cloud-based automated assistant components 119. For example, assume that the assistant non-input device 1851 is a smart doorbell IoT device. In response to an individual pressing a button on the doorbell IoT device, the doorbell IoT device can transmit corresponding data to one of the assistant non-input systems 180 (e.g., one of the assistant non-input systems managed by the manufacturer of the doorbell, which can be a 1P system or a 3P system). One of the assistant non-input systems 180 can determine a change in the status of the doorbell IoT device based on such data. For example, one of the assistant non-input systems 180 can determine a change in the doorbell from an inactive state (e.g., the button has not been pressed recently) to an active state (the button has been pressed recently), and the change in the doorbell status can be (e.g., via Figure 1 the network 110) transmitted to one or more cloud-based automated assistant components 119 and / or one or more assistant input devices 106. It is noted that although user input (e.g., pressing a button on the doorbell) is received at the assistant non-input device 1851, the user input does not point to the automated assistant 120 (hence the term "assistant non-input device"). As another example, assume that the assistant non-input device 1851 is a smart thermostat IoT device with a microphone, but the smart thermostat does not include an automated assistant client 118. An individual can interact with the smart thermostat (e.g., using touch input or voice input) to change the temperature, set a specific value as a set point for controlling the HVAC system via the smart thermostat, etc. However, the individual cannot communicate directly with the automated assistant 120 via the smart thermostat unless the smart thermostat includes an automated assistant client 118.

[0035] In various embodiments, one or more cloud-based automated assistant components 119 can further include various engines. For example, as Figure 1 shown, one or more cloud-based automated assistant components 119 can further include an event detection engine 130, a device identification engine 140, an event processing engine 150, a semantic tagging engine 160, and a query / command processing engine 170. Although in Figure 1These various engines are depicted as one or more cloud-based automated assistant components 119, but it should be understood that this is for purposes of example and is not meant to be limiting. For example, assistant input device 106 and / or assistant non-input device 185 can include one or more of these various engines. As another example, these various engines can be distributed across assistant input device 106, assistant non-input device 185 can include one or more of these various engines, and / or one or more cloud-based automated assistant components 119.

[0036] In some embodiments, event detection engine 130 is capable of detecting various events that occur in the ecosystem. In some versions of these embodiments, event detection engine 130 is capable of determining when a given assistant input device in assistant input device 106 and / or a given assistant non-input device in assistant non-input device 185 (e.g., a given assistant device in the assistant devices) has been newly added to the ecosystem or has moved locations within the ecosystem. For example, event detection engine 130 can determine when a given assistant device in the assistant devices has been newly added to the ecosystem based on one or more wireless signals detected via network 110 and through device identification engine 140. For example, when a given assistant device in the assistant devices is newly connected to one or more networks 110, the given assistant device in the assistant devices can broadcast a signal indicating that it has been newly added to network 110. As another example, event detection engine 130 can determine when a given assistant device in the assistant devices has moved locations within the ecosystem based on one or more wireless signals detected via network 110. In these examples, device identification engine 140 can process the signals to determine that a given assistant device in the assistant devices has been newly added to network 110 and / or to determine that a given assistant device in the assistant devices has moved locations within the ecosystem. The one or more wireless signals detected by device identification engine 140 can be, for example, network signals and / or acoustic signals that are not perceptible to humans and optionally include corresponding unique identifiers of a given assistant device in the assistant devices and / or other assistant devices in proximity to the given assistant device in location. For example, when a given assistant device in the assistant devices has moved locations within the ecosystem, device identification engine 140 can detect one or more wireless signals transmitted by other assistant devices in proximity to the given assistant device in location. These signals can be processed to determine that one or more other assistant devices in proximity to the given assistant device in location are different from one or more assistant devices that were previously in proximity to the given assistant device in location.

[0037] In some other versions of these embodiments, the automated assistant 120 can cause a given assistant device among the assistant devices newly added to the ecosystem or relocated within the ecosystem to be assigned to an assistant device group (e.g., in the device topology representation of the ecosystem stored in the device topology database 193). For example, in an embodiment where a given assistant device among the assistant devices is newly added to the ecosystem, the given assistant device among the assistant devices can be added to an existing assistant device group, or a new assistant device group including the given assistant device among the assistant devices can be created. For example, if a given assistant device among the assistant devices is in proximity to multiple assistant devices belonging to the "kitchen" group (e.g., a smart oven, a smart coffee maker, an interactive standalone speaker associated with a unique identifier or label indicating its location in the kitchen, and / or other auxiliary devices), the given assistant device among the assistant devices can be added to the "kitchen" group, or a new group can be created. As another example, in an embodiment where a given assistant device among the assistant devices has been relocated within the ecosystem, the given assistant device among the assistant devices can be added to an existing assistant device group, or a new assistant device group including the given assistant device among the assistant devices can be created. For example, if a given assistant device among the assistant devices was in proximity to multiple assistant devices belonging to the aforementioned "kitchen" group but is now in proximity to multiple assistant devices belonging to the "garage" group (e.g., a smart garage door, a smart door lock, and / or other assistant devices), the given assistant device among the assistant devices can be removed from the "kitchen" group and added to the "garage" group.

[0038] In some additional or alternative versions of these embodiments, the event detection engine 130 can detect the occurrence of an acoustic event. The occurrence of an acoustic event can be detected based on audio data received at one or more assistant input devices 106 and / or one or more assistant non-input devices 185 (e.g., one or more assistant devices). The audio data received at one or more assistant devices can be processed by an event detection model stored in the ML model database 192. In these embodiments, each of the one or more assistant devices that detect the occurrence of an acoustic event includes a corresponding microphone.

[0039] In some other versions of these embodiments, the occurrence of an acoustic event can include ambient noise captured in audio data at one or more assistant devices (and optionally only the occurrence of ambient noise detected when voice reception is active at one or more assistant devices). The ambient noise detected at each of the one or more assistant devices can be stored in the device activity database 191. In these embodiments, the event processing engine 150 can use ambient noise detection models (e.g., stored in the ML model database 192) to process the ambient noise detected at one or more assistant devices, and these ambient noise detection models are trained to classify the ambient noise into one or more semantic categories of a plurality of different semantic categories based on metrics generated when processing the ambient noise using the ambient noise detection models. The plurality of different categories can include, for example, a movie or TV sound category, a cooking sound category, a music sound category, a garage or workshop sound category, a yard sound category, and / or other different sound categories that are semantically meaningful. For example, if the event processing engine 150 determines that the ambient noise processed using the ambient noise detection model includes sounds corresponding to a microwave oven making a sound, food sizzling on a frying pan, a food processor processing food, etc., then the event processing engine 150 can classify the ambient noise into the cooking sound category. As another example, if the event processing engine 150 determines that the ambient noise processed using the ambient noise detection model includes sounds corresponding to a saw buzzing, a hammer striking, etc., then the event processing engine 150 can classify the ambient noise into the garage or workshop category. The classification of the ambient noise detected at a particular device can also be used as device-specific signals that are used to infer semantic labels for the assistant device (e.g., as described with respect to the semantic tagging engine 160).

[0040] In some additional or alternative versions of those further embodiments, the occurrence of a sound event can include a hotword or a specific sound detected at one or more assistant devices. In these embodiments, the event processing engine 150 can use a hotword detection model to process audio data detected at one or more assistant devices, and the hotword detection model is trained to determine whether the audio data includes a specific word or phrase that invokes the automated assistant 120 based on a metric generated when processing the audio data using the hotword detection model. For example, the event processing engine 150 can process the audio data to determine whether the audio data captures a spoken utterance of a user that includes "Assistant", "Hey Assistant", "Okay, Assistant", and / or any other word or phrase that invokes the automated assistant. Additionally, the metric generated using the hotword detection model can include a corresponding confidence level or probability indicating whether the audio data includes a word or phrase that invokes the automated assistant 120. In some versions of these embodiments, if the metric meets a threshold, the event processing engine 150 can determine that the audio data captures the word or phrase. For example, if the event processing engine 150 generates a metric of 0.70 associated with audio data that captures a word or phrase that invokes the automated assistant 120 and the threshold is 0.65, the event processing engine 150 can determine that the audio data captures a word or phrase that invokes the automated assistant 120.

[0041] In these embodiments, the event processing engine 150 can additionally or alternatively use a sound detection model to process audio data detected at one or more assistant devices, and the sound detection model is trained to determine whether the audio data includes a specific sound based on a metric generated when processing the audio data using the sound detection model. The specific sound can include, for example, glass breaking, a dog barking, a cat meowing, a doorbell ringing, a smoke alarm going off, or a carbon monoxide detector going off. For example, the event processing engine 150 can process the audio data to determine whether the audio data captures any of these specific sounds. In this example, a single sound detection model can be trained to determine whether multiple specific sounds are captured in the audio data, or multiple sound detection models can be trained to determine whether a given specific sound is captured in the audio data. Additionally, the metric generated using the sound detection model can include a corresponding confidence level or probability indicating whether the audio data includes a specific sound. In some versions of these embodiments, if the metric meets a threshold, the event processing engine 150 can determine that the audio data captures the specific sound. For example, if the event processing engine 150 generates a metric of 0.70 associated with audio data that captures the sound of glass breaking and the threshold is 0.65, the event processing engine 150 can determine that the audio data captures the sound of glass breaking.

[0042] In various embodiments, the occurrence of an acoustic event can be captured by multiple assistant devices in an ecosystem. For example, multiple assistant devices in the environment can capture audio data that is temporally corresponding (e.g., temporally corresponding in that the corresponding audio data is detected at the multiple assistant devices simultaneously or within a threshold duration). In these embodiments, and in response to a given assistant device detecting audio data in the ecosystem, the device identification engine 140 is capable of identifying one or more additional assistant devices that should also have detected audio data that is temporally corresponding to the captured acoustic event. For example, the device identification engine 140 can identify one or more additional assistant devices that should also have detected audio data that is temporally corresponding to the captured acoustic event based on one or more additional assistant devices that have historically detected audio data that is temporally corresponding to the captured acoustic event. In other words, the device identification engine 140 can anticipate that one or more additional assistant devices should also capture audio data including the acoustic event because the given assistant device and the one or more additional assistant devices have historically captured audio data that is temporally corresponding to the same acoustic.

[0043] In various embodiments, one or more device-specific signals generated or detected by a corresponding assistant device can be stored in the device activity database 191. In some embodiments, the device activity database 191 can correspond to a portion of the memory dedicated to the device activity of that particular assistant device. In some additional or alternative embodiments, the device activity database 191 can correspond to (e.g., via Figure 1 the network 110) the memory of a remote system that communicates with the assistant device. This device activity can be used to generate candidate semantic tags for a given assistant device in the assistant device (e.g., as described with respect to the semantic tagging engine 160). The device activity can include, for example, queries or requests received at the corresponding assistant device (and / or the semantic categories associated with each of the multiple queries or requests), commands executed at the corresponding assistant device (and / or the semantic categories associated with each of the multiple commands), environmental noise detected at the corresponding assistant device (and / or the semantic categories associated with various instances of the environmental noise), the unique identifier or tag of any assistant device located in proximity to the given assistant device (e.g., identified via the event detection engine 140), user preferences of a user associated with the ecosystem, and / or any other data received, generated, and / or executed by the corresponding assistant device, where the user preferences are determined based on user interactions with multiple assistant devices in the ecosystem (e.g., browsing history, search history, purchase history, music history, movie or TV history, and / or any other user interactions associated with the multiple assistant devices).

[0044] In some embodiments, the semantic tagging engine 160 is capable of processing one or more device-specific signals to generate candidate semantic tags for a given assistant device in an assistant device (e.g., a given assistant input device in the assistant input device 106 and / or a given assistant non-input device in the assistant non-input device 185) based on the one or more device-specific signals. In some versions of those embodiments, the given assistant device for which candidate semantic tags are generated can be identified in response to determining that the given assistant device has been newly added to the ecosystem and / or has moved locations within the ecosystem. In some additional or alternative versions of those embodiments, the given assistant device for which candidate semantic tags are generated can be identified periodically (e.g., once a month, once every six months, once a year, etc.). In some additional or alternative versions of those embodiments, the given assistant device for which candidate semantic tags are generated can be identified in response to determining that the part of the ecosystem in which the given assistant device is located has been repurposed (e.g., a room in the primary residence of the ecosystem has been repurposed from a study to a bedroom). In these embodiments, the event detection engine 130 can be utilized to identify the given assistant device. Identifying the given assistant device in these and other ways will be described with reference to Figure 2A and Figure 2B as follows.

[0045] In some embodiments, the semantic tagging engine 160 is capable of selecting a given semantic tag for a given assistant device from among the candidate semantic tags based on one or more device-specific signals. Generating candidate semantic tags for a given assistant device based on one or more device-specific tags and selecting the given semantic tag from among the candidate semantic tags are described below (e.g., with reference to Figure 2A and Figure 2B ).

[0046] In an implementation where candidate semantic tags for a given assistant device are generated based on queries, requests, and / or commands (or corresponding text thereof) stored in the device activity database 191, a semantic classifier (e.g., stored in the ML model database 192) can be used to process the queries, requests, and / or commands to index the device activities for the given assistant device into one or more different semantic categories corresponding to different types of queries, requests, and / or commands. Candidate semantic tags can be generated based on the semantic categories into which the queries, commands, and / or requests are classified, and a given semantic tag selected for the given assistant device can be selected based on the number of queries, requests, and / or commands classified in the given semantic category. For example, assume that a given assistant device has previously received nine queries related to obtaining cooking recipes and two commands related to controlling smart lights in an ecosystem. In this example, the candidate semantic tags can include, for example, a first semantic tag "kitchen device" and a second semantic tag "control smart light device". Additionally, the semantic tagging engine 160 can select the first semantic tag "kitchen device" as the given semantic tag for the given assistant device because the historical usage of the given assistant device indicates that it is primarily used for cooking-related activities.

[0047] In some embodiments, the semantic classifier stored in the ML model database 192 can be a natural language understanding engine (e.g., implemented by the NLP module 122 described below). An intent determined based on processing a query, command, and / or request previously received at the assistant device can be mapped to one or more semantic categories. Notably, the multiple different semantic categories described herein can be defined with various levels of granularity. For example, a semantic category can be associated with a genus category of smart device commands and / or a species category for that genus, such as a category of smart lighting commands, a category of smart thermostat commands, and / or a category of smart camera commands. In other words, each category can have a unique set of intents associated with that category determined by the semantic classifier, although some intents of a category can also be associated with other categories. In some additional or alternative embodiments, the semantic classifier stored in the ML model database 192 can be used to generate a text embedding (e.g., a lower-dimensional representation, such as a word2vec representation) of the text corresponding to a query, command, and / or request. These embeddings can be points in an embedding space where semantically similar words or phrases are associated with the same or similar parts of the embedding space. Moreover, these parts of the embedding space can be associated with one or more of the multiple different semantic categories, and a given embedding in the embedding can be classified as a given semantic category in the semantic category if a distance metric between the given embedding in the embedding and one or more parts of the embedding space meets a distance threshold. For example, cooking-related words or phrases can be associated with a first part of the embedding space associated with the "cooking" semantic label, weather-related words or phrases can be associated with a second part of the embedding space associated with the "weather" semantic label, and so on.

[0048] In embodiments where one or more device-specific signals additionally or alternatively include ambient noise activity, an ambient noise detection model (e.g., stored in the ML model database 192) can be used to process instances of ambient noise to index device activity for a given assistant device into one or more different semantic categories corresponding to different types of ambient noise. Candidate semantic labels can be generated based on the semantic categories into which instances of ambient noise are classified, and a given semantic label selected for a given assistant device can be selected based on the number of instances of the environment classified in the given semantic category. For example, assume that the ambient noise detected at a given assistant device (and optionally only when speech recognition activity is occurring) primarily includes ambient noise classified as cooking sounds. In this example, the semantic tagging engine 160 can select the semantic label "kitchen equipment" as the given semantic label for the given assistant device because the ambient noise captured in the audio data indicates that the device is located near cooking-related activities.

[0049] In some embodiments, an environmental noise detection model stored in the ML model database 192 can be trained to detect a specific sound, and based on the output generated across the environmental noise detection models, it can be determined whether an instance of environmental noise includes the specific sound. The environmental noise detection model can be trained using, for example, supervised learning techniques. For example, a plurality of training instances can be obtained. Each training instance can include: a training instance input including environmental noise, and a corresponding training instance output including an indication of whether the training instance input includes the specific sound that the environmental noise detection model is being trained to detect. For example, if the environmental noise detection model is being trained to detect the sound of glass breaking, a label (e.g., "yes") or value (e.g., "1") can be assigned to the training instance including the sound of glass breaking, and a different label (e.g., "no") or value (e.g., "0") can be assigned to the training instance not including the sound of glass breaking. In some additional or alternative embodiments, the environmental noise detection model stored in the ML model database 192 can be used to generate an audio embedding (e.g., a lower-dimensional representation of an instance of environmental noise) based on an instance of environmental noise (or its acoustic features, such as mel-Cepstral frequency coefficients, raw audio waveforms, and / or other acoustic features). These embeddings can be points within an embedding space where similar sounds (or acoustic features capturing the sounds) are associated with the same or similar portions of the embedding space. Additionally, these portions of the embedding space can be associated with one or more semantic categories among a plurality of different semantic categories, and if a distance metric between a given embedding in the embedding and one or more portions of the embedding space meets a distance threshold, the given embedding in the embedding can be classified as a given semantic category among the semantic categories. For example, an instance of glass breaking can be associated with a first portion of the embedding space associated with the "glass breaking" sound, an instance of a doorbell ringing can be associated with a second portion of the embedding space associated with the "doorbell" sound, and so on.

[0050] In embodiments where one or more device - specific signals additionally or alternatively include unique identifiers or tags of additional assistant devices that are in position proximity to a given assistant device, candidate semantic tags can be generated based on those unique identifiers or tags, and a given semantic tag selected for the given assistant device can be selected based on one or more unique identifiers or tags of the additional assistant devices. For example, assume that a first tag, "smart oven", is associated with a first assistant device in position proximity to a given assistant device, and a second tag, "smart coffee maker", is associated with a second assistant device in position proximity to the given assistant device. In this example, the semantic tagging engine 160 can select the semantic tag "kitchen appliances" as the given semantic tag for the given assistant device because the tags associated with the additional assistant devices in position proximity to the given assistant device are cooking - related. The unique identifiers or tags can be processed using the semantic classifier stored in the ML model database 192 in the same or a similar manner as described above for processing queries, commands, and / or requests.

[0051] In embodiments where candidate semantic tags for a given assistant device are generated based on user preferences, a semantic classifier (e.g., stored in the ML model database 192) can be used to process the user preferences to index the user preferences into one or more different semantic categories corresponding to different types of user preferences. The candidate semantic tags can be generated based on the semantic categories into which the user preferences are classified, and the given semantic tag selected for the given assistant device can be selected based on the given semantic category related to the given assistant device into which the user preferences are classified. For example, assume that user preferences indicate that a user associated with an ecosystem likes cooking and likes a virtual chef named Johnny Flay. In this example, the candidate semantic tags can include, for example, a first semantic tag "cooking appliances" and a second candidate semantic tag "Johnny Flay appliances". In some versions of these embodiments, using the user preferences as a device - specific signal for generating one or more candidate semantic tags can be in response to receiving a user input to assign semantic tags to an assistant device based on user preferences.

[0052] In some embodiments, the semantic tagging engine 160 can automatically assign a given semantic tag to a given assistant device in a device topology representation of the ecosystem (e.g., stored in the device topology database 193). In some additional or alternative embodiments, the semantic tagging engine 160 can cause the automated assistant 120 to generate a prompt that includes candidate semantic tags. The prompt can request a selection from a user associated with the ecosystem of one of the candidate tags as the given semantic tag. Additionally, the prompt can be presented visually and / or audibly at a given assistant device in the assistant devices (which may or may not be the given assistant device to which the given semantic tag is assigned) and / or at the client device of the user (e.g., a mobile device). In response to receiving a selection of one of the candidate tags as the given semantic tag, the selected given semantic tag can be assigned to the given assistant device in the device topology representation of the ecosystem (e.g., stored in the device topology database 193). In some versions of these embodiments, the given semantic tag assigned to the given assistant device can be added to a list of semantic tags for the given assistant device. In other words, multiple semantic tags can be associated with the given assistant device. In other versions of these embodiments, the given semantic tag assigned to the given assistant device can replace any other semantic tag used for the given assistant device. In other words, only a single semantic tag can be associated with the given assistant device.

[0053] In some embodiments, the query / command processing engine 170 can process queries, requests, or commands that are directed to the automated assistant 120 and received via one or more assistant input devices 106. The query / command processing engine 170 can process the queries, requests, or commands to select one or more assistant devices to satisfy the query or command. It is noted that the one or more assistant devices selected to satisfy the query or command can be different from the one or more assistant input devices 106 that received the query or command. The query / command processing engine 170 can select one or more assistant devices to satisfy the spoken utterance based on one or more criteria. The one or more criteria can include, for example, the proximity of one or more devices to the user providing the spoken utterance (e.g., determined using the presence sensor 105 described below), the device capabilities of one or more devices in the ecosystem, the semantic tags assigned to one or more assistant devices, and / or other criteria for selecting assistant devices to satisfy the spoken utterance.

[0054] For example, assume that a display device is required to satisfy an oral utterance. In this example, candidate assistant devices considered when selecting a given assistant device to satisfy an oral utterance can be restricted to those that include a display device. If multiple assistant devices in the ecosystem include a display device, the given assistant device that includes the display device and is closest to the user can be selected to satisfy the utterance. In contrast, in an implementation where only a speaker is required to satisfy an oral utterance (e.g., a display device is not required to satisfy an oral utterance), candidate assistant devices considered when selecting a given assistant device to satisfy an oral utterance can include those that have a speaker, regardless of whether they include a display device.

[0055] As another example, assume that an oral utterance includes semantic attributes that match a semantic label assigned to a given assistant device. The query / command processing engine 170 can determine that the semantic attributes of the spoken utterance match the semantic label assigned to the given assistant device (e.g., it is an exact match or a soft match) by generating a first embedding corresponding to one or more terms (or text corresponding thereto) of the spoken utterance and a second embedding corresponding to one or more terms of the semantic label assigned to the given assistant device and comparing these embeddings to determine whether a distance metric between the embeddings meets a distance threshold indicating an embedding match. In this example, the query / command processing engine 170 can select the given assistant device to satisfy the oral utterance based on the oral utterance matching the semantic label (and optionally, in addition to or instead of the proximity of the user providing the oral utterance to the given assistant device). In this way, the selection of the assistant device to satisfy the oral utterance can be biased towards the semantic label assigned to the assistant device as described herein.

[0056] In various implementations, one or more assistant input devices 106 can include one or more corresponding presence sensors 105 1-N(Also referred to herein simply as "presence sensor 105"), which is configured to provide a signal indicating detected presence (especially human presence) with approval from the corresponding user. In some of these embodiments, the automated assistant 120 is capable of identifying one or more assistant input devices 106 to satisfy an oral utterance from a user associated with the ecosystem at least in part based on the presence of the user at one or more assistant input devices 106. The oral utterance can be satisfied by rendering response content at one or more assistant input devices 106 (e.g., audibly and / or visually), by causing one or more assistant input devices 106 to be controlled based on the oral utterance, and / or by causing one or more assistant input devices 106 to perform any other action to satisfy the oral utterance. As described herein, the automated assistant 120 can utilize data determined based on the corresponding presence sensor 105 to determine those assistant input devices 106 based on where the user is nearby or was recently nearby, and provide corresponding commands only to those assistant input devices 106. In some additional or alternative embodiments, the automated assistant 120 can utilize data determined based on the corresponding presence sensor 105 to determine whether any user (any user or a specific user) is currently near any assistant input device 106, and can optionally suppress the provision of commands based on determining that no user (any user or a specific user) is near any assistant input device 106.

[0057] The corresponding presence sensor 105 can take various forms. Some assistant input devices 106 can be equipped with one or more digital cameras, which are configured to capture and provide a signal indicating detected movement in their field of view. Additionally or alternatively, some assistant input devices 106 can be equipped with other types of light-based presence sensors 105, such as passive infrared ("PIR") sensors that measure infrared ("IR") light radiated from objects within their field of view. Additionally or alternatively, some assistant input devices 106 can be equipped with presence sensors 105 that detect acoustic (or pressure) waves, such as one or more microphones. Furthermore, in addition to assistant input devices 106, one or more assistant non-input devices 185 can additionally or alternatively include the corresponding presence sensors 105 described herein, and signals from such sensors can additionally be used by the automated assistant 120 to determine whether and / or how to satisfy an oral utterance according to the embodiments described herein.

[0058] Additionally or alternatively, in some embodiments, the sensor 105 can be configured to detect other phenomena associated with the presence of humans or devices in the ecosystem. For example, in some embodiments, a given assistant device in the assistant devices can be equipped with a presence sensor 105 that detects the presence of various types of wireless signals (e.g., waves such as radio, ultrasound, electromagnetic, etc.) transmitted by other assistant devices (e.g., mobile devices, wearable computing devices, etc.) carried / operated by a specific user and / or other assistant devices in the ecosystem (e.g., as described with respect to the event detection engine 130). For example, some assistant devices can be configured to transmit waves that are imperceptible to humans, such as ultrasonic or infrared waves, which can be detected by one or more assistant input devices 106 (e.g., via an ultrasonic / infrared receiver such as a microphone with ultrasonic capabilities).

[0059] Additionally or alternatively, various assistant devices can transmit other types of waves that are imperceptible to humans, such as radio waves (e.g., Wi-Fi, Bluetooth, cellular, etc.), which can be detected by other assistant devices (e.g., mobile devices, wearable computing devices, etc.) carried / operated by a specific user and used to determine the specific location of the operating user. In some embodiments, Wi-Fi triangulation can be used, for example, to detect a person's location based on Wi-Fi signals to / from the assistant device. In other embodiments, other wireless signal characteristics such as time of flight, signal strength, etc. can be used by various assistant devices individually or in combination to determine the location of a specific person based on signals transmitted by other assistant devices carried / operated by a specific user.

[0060] Additionally or alternatively, in some embodiments, one or more assistant input devices 106 can perform speech recognition to identify the user from the user's speech. For example, some instances of the automated assistant 120 can be configured to match the speech to the user's profile, e.g., for the purpose of providing / limiting access to various resources. In some embodiments, the movement of the speaker can then be determined, for example, by the presence sensor 105 of the assistant device. In some embodiments, based on this detected movement, the location of the user can be predicted, and when any content is rendered at the assistant device at least partially based on the proximity of the assistant device to the user's location, that location can be assumed to be the user's location. In some embodiments, the user can simply be assumed to be at the last location where he or she interacted with the automated assistant 120, especially if not much time has passed since the last interaction.

[0061] Each of the assistant input devices 106 further includes a corresponding user interface component 107 1-N(also referred to herein as “user interface component 107” for short), each of which can include one or more user interface input devices (e.g., microphone, touch screen, keyboard) and / or one or more user interface output devices (e.g., display, speaker, projector). As an example, the user interface component 1071 of the assistant input device 1061 can include only a speaker and a microphone, while the user interface component 107 of the assistant input device 106 N can include a speaker, a touch screen, and a microphone. Additionally or alternatively, in some embodiments, the assistant non-input device 185 can include one or more user interface input devices and / or one or more user interface output devices of the user interface component 107, but the user input devices (if any) for the assistant non-input device 185 may not allow the user to directly interact with the automated assistant 120. N Each of the assistant input devices 106 and / or any other computing device operating one or more cloud-based automated assistant components 119 can include one or more memories for storing data and software applications, one or more processors for accessing data and executing applications, and other components for facilitating communication over a network. Operations performed by one or more assistant input devices 106 and / or by the automated assistant 120 can be distributed across multiple computer systems. The automated assistant 120 can be implemented as, for example, a computer program running on one or more computers at one or more locations coupled to each other via a network (e.g.,

[0062] any of the networks 110). Figure 1 As described above, in various embodiments, each of the assistant input devices 106 can operate a corresponding automated assistant client 118. In various embodiments, each automated assistant client 118 can include a corresponding voice capture / text-to-speech (TTS) / speech-to-text (STT) module 114

[0063] (also referred to herein as “voice capture / TTS / STT module 114” for short). In other embodiments, one or more aspects of the corresponding voice capture / TTS / STT module 114 can be implemented separately from the corresponding automated assistant client 118. 1-N

[0064] Each respective voice capture / TTS / STT module 114 may be configured to perform one or more functions, including, for example: capturing a user's voice (voice capture, e.g., via a respective microphone (in some cases, the microphone may include a presence sensor 105)); converting the captured audio to text and / or other representations or embeddings (STT) using a speech recognition model stored in the ML model database 192; and / or converting text to speech (TTS) using a speech synthesis model stored in the ML model database 192. Instances of these models may be stored locally at each of the respective assistant input devices 106 and / or may be accessible by the assistant input devices (e.g., via the network 110 of Figure 1 ). In some embodiments, because one or more of the assistant input devices 106 may be relatively limited in terms of computing resources (e.g., processor cycles, memory, battery, etc.), the respective voice capture / TTS / STT modules 114 local to each of the assistant input devices 106 may be configured to use a speech recognition model to convert a limited number of different spoken phrases to text (or other forms, e.g., lower-dimensional embeddings). Other voice inputs may be sent to one or more cloud-based automated assistant components 119, which may include a cloud-based TTS module 116 and / or a cloud-based STT module 117. Figure 1 The cloud-based STT module 117 may be configured to utilize the virtually infinite resources of the cloud to convert audio data captured by the voice capture / TTS / STT module 114 to text (which may then be provided to the natural language processor module 122) using a speech recognition model stored in the ML model database 192. The cloud-based TTS module 116 may be configured to utilize the virtually infinite resources of the cloud to convert text data (e.g., text formulated by the automated assistant 120) to computer-generated voice output using a speech synthesis model stored in the ML model database 192. In some embodiments, the cloud-based TTS module 116 may provide the computer-generated voice output to one or more assistant devices for direct output, e.g., using the respective speakers of the respective assistant devices. In other embodiments, the text data generated by the automated assistant 120 using the cloud-based TTS module 116 (e.g., including client device notifications in a command) may be provided to the voice capture / TTS / STT module 114 of the respective assistant device, which may then convert the text data to computer-generated voice locally using the speech synthesis model and cause the computer-generated voice to be rendered via the local speaker of the respective assistant device.

[0065]

[0066] ​The automated assistant 120 (and in particular, one or more cloud-based automated assistant components 119) can include a natural language processing (NLP) module 122, the aforementioned cloud-based TTS module 116, the aforementioned cloud-based STT module 117, and other components, some of which are described in more detail below. In some embodiments, one or more engines and / or modules of the automated assistant 120 can be omitted, combined, and / or implemented in components separate from the automated assistant 120. Instances of the NLP module 122 can alternatively or additionally be implemented locally at the assistant input device 106.

[0067] In some embodiments, the automated assistant 120 generates response content in response to various inputs generated by a user in a human-machine dialogue session with the automated assistant 120 via one of the assistant input devices 106. The automated assistant 120 can provide the response content via the assistant input device 106 and / or the assistant non-input device 185 (e.g., when separated from the assistant device, via Figure 1 one or more of the networks 110) for presentation to the user as part of the dialogue session. For example, the automated assistant 120 can generate response content in response to a free-form natural language input provided via one of the assistant input devices 106. As used herein, a free-form input is an input formulated by the user and not limited to a set of options presented for the user to select.

[0068] The NLP module 122 of the automated assistant 120 processes the natural language input generated by the user via the assistant input device 106 and can generate an annotation output for use by one or more other components of the automated assistant 120, the assistant input device 106, and / or the assistant non-input device 185. For example, the NLP module 122 can process a natural language free-form input generated by the user via one or more corresponding user interface input devices of the assistant input device 106. The annotation output generated based on processing the natural language free-form input can include one or more annotations of the natural language input and optionally one or more (e.g., all) terms of the natural language input.

[0069] In some embodiments, the NLP module 122 is configured to identify and annotate various types of syntactic information in natural language input. For example, the NLP model 122 can include a part-of-speech tagger that is configured to annotate terms with their syntactic roles. In some embodiments, the NLP model 122 can additionally and / or alternatively include an entity tagger (not depicted) that is configured to annotate entity references in one or more segments, such as references to people (including, for example, literary characters, celebrities, public figures, etc.), organizations, locations (real and fictional), and the like. In some embodiments, data about entities can be stored in one or more databases, such as a knowledge graph (not depicted). In some embodiments, the knowledge graph can include nodes representing known entities (and in some cases, entity attributes) and edges connecting the nodes and representing relationships between the entities.

[0070] The entity tagger of the NLP module 122 can annotate references to entities at a higher granularity level (e.g., such that all references to an entity class such as a person can be identified) and / or at a lower granularity level (e.g., such that all references to a specific entity such as a specific person can be identified). The entity tagger can rely on the content of the natural language input to disambiguate specific entities and / or can optionally communicate with a knowledge graph or other entity database to disambiguate specific entities.

[0071] In some embodiments, the NLP module 122 can additionally and / or alternatively include a coreference resolver (not depicted) that is configured to group or "cluster" references to the same entity based on one or more context clues. For example, based on the "front door lock" mentioned in a client device notification rendered immediately before receiving the natural language input "lock it", the coreference resolver can resolve the term "it" in the natural reference input "lock it" to "front door lock".

[0072] In some embodiments, one or more components of the NLP module 122 can rely on annotations from one or more other components of the NLP module 122. For example, in some embodiments, a named entity tagger can rely on annotations from a coreference resolver and / or a dependency parser to annotate all mentions of a specific entity. Similarly, for example, in some embodiments, when clustering references to the same entity, the coreference resolver can rely on annotations from a dependency parser. In some embodiments, when processing a specific natural language input, one or more components of the NLP module 122 can use relevant data outside of the specific natural language input to determine one or more annotations - such as an assistant input device notification rendered immediately before receiving the natural language input on which the assistant input device notification is based.

[0073] Although Figure 1 depicted as having a specific configuration with components implemented by an assistant device and / or a server and depicted as having an assistant device and / or a server communicating via a specific network, it should be understood that this is not meant to be limiting for example purposes. For example, the assistant input device 106 and the assistant non-input device can be directly communicatively coupled to each other via one or more networks (not depicted). As another example, the operation of one or more cloud-based automated assistant components 119 can be implemented locally at one or more assistant input devices 106 and / or one or more assistant non-input devices. As yet another example, instances of various ML models stored in the ML model database 192 can be stored locally at the assistant device, and / or instances of the device topology representation of the ecosystem stored in the device topology database 193 can be stored locally at the assistant input device. In addition, in an implementation where data (e.g., device activity, audio data or the recognized text corresponding thereto, device topology representation, and / or any other data described herein) is transmitted over any one of the one or more networks 110 via Figure 1 , the data can be encrypted, filtered, or otherwise protected in any manner to ensure user privacy.

[0074] By using the techniques described herein to infer semantic tags and assign semantic tags to assistant devices in the ecosystem, the device topology representation of the ecosystem can be kept up to date. In addition, the semantic tags assigned to the assistant devices are semantically meaningful to the user because the semantic tags assigned to the corresponding assistant devices are selected based on the use of the corresponding assistant devices and / or the corresponding parts of the ecosystem in which the corresponding assistant devices are located. Thus, when an utterance is received at one or more assistant devices in the ecosystem, the automated assistant can more accurately select one or more assistant devices most suitable to satisfy the utterance. Thus, the number of user inputs received by one or more assistant devices in the ecosystem can be reduced because the user associated with the ecosystem does not need to specify a particular device to satisfy the utterance or repeat the utterance in the case where an incorrect device is selected to satisfy the utterance, thereby saving computing resources and / or network resources at the assistant device by reducing network traffic. In addition, the number of user inputs received by one or more assistant devices in the ecosystem can be reduced because when an assistant device is newly added to the ecosystem or moved within the ecosystem, the user does not need to manually update the device topology representation via a software application associated with the ecosystem.

[0075] Now refer to Figure 2A and Figure 2B for additional description of the various components of Figure 1 . In Figure 2A and Figure 2BA floor plan of a home is depicted. The depicted floor plan includes multiple rooms 250 to 262. Multiple assistant input devices 106 1-5 are deployed throughout at least some of the rooms. Each of the assistant input devices 106 1-5 can implement an instance of an automated assistant client 118 configured with selected aspects of the present disclosure, and can include one or more input devices, such as a microphone capable of capturing the words spoken by a nearby person. For example, a first assistant input device 1061 in the form of an interactive stand-alone speaker and display device (e.g., a display screen, a projector, etc.) is deployed in Figure 2A room 250 (which is a kitchen in this example) and in Figure 2B room 256 (which is a living room in this example). A second assistant input device 1062 in the form of a so-called "smart" TV (e.g., a networked TV having one or more processors implementing a corresponding instance of the automated assistant client 118) is deployed in room 252 (which is a study in this example). A third assistant input device 1063 in the form of an interactive stand-alone speaker without a display is deployed in room 254 (which is a bedroom in this example). A fourth assistant input device 1064 in the form of another interactive stand-alone speaker is deployed in room 256 (which is a living room in this example). A fifth assistant input device 1065 also in the form of a smart TV is also deployed in room 250 (which is a kitchen in this example).

[0076] Although not depicted in Figure 2A and Figure 2B the multiple assistant input devices 106 1-4 can be communicatively coupled to each other and / or to other resources (e.g., the Internet) via one or more wired or wireless WANs and / or LANs (e.g., via Figure 1 network 110). Additionally, there can also be other assistant input devices - particularly mobile devices such as smart phones, tablets, laptops, wearable devices, etc., e.g., carried by one or more people in the home, and the other assistant input devices may or may not also be connected to the same WAN and / or LAN. It should be understood that Figure 2A and Figure 2B the configuration of the assistant input devices depicted is only one example; more or fewer and / or different assistant input devices 106 can be deployed across any number of other rooms and / or areas in the home, and / or at locations other than a residential home (e.g., an enterprise, a hotel, a public place, an airport, a vehicle, and / or other locations or spaces).

[0077] In Figure 2A and Figure 2BFurther depicted in [the figure] are a plurality of assistant non-input devices 185 1-5 . For example, a first assistant non-input device 1851 in the form of a smart doorbell is deployed outside the home, near the front door of the home. A second assistant non-input device 1852 in the form of a smart lock is deployed outside the home, on the front door of the home. A third assistant non-input device 1853 in the form of a smart washing machine is deployed in room 262 (which is a laundry room in this example). A fourth assistant non-input device 1854 in the form of a door open / close sensor is deployed near the back door in room 262 and detects whether the back door is open or closed. A fifth assistant non-input device 1855 in the form of a smart thermostat is deployed in room 252 (which is a study in this example).

[0078] Each of the assistant non-input devices 185 is capable of (e.g., via Figure 1 network 110) communicating with a corresponding assistant non-input system 180 (shown in Figure 1 ) to provide data to the corresponding assistant non-input system 180 and optionally to be controlled based on commands provided by the corresponding assistant non-input system 180. One or more of the assistant non-input devices 185 can additionally or alternatively (e.g., via Figure 1 network 110) communicate directly with one or more assistant input devices 106 to provide data to the one or more assistant input devices 106 and optionally to be controlled based on commands provided by the one or more assistant input devices 106. It should be understood that Figure 2A and Figure 2B the configuration of the assistant non-input devices 185 depicted in [the figure] is merely an example; more or fewer and / or different assistant non-input devices 185 can be deployed across any number of other rooms and / or areas in the home and / or at locations other than a residential home (e.g., an enterprise, a hotel, a public place, an airport, a vehicle, and / or other locations or spaces).

[0079] In various embodiments, semantic tags can be assigned to a given assistant device (e.g., a given one of assistant input device 106 or assistant non-input device 185) based on the processing of one or more device-specific signals associated with the corresponding assistant device. The one or more device-specific signals can be detected and / or generated by the given assistant device. The one or more device-specific signals can include, for example, one or more queries previously received at the given assistant device (if any), one or more commands previously executed at the given assistant device (if any), instances of ambient noise previously detected at the given assistant device (and optionally only when voice reception is active at the given assistant device), unique identifiers (or tags) for the corresponding assistant device that is in a location proximate to the given assistant device, and / or user preferences of a user associated with the ecosystem. Each of the one or more device-specific signals associated with the given assistant device can be processed to classify each of them into one or more semantic categories out of a plurality of different semantic categories.

[0080] In addition, one or more candidate semantic tags can be generated based on the one or more device-specific signals. The candidate semantic tags can be generated using one or more rules (which are optionally heuristically defined) or machine learning models (e.g., stored in the ML model database 192). For example, one or more heuristically defined rules can indicate that candidate semantic tags should be generated that are associated with each of the semantic categories into which the one or more device-specific signals are classified. For example, assume that the device-specific signals are classified into the "kitchen" category, the "cooking" category, the "bedroom" category, and the "living room" category. In this example, the candidate semantic tags can include a first candidate semantic tag "kitchen assistant device", a second candidate semantic tag "cooking assistant device", a third candidate semantic tag "bedroom assistant device", and a fourth semantic tag "living room assistant device". As another example, a machine learning model trained to generate candidate semantic tags can be used to process the one or more device-specific signals (or the one or more semantic categories corresponding thereto). For example, the machine learning model can be trained based on a plurality of training instances. Each of the training instances can include a training instance input and a corresponding training instance output. The training instance input can include, for example, one or more device-specific signals and / or one or more semantic categories, and the corresponding training instance output can include, for example, a ground truth output corresponding to the semantic tag that should be assigned based on the training instance input.

[0081] In addition, a semantic tag assigned to a given assistant device can be selected from among one or more candidate semantic tags. The semantic tag to be assigned to a given assistant device can be selected from among one or more candidate semantic tags based on a confidence level associated with each of the one or more candidate semantic tags. In some embodiments, the semantic tag assigned to a given assistant device can be automatically assigned to the given assistant device, while in additional or alternative embodiments, a user associated with the Figure 2A and Figure 2B ecosystem can be prompted to select, from a list of one or more candidate semantic tags, the semantic tag to be assigned to the given assistant device (e.g., as described with respect to Figure 3 ). In some additional or alternative embodiments, the semantic tag can be automatically assigned to the given assistant device if the given semantic tag is unique (relative to other assistant devices in the ecosystem that are in a location proximate to the given assistant device).

[0082] In some versions of those embodiments, the given assistant device to which the semantic tag is assigned can be identified in response to determining that the given assistant device has been newly added to the ecosystem (e.g., via an event detection engine 130 and / or a device identification engine 140 of Figure 1 ). For example, and with specific reference to Figure 2A , assume that a first assistant input device 1061 in the form of an interactive standalone speaker and device is newly deployed in room 250 (which is the kitchen in this example). With respect to one or more device-specific signals associated with the first assistant input device 1061 in Figure 2A , assume that no previous queries or commands have been received at or executed by the first assistant input device 1061 (other than configuring the first assistant input device 1061), since the first assistant input device 1061 has been newly added to the ecosystem, assume that several instances of ambient noise have been captured while the user of the ecosystem is configuring the first assistant input device 1061 (e.g., when voice reception is active while the user provides an oral utterance including, for example, a name, a test utterance to establish a voice embedding for the user, etc.), assume that a corresponding unique identifier (or tag) associated with a fifth assistant input device 1065 in the form of a smart TV in room 250 (e.g., "kitchen TV") and a fifth assistant non-input device 1855 in the form of a smart thermostat in room 252 (e.g., "thermostat") is detected at the first assistant input device 1061, and assume that the user preferences of the user associated with the ecosystem are known.

[0083] In this example, an ambient noise detection model can be used to process instances of ambient noise (if any) to classify the ambient noise into one or more semantic categories. For example, assume that instances of ambient noise capture water droplets falling into a sink in room 250, a microwave oven or an oven making sounds in room 250, food sizzling on a pan on a stove in room 250, etc. These instances of ambient noise can be classified into a "kitchen" semantic category, a "cooking" semantic category, and / or other semantic categories related to the noises commonly encountered in a kitchen. Additionally or alternatively, the ambient noise can capture a movie, a TV show, or an advertisement visually and / or audibly rendered via a fifth assistant input device 1065 in the form of a smart TV in room 250. These instances of ambient noise can be classified into a "TV" semantic category, a "movie" semantic category, and / or other semantic categories related to the noises commonly encountered in a smart TV. Additionally or alternatively, the unique identifiers (or tags) of the fifth assistant input device 1065 (e.g., "kitchen TV") and the fifth assistant non-input device 1855 ("thermostat") can be processed to generate one or more semantic tags. These unique identifiers (or tags) can be classified into a "kitchen" semantic category, a "smart device" semantic category, and / or other semantic categories related to assistant devices that are physically close to the first assistant input device 1061 in the Figure 2A ecosystem. Additionally or alternatively, assume that user preferences indicate that the user is interested in cooking and a virtual chef named Johnny Flay. Based on being classified into a "cooking" category or a "kitchen" category by one or more other device-specific signals, and the user preferences for cooking and Johnny Flay also being classified into a "cooking" category or a "kitchen" category, these user preferences can be recognized as being related to the first assistant input device 1061. Thus, the candidate semantic tags in this example can include "kitchen speaker device", "cooking speaker device", "TV speaker device", "movie speaker device", "Johnny Flay device", and / or other candidate semantic tags based on one or more device-specific tags associated with the first assistant input device 1061. Further, in this example, a given semantic tag from among the semantic tags can be automatically assigned to the first assistant input device 1061, or the user associated with the ecosystem can be prompted to select one or more candidate semantic tags to assign the given semantic tag to the first assistant input device 1061 (e.g., while configuring the first assistant input device 1061).

[0084] In some additional or alternative embodiments, it can be done periodically (e.g., once a week, once a month, once every six months, and / or once at any other time period via Figure 1The event detection engine 130 and / or the device identification engine 140) identify a given assistant device to which a semantic tag is to be assigned. For example, and with specific reference to Figure 2A , assume that a first assistant input device 1061 in the form of an interactive standalone speaker and display device has been deployed in room 250 (which is the kitchen in this example) for six months. Regarding one or more device-specific signals associated with the first assistant input device 1061 in Figure 2A , assume that a previous query or command has been received or executed at the first assistant input device 1061, assume that an instance of ambient noise has been captured at the first assistant input device 1061, and assume that corresponding unique identifiers (or tags) associated with a fifth assistant input device 1065 in the form of a smart TV (e.g., "kitchen TV") in room 250 and a fifth assistant non-input device 1855 in the form of a smart thermostat (e.g., "thermostat") in room 252 are still detected at the first assistant input device 1061.

[0085] In this example, a semantic classifier can be used to process queries and commands (or their corresponding text) to classify the queries and commands into one or more semantic categories. For example, previous queries and commands received at the first assistant input device 1061 can include queries related to requesting cooking recipes, commands related to setting a timer, commands related to controlling any smart device in the kitchen, and / or other queries or commands. These instances of queries and commands can be classified into a "cooking" semantic category, a "control smart device" category, and / or other semantic categories based on the queries and commands received at the first assistant input device 1061. As described above, candidate semantic tags can additionally or alternatively be determined based on instances of ambient noise and / or unique identifiers (or tags) located in proximity to the first assistant input device 1061. Thus, candidate semantic tags in this example can include "kitchen device", "timer device", "thermostat display device", "cooking display device", "TV device", "movie device", "Johnny Flay recipe device", and / or other candidate semantic tags based on one or more device-specific tags associated with the first assistant input device 1061. Additionally, in this example, a given semantic tag from among the semantic tags can be automatically assigned to the first assistant input device 1061, or the user associated with the ecosystem can be prompted to select one or more candidate semantic tags to assign the given semantic tag to the first assistant input device 1061 (e.g., while configuring the first assistant input device 1061).

[0086] In some additional or alternative embodiments, in response to (e.g., via Figure 1The event detection engine 130 and / or the device identification engine 140) determine that a given assistant device has moved locations within the ecosystem to identify the given assistant device to which a semantic tag is to be assigned. For example, and now specifically referring to Figure 2B , assume that a first assistant input device 1061 in the form of an interactive standalone speaker and display device moves from a room 250, which is a kitchen in this example, to a room 256, which is a living room in this example. Regarding one or more device-specific signals associated with the first assistant input device 1061 in Figure 2B , assume that a previous query or command has been received at or executed by the first assistant input device 1061, assume that several instances of ambient noise have been captured, and assume that a corresponding unique identifier (or tag) associated with a fourth assistant input device 1064 in the form of another interactive standalone speaker (e.g., "living room speaker device") in room 256 has been detected at the first assistant input device 1061, and assume that the user preferences of the user associated with the ecosystem are known. In this example, one or more device-specific signals can be restricted to those generated or received after the first assistant input device 1061 has moved locations within the ecosystem (less user preferences).

[0087] In this example, a semantic classifier can be used to process queries and commands (or their corresponding text) to classify the queries and commands into one or more semantic categories. For example, queries and commands previously received at the first assistant input device 1061 can include queries related to requesting weather or traffic information, commands related to planning a vacation, and / or other queries or commands. These instances of queries and commands can be classified as an "information" semantic category (or more specifically, a "weather information" category and a "traffic information" category), a "planning" category, and / or other semantic categories based on the queries and commands received at the first assistant input device 1061. Additionally or alternatively, an ambient noise detection model can be used to process instances of ambient noise to classify the ambient noise into one or more semantic categories. For example, the ambient noise can capture music or podcasts audibly rendered by the fourth assistant input device 1064, people talking on a couch depicted in room 256, movies or TV shows audibly rendered by a computing device in the ecosystem, etc. These instances of ambient noise can be classified as a "music" semantic category, a "talking" semantic category, a "movie" semantic category, a "TV show" semantic category, and / or other semantic categories related to the noise typically encountered in a kitchen. Additionally or alternatively, the unique identifier (or tag) of the fourth assistant input device 1064 (e.g., "living room speaker device") can be processed to generate one or more semantic tags. The unique identifier (or tag) can be classified as, for example, a "living room" semantic category and / or related to theFigure 2B Other semantic categories related to assistant devices that are located in the ecosystem in proximity to the first assistant input device 1061. Additionally or alternatively, assume that user preferences indicate that the user is interested in cooking and a virtual movie called Vehicles and a specific character in the movie named Thunder McKing. These user preferences can be classified as "movie" based on one or more other device-specific signals, and the user preferences for the movie Vehicles and Thunder McKing are also classified as the "movie" category and are identified as being related to the first assistant input device 1061. Thus, the candidate semantic tags in this example can include "living room device", "planning device", "vehicle device", "Thunder McKing device", and / or other candidate semantic tags based on one or more device-specific tags associated with the first assistant input device 1061. Additionally, in this example, a given semantic tag from among the semantic tags can be automatically assigned to the first assistant input device 1061, or the user associated with the ecosystem can be prompted to select one or more candidate semantic tags to assign the given semantic tag to the first assistant input device 1061.

[0088] In some additional or alternative embodiments, a given assistant device to which a semantic tag is to be assigned can be identified in response to (e.g., via Figure 1 the event detection engine 130 and / or the device identification engine 140) determining that the portion of the ecosystem in which the given assistant device is located has been repurposed. For example, and now specifically referring to Figure 2B , assume that the first assistant input device 1061 in the form of an interactive standalone speaker and display device is located in room 256 (which is the living room in this example), but the living room has been repurposed as a bedroom. Regarding one or more device-specific signals associated with the first assistant input device 1061 in Figure 2B , assume that previous queries or commands have been received at or executed by the first assistant input device 1061, assume that several instances of ambient noise have been captured, and assume that a corresponding unique identifier (or tag) associated with a fourth assistant input device 1064 in the form of another interactive standalone speaker (e.g., "living room speaker device") in room 256 has been detected at the first assistant input device 1061. In this example, the one or more device-specific signals can be limited to those generated or received after room 256 has been repurposed.

[0089] In this example, a semantic classifier can be used to process queries and commands (or their corresponding text) to classify the queries and commands into one or more semantic categories. For example, the queries and commands previously received at the first assistant input device 1061 can include commands related to setting an alarm, commands related to a good morning or good night routine, and / or other queries or commands. These instances of commands can be classified into an "alarm" semantic category, a "routine" category (or more specifically, a "morning routine" category or a "night routine" category), and / or other semantic categories based on the queries and commands received at the first assistant input device 1061. Additionally or alternatively, an ambient noise detection model can be used to process instances of ambient noise to classify the ambient noise into one or more semantic categories. For example, the ambient noise can capture the snoring of one or more users, human conversations, etc. These instances of ambient noise can be classified into a "bedroom" semantic category, a "conversation" semantic category, and / or other semantic categories related to the noise typically encountered in a bedroom. Additionally or alternatively, the unique identifier (or label) of the fourth assistant input device 1064 (e.g., "living room speaker device") can be processed to generate one or more semantic labels. This unique identifier (or label) can be classified into a "living room" semantic category and / or other semantic categories related to assistant devices that are physically close to the first assistant input device 1061 in the Figure 2B ecosystem. Thus, the candidate semantic labels in this example can include "living room display device", "bedroom display device", and / or other candidate semantic labels based on one or more device-specific labels associated with the first assistant input device 1061. Further, in this example, a given semantic label from among the semantic labels can be automatically assigned to the first assistant input device 1061, or the user associated with the ecosystem can be prompted to select one or more candidate semantic labels to assign the given semantic label to the first assistant input device 1061. It is noted that in this example, the system can determine that the given semantic label corresponds to a "bedroom display device" even though the unique identifier (or label) of the fourth assistant input device 1064 corresponds to a "living room" category because the use of the first assistant input device 1061 indicates that it is located in the bedroom.

[0090] Although in this document 2A is described with respect to a given assistant device to which a semantic label is assigned being an assistant input device (e.g., the first assistant input device 1061) Figure 2B, but it should be understood that this is for illustrative purposes and is not meant to be limiting. For example, the techniques described herein can also be used to assign corresponding semantic tags to assistant non-input devices 185. For example, assume that a smart light (which may be an assistant non-input device without any microphones) is newly added to room 252 (which is a bedroom in this example). In this example, the unique identifier or tag associated with a third assistant input device 1063 in the form of an interactive stand-alone speaker without a display (e.g., a "bedroom speaker device") can be used to infer the semantic tag of the newly added smart light, "bedroom smart light", using the techniques described herein. Additionally, assume that the smart light is moved from room 254 to room 262 (which is a laundry room in this example). In this example, the unique identifier or tag associated with a third assistant non-input device 1853 in the form of a smart washing machine can be used to infer the semantic tag of the recently moved smart light, "laundry room smart light", using the techniques described herein.

[0091] Now turning to Figure 3 , a flowchart of an example method 300 for illustrating the assignment of a given semantic tag to a given assistant device in an ecosystem is described. For convenience, the operations of method 300 are described with reference to the system performing the operations. The system of method 300 includes one or more processors and / or other components of a computing device. For example, the system of method 300 can be implemented by Figure 1 , Figure 2A or Figure 2B an assistant input device 106 of Figure 1 , Figure 2A or Figure 2B an assistant non-input device 185 of Figure 5 , a computing device 510 of

[0092] Figure 1 one of the assistant input devices 106 of Figure 1 Figure 1 one of the assistant non-input devices 185 of Figure 1as described by the event detection engine 130). In some additional or alternative embodiments, a given assistant device can be identified periodically (e.g., once a month, once every six months, once a year, etc.).

[0093] At block 354, the system obtains a device-specific signal associated with a given assistant device. The device-specific signal can be detected and / or generated by the given assistant device. In some embodiments, block 354 can include one or more of optional sub-blocks 354A, sub-block 354B, sub-block 354C, or sub-block 354D. If included, at sub-block 354A, the system obtains the plurality of queries or commands (if any) previously received at the given assistant device. If included, at sub-block 354B, the system additionally or alternatively obtains an instance of ambient noise previously detected at the given assistant device (and optionally, only when voice reception is active at the given assistant device (e.g., after receiving a particular word or phrase that invokes the automated assistant), or via a digital signal processor (DSP) when voice reception is inactive). In some embodiments, the ambient noise obtained is limited to the ambient noise detected when voice reception is active at the given assistant device. If included, at sub-block 354C, the system additionally or alternatively obtains (e.g., using Figure 1 the unique identifier (or tag) of the corresponding assistant device that is in a location close to the given assistant device as determined by the device identification engine 140). If included, at sub-block 354D, the system additionally or alternatively obtains the user preferences of the user associated with the ecosystem.

[0094] At block 356, the system processes device-specific signals to generate candidate semantic tags for a given assistant device. In embodiments in which one or more of the device-specific signals include multiple queries or commands previously received at the given assistant device, a semantic classifier can be used to process the multiple queries or commands (or the text corresponding thereto) to classify each of the multiple queries or commands into one or more different semantic categories. For example, queries related to cooking recipes and commands related to controlling a smart oven or a smart coffee maker in a kitchen can be classified into a cooking category or a kitchen category, queries related to weather can be classified into a weather category, commands related to controlling lights can be classified into a lights category, and so on. In embodiments in which one or more of the device-specific signals additionally or alternatively include instances of ambient noise detected at the given assistant device, an ambient noise detection model can be used to process the instances of ambient noise to classify each of the instances of ambient noise into one or more different semantic categories. For example, if an instance of ambient noise is determined to correspond to a microwave oven making a sound, food sizzling in a frying pan, a food processor processing food, etc., the instance of ambient noise can be classified into a cooking category. As another example, if an instance of ambient noise is determined to correspond to a saw buzzing, a hammer pounding, etc., the instance of ambient noise can be classified into a garage category and a workshop category. In embodiments in which one or more of the device-specific signals additionally or alternatively include unique identifiers (or tags) of corresponding assistant devices that are in a location close to the given assistant device, the unique identifiers (or tags) for the corresponding assistant devices can be classified into one or more different semantic categories. For example, if the unique identifier (or tag) for a corresponding assistant device corresponds to "coffee maker", "oven", "microwave oven", the unique identifier can be classified into a kitchen category or a cooking category. As another example, if the unique identifier (or tag) for a corresponding assistant device corresponds to "bedroom light" and "bedroom projection device", the unique identifier can be classified into a bedroom category. In embodiments in which one or more of the device-specific signals additionally or alternatively include user preferences of a user associated with an ecosystem, the user preferences can be classified into one or more different semantic categories. For example, if the user preferences indicate that the user is interested in cooking, cooking shows, a particular chef, and / or other cooking-related interests, the user preferences can be classified into a kitchen category, a cooking category, or a category associated with the particular chef.

[0095] Candidate semantic tags can be generated based on the processing of one or more device-specific signals. For example, assume that one or more device-specific signals indicate that a given assistant device is located in the kitchen or living room of a user's primary home associated with an ecosystem. Further assume that the given assistant device is an interactive stand-alone speaker device with a display. In this example, a first candidate semantic tag "Kitchen display device" and a second candidate semantic tag "Living room display device" can be generated. In this document (e.g., refer to Figure 2A and Figure 2B ), generating candidate semantic tags based on one or more device-specific signals is described in more detail.

[0096] At block 358, the system determines whether to prompt a user associated with the ecosystem to request a selection of a given semantic tag from among the candidate semantic tags. The system can determine whether to prompt the user to request a selection of a given semantic tag based on whether the corresponding confidence level associated with the candidate semantic tag meets a threshold confidence level. The corresponding confidence level can be determined based on, for example, the number of one or more device-specific signals classified as a given semantic category. For example, assume that the given assistant device identified at block 352 is an interactive stand-alone speaker that implements an instance of an automated assistant. Further assume that each of the one or more device-specific signals indicates that the interactive stand-alone speaker is located in a bedroom. For example, based on previous queries or commands received at the interactive stand-alone speaker associated with setting an alarm or a good night routine, based on instances of ambient noise including snoring, and / or based on other assistant devices having unique identifiers (or tags) of "bedroom light" and "bedroom projection device". In this example, since the interactive stand-alone speaker is associated with the bedroom in the user's primary home associated with the ecosystem, the system can be highly confident in the semantic tag "Bedroom speaker device". However, if some of the queries or commands received at the interactive stand-alone speaker are associated with cooking recipes, the system may be less confident in the semantic tag "Bedroom speaker device".

[0097] If, at an iteration of block 358, the system determines not to prompt the user to request a selection of a given semantic tag, the system can proceed to block 360. At block 360, the system automatically assigns the given semantic tag from among the candidate semantic tags to a given assistant device in the device topology representation in the ecosystem. The device topology representation of the ecosystem can be stored locally at one or more assistant devices in the ecosystem and / or at a remote system communicating with one or more assistant devices in the ecosystem. In some embodiments, the given semantic tag can be the only semantic tag associated with the given assistant device (and optionally replace other unique identifiers or tags assigned to the given assistant device), while in other embodiments, the given semantic tag can be added to a list of unique identifiers or tags assigned to the given assistant device. The system can then return to block 352 to identify another given assistant device from among the plurality of assistant devices in the ecosystem, thereby generating another given semantic tag and assigning it to the another given assistant device.

[0098] If at the iteration of box 358, the system determines that the user is prompted to request the selection of a given semantic tag, the system can proceed to box 362. At box 362, the system generates a prompt requesting the selection of a given semantic tag from the user of the client device based on the candidate semantic tags. At box 364, the system causes the prompt to be rendered at the user's client device. The prompt can be rendered visually and / or audibly at the user's client device, and can optionally be based on the capabilities of the client device at which the prompt is rendered. For example, the prompt can be rendered via a software application accessible at the client device (e.g., a software application associated with the ecosystem or one or more assistant devices included in the ecosystem). In this example, if the client device includes a display, the prompt can be rendered visually via the software application (or the home screen of the client device), and / or audibly via the speaker of the client device. However, if the client device does not include a display, the prompt can only be rendered audibly via the speaker of the client device. At box 366, the system receives a selection of a given semantic tag in response to the prompt. The user's client device can be, for example, the given assistant identified at box 352 or a different client device (e.g., the user's mobile device, or any other assistant device in the ecosystem that is capable of rendering the prompt). For example, assume that the system generates candidate semantic tags "bedroom speaker device" and "kitchen speaker device". In this example, the system can generate a prompt that includes selectable elements associated with both of these semantic tags, and request the user to provide input (e.g., touch or verbal) to select one of the candidate semantic tags to assign to the given assistant device as the given semantic tag to be assigned to the given assistant device. At box 368, the system assigns the given semantic tag to the given assistant device in the device topology representation of the ecosystem in a manner similar to that described above with respect to box 360.

[0099] Now go to Figure 4 , depicts a flow chart illustrating an example method 400 of using assigned semantic tags to satisfy a query or command received at an assistant device in an ecosystem. For convenience, the operations of method 400 are described with reference to a system performing the operations. The system of method 400 includes one or more processors and / or other components of a computing device. For example, the system of method 400 can be Figure 1 , Figure 2A or Figure 2B assistant input device 106, Figure 1 , Figure 2A or Figure 2B Assistant non-input device 185, Figure 5implemented by the computing device 510, one or more servers, other computing devices, and / or any combination thereof. Additionally, although the operations of method 400 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, and / or added.

[0100] At block 452, the system receives audio data corresponding to a user's spoken utterance via a respective microphone of a respective assistant device in an ecosystem that includes a plurality of assistant devices. The user may be associated with the ecosystem.

[0101] At block 454, the system processes the audio data to identify semantic attributes of a query or command included in the spoken utterance. The semantic attributes of the query or command may correspond to language units, such as words or phrases, that define related fields or related sets of words and / or phrases. In some embodiments, the system is capable of using a speech recognition model to process the audio data to convert the spoken utterance captured in the audio data into text, and is capable of using a semantic classifier to identify semantic attributes based on the recognized text. In additional or alternative embodiments, the system is capable of using a semantic classifier to process the audio data and is capable of directly identifying semantic attributes based on the audio data. For example, assume that the spoken utterance received at block 452 is "Show me a chili recipe." In this example, the semantic classifier can be used to process the spoken utterance (or the text corresponding thereto) to identify the semantic attributes "chili," "food," "kitchen," and / or "cooking." It is noted that the semantic attributes identified at block 454 may include one or more terms or phrases included in the spoken utterance, or may include a given semantic category to which the spoken utterance is classified.

[0102] At block 456, the system determines whether the spoken utterance designates a given assistant device for fulfilling the query or command included in the spoken utterance. If, at an iteration of block 456, the system determines that the spoken utterance designates a given assistant device for fulfilling the query or command, the system may proceed to block 466. Block 466 is described below. For example, assume that the spoken utterance received at block 452 is "Show me a chili recipe on the kitchen display device". In this example, the system can utilize the "kitchen display device" because the user has designated that the "kitchen display device" should be used to present the "chili recipe" to the user who provided the spoken utterance. Thus, the system can select the "kitchen display device" to fulfill the spoken utterance. If, at an iteration of block 456, the system determines that the spoken utterance does not designate a given assistant device for fulfilling the query or command, the system may proceed to block 458. For example, assume that the spoken utterance received at block 452 is simply "Show me a chili recipe" without designating any assistant device to fulfill the spoken utterance. In this example, the system can determine that the spoken utterance does not designate a given assistant device to fulfill the spoken utterance. Thus, the system needs to determine which (if any) assistant device in the ecosystem should be utilized to fulfill the spoken utterance.

[0103] At block 458, the system determines whether the semantic properties identified at block 454 match a given semantic label assigned to a given assistant device from among multiple assistant devices. The system can generate an embedding of one or more terms corresponding to the semantic properties and can compare the embedding of the semantic properties to one or more corresponding embeddings of one or more terms corresponding to the respective semantic labels assigned to one or more of the multiple assistant devices in the ecosystem. Additionally, the system can determine whether the embedding of the semantic properties matches any of the one or more corresponding embeddings of the respective semantic labels assigned to one or more of the multiple assistant devices in the ecosystem. For example, the system can determine whether a distance metric between the embedding of the semantic properties and each of the one or more corresponding embeddings of the respective semantic labels assigned to one or more of the multiple assistant devices in the ecosystem is. Additionally, the system can determine whether the distance metric meets a distance threshold (e.g., identifying an exact match or a soft match).

[0104] If, at an iteration of block 458, the system determines that the semantic attributes identified at block 454 match a given semantic label assigned to a given assistant device, the system can proceed to block 460. At block 460, the system causes a given client device associated with the given semantic label in the device topology representation of the ecosystem to satisfy the query or command included in the spoken utterance. The system can then return to block 452 to monitor additional audio via the respective microphones of the plurality of assistant devices in the ecosystem. More specifically, the system is capable of causing a given assistant device to perform one or more actions to satisfy the utterance. For example, assume that the spoken utterance received at block 452 has identified semantic attributes of "peppers", "food", "kitchen", and / or "cooking" of "show me a pepper recipe". Further assume that the semantic label assigned to the given assistant device is "kitchen display device". In this example, even though the spoken utterance was not received at this assistant device, the system is capable of selecting the given assistant device assigned the semantic label "kitchen display device". Additionally, the system is capable of causing the given assistant device assigned the semantic label "kitchen display device" to visually render a pepper recipe in response to the spoken utterance received by the respective microphone of the respective assistant device.

[0105] In various embodiments, semantic labels that match semantic attributes identified based on a spoken utterance can be assigned to the plurality of assistant devices in the ecosystem. Notably, and although not depicted in Figure 4 for clarity, in addition to or instead of using proximity information (e.g., as described above with respect to Figure 1 the query / command processing engine 170), when selecting one or more assistant devices to satisfy a spoken utterance, the system can determine whether the semantic attributes identified at block 454 match a given semantic label assigned to a given assistant device from among the plurality of assistant devices. Continuing the above example, assume that there are a plurality of assistant devices assigned the semantic label "kitchen display device". In this example, the assistant device among the plurality of assistant devices assigned the semantic label "kitchen display device" that is closest to the user in the ecosystem can be used to satisfy the spoken utterance. Additionally, and also not depicted in Figure 4 for clarity, in addition to or instead of using device capability information (e.g., as described above with respect to Figure 1As described by the query / command processing engine 170, when selecting one or more assistant devices to satisfy an oral utterance, the system can determine whether the semantic attributes identified at block 454 match a given semantic tag assigned to a given assistant device from among multiple assistant devices. Continuing with the above example, assume that a first assistant device with a display device is assigned the semantic tag "kitchen display device", and a second assistant device without a display device is assigned the semantic tag "kitchen speaker device". In this example, the first assistant device assigned the semantic tag "kitchen display device" can be selected over the second assistant device assigned the semantic tag "kitchen speaker device" to satisfy the oral utterance because the oral utterance specifies "show me" a chili recipe, and the first assistant device can display the chili recipe in response to the oral utterance, while the second assistant device cannot display the chili recipe.

[0106] If, at the iteration of block 458, the system determines that the semantic attributes identified at block 454 do not match any semantic tag assigned to any assistant device in the ecosystem, the system can proceed to block 462. At block 462, the system identifies a given assistant device that is close to the user. For example, the system can identify the given assistant device that is closest to the user in the ecosystem (e.g., as described with respect to Figure 1 the presence sensor 105).

[0107] At block 464, the system determines whether a given assistant device identified at block 462 is able to satisfy the query or command. The capabilities of each assistant device can be stored in a device topology representation of the ecosystem and associated with the corresponding assistant device (e.g., as a device attribute for the corresponding assistant device). If, at an iteration of block 464, the system determines that the given assistant device identified at block 462 cannot satisfy the query or command, the system can return to block 462 to identify another given assistant device that is also near the user. The system can then proceed again to block 464 to determine whether the other given assistant device identified at a subsequent iteration of block 462 is able to satisfy the query or command. The system can repeat this process until an assistant device that can satisfy the query or command is identified. For example, assume that the assistant device identified at block 462 is a standalone speaker device that lacks a display but requires a display to satisfy the spoken utterance. In this example, the system can determine at block 464 that the assistant device identified at block 462 cannot satisfy the spoken utterance and can return to block 462 to identify another assistant device near the user at block 462. If, at an iteration of block 464, the system determines that the given assistant device identified at block 462 is able to satisfy the query or command, the system can proceed to block 466. Continuing the above example, further assume that the given assistant device identified at the first iteration of block 462 is a standalone speaker device that includes a display, or that the other assistant device identified at another iteration of block 462 is a standalone speaker device that has a display. In this example, the system can determine that the assistant device identified at a subsequent iteration of block 462 can satisfy the spoken utterance at block 464, and the system can proceed to block 466.

[0108] At block 466, the system causes the given assistant device to satisfy the query or command included in the spoken utterance. The system can satisfy the spoken utterance in a similar manner as described above with respect to block 460. The system can then return to block 452 to monitor additional audio via the respective microphones of the plurality of assistant devices in the ecosystem.

[0109] Figure 5 is a block diagram of an example computing device 510 that can optionally be used to perform one or more aspects of the techniques described herein. In some embodiments, one or more assistant input devices, one or more cloud-based automated assistant components, one or more assistant non-input systems, one or more assistant non-input devices, and / or other components can include one or more components of the example computing device 510.

[0110] The computing device 510 generally includes at least one processor 514 that communicates with a plurality of peripheral devices via a bus subsystem 512. These peripheral devices can include a storage subsystem 524 (including, for example, a memory subsystem 525 and a file storage subsystem 526), a user interface output device 520, a user interface input device 522, and a network interface subsystem 516. The input and output devices allow a user to interact with the computing device 510. The network interface subsystem 516 provides an interface to an external network and is coupled to corresponding interface devices in other computing devices.

[0111] The user interface input device 522 can include a keyboard, a pointing device such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touchscreen incorporated into a display, an audio input device such as a speech recognition system, microphone, and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways of inputting information into the computing device 510 or onto a communication network.

[0112] The user interface output device 520 can include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem can include a cathode ray tube display (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for generating a visible image. The display subsystem can also provide a non-visual display such as via an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and ways of outputting information from the computing device 510 to a user or to another machine or computing device.

[0113] The storage subsystem 524 stores programs and data structures that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 524 can include logic that performs selected aspects of the methods described herein and implements Figure 1 the various components depicted in

[0114] These software modules are typically executed by the processor 514 alone or in combination with other processors. The memory 525 used in the storage subsystem 524 can include multiple memories, and the multiple memories include a main random access memory (RAM) 530 for storing instructions and data during program execution, and a read-only memory (ROM) 532 in which fixed instructions are stored. The file storage subsystem 526 can provide persistent storage for program and data files, and can include a hard disk drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical disk drive or a removable media cartridge. The modules implementing the functions of certain embodiments can be stored in the storage subsystem 524 through the file storage subsystem 526, or stored in other machines accessible by one or more processors 514.

[0115] The bus subsystem 512 provides a mechanism for the various components and subsystems of the computing device 510 to communicate with each other in the expected manner. Although the bus subsystem 512 is schematically shown as a single bus, alternative embodiments of the bus subsystem can use multiple buses.

[0116] The computing device 510 can be of various types, and the various types include workstations, servers, computing clusters, blade servers, server farms or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, Figure 5 the description of the computing device 510 depicted herein is only intended as a specific example for the purpose of illustrating some embodiments. Compared with the computing device depicted herein, many other configurations of the computing device 510 with more or fewer components are possible. Figure 5

[0117] In certain embodiments discussed herein where personal information about a user may be collected or used (e.g., user data extracted from other electronic communications, information about the user's social network, the user's location, the user's time, the user's biometric information, and the user's activities and demographic information, relationships between users, etc.), the user is provided with one or more opportunities to control whether information is collected, whether personal information is stored, whether personal information is used, and how the information collected, stored, and used about the user is used. That is, the systems and methods discussed herein collect, store, and / or use user personal information only after receiving explicit authorization from the relevant user.

[0118] For example, a user is provided with control over whether a program or feature collects user information about that particular user or other users associated with the program or feature. Each user whose personal information is to be collected is provided with one or more options to allow control over the information collection related to that user, thereby providing permission or authorization regarding whether to collect the information and which portions of the information are to be collected. For example, one or more such control options can be provided to the user via a communication network. Additionally, before storing or using certain data, the certain data can be processed in one or more ways such that personally identifiable information is removed. As an example, a user's identity can be processed such that personally identifiable information cannot be determined. As another example, a user's geographical location can be generalized to a larger area such that the user's specific location cannot be determined.

[0119] In some embodiments, a method implemented by one or more processors is provided and includes: identifying a given assistant device among a plurality of assistant devices in an ecosystem; obtaining one or more device-specific signals associated with the given assistant device, the one or more device-specific signals being generated or received by the given assistant device; processing one or more of the device-specific signals to generate one or more candidate semantic tags for the given assistant device; selecting a given semantic tag for the given assistant device from among the one or more candidate semantic tags; and assigning the given semantic tag to the given assistant device in a device topology representation of the ecosystem. Assigning the given semantic tag to the given assistant device includes automatically assigning the given semantic tag to the given assistant device.

[0120] These and other embodiments of the techniques disclosed herein may include one or more of the following features.

[0121] In some embodiments, one or more device-specific signals can at least include device activities associated with a given assistant device, and the device activities associated with the given assistant device can include a plurality of queries or commands previously received at the given assistant device. In some versions of those embodiments, processing one or more of the device-specific signals to generate one or more candidate semantic labels for the given assistant device can include: using a semantic classifier to process the device activities associated with the given assistant device to classify each of the plurality of queries or commands into one or more of a plurality of different categories; and generating one or more of the candidate semantic labels based on the one or more of the plurality of different categories into which each of the plurality of queries or commands is classified. In some other versions of those embodiments, selecting a given semantic label for the given assistant device can include selecting the given semantic label for the given assistant device based on the number of the plurality of queries or commands classified into a given one of the plurality of different categories, the given semantic label being associated with the given semantic label.

[0122] In some embodiments, one or more of the device-specific signals can at least include device activities associated with a given assistant device, wherein the device activities associated with the given assistant device include environmental noise previously detected at the given assistant device, and the environmental noise can be environmental noise previously detected at the given assistant device during a voice reception activity. In some versions of those embodiments, the method can further include: using an environmental noise detection model to process the environmental noise previously detected at the given assistant device to classify the environmental noise into one or more of a plurality of different categories; and generating one or more of the candidate semantic labels based on the one or more of the plurality of different categories into which the environmental noise is classified. In some other versions of those embodiments, selecting a given semantic label for the given assistant device can include selecting the given semantic label for the given assistant device based on the environmental noise classified into a given one of the plurality of different categories, the given semantic label being associated with the given semantic label.

[0123] In some embodiments, one or more of the device-specific signals can include the respective unique identifiers of one or more of a plurality of assistant devices that are in the ecosystem and are in a location close to a given assistant device. In some versions of those embodiments, processing one or more of the device-specific signals to generate one or more of the candidate semantic labels for the given assistant device can include: identifying, based on one or more wireless signals, one or more of a plurality of assistant devices that are in the ecosystem and are in a location close to the given assistant device; obtaining the respective unique identifiers of one or more of a plurality of assistant devices that are in a location close to the given assistant device; and generating one or more of the candidate semantic labels based on the respective unique identifiers of one or more of a plurality of assistant devices that are in a location close to the given assistant device. In some other versions of those embodiments, selecting a given semantic label for the given assistant device can include selecting the given semantic label for the given assistant device based on the attributes of the respective unique identifiers of one or more of a plurality of assistant devices that are in a location close to the given assistant device.

[0124] In some embodiments, a given assistant device can be identified in response to determining that the given assistant device has been newly added to the ecosystem, or in response to determining that the given assistant device has moved locations within the ecosystem.

[0125] In some versions of those embodiments, a given assistant device can be identified in response to determining that the given assistant device has moved locations within the ecosystem. In some other versions of those embodiments, assigning a given semantic label to the given assistant device can include: adding the given semantic label to a list of semantic labels associated with the given assistant device; or replacing an existing semantic label associated with the given assistant device with the given semantic label. In some additional or alternative versions of those embodiments, determining that the given assistant device has moved locations within the ecosystem can include identifying, based on one or more wireless signals, that a current subset of a plurality of assistant devices that are in the ecosystem and are in a location close to the given assistant device is different from a stored subset of a plurality of assistant devices stored in association with the given assistant device. In other versions of those embodiments, the method can further include: switching the given assistant device from an existing group of assistant devices that includes one or more of a plurality of assistant devices to another existing group of assistant devices that includes one or more of a plurality of assistant devices, or creating a new group of assistant devices that includes at least the given assistant device.

[0126] In some versions of those embodiments, a given assistant device can be identified in response to determining that the given assistant device has been newly added to the ecosystem. In some other versions of those embodiments, assigning a given semantic tag to a given assistant device can include adding the given semantic tag to a list of semantic tags associated with the given assistant device. In some additional or alternative versions of those further embodiments, determining that a given assistant device has been newly added to the ecosystem can include identifying, based on one or more wireless signals, that the given assistant device has been added to a wireless network associated with the ecosystem. In some other versions of those embodiments, the method can further include: adding the given assistant device to an existing group of assistant devices that includes one or more assistant devices among a plurality of assistant devices, or creating a new group of assistant devices that includes at least the given assistant device.

[0127] In some embodiments, a given assistant device can be identified periodically to verify whether the existing semantic tag assigned to the given assistant device is correct.

[0128] In some embodiments, the method can further include, after assigning a given semantic tag to a given assistant device in a device topology representation of the ecosystem: receiving, via one or more respective microphones of one of the plurality of assistant devices in the ecosystem, audio data corresponding to an oral utterance from a user associated with the ecosystem, the oral utterance including a query or command; processing the audio data corresponding to the oral utterance to determine semantic attributes of the query or command; determining that the semantic attributes of the query or command match the given semantic tag assigned to the given assistant device; and, in response to determining that the semantic attributes of the query or command match the given semantic tag assigned to the given assistant device, causing the given assistant device to satisfy the query or command.

[0129] In some embodiments, one or more device-specific signals can include at least user preferences of a user associated with an ecosystem, and the user preferences can be determined based on user interactions with a plurality of assistant devices in the ecosystem. In some versions of those embodiments, processing one or more of the device-specific signals to generate one or more candidate semantic labels for a given assistant device can include: using a semantic classifier to process the user preferences to identify at least one semantic category associated with the user preferences among a plurality of different semantic categories; and generating one or more of the candidate semantic labels based on the given semantic category. In yet another version of those embodiments, selecting a given semantic label for a given assistant device can include: determining that the given semantic label among a plurality of candidate semantic labels is relevant to the given assistant device; and selecting the given semantic label for the given assistant device in response to determining that the given semantic label is relevant to the given assistant device. Determining that the given semantic label can be relevant to the given assistant device is based on one or more additional device-specific signals associated with the given assistant device. In some additional or alternative versions of those further embodiments, processing the user preferences to identify at least one semantic category can be in response to receiving user input that assigns one or more semantic labels to at least the given assistant device.

[0130] In some embodiments, a method implemented by one or more processors is provided and includes: identifying a given assistant device among a plurality of assistant devices in an ecosystem; obtaining one or more device-specific signals associated with the assistant device, the one or more device-specific signals being generated or received by the given assistant device; processing one or more of the device-specific signals to generate one or more candidate semantic labels for the given assistant device; generating a prompt requesting a selection of a given semantic label from a candidate semantic label among the one or more candidate semantic labels from a user of a client device, selecting one or more of the candidate semantic labels; causing the prompt to be rendered at the user's client device; and in response to receiving a selection of the given semantic label in response to the prompt, assigning the given semantic label to the given assistant device in a device topology representation of the ecosystem.

[0131] These and other embodiments of the techniques disclosed herein can include one or more of the following features.

[0132] In some embodiments, the one or more device-specific signals may include two or more of the following: device activity associated with a given assistant device, the device activity associated with the given assistant device including multiple queries or commands previously received at the given assistant device; ambient noise previously detected at the given assistant device, the ambient noise being previously detected at the given assistant device when voice reception activity occurred; or corresponding unique identifiers of one or more assistant devices among multiple assistant devices that are located proximate to the given assistant device in the ecosystem.

[0133] In some embodiments, a method implemented by one or more processors is provided and includes: identifying a given assistant device from among multiple assistant devices in an ecosystem; obtaining one or more device-specific signals associated with the given assistant device, the one or more device-specific signals being generated or received by the given assistant device; determining a given semantic tag for the given assistant device based on one or more of the device-specific signals; assigning the given semantic tag to the given assistant device in a device topology representation of the ecosystem; and after assigning the given semantic tag to the given assistant device in the device topology representation of the ecosystem: receiving audio data corresponding to a spoken utterance from a user associated with the ecosystem via one or more corresponding microphones of an assistant device among the multiple assistant devices in the ecosystem, the spoken utterance including a query or a command; processing the audio data corresponding to the spoken utterance to determine semantic attributes of the query or command; determining that the semantic attributes of the query or command match the given semantic tag assigned to the given assistant device; and in response to determining that the semantic attributes of the query or command match the given semantic tag assigned to the given assistant device, satisfying the query or command with the given assistant device.

[0134] In addition, some embodiments include one or more processors (e.g., central processing units (CPUs), graphics processing units (GPUs), and / or tensor processing units (TPUs)) of one or more computing devices, wherein the one or more processors are operable to execute instructions stored in an associated memory, and wherein the instructions are configured to cause any of the foregoing methods to be performed. Some embodiments also include one or more non-transitory computer-readable storage media storing computer instructions, which are executable by one or more processors to perform any of the foregoing methods. Some embodiments also include a computer program product comprising instructions executable by one or more processors to perform any of the foregoing methods.

[0135] It should be appreciated that all combinations of the above concepts and additional concepts described in more detail herein are contemplated as being part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are contemplated as being part of the subject matter disclosed herein.

Claims

1. A method implemented by one or more processors, the method comprising: Identifying a given assistant device among a plurality of assistant devices in an ecosystem; Obtaining one or more device-specific signals associated with the given assistant device, the one or more device-specific signals being generated or received by the given assistant device, wherein one or more of the device-specific signals at least include device activities associated with the given assistant device, and wherein the device activities associated with the given assistant device include environmental noise previously detected at the given assistant device during a voice reception activity; Processing one or more of the device-specific signals to generate one or more candidate semantic tags for the given assistant device, wherein processing one or more of the device-specific signals to generate the one or more candidate semantic tags for the given assistant device includes: Processing the environmental noise previously detected at the given assistant device during a voice reception activity to generate one or more of the candidate semantic tags for the given assistant device; Selecting a given semantic tag for the given assistant device from among the one or more candidate semantic tags; and Assigning the given semantic tag to the given assistant device in a device topology representation of the ecosystem, wherein assigning the given semantic tag to the given assistant device includes automatically assigning the given semantic tag to the given assistant device.

2. The method according to claim 1, wherein Processing the environmental noise previously detected at the given assistant device during a voice reception activity to generate one or more of the candidate semantic tags for the given assistant device includes: Using an environmental noise detection model to process the environmental noise previously detected at the given assistant device to classify the environmental noise into one or more of the plurality of different categories; and Generating one or more of the candidate semantic tags based on the one or more of the plurality of different categories into which the environmental noise is classified.

3. The method according to claim 2, wherein Selecting the given semantic tag for the given assistant device includes: Selecting the given semantic tag for the given assistant device based on the environmental noise classified into a given category of the plurality of different categories, the given semantic tag being associated with the given semantic tag.

4. The method according to claim 3, wherein, The one or more device-specific signals further include corresponding unique identifiers of one or more of the plurality of assistant devices that are physically close to the given assistant device in the ecosystem.

5. The method according to claim 4, wherein, Processing one or more of the device-specific signals to generate one or more of the candidate semantic tags for the given assistant device further includes: Identifying, based on one or more wireless signals, one or more of the plurality of assistant devices that are in a location close to the given assistant device in the ecosystem; Obtaining the respective unique identifiers of one or more of the plurality of assistant devices that are in a location close to the given assistant device; and Generating one or more of the candidate semantic labels further based on the respective unique identifiers of one or more of the plurality of assistant devices that are in a location close to the given assistant device.

6. The method according to claim 5, wherein, Selecting the given semantic label for the given assistant device further includes: Selecting the given semantic label for the given assistant device further based on the attributes of the respective unique identifiers of one or more of the plurality of assistant devices that are in a location close to the given assistant device.

7. The method according to claim 1, wherein Identifying the given assistant device in response to determining that the given assistant device has been newly added to the ecosystem, or in response to determining that the given assistant device has moved locations within the ecosystem.

8. The method according to claim 7, wherein, Identifying the given assistant device in response to determining that the given assistant device has moved locations within the ecosystem.

9. The method according to claim 8, wherein Assigning the given semantic label to the given assistant device includes: Adding the given semantic label to a list of semantic labels associated with the given assistant device; or Replacing an existing semantic label associated with the given assistant device with the given semantic label.

10. The method according to claim 8, wherein, Determining that the given assistant device has moved locations within the ecosystem includes: Identifying, based on one or more wireless signals, that a current subset of the plurality of assistant devices that are in a location close to the given assistant device in the ecosystem is different from a stored subset of the plurality of assistant devices stored in association with the given assistant device.

11. The method according to claim 10, further comprising: Switching the given assistant device from an existing group of assistant devices that includes one or more of the plurality of assistant devices to another existing group of assistant devices that includes one or more of the plurality of assistant devices, or Creating a new group of assistant devices that includes at least the given assistant device.

12. The method according to claim 7, wherein Identifying the given assistant device in response to determining that the given assistant device has been newly added to the ecosystem.

13. The method according to claim 12, wherein, Assigning the given semantic label to the given assistant device includes: Adding the given semantic label to a list of semantic labels associated with the given assistant device.

14. The method according to claim 12, wherein, Determining that the given assistant device has been newly added to the ecosystem includes: Identifying, based on one or more wireless signals, that the given assistant device has been added to a wireless network associated with the ecosystem.

15. The method according to claim 1, wherein The device activities associated with the given assistant device further include a plurality of queries or commands previously received at the given assistant device.

16. The method according to claim 15, wherein, Processing one or more of the device-specific signals to generate one or more of the candidate semantic labels for the given assistant device further includes: Using a semantic classifier to process the device activities associated with the given assistant device to classify each of the plurality of queries or commands into one or more of a plurality of different categories; and Further generating one or more candidate semantic labels among the candidate semantic labels based on one or more of the plurality of different categories into which each of the plurality of queries or commands is classified.

17. The method according to claim 16, wherein, Selecting the given semantic label for the given assistant device further includes: Further selecting the given semantic label for the given assistant device based on the number of the plurality of queries or commands classified into a given category among the plurality of different categories, the given semantic label being associated with the given assistant device.

18. A method implemented by one or more processors, the method comprising: Identifying a given assistant device among a plurality of assistant devices in an ecosystem; Obtaining one or more device-specific signals associated with the assistant device, the one or more device-specific signals being generated or received by the given assistant device, wherein one or more of the device-specific signals at least include device activities associated with the given assistant device, and wherein the device activities associated with the given assistant device include ambient noise previously detected at the given assistant device during a voice reception activity; Processing one or more of the device-specific signals to generate one or more candidate semantic labels for the given assistant device, wherein processing one or more of the device-specific signals to generate one or more of the candidate semantic labels for the given assistant device includes: Processing the ambient noise previously detected at the given assistant device during a voice reception activity to generate one or more of the candidate semantic labels for the given assistant device; Generating a prompt requesting a selection of a given semantic label from a user of a client device based on one or more of the candidate semantic labels, the selection being from one or more of the candidate semantic labels; Causing the prompt to be rendered at the client device of the user; and Assigning the given semantic label to the given assistant device in a device topology representation of the ecosystem in response to receiving the selection of the given semantic label in response to the prompt.

19. The method according to claim 18, wherein Processing the ambient noise previously detected at the given assistant device during a voice reception activity to generate one or more of the candidate semantic labels for the given assistant device includes: Using an ambient noise detection model to process the ambient noise previously detected at the given assistant device to classify the ambient noise into one or more of the plurality of different categories; and Generating one or more of the candidate semantic labels based on one or more of the plurality of different categories into which the ambient noise is classified; and Among them, selecting the given semantic tag for the given assistant device includes: selecting the given semantic tag for the given assistant device based on the environmental noise classified as a given category among the multiple different categories, where the given semantic tag is associated with the given semantic tag.

20. A method implemented by one or more processors, the method including: identifying a given assistant device among multiple assistant devices in an ecosystem; obtaining one or more device-specific signals associated with the given assistant device, the one or more device-specific signals being generated or received by the given assistant device, where one or more of the device-specific signals at least include device activities associated with the given assistant device, and where the device activities associated with the given assistant device include environmental noise previously detected at the given assistant device during a voice reception activity; determining a given semantic tag for the given assistant device based on one or more of the device-specific signals, where determining the given semantic tag for the given assistant device is at least based on the environmental noise previously detected at the given assistant device during a voice reception activity; assigning the given semantic tag to the given assistant device in a device topology representation of the ecosystem; and after assigning the given semantic tag to the given assistant device in the device topology representation of the ecosystem: receiving, via one or more corresponding microphones of one of the multiple assistant devices in the ecosystem, audio data corresponding to an oral utterance from a user associated with the ecosystem, the oral utterance including a query or a command; processing the audio data corresponding to the oral utterance to determine semantic attributes of the query or the command; determining that the semantic attributes of the query or the command match the given semantic tag assigned to the given assistant device; and in response to determining that the semantic attributes of the query or the command match the given semantic tag assigned to the given assistant device, causing the given assistant device to satisfy the query or the command.

21. At least one computing device, including: at least one processor; and a memory storing instructions that, when executed, cause the at least one processor to perform the method according to any one of claims 1 to 20.

22. A non-transitory computer-readable storage medium storing instructions that, when executed, cause at least one processor to perform the method according to any one of claims 1 to 20.

Citation Information

Patent Citations

  • Automatically determining language for speech recognition of spoken utterance received via an automated assistant interface

    CN110998717A

  • Providing custom names for headless devices

    US20150052231A1