Inferring semantic indicators for assistant devices based on device-specific signals

Semantic labeling of assistant devices using device-specific signals addresses the inefficiencies of manual labeling by automatically updating device topology, enhancing accuracy and reducing user input and resource consumption.

JP7818056B2Active Publication Date: 2026-02-19GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024164930
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-10-29
Filing Date
2024-09-24
Publication Date
2026-02-19
Estimated Expiration
2040-12-14

AI Technical Summary

Technical Problem

Existing techniques require users to manually assign or infer device labels for assistant devices, which can be inaccurate or require user intervention when devices are added or moved, leading to inefficiencies and resource waste.

Method used

Assign semantic labels to assistant devices based on device-specific signals, including queries, commands, ambient noise, and user preferences, automatically updating the device topology representation without user input.

Benefits of technology

Enables accurate and efficient device selection by the automated assistant, reducing user input and conserving computational and network resources, while maintaining a semantically meaningful device labeling system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007818056000001
    Figure 0007818056000001
  • Figure 0007818056000002
    Figure 0007818056000002
  • Figure 0007818056000003
    Figure 0007818056000003
Patent Text Reader

Abstract

To provide a method for inferring a semantic label for an assistant device based on a device-specific signal, and a storage medium.SOLUTION: Implementations can identify a given assistant device from among a plurality of assistant devices in an ecosystem, obtain device-specific signal(s) that are generated by the given assistant device, process the device-specific signal(s) to generate candidate semantic label(s), select a semantic label for the given semantic device from among the candidate semantic label(s), and assign, in a device topology representation of the ecosystem, the given semantic label to the given assistant device. Implementations can optionally receive a spoken utterance that includes a query or command at the assistant device(s), determine a semantic property of the query or command matches the given semantic label to the given assistant device, and assign the given semantic label to the given assistant device.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] A person can engage in human-to-computer interactions with interactive software applications referred to herein as “automated assistants” (also referred to as “chatbots,” “interactive personal assistants,” “intelligent personal assistants,” “personal voice assistants,” “conversational agents,” etc.). For example, a person (who may be referred to as a “user” when interacting with an automated assistant) may provide input (e.g., commands, queries, and / or requests) to the automated assistant, which may cause the automated assistant to generate and provide a responsive output, control one or more Internet of things (IoT) devices, and / or perform one or more other functions. The input provided by the user may be, for example, spoken natural language input (i.e., utterances) and / or typed natural language input, which may in some cases be converted to text (or other semantic representation) and then further processed.

[0002] In some cases, an automated assistant may include an automated assistant client that runs locally on an assistant device and with which a user interacts directly, as well as a cloud-based counterpart that leverages the virtually limitless resources of the cloud to help the automated assistant respond to user input. For example, the automated assistant may provide its cloud-based counterpart with an audio recording of the user's speech (or a text transcription thereof) and, optionally, data indicating the user's identity (e.g., a certificate). The cloud-based counterpart may perform various processing on the query and return results to the automated assistant client, which may then provide corresponding output to the user.

[0003] Many users may interact with an automated assistant using multiple assistant devices. For example, some users may have a coordinated “ecosystem” of assistant devices that can receive user input directed to the automated assistant and / or be controlled by the automated assistant, such as one or more smartphones, one or more tablet computers, one or more vehicle computing systems, one or more wearable computing devices, one or more smart televisions, one or more interactive standalone speakers, and / or one or more IoT devices, among other assistant devices. A user may use any of these assistant devices (assuming the automated assistant client is installed and the assistant device is capable of receiving input) to engage in human-computer interactions with the automated assistant. In some cases, these assistant devices may be scattered around a user's primary residence, secondary residence, workplace, and / or other structure. For example, mobile assistant devices such as smartphones, tablets, and smartwatches may be worn by the user and / or may be where the user last left them. Other assistant devices, such as traditional desktop computers, smart televisions, interactive standalone speakers, and IoT devices, may be more stationary but may still be located in various places (e.g., rooms) in the user's home or workplace. Summary of the Invention [Problem to be solved by the invention]

[0004] Techniques exist for enabling a user (e.g., a single user, multiple users in a household, coworkers, housemates, etc.) to manually assign labels to assistant devices in an ecosystem of assistant devices and subsequently interact with or control any one of the assistant devices using an automated assistant client on any one of the assistant devices. For example, a user can issue a spoken command to the automated assistant client on an assistant device, such as "Show me chili recipes on my kitchen device," to cause the assistant device (or another assistant device in the ecosystem) to retrieve search results for chili recipes and present the search results to the user via the kitchen device. However, such techniques require the user to specify a particular assistant device by a previously assigned label (e.g., "kitchen device") that the user may have forgotten, or require the automated assistant to infer the "best" device (e.g., the device closest to the user) to provide search results. Furthermore, if a particular assistant device is new to the ecosystem or moves within the ecosystem, the label assigned to a particular assistant device by the user may not represent the particular assistant device. [Means for solving the problem]

[0005] Implementations described herein relate to assigning semantic labels to each assistant device in a device topology representation of an ecosystem including multiple assistant devices. The semantic label assigned to each assistant device can be inferred based on one or more device-specific signals associated with the respective assistant device. The one or more device-specific signals can include, for example, one or more queries (if any) previously received at the respective assistant device, one or more commands (if any) previously executed at the respective assistant device, instances of ambient noise previously detected at the respective assistant device (and optionally only when speech reception was active at the respective assistant device), unique identifiers (or labels) of any other assistant devices positionally proximate to the respective assistant device, and / or user preferences of a user associated with the ecosystem determined based on user interactions with multiple assistant devices in the ecosystem. Each of the one or more device-specific signals associated with each assistant device can be processed to classify each of them into one or more semantic categories of a plurality of heterogeneous semantic categories. One or more candidate semantic labels can be generated for each assistant device based on the semantic category into which one or more of the device-specific signals are classified. Further, a given semantic indicator from among one or more candidate semantic indicators and for a given one of the respective assistant devices may be selected and assigned to the given one of the respective assistant devices in the device topology representation of the ecosystem.

[0006] For example, assume that a given assistant device is a two-way standalone speaker device with a display located at a primary residence of a user associated with the ecosystem. Further assume that multiple queries related to searching for food recipes have been received and executed at the given assistant device and / or multiple commands related to setting timers have been received and executed at the given assistant device, that instances of ambient noise have been detected at the client device, that a unique identifier (or indicator) associated with an additional assistant device in the ecosystem corresponding to a "smart oven" is detected at the given assistant device, and that user preferences of a user associated with the ecosystem indicate that the user prefers a fictional chef named Johnny Flay. In this example, assume that a query related to searching for food recipes is categorized into the “recipe,” “kitchen,” and / or “kitchen” categories, and that a command related to setting a timer is categorized into the “timing” and / or “cooking” categories; further assume that instances of ambient noise are categorized into the “kitchen” and / or “cooking” categories based on ambient noise capturing cooking sounds (e.g., food frying in a frying pan, a knife cutting food, a microwave in use, etc.); further assume that a unique identifier (or indicator) of a “smart oven” associated with an additional assistant device is categorized into the “kitchen” and / or “cooking” categories; and further assume that a fictional chef is categorized into the “cooking” category (or the more specific category of “Johnny Flay”). As a result, candidate semantic indicators of “recipe display device,” “kitchen display device,” “cooking display device,” “timing display device,” and “Johnny Flay device” may be generated for the interactive standalone speaker with a display.Furthermore, from among the candidate semantic labels, a given semantic label may be assigned to a two-way standalone speaker device having a display in a device topology representation of the ecosystem for the user's main residence.

[0007] In some implementations, a given semantic label may be automatically assigned to a given assistant device in the device topology representation of the ecosystem. For example, if a reliability level associated with the given semantic label meets a threshold reliability level, the given semantic label may be automatically assigned to a given assistant device in the device topology representation of the ecosystem. The reliability level associated with the given semantic label may be determined while processing one or more device-specific signals associated with the given assistant device. For example, the reliability level associated with the given assistant device may be based on the amount of one or more device-specific signals that fall into one or more semantic categories. For example, if nine queries related to searching for food recipes have been received at a given assistant device and only one query related to searching for weather information has been received at a given assistant device, the semantic label "cooking display device" or "recipe display device" may be automatically assigned to the given assistant device in the device topology representation of the ecosystem (even if the given assistant device is not located in the user's kitchen). For example, the confidence level associated with a given assistant device may be based on a measure determined based on output generated using a semantic classifier and / or an ambient noise detection model to process one or more device-specific signals. For example, previously received queries or commands (or their corresponding text) may be processed using a semantic classifier to classify each of the queries or commands into one or more of the semantic categories, the instances of ambient noise may be processed using an ambient noise detection model to classify each of the instances of ambient noise into one or more of the semantic categories, and the unique identifiers (or indicators) may be processed using a semantic classifier to classify each of the unique identifiers (or indicators) into one or more of the semantic categories along with a respective measure.As another example, if a given semantic indicator is unique (with respect to other assistant devices that are positionally close to the given assistant device in the ecosystem), the given semantic indicator can be automatically assigned to the given assistant device in the device topology representation of the ecosystem.

[0008] In some additional or alternative implementations, in response to receiving a user input to assign a given semantic label to a given assistant device, the given semantic label may be assigned to the given assistant device in the device topology representation of the ecosystem. For example, a prompt may be generated to prompt a user associated with the ecosystem to select a given semantic label from one or more candidate semantic labels. The prompt may be rendered on the user's client device (e.g., the user's given assistant device or another client device (e.g., a mobile phone)), and in response to receiving the selection of the given semantic label, the given semantic label may be assigned to the given assistant device. For example, assume that one or more candidate semantic labels include a "cooking display device," a "recipe display device," and a "weather display device." In this case, the prompt may include each of the candidate semantic labels and request the user to select a given semantic label from the candidate semantic labels to be assigned to the given assistant device (and optionally replace an existing semantic label). While the above examples are described with reference to a single semantic indicator being assigned to a given assistant device, it should be understood that this is for illustrative purposes and is not intended to be limiting. For example, the assistant devices described herein may be assigned multiple semantic indicators, such that each assistant device is stored in association with a list of semantic indicators.

[0009] In various implementations, after one or more semantic labels are assigned to each assistant device in the device topology representation of the ecosystem, the semantic labels assigned to the assistant devices according to the techniques described herein may also be used to process utterances received at one or more assistant devices in the ecosystem. For example, audio data corresponding to the utterance may be processed to identify semantic features included in the utterance. Furthermore, an embedding (e.g., a word2vec representation) of the identified semantic features may be generated and compared with multiple embeddings (e.g., each word2vec representation) of each semantic label assigned to the assistant device in the ecosystem. Furthermore, based on the comparison, it may be determined that the semantic features match a given embedding of the multiple embeddings of each semantic label. For example, assume that the embedding is a word2vec representation. In this example, the cosine distance between the word2vec representation of the semantic feature and each of the word2vec representations of the respective semantic indicators can be determined, and the given semantic indicators associated with each cosine distance that meets the threshold distance can be used to determine whether the semantic feature of the utterance matches the given semantic indicator (e.g., an exact match or a rough match). As a result, a given assistant device associated with the given semantic indicator can be selected to fulfill the utterance. Additionally or alternatively, the user's proximity to the given assistant device and / or the device capabilities of the given assistant device can be considered when selecting a given assistant device to fulfill the utterance.

[0010] By inferring and assigning semantic labels to assistant devices in an ecosystem using the techniques described herein, the device topology representation of the ecosystem can be kept up to date without requiring diverse user interface inputs (or any user interface inputs) to do so. Furthermore, the semantic labels assigned to assistant devices are semantically useful to users in that the semantic labels assigned to each assistant device are selected based on the use of each assistant device and / or the respective portion of the ecosystem in which each assistant device is located. Thus, when an utterance is received at one or more of the assistant devices in the ecosystem, the automated assistant can more robustly and / or accurately select one or more of the assistant devices that are best suited to satisfy the utterance. As a result, because a user associated with the ecosystem does not need to specify a particular device to satisfy the utterance or repeat the utterance when the wrong device is selected to satisfy the utterance, the amount and / or length of user inputs received by one or more of the assistant devices in the ecosystem can be reduced, thereby conserving computational and / or network resources at the assistant device by reducing network traffic. Additionally, when an assistant device is newly added to the ecosystem, moved within the ecosystem, or located in a repurposed portion of the ecosystem (e.g., a room in a user's main residence that is converted from a study to a bedroom), the user does not need to manually update the device topology representation via a software application associated with the ecosystem, thereby reducing the amount of user input received by one or more of the assistant devices in the ecosystem.

[0011] The above description is provided as a summary of only some implementations of the present disclosure. Further description of these and other implementations is set forth in more detail herein. As one non-limiting example, various implementations are set forth in more detail in the claims contained herein.

[0012] Additionally, some implementations include one or more processors of one or more computing devices, the one or more processors operable to execute instructions stored in associated memory, the instructions configured to cause performance of any of the methods described herein. Some implementations also include one or more non-transitory computer-readable storage media that store computer instructions executable by the one or more processors to perform any of the methods described herein.

[0013] It is understood that all combinations of the foregoing concepts, and additional concepts more fully described herein, are considered to be part of the subject matter disclosed herein, for example, all combinations of claimed subject matter appearing at the end of this disclosure are considered to be part of the subject matter disclosed herein. [Brief explanation of the drawings]

[0014] [Figure 1] FIG. 1 is a block diagram of an example environment in which implementations disclosed herein may be implemented. [Figure 2A] FIG. 10 illustrates examples related to assigning a given semantic indicator to a given assistant device that is newly added to an ecosystem of assistant devices and / or moves within an ecosystem of assistant devices, according to various implementations. [Figure 2B]FIG. 10 illustrates examples related to assigning a given semantic indicator to a given assistant device that is newly added to an ecosystem of assistant devices and / or moves within an ecosystem of assistant devices, according to various implementations. [Figure 3] 1 is a flowchart illustrating an example method for assigning a given semantic indicator to a given assistant device in an ecosystem, according to various implementations. [Figure 4] 10 is a flowchart illustrating an example method for using assigned semantic indicators in satisfying a query or command received at an assistant device in an ecosystem, according to various implementations. [Figure 5] FIG. 1 illustrates an exemplary architecture of a computing device, according to various implementations. DETAILED DESCRIPTION OF THE INVENTION

[0015] There has been a proliferation of smart, multi-sensing, network-connected devices (also referred to herein as assistant devices), such as smartphones, tablet computers, vehicle computing systems, wearable computing devices, smart televisions, interactive standalone speakers (e.g., with or without displays), sound speakers, home alarms, door locks, cameras, lighting systems, treadmills, thermostats, scales, smart beds, watering systems, garage door openers, home appliances, baby monitors, fire alarms, moisture meters, etc. Often, multiple assistant devices are located within the confines of a structure such as a home, or within various associated structures, such as a user's primary residence and a user's secondary residence, a user's vehicle, and / or a user's workplace.

[0016] Additionally, there is a proliferation of assistant devices (also referred to herein as assistant input devices) that each include an automated assistant client that can form a logical instance of an automated assistant. These assistant input devices may be dedicated to assistant functionality (e.g., an interactive standalone speaker and / or standalone audio / visual device that includes only an assistant client and associated interface and is dedicated to assistant functionality) or may perform assistant functionality in addition to other functions (e.g., a mobile phone or tablet that includes an assistant client as one of its various applications). Moreover, some IoT devices may also be assistant input devices. For example, some IoT devices may include an automated assistant client and at least a speaker and / or microphone that serve (at least in part) as a user interface output and / or input device for the automated assistant client's assistant interface. Some assistant devices may not implement an automated assistant client or may not have means for interfacing with a user (e.g., a speaker and / or microphone), but they may still be controlled by an automated assistant (also referred to herein as assistant non-input devices). For example, a smart light bulb may not include an automation assistant client, a speaker, and / or a microphone, but commands and / or requests can be sent to the smart light bulb via the automation assistant to control the functionality of the smart lighting (e.g., turn lights on / off, dim, change color, etc.).

[0017] Various techniques have been proposed for labeling and / or grouping assistant devices (including both assistant-input devices and assistant-non-input devices) in an ecosystem of assistant devices. For example, when a new assistant device is added to the ecosystem, a user associated with the ecosystem can manually assign a label (or a unique identifier) ​​to the new assistant device in a device topology representation of the ecosystem and / or manually add the new assistant device to a group of assistant devices in the ecosystem via a software application (e.g., via an automated assistant application, a software application associated with the ecosystem, a software application associated with the new assistant device, etc.). As described herein, the label originally assigned to the assistant device may be forgotten by the user or may not be semantically meaningful with respect to how the assistant device is utilized or where the assistant device is located in the ecosystem. Furthermore, when an assistant device moves within the ecosystem, the user may be required to manually change the label assigned to the assistant device and / or manually change the group to which the assistant device is assigned via a software application. Otherwise, the label assigned to the assistant device and / or the group to which the assistant device is assigned may not accurately reflect the location or use of the assistant device and / or may not be semantically meaningful for the assistant device. For example, if a smart speaker labeled "living room speaker" is located in the living room of a user's primary residence, but the smart speaker is moved to the kitchen of the user's primary residence, the smart speaker may still be labeled as "living room speaker" unless the user manually changes the label in the device topology representation for the user's primary residence ecosystem, even though the label does not represent the location of the assistant device.

[0018] The device topology representation may include a label (or unique identifier) ​​associated with each assistant device. Additionally, the device topology representation can specify a label (or unique identifier) ​​associated with each assistant device. The device attributes for a given assistant device can, for example, indicate one or more input and / or output modalities supported by the respective assistant device. For example, a device attribute for a standalone speaker-only assistant client device can indicate that it is capable of providing auditory output but not visual output. The device attributes for a given assistant device can additionally or alternatively, for example, identify one or more controllable states of the given assistant device, identify a party (e.g., a first party (1P) or a third party (3P)) that manufactures, distributes, and / or creates firmware for the assistant device, and / or identify a unique identifier for the given assistant device, such as an immutable identifier provided by a 1P or 3P, or a label assigned to the given assistant device by a user. According to various implementations disclosed herein, the device topology representation can optionally further specify which smart devices can be locally controlled by which assistant devices, local addresses of the locally controllable assistant devices (or local addresses of hubs that can directly and locally control those assistant devices), local signal strengths, and / or other priority indicators between the respective assistant devices. Furthermore, according to various implementations disclosed herein, the device topology representation (or a variant thereof) can be stored locally on each of the multiple assistant devices for use in locally controlling and / or locally assigning labels to the assistant devices. Furthermore, the device topology representation can specify groups associated with each assistant device, which can be defined at various levels of granularity.For example, multiple smart lights in the living room of a user's primary residence may be considered to belong to the "living room lights" group. Additionally, if the living room of the primary residence also includes a smart speaker, all of the assistant devices located in the living room may be considered to belong to the "living room assistant devices" group.

[0019] The automation assistant can detect various events occurring in the ecosystem based on one or more signals generated by one or more of the assistant devices. For example, the automation assistant can process one or more of the signals to detect these events using an event detection model or an event detection rule. Furthermore, the automation assistant can cause one or more actions to be performed based on an output generated based on one or more of the signals for an event occurring in the ecosystem. In some implementations, the detected event can be a device-related event associated with one or more of the assistant devices (e.g., assistant input devices and / or assistant non-input devices). For example, a given one of the assistant devices can detect when it is newly added to the ecosystem based on one or more wireless signals generated by the given one of the assistant devices (and optionally a unique identifier associated with the given one of the assistant devices included in one or more of the wireless signals). As another example, a given one of the assistant devices can detect when it moves within the ecosystem based on being surrounded by one or more different assistant devices that previously surrounded it (and optionally determined based on the unique identifiers of each of the one or more different assistant devices). In these implementations, one or more of the actions performed by the automated assistant can include, for example, determining a semantic indicator for the given one of the assistant devices in response to determining that it is newly introduced to the ecosystem or has moved within the ecosystem, and causing the semantic indicator to be assigned to the given one of the assistant devices in a device topology representation of the ecosystem.

[0020] In some additional or alternative implementations, the detected event may be an acoustic event captured through a respective microphone of one or more assistant devices. The automated assistant may have audio data capturing the acoustic event processed using an acoustic event model. The acoustic event detected by the acoustic event model may include, for example, using a hotword detection model to detect a hotword included in an utterance that invokes the automated assistant, using an ambient noise detection model to detect ambient noise in the ecosystem (and optionally while speech acceptance is active on a given one of the assistant devices), using a sound detection model to detect specific sounds in the ecosystem (e.g., glass breaking, a dog barking, a cat meowing, a doorbell ringing, a fire alarm going off, or a carbon monoxide detector going off), and / or other acoustic-related events that may be detected using the respective acoustic event detection model. For example, assume that audio data is detected through at least one respective microphone of the assistant device. In this example, the automated assistant may have audio data processed by at least one hotword detection model on the assistant device to determine whether the audio data captures a hotword for invoking the automated assistant. Additionally, the automated assistant may additionally or alternatively cause the audio data to be processed by at least one ambient noise detection model of the assistant device to classify any ambient (or background) noise captured in the audio data into one or more disparate semantic categories of ambient noise (e.g., movie or television sounds, cooking sounds, and / or other disparate sound categories). Moreover, the automated assistant may additionally or alternatively cause the audio data to be processed by at least one sound detection model of the assistant device to determine whether any particular sound is captured in the audio data.

[0021] Implementations described herein relate to inferring semantic indicators for assistant devices based on one or more signals generated by each of the respective devices. The implementations further relate to assigning semantic indicators to assistant devices in a device topology representation of an ecosystem. The semantic indicators may be automatically assigned to assistant devices or may be presented to a user associated with the ecosystem to prompt for selection of one or more semantic indicators to be assigned to the assistant devices. The implementations further relate to subsequently using the semantic indicators when processing an utterance to determine whether the utterance includes a word or phrase that matches any of the semantic indicators, and, when the utterance is determined to include a word or phrase that matches one of the semantic indicators, using an assistant device associated with the matching one of the semantic indicators to satisfy the utterance.

[0022] 1, an example environment in which the techniques disclosed herein may be implemented is shown. The example environment includes multiple assistant input devices 106 1-N (also referred to herein simply as “assistant input device(s) 106”), one or more cloud-based automated assistant components 119, one or more assistant non-input systems 180, one or more assistant non-input devices 185 1-N (also referred to herein simply as "Assistant non-input devices 185"), device activity database 191, machine learning ("ML") model database, and device topology database 193. Assistant input devices 106 and Assistant non-input devices 185 of FIG. 1 may also be collectively referred to herein as "Assistant devices."

[0023] One or more (e.g., all) of the assistant input devices 106 may be connected to a respective automated assistant client 118 1-NHowever, in some implementations, one or more of the assistant input devices 106 may optionally run a respective automated assistant client 118. 1-N The automated assistant client 118 may lack an instance of the automated assistant client 118 and still include engines and hardware components (e.g., microphone, speaker, speech recognition engine, natural language processing engine, speech synthesis engine, etc.) for receiving and processing user input directed to the automated assistant. 1-N An instance of may be an application separate from (e.g., installed "on top of") the operating system of each assistant input device 106, or alternatively may be implemented directly by the operating system of each assistant input device 106. As described further below, automated assistant client 118 1-N Each instance of optionally 1-N In responding to various requests provided by the cloud-based automated assistant component 119, the cloud-based automated assistant component 119 may interact with one or more of the cloud-based automated assistant components 119. Additionally, as also described below, other engines of the assistant input device 106 may optionally interact with one or more of the cloud-based automated assistant components 119.

[0024] One or more cloud-based automation assistant components 119 may be implemented on one or more computing systems (e.g., servers collectively referred to as "cloud" or "remote" computing systems) that are communicatively coupled to respective assistant input devices 106 via one or more local area networks ("LANs" including Wi-Fi LANs, Bluetooth networks, near field communications networks, mesh networks, etc.) and / or wide area networks ("WANs" including the Internet, etc.). The communicative coupling of cloud-based automation assistant components 119 with assistant input devices 106 is indicated generally by 1101 in FIG. 1 . Also, in some embodiments, assistant input devices 106 may be communicatively coupled to each other via one or more networks (e.g., LANs and / or WANs), indicated generally by 1102 in FIG. 1 .

[0025] One or more cloud-based automation assistant components 119 may also be communicatively coupled to one or more assistant non-input systems 180 via one or more networks (e.g., LAN and / or WAN). The communicative coupling of the cloud-based automation assistant components 119 with the assistant non-input systems 180 is generally indicated by 1103 in FIG. 1 . Furthermore, the assistant non-input systems 180 may each be communicatively coupled to one or more (e.g., groups) of assistant non-input devices 185 via one or more networks (e.g., LAN and / or WAN). For example, a first assistant non-input system 180 may be communicatively coupled to and receive data from one or more first groups of assistant non-input devices 185, a second assistant non-input system 180 may be communicatively coupled to and receive data from one or more second groups of assistant non-input devices 185, and so on. The communicative coupling of the assistant non-input systems 180 with the assistant non-input devices 185 is generally indicated by 1104 in FIG. 1 .

[0026] An instance of an automated assistant client 118, through its interaction with one or more cloud-based automated assistant components 119, may form what appears to a user to be a logical instance of an automated assistant 120 through which the user may engage in human-to-computer interactions. Two instances of such an automated assistant are shown in FIG. 1. A first automated assistant 120A, surrounded by dashed lines, includes an automated assistant client 1181 of an assistant input device 1061 and one or more cloud-based automated assistant components 119. A second automated assistant 120B, surrounded by dash-dash-dot lines, includes an automated assistant client 1181 of an assistant input device 1061 and one or more cloud-based automated assistant components 119. N Automation Assistant Client 118 Nand one or more cloud-based automated assistant components 119. Thus, it should be understood that each user interacting with an automation assistant client 118 executing on one or more of the assistant input devices 106 can effectively interact with a user-specific logical instance of the automation assistant 120 (or a logical instance of the automation assistant 120 shared among a household or other group of users). For brevity and simplicity, the term "automation assistant" as used herein refers to the combination of an automation assistant client 118 executing on each one of the assistant input devices 106 and one or more cloud-based automated assistant components 119 (which may be shared among various automation assistant clients 118). While only multiple assistant input devices 106 are shown in FIG. 1 , it is understood that the cloud-based automated assistant component 119 can additionally serve many additional groups of assistant input devices.

[0027] Assistant input devices 106 may include, for example, one or more of a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a computing device in a user's vehicle (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), an interactive standalone speaker (e.g., with or without a display), a smart appliance such as a smart television, a user's wearable device including a computing device (e.g., a user's watch with a computing device, a user's eyeglasses with a computing device, a virtual reality or augmented reality computing device), and / or any IoT device capable of receiving user input directed to automation assistant 120. Additional and / or alternative assistant input devices may be provided. Assistant non-input devices 185 may include many of the same devices as assistant input devices 106, but are not capable of receiving user input directed to automation assistant 120 (e.g., do not include a user interface input component). Although assistant non-input devices 185 do not receive user input directed to automation assistant 120, assistant non-input devices 185 may still be controlled by automation assistant 120.

[0028] In some implementations, multiple assistant input devices 106 and assistant non-input devices 185 can be associated with each other in various ways to facilitate the execution of the techniques described herein. For example, in some implementations, multiple assistant input devices 106 and assistant non-input devices 185 can be associated with each other by being communicatively coupled via one or more networks (e.g., via network 110 of FIG. 1 ). This can be the case, for example, when multiple assistant input devices 106 and assistant non-input devices 185 are deployed throughout a particular area or environment, such as a home, a building, etc. Additionally or alternatively, in some implementations, multiple assistant input devices 106 and assistant non-input devices 185 can be associated with each other by being members of a collaborative ecosystem that is at least selectively accessible by one or more users (e.g., an individual, a family member, employees of an organization, other predetermined groups, etc.). In some of these implementations, the ecosystem of multiple assistant input devices 106 and assistant non-input devices 185 can be manually and / or automatically associated with each other in a device topology representation of the ecosystem stored in device topology database 193.

[0029] Assistant-non-input systems 180 may include one or more first-party (1P) systems and / or one or more third-party (3P) systems. A 1P system refers to a system controlled by a party that is the same as the party that controls the automated assistant 120 referred to herein. A 3P system, as used herein, refers to a system controlled by a party other than the party that controls the automated assistant 120 referred to herein.

[0030] The Assistant non-input system 180 can receive data from the Assistant non-input device 185 and / or one or more cloud-based automation assistant components 119 communicatively coupled thereto (e.g., via the network 110 of FIG. 1 ) and selectively transmit data (e.g., status, state changes, and / or other data) to the Assistant non-input device 185 and / or one or more cloud-based automation assistant components 119. For example, assume that the Assistant non-input device 1851 is a smart doorbell IoT device. In response to an individual pressing a button on the doorbell IoT device, the doorbell IoT device can transmit data corresponding to one of the Assistant non-input systems 180 (e.g., one of the Assistant non-input systems managed by the doorbell manufacturer, which may be a 1P system or a 3P system). One of the Assistant non-input systems 180 can determine a change in the state of the doorbell IoT device based on such data. For example, one of the assistant non-input systems 180 can determine a doorbell change from an inactive state (e.g., no recent button press) to an active state (recent button press), and the change in doorbell state can be transmitted (e.g., via network 110 of FIG. 1 ) to one or more cloud-based automation assistant components 119 and / or one or more of the assistant input devices 106. In particular, user input is received at assistant non-input device 1851 (e.g., a button press on the doorbell), but the user input is not directed to the automation assistant 120 (hence the term “assistant non-input device”). As another example, assume that assistant non-input device 1851 is a smart thermostat IoT device with a microphone, but the smart thermostat does not include the automation assistant client 118. An individual can operate the smart thermostat (e.g., using touch input or spoken input) to change the temperature, set a particular value as a setpoint for controlling an HVAC system via the smart thermostat, etc.However, unless the smart thermostat includes an automation assistant client 118, an individual cannot directly communicate with the automation assistant 120 via the smart thermostat.

[0031] In various implementations, one or more cloud-based automation assistant components 119 may further include various engines. For example, as shown in FIG. 1 , one or more cloud-based automation assistant components 119 may further include an event detection engine 130, a device identification engine 140, an event processing engine 150, a semantic signature engine 160, and a query / command processing engine 170. It should be understood that while these various engines are shown as one or more cloud-based automation assistant components 119 in FIG. 1 for purposes of illustration and not limitation, they are not intended to be limiting. For example, assistant input device 106 and / or assistant non-input device 185 may include one or more of these various engines. As another example, these various engines may be distributed across assistant input device 106, and assistant non-input device 185 may include one or more of these various engines and / or one or more cloud-based automation assistant components 119.

[0032] In some implementations, event detection engine 130 can detect various events occurring in the ecosystem. In some versions of these implementations, event detection engine 130 can detect when a given one of assistant input devices 106 and / or a given one of assistant non-input devices 185 (e.g., a given one of assistant devices) is newly added to the ecosystem or moves within the ecosystem. For example, event detection engine 130 can determine when a given one of assistant devices is newly added to the ecosystem based on one or more wireless signals detected over network 110 and via device identification engine 140. For example, when a given one of assistant devices is newly connected to one or more of networks 110, the given one of assistant devices can broadcast a signal indicating that it is newly added to network 110. As another example, event detection engine 130 can determine when a given one of assistant devices has moved within the ecosystem based on one or more wireless signals detected over network 110. In these examples, device identification engine 140 can process the signals to determine that a given one of the assistant devices is newly added to network 110 and / or that a given one of the assistant devices has moved within the ecosystem. The one or more wireless signals detected by device identification engine 140 can be, for example, network signals and / or acoustic signals that are imperceptible to humans and optionally include a unique identifier for the given one of the assistant devices and / or each of the other assistant devices that are positionally close to the given one of the assistant devices. For example, when a given one of the assistant devices moves within the ecosystem, device identification engine 140 can detect one or more wireless signals being transmitted by other assistant devices that are positionally close to the given one of the assistant devices.These signals can be processed to determine that one or more other assistant devices that are positionally close to a given one of the assistant devices are different from one or more assistant devices that were previously positionally close to the given one of the assistant devices.

[0033] In some further versions of these implementations, automation assistant 120 can cause a given one of the assistant devices that is newly added to the ecosystem or moved within the ecosystem to be assigned to a group of assistant devices (e.g., in the device topology representation of the ecosystem stored in device topology database 193). For example, in an implementation in which a given one of the assistant devices is newly added to the ecosystem, the given one of the assistant devices may be added to an existing group of assistant devices, or a new group of assistant devices including the given one of the assistant devices may be created. For example, if the given one of the assistant devices is positionally close to multiple assistant devices that belong to a “kitchen” group (e.g., a smart oven, a smart coffee maker, a two-way standalone speaker associated with a unique identifier or sign indicating that it is located in the kitchen, and / or other assistant devices), the given one of the assistant devices may be added to the “kitchen” group, or a new group may be created. As another example, in an implementation in which a given one of the assistant devices moves within the ecosystem, the given one of the assistant devices may be added to an existing group of assistant devices, or a new group of assistant devices including the given one of the assistant devices may be created. For example, if a given one of the assistant devices was geographically close to multiple assistant devices belonging to the aforementioned "kitchen" group, but is now geographically close to multiple assistant devices (e.g., smart garage doors, smart door locks, and / or other assistant devices) belonging to the "garage" group, the given one of the assistant devices may be removed from the "kitchen" group and added to the "garage" group.

[0034] In some additional or alternative versions of these implementations, event detection engine 130 can detect the occurrence of an acoustic event. The occurrence of an acoustic event can be detected based on audio data received at one or more of assistant input devices 106 and / or one or more of assistant non-input devices 185 (e.g., one or more of the assistant devices). The audio data received at one or more of the assistant devices can be processed by an event detection model stored in ML model database 192. In these implementations, each of the one or more assistant devices that detect the occurrence of an acoustic event includes a respective microphone.

[0035] In some further versions of these implementations, the occurrence of an acoustic event may include ambient noise captured in audio data at one or more of the assistant devices (and optionally may include only occurrences of ambient noise detected when speech acceptance is active at one or more of the assistant devices). The ambient noise detected at each of the one or more assistant devices may be stored in device activity database 191. In these implementations, event processing engine 150 can process the ambient noise detected at one or more assistant devices using an ambient noise detection model (e.g., stored in ML model database 192) that is trained to classify the ambient noise into one or more of a plurality of heterogeneous semantic categories based on a measure generated when processing the ambient noise using the ambient noise detection model. The plurality of heterogeneous categories may include, for example, a movie or television sound category, a cooking sound category, a music sound category, a garage or workshop sound category, a patio sound category, and / or other semantically meaningful heterogeneous sound categories. For example, if event processing engine 150 determines that the ambient noise processed using the ambient noise detection model includes sounds corresponding to a microwave humming, food frying in a frying pan, a food processor processing food, etc., event processing engine 150 can classify the ambient noise into a cooking sounds category. As another example, if event processing engine 150 determines that the ambient noise processed using the ambient noise detection model includes sounds corresponding to a circular saw operating, hammering, etc., event processing engine 150 can classify the ambient noise into a garage or workshop category. The classification of the ambient noise detected at a particular device can also be used as a device-specific signal to be utilized in inferring a semantic indicator (e.g., as described with respect to semantic indicator engine 160) for the assistant device.

[0036] In some additional or alternative versions of these further implementations, the occurrence of the acoustic event may include a hot word or a particular sound detected on one or more of the assistant devices. In these implementations, event processing engine 150 can process audio data detected on one or more assistant devices using a hot word detection model trained to determine whether the audio data includes a particular word or phrase that invokes automated assistant 120 based on a measure generated when processing the audio data using the hot word detection model. For example, event processing engine 150 can process the audio data to determine whether the audio data captures a user utterance that includes "assistant," "hey assistant," "OK assistant," and / or any other word or phrase that invokes an automated assistant. Furthermore, the measure generated using the hot word detection model may include a respective confidence level or probability indicating whether the audio data includes a word or phrase that invokes automated assistant 120. In some versions of these implementations, event processing engine 150 can determine that the audio data captures a word or phrase if the measure meets a threshold. For example, if the event processing engine 150 generates a measure of 0.70 associated with audio data that captures a word or phrase that invokes the automated assistant 120, and the threshold value is 0.65, the event processing engine 150 may determine that the audio data captures a word or phrase that invokes the automated assistant 120.

[0037] In these implementations, event processing engine 150 can additionally or alternatively process audio data detected at one or more assistant devices using a sound detection model trained to determine whether the audio data contains a specific sound based on a measure generated when processing the audio data using the sound detection model. The specific sounds may include, for example, glass breaking, a dog barking or a cat meowing, a doorbell ringing, a fire alarm going off, or a carbon monoxide detector going off. For example, event processing engine 150 can process the audio data to determine whether the audio data captures any of these specific sounds. In this example, a single sound detection model may be trained to determine whether multiple specific sounds are captured in the audio data, or multiple sound detection models may be trained to determine whether a given specific sound is captured in the audio data. Furthermore, the measure generated using the sound detection model may include a respective confidence level or probability indicating whether the audio data contains a specific sound. In some versions of these implementations, event processing engine 150 can determine that the audio data captures a specific sound if the measure meets a threshold. For example, if the event processing engine 150 generates a measure of 0.70 associated with audio data capturing the sound of glass breaking, and the threshold value is 0.65, the event processing engine 150 may determine that the audio data captures the sound of glass breaking.

[0038] In various implementations, the occurrence of an acoustic event may be detected by various assistant devices in the ecosystem. For example, various assistant devices in an environment may capture time-corresponding audio data (e.g., time-corresponding in that the respective audio data is detected at the various assistant devices at the same time or within a threshold length of time). In these implementations, in response to a given assistant device detecting audio data in the ecosystem, device identification engine 140 can identify one or more additional assistant devices that would have similarly detected time-corresponding audio data capturing the acoustic event. For example, device identification engine 140 can identify one or more additional assistant devices that would have similarly detected time-corresponding audio data capturing the acoustic event based on the one or more additional assistant devices having previously detected time-corresponding audio data capturing the acoustic event. In other words, because a given assistant device and one or more additional assistant devices have previously captured time-corresponding audio data containing the same sound, device identification engine 140 can predict that one or more additional assistant devices will similarly capture audio data containing the acoustic event.

[0039] In various implementations, one or more device-specific signals generated or detected by each assistant device can be stored in device activity database 191. In some implementations, device activity database 191 can correspond to a portion of memory dedicated to device activity for that particular assistant device. In some additional or alternative implementations, device activity database 191 can correspond to memory of a remote system in communication with the assistant device (e.g., via network 110 of FIG. 1 ). This device activity can be utilized in generating candidate semantic indicators for a given one of the assistant devices (e.g., as described with respect to semantic indicator engine 160). Device activity can include, for example, queries or requests received at each assistant device (and / or semantic categories associated with each of the multiple queries or requests), commands executed at each assistant device (and / or semantic categories associated with each of the multiple commands), ambient noise detected at each assistant device (and / or semantic categories associated with various instances of ambient noise), unique identifiers or indicators (e.g., identified via event detection engine 140) of any assistant devices positionally proximate to a given assistant device, user preferences of a user associated with the ecosystem determined based on user interactions with multiple assistant devices in the ecosystem (e.g., browsing history, search history, purchase history, music history, movie or TV history, and / or any other user interactions associated with multiple assistant devices), and / or any other data received, generated, and / or executed by each assistant device.

[0040] In some implementations, semantic label engine 160 can process one or more device-specific signals and generate a semantic label candidate for a given one of the assistant devices (e.g., a given one of assistant input devices 106 and / or a given one of assistant non-input devices 185) based on the one or more device-specific signals. In some versions of these implementations, a given assistant device for which a semantic label candidate is generated can be identified in response to determining that a given assistant device is newly added to the ecosystem and / or moves within the ecosystem. In some additional or alternative versions of these implementations, a given assistant device for which a semantic label candidate is generated can be identified periodically (e.g., once a month, once every six months, once a year, etc.). In some additional or alternative versions of these implementations, a given assistant device for which a candidate semantic label is generated can be identified in response to determining that a portion of the ecosystem in which the given assistant device is located has been repurposed (e.g., a room in the ecosystem's main residence has been repurposed from a study to a bedroom). In these implementations, the given assistant device can be identified utilizing event detection engine 130. Identifying a given assistant device in these and other ways is described with respect to FIGS. 2A and 2B.

[0041] In some implementations, the semantic indicator engine 160 can select a given semantic indicator for a given assistant device from among the candidate semantic indicators based on one or more device-specific signals. Generating candidate semantic indicators for a given assistant device and selecting a given semantic indicator from among the candidate semantic indicators based on one or more device-specific indicators is described below (e.g., in connection with FIGS. 2A and 2B).

[0042] In implementations in which candidate semantic labels for a given assistant device are generated based on queries, requests, and / or commands stored in device activity database 191 (or corresponding text), the queries, requests, and / or commands can be processed using a semantic classifier (e.g., stored in ML model database 192) to index device activity for the given assistant device into one or more different semantic categories corresponding to disparate queries, requests, and / or commands. The candidate semantic labels can be generated based on the semantic categories into which the queries, commands, and / or requests are classified, and a given semantic label selected for a given assistant device can be selected based on the amount of queries, requests, and / or commands classified in a given semantic category. For example, assume that a given assistant device has previously received nine queries related to obtaining cooking recipes and two commands related to controlling smart lighting in an ecosystem. In this example, the candidate semantic indicators may include, for example, a first semantic indicator of "kitchen device" and a second semantic indicator of "smart lighting control device." Further, the semantic indicator engine 160 may select the first semantic indicator of "kitchen device" as the given semantic indicator for the given assistant device because past use of the given assistant device indicates that the given assistant device is primarily used for cooking-related activities.

[0043] In some implementations, the semantic classifier stored in ML model database 192 may be a natural language understanding engine (e.g., implemented by NLP module 122, described below). An intent determined based on processing a query, command, and / or request previously received at the assistant device may be mapped to one or more of the semantic categories. In particular, the heterogeneous semantic categories described herein may be defined at various levels of granularity. For example, a semantic category may be associated with a class of smart device commands and / or a type category of the class, such as a smart lighting command category, a smart thermostat command category, and / or a smart camera command category. In other words, each category may have a unique set of intents associated with it as determined by the semantic classifier, although some intents of a category may also be associated with additional categories. In some additional or alternative implementations, semantic classifiers stored in the ML model database 192 may be utilized to generate text embeddings (e.g., low-dimensional representations, such as word2vec representations) corresponding to query, command, and / or request text. These embeddings may be points in an embedding space such that semantically similar words or phrases are associated with the same or similar portions of the embedding space. Furthermore, these portions of the embedding space may be associated with one or more of a plurality of heterogeneous semantic categories, and a given one of the embeddings may be classified into a given one of the semantic categories if a distance measure between the given one of the embeddings and one or more of the portions of the embedding space satisfies a distance threshold. For example, a word or phrase related to cooking may be associated with a first portion of the embedding space associated with the semantic label “cooking,” a word or phrase related to weather may be associated with a second portion of the embedding space associated with the semantic label “weather,” and so on.

[0044] In implementations in which one or more device-specific signals additionally or alternatively include ambient noise activity, the instances of ambient noise can be processed using an ambient noise detection model (e.g., stored in ML model database 192) to index device activity for a given assistant device into one or more different semantic categories corresponding to different types of ambient noise. Candidate semantic labels can be generated based on the semantic categories into which the instances of ambient noise are classified, and a given semantic label selected for a given assistant device can be selected based on the amount of ambient noise instances classified in a given semantic category. For example, assume that the ambient noise detected at a given assistant device (and optionally only when speech recognition is active) primarily includes ambient noise classified as cooking sounds. In this example, semantic label engine 160 can select the semantic label “kitchen device” as the given semantic label for the given assistant device because the ambient noise captured in the audio data indicates that the device is located near cooking-related activity.

[0045] In some implementations, an ambient noise detection model stored in the ML model database 192 can be trained to detect a particular sound, and whether an instance of ambient noise includes a particular sound can be determined based on an output generated across the ambient noise detection model. The ambient noise detection model can be trained, for example, using supervised learning techniques. For example, multiple training instances can be obtained. Each of the training instances can include a training instance input including ambient noise and a corresponding training instance output including an indication of whether the training instance input includes the particular sound that the ambient noise detection model is trained to detect. For example, if the ambient noise detection model is trained to detect the sound of glass breaking, a training instance including the sound of glass breaking can be assigned a label (e.g., "yes") or value (e.g., "1"), and a training instance that does not include the sound of glass breaking can be assigned a different label (e.g., "no") or value (e.g., "0"). In some additional or alternative implementations, the ambient noise detection models stored in the ML model database 192 may be utilized to generate audio embeddings (e.g., low-dimensional representations of instances of ambient noise) based on the instances of ambient noise (or their acoustic features, such as Mel-frequency cepstral coefficients, raw audio waveforms, and / or other acoustic features). These embeddings may be points in an embedding space such that similar sounds (or acoustic features that capture sounds) are associated with the same or similar portions of the embedding space. Furthermore, these portions of the embedding space may be associated with one or more of a plurality of heterogeneous semantic categories, and a given one of the embeddings may be classified into a given one of the semantic categories if a distance measure between the given one of the embeddings and one or more of the portions of the embedding space satisfies a distance threshold. For example, an instance of glass breaking may be associated with a first portion of the embedding space related to the sound of “glass breaking,” an instance of doorbell ringing may be associated with a second portion of the embedding space related to the sound of “doorbell,” and so on.

[0046] In an implementation in which the one or more device-specific signals additionally or alternatively include unique identifiers or indicators of additional assistant devices that are close to the given assistant device, candidate semantic indicators may be generated based on the unique identifiers or indicators, and a given semantic indicator selected for the given assistant device may be selected based on one or more of the unique identifiers or indicators of the additional assistant devices. For example, assume that a first indicator, "smart oven," is associated with a first assistant device that is close to the given assistant device, and a second indicator, "smart coffee maker," is associated with a second assistant device that is close to the given assistant device. In this example, the semantic indicator engine 160 may select the semantic indicator "kitchen device" as the given semantic indicator for the given assistant device because the indicators associated with the additional assistant devices that are close to the given assistant device are related to cooking. The unique identifier or indicator may be processed using a semantic classifier stored in the ML model database 192 in the same or similar manner as described above in connection with processing queries, commands, and / or requests.

[0047] In an implementation in which candidate semantic labels for a given assistant device are generated based on user preferences, the user preferences can be processed using a semantic classifier (e.g., stored in ML model database 192) to index the user preferences into one or more different semantic categories corresponding to the disparate user preferences. Candidate semantic labels can be generated based on the semantic categories into which the user preferences are categorized, and a given semantic label selected for a given assistant device can be selected based on the given semantic category into which the user preferences associated with the given assistant device are categorized. For example, assume that user preferences indicate that a user associated with the ecosystem likes cooking and prefers a fictional chef named Johnny Flay. In this example, the candidate semantic labels can include, for example, a first semantic label called "cooking device" and a second semantic label called "Johnny Flay device." In some versions of those implementations, utilizing the user preferences as a device-specific signal to generate one or more candidate semantic indicators can be in response to receiving user input to assign semantic indicators to the assistant device based on the user preferences.

[0048] In some implementations, semantic sign engine 160 can automatically assign a given semantic sign to a given assistant device in a device topology representation of the ecosystem (e.g., stored in device topology database 193). In some additional or alternative implementations, semantic sign engine 160 can cause automated assistant 120 to generate a prompt including candidate semantic signs. The prompt can ask a user associated with the ecosystem to select one of the candidate signs as the given semantic sign. Further, the prompt can be visually and / or audibly rendered on a given one of the assistant devices (which may or may not be the given assistant device to which the given semantic sign is assigned) and / or on the user's client device (e.g., a mobile device). In response to receiving a selection of one of the candidate signs as the given semantic sign, the selected given semantic sign can be assigned to the given assistant device in the device topology representation of the ecosystem (e.g., stored in device topology database 193). In some versions of these implementations, a given semantic indicator assigned to a given assistant device may be added to a list of semantic indicators for the given assistant device. In other words, multiple semantic indicators may be associated with a given assistant device. In other versions of these implementations, a given semantic indicator assigned to a given assistant device may replace any other semantic indicator for the given assistant device. In other words, only a single semantic indicator may be associated with a given assistant device.

[0049] In some implementations, query / command processing engine 170 can process queries, requests, or commands directed to automated assistant 120 and received via one or more of assistant input devices 106. Query / command processing engine 170 can process the queries, requests, or commands to select one or more assistant devices to satisfy the query or command. In particular, one or more of the assistant devices selected to satisfy the query or command may be different from one or more of the assistant input devices 106 that received the query or command. Query / command processing engine 170 can select one or more assistant devices to satisfy an utterance based on one or more criteria. The one or more criteria may include, for example, one or more proximity of the device to the user who provided the utterance (e.g., determined using presence sensor 105 described below), one or more device capabilities of the devices in the ecosystem, semantic labels assigned to one or more assistant devices, and / or other criteria for selecting an assistant device to satisfy the utterance.

[0050] For example, assume that a display device is required to satisfy the utterance. In this example, the assistant device candidates considered when selecting a given assistant device to satisfy the utterance may be limited to those that include a display device. If multiple assistant devices in the ecosystem include a display device, the given assistant device that includes a display device and is closest to the user may be selected to satisfy the utterance. In contrast, in an implementation in which only a speaker is required to satisfy the utterance (e.g., a display device is not required to satisfy the utterance), the assistant device candidates considered when selecting a given assistant device to satisfy the utterance may include those that have a speaker, regardless of whether they include a display device.

[0051] As another example, assume that an utterance includes semantic characteristics that match a semantic indicator assigned to a given assistant device. The query / command processing engine 170 can determine that the semantic characteristics of the utterance match the semantic indicator assigned to the given assistant device by generating a first embedding corresponding to one or more words of the utterance (or corresponding text) and a second embedding corresponding to one or more words of the semantic indicator assigned to the given assistant device, comparing the embeddings, and determining whether a distance measure between the embeddings satisfies a distance threshold that indicates that the embeddings match (e.g., whether it is an exact match or a rough match). In this example, the query / command processing engine 170 can select a given assistant device to satisfy the utterance based on the utterance's match with the semantic indicator (optionally in addition to or instead of the proximity of the user who provided the utterance to the given assistant device). In this way, the selection of an assistant device to satisfy an utterance can be biased toward semantic indicators assigned to the assistant device as described herein.

[0052] In various implementations, one or more of the assistant input devices 106 includes one or more respective presence sensors 105 configured to provide a signal indicative of a detected presence, particularly a human presence, upon acknowledgement from a corresponding user. 1-N(also referred to herein simply as "presence sensors 105"). In some of these implementations, automated assistant 120 can identify one or more of assistant input devices 106 based at least in part on the presence of a user at one or more of assistant input devices 106 to satisfy an utterance from a user associated with the ecosystem. The utterance can be satisfied by rendering responsive content (e.g., audibly and / or visually) at one or more of assistant input devices 106, by causing one or more of assistant input devices 106 to be controlled based on the utterance, and / or by causing one or more of assistant input devices 106 to perform any other action to satisfy the utterance. As described herein, automated assistant 120 can utilize data determined based on each presence sensor 105 in determining those assistant input devices 106 based on where the user is or has recently been, and provide corresponding commands to only those assistant input devices 106. In some additional or alternative implementations, the automated assistant 120 can utilize data determined based on each presence sensor 105 in determining whether a user (any user or a specific user) is currently near any of the assistant input devices 106, and can optionally suppress providing commands based on determining that a user (any user or a specific user) is not near any of the assistant input devices 106.

[0053] Each presence sensor 105 may take a variety of forms. Some assistant input devices 106 may be equipped with one or more digital cameras configured to capture and provide signals indicative of detected movement in their field of view. Additionally or alternatively, some assistant input devices 106 may be equipped with other types of light-based presence sensors 105, such as passive infrared ("PIR") sensors that measure infrared ("IR") light emanating from objects within their field of view. Additionally or alternatively, some assistant input devices 106 may be equipped with presence sensors 105 that detect sound waves (or pressure waves), such as one or more microphones. Moreover, in addition to assistant input devices 106, one or more of assistant non-input devices 185 may additionally or alternatively include respective presence sensors 105 described herein, and signals from such sensors may be additionally utilized by automated assistant 120 when determining whether and / or how to satisfy an utterance according to implementations described herein.

[0054] Additionally or alternatively, in some implementations, presence sensor 105 can be configured to detect other phenomena related to the presence of people or devices in the ecosystem. For example, in some embodiments, a given one of the assistant devices can be equipped with presence sensor 105 that detects various types of wireless signals (e.g., waves such as radio waves, ultrasound, electromagnetic waves, etc.) emitted by, for example, other assistant devices (e.g., mobile devices, wearable computing devices, etc.) carried / operated by a particular user and / or other assistant devices in the ecosystem (e.g., as described with respect to event detection engine 130). For example, some of the assistant devices can be configured to emit waves imperceptible to humans, such as ultrasound or infrared waves, that can be detected by one or more of assistant input devices 106 (e.g., via an ultrasound / infrared receiver such as an ultrasound-enabled microphone).

[0055] Additionally or alternatively, various assistant devices may emit other types of waves imperceptible to humans, such as radio waves (e.g., Wi-Fi, Bluetooth, cellular, etc.), that can be detected by other assistant devices (e.g., mobile devices, wearable computing devices, etc.) carried / operated by a particular user and used to determine the specific location of the operating user. In some implementations, Wi-Fi triangulation can be used to detect a person's location based on Wi-Fi signals to / from the assistant device, for example. In other implementations, other wireless signal characteristics, such as time-of-flight, signal strength, etc., can be used by various assistant devices, alone or collectively, to determine a person's location based on signals emitted by other assistant devices carried / operated by a particular user.

[0056] Additionally or alternatively, in some implementations, one or more of the assistant input devices 106 may perform voice recognition to recognize a user from their voice. For example, some instances of the automated assistant 120 may be configured to match the voice with a user profile, e.g., for the purpose of providing / restricting access to various resources. In some implementations, the speaker's movements may then be determined, e.g., by a presence sensor 105 on the assistant device. In some implementations, based on such detected movements, the user's location may be predicted, and this location may be considered the user's location when any content is rendered on the assistant devices based at least in part on the assistant devices' proximity to the user's location. In some implementations, the user may simply be considered to be in the location where the user last engaged with the automated assistant 120, especially if not much time has passed since that last engagement.

[0057] Each of the assistant input devices 106 further includes a respective user interface component 107 1-N (also referred to herein simply as "user interface component 107"), each of which may include one or more user interface input devices (e.g., microphone, touch screen, keyboard) and / or one or more user interface output devices (e.g., display, speaker, projector). As an example, user interface component 1071 of assistant input device 1061 may include only a speaker and microphone, while assistant input device 106 N User Interface Components 107 N may include a speaker, a touchscreen, and a microphone. Additionally or alternatively, in some implementations, assistant non-input devices 185 may include one or more user interface input devices and / or one or more user interface output devices of user interface component 107, although the user input devices (if any) for assistant non-input devices 185 may not enable a user to directly interact with automated assistant 120.

[0058] Assistant input devices 106 and / or any other computing devices operating one or more of cloud-based automation assistant components 119 may each include one or more memories for storage of data and software applications, one or more processors for accessing data and executing applications, and other components that facilitate communication over a network. Operations performed by one or more of assistant input devices 106 and / or by automation assistant 120 may be distributed among multiple computer systems. Automation assistant 120 may be implemented, for example, as a computer program running on one or more computers at one or more locations coupled to each other through a network (e.g., any of networks 110 of FIG. 1 ).

[0059] As mentioned above, in various implementations, each of the assistant input devices 106 can operate a respective automated assistant client 118. In various implementations, each automated assistant client 118 operates a respective speech capture / text-to-speech (TTS) / speech-to-text (STT) module 114. 1-N (also referred to herein simply as "speech capture / TTS / STT module 114"). In other implementations, one or more aspects of each speech capture / TTS / STT module 114 may be implemented separately from each automated assistant client 118.

[0060] Each respective speech capture / TTS / STT module 114 can be configured to perform one or more functions including, for example, capturing a user's speech (speech capture, e.g., via a respective microphone (which in some cases may comprise presence sensor 105)), converting the captured audio to text and / or other representations or embeddings (STT) using speech recognition models stored in ML model database 192, and / or converting text to speech (TTS) using speech synthesis models stored in ML model database 192. Instances of these models can be stored locally on each respective assistant input device 106 and / or accessible by the assistant input device (e.g., via network 110 of FIG. 1 ). In some implementations, one or more of assistant input devices 106 may have relatively limited computational resources (e.g., processor cycles, memory, battery, etc.), so a respective speech capture / TTS / STT module 114 local to each of assistant input devices 106 may be configured to convert a finite number of different spoken phrases into text (or other formats, such as low-dimensional embeddings) using a speech recognition model. Other speech inputs may be sent to one or more of cloud-based automated assistant components 119, which may include cloud-based TTS module 116 and / or cloud-based STT module 117.

[0061] Cloud-based STT module 117 can be configured to leverage the virtually limitless resources of the cloud to convert audio data captured by speech capture / TTS / STT module 114 into text (which can then be provided to natural language processor module 122) using speech recognition models stored in ML model database 192. Cloud-based TTS module 116 can be configured to leverage the virtually limitless resources of the cloud to convert text data (e.g., text compiled by automated assistant 120) into computer-generated speech output using speech synthesis models stored in ML model database 192. In some implementations, cloud-based TTS module 116 can provide computer-generated speech output to one or more of the assistant devices, for example, to be output directly using the respective speakers of the respective assistant devices. In other implementations, text data generated by automated assistant 120 using cloud-based TTS module 116 (e.g., client device notifications included in commands) may be provided to speech capture / TTS / STT module 114 of the respective assistant device, which may then use a speech synthesis model to locally convert the text data into computer-generated speech that is rendered through a local speaker on the respective assistant device.

[0062] The automation assistant 120 (and in particular the one or more cloud-based automation assistant components 119) may include a natural language processing (NLP) module 122, the aforementioned cloud-based TTS module 116, the aforementioned cloud-based STT module 117, and other components, some of which are described in more detail below. In some implementations, one or more of the engines and / or modules of the automation assistant 120 may be omitted, combined, and / or implemented in a component separate from the automation assistant 120. An instance of the NLP module 122 may additionally or alternatively be implemented locally in the assistant input device 106.

[0063] In some implementations, the automation assistant 120 generates response content in response to various inputs generated by a user of one of the assistant input devices 106 during a human-computer interaction session with the automation assistant 120. The automation assistant 120 can provide the response content (e.g., via one or more of the networks 110 of FIG. 1 when separate from the assistant device) for presentation to the user as part of the interaction session via the assistant input device 106 and / or the assistant non-input device 185. For example, the automation assistant 120 can generate response content in response to free-form natural language input provided via one of the assistant input devices 106. As used herein, free-form input is input organized by the user that is not constrained to a group of options presented for selection by the user.

[0064] NLP module 122 of automation assistant 120 may process natural language input generated by a user via assistant input devices 106 and generate annotated output for use by one or more other components of automation assistant 120, assistant input devices 106, and / or assistant non-input devices 185. For example, NLP module 122 may process free-form natural language input generated by a user via one or more respective user interface input devices of assistant input devices 106. The annotated output generated based on processing the free-form natural language input may include one or more annotations of the natural language input and, optionally, one or more (e.g., all) of the terms of the natural language input.

[0065] In some implementations, the NLP module 122 is configured to identify and annotate various types of grammatical information in the natural language input. For example, the NLP module 122 may include a portion of an utterance tagger configured to annotate words with their grammatical roles. In some implementations, the NLP module 122 may additionally and / or alternatively include an entity tagger (not shown) configured to annotate references to entities in one or more segments, such as references to people (e.g., including literary characters, celebrities, public figures, etc.), organizations, locations (real and imaginary), etc. In some implementations, data about entities may be stored in one or more databases, such as a knowledge graph (not shown). In some implementations, the knowledge graph may include nodes representing known entities (and possibly entity attributes) and edges connecting the nodes to represent relationships between the entities.

[0066] The entity tagger of NLP module 122 may annotate references to an entity at a high level of granularity (e.g., to enable identification of all references to an entity class, such as people) and / or at a low level of granularity (e.g., to enable identification of all references to a particular entity, such as a particular person). The entity tagger may rely on the content of the natural language input to resolve for a particular entity, and / or may optionally communicate with a knowledge graph or other entity database to resolve for a particular entity.

[0067] In some implementations, the NLP module 122 may additionally and / or alternatively include a coreference resolver (not shown) configured to group or "cluster" mentions into the same entity based on one or more contextual cues. For example, the coreference resolver may be utilized to resolve the word "it" in the natural language input "lock it" to "front door lock" based on the "front door lock" being mentioned in a client device notification that was rendered immediately prior to receiving the natural language input "lock it."

[0068] In some implementations, one or more components of NLP module 122 may rely on annotations from one or more other components of NLP module 122. For example, in some implementations, a designated entity tagger may rely on annotations from a coreference resolver and / or a dependency analyzer when annotating all references to a particular entity. Also, for example, in some implementations, a coreference resolver may rely on annotations from a dependency analyzer when clustering references to the same entity. In some implementations, when processing a particular natural language input, one or more components of NLP module 122 may determine one or more annotations using relevant data outside of the particular natural language input, such as an assistant input device notification that was rendered immediately before receiving the natural language input on which the assistant input device notification is based.

[0069] 1 is illustrated as having a particular configuration of components implemented by the assistant device and / or server and as having the assistant device and / or server communicating over a particular network, it should be understood that this is for illustrative purposes only and is not intended to be limiting. For example, assistant input device 106 and assistant non-input devices may be directly and communicatively coupled to each other via one or more networks (not shown). As another example, the operation of one or more cloud-based automated assistant components 119 may be implemented locally on one or more of assistant input devices 106 and / or one or more of the assistant non-input devices. As yet another example, instances of various ML models stored in ML model database 192 may be stored locally on the assistant device, and / or instances of the device topology representation of the ecosystem stored in device topology database 193 may be stored locally on the assistant input device. Additionally, in implementations in which data (e.g., device activity, corresponding audio data or recognized text, device topology representations, and / or any other data described herein) is transmitted over any of one or more networks 110 of FIG. 1, the data may be encrypted, filtered, or otherwise protected in any manner to ensure user privacy.

[0070] Using the techniques described herein, semantic labels can be inferred and assigned to assistant devices in the ecosystem to maintain an up-to-date device topology representation of the ecosystem. Furthermore, the semantic labels assigned to the assistant devices are semantically meaningful to the user in that they are selected based on the use of the respective assistant devices and / or the respective parts of the ecosystem in which the respective assistant devices are located. Thus, when an utterance is received by one or more of the assistant devices in the ecosystem, the automated assistant can more accurately select one or more of the assistant devices best suited to satisfy the utterance. As a result, because a user associated with the ecosystem does not need to specify a particular device to satisfy the utterance or repeat the utterance when the wrong device is selected to satisfy the utterance, the amount of user input received by one or more of the assistant devices in the ecosystem can be reduced, thereby conserving computational and / or network resources at the assistant device by reducing network traffic. Additionally, when an assistant device is newly added to or moved within the ecosystem, the amount of user input received by one or more of the assistant devices in the ecosystem can be reduced because a user does not need to manually update the device topology representation via a software application associated with the ecosystem.

[0071] Further description of the various components of Figure 1 will now be provided with reference to Figures 2A and 2B. A floor plan of a house is illustrated in Figures 2A and 2B. The illustrated floor plan includes multiple rooms 250-262. Multiple assistant input devices 106 1-5 are deployed throughout at least some of the rooms. 1-5Each of the rooms may implement an instance of automated assistant client 118 configured using selected aspects of the present disclosure and may include one or more input devices, such as a microphone, capable of capturing speech spoken by nearby people. For example, a first assistant input device 1061 in the form of an interactive standalone speaker and a display device (e.g., a display screen, a projector, etc.) are deployed in room 250 of FIG. 2A , which in this example is a kitchen, and in room 256 of FIG. 2B , which in this example is a living room. A second assistant input device 1062 in the form of a so-called “smart” television (e.g., a network-connected television with one or more processors implementing respective instances of automated assistant client 118) is deployed in room 252, which in this example is a study. A third assistant input device 1063 in the form of an interactive standalone speaker without a display is deployed in room 254, which in this example is a bedroom. A fourth assistant input device 1064 in the form of another interactive standalone speaker is deployed in room 256, which in this example is a living room. A fifth assistant input device 1065, also in the form of a smart television, is also located in room 250, in this example the kitchen.

[0072] Although not shown in FIGS. 2A and 2B, multiple assistant input devices 106 1-4may be communicatively coupled with each other and / or other resources (e.g., the Internet) via one or more wired or wireless WANs and / or LANs (e.g., via network 110 of FIG. 1 ). In addition, other assistant input devices, particularly certain mobile devices such as smartphones, tablets, laptops, wearable devices, etc., may also be present and may, for example, be carried by one or more people in the home, which may or may not be connected to the same WAN and / or LAN. The configuration of assistant input devices illustrated in FIGS. 2A and 2B is just one example; more or fewer and / or different assistant input devices 106 may be deployed throughout any number of other rooms and / or areas of the home and / or in locations other than the home (e.g., businesses, hotels, public places, airports, vehicles, and / or other locations or spaces).

[0073] Multiple Assistant Non-Input Devices 185 1-5 are further shown in Figures 2A and 2B. For example, a first assistant non-input device 1851 in the form of a smart doorbell is deployed outside the home near the front door. A second assistant non-input device 1852 in the form of a smart lock is deployed outside the home adjacent to the front door. A third assistant non-input device 1853 in the form of a smart washing machine is deployed in room 262, which in this example is the laundry room. A fourth assistant non-input device 1854 in the form of a door open / close sensor is deployed near the back door of room 262 and detects whether the back door is open or closed. A fifth assistant non-input device 1855 in the form of a smart thermostat is deployed in room 262, which in this example is a study.

[0074] Each of Assistant non-input devices 185 can communicate with a respective Assistant non-input system 180 (shown in FIG. 1 ) (e.g., via network 110 of FIG. 1 ) to provide data to the respective Assistant non-input system 180, optionally so that the data can be controlled based on commands provided by the respective Assistant non-input system 180. One or more of Assistant non-input devices 185 can additionally or alternatively communicate directly with one or more of Assistant input devices 106 (e.g., via network 110 of FIG. 1 ) to provide data to one or more of Assistant input devices 106, optionally so that the data can be controlled based on commands provided by one or more of Assistant input devices 106. The configuration of assistant non-input devices 185 shown in Figures 2A and 2B is merely an example, and more or fewer and / or different assistant non-input devices 185 may be deployed throughout any number of other rooms and / or areas of the home and / or in locations other than the home (e.g., businesses, hotels, public places, airports, vehicles, and / or other locations or spaces).

[0075] In various implementations, a semantic indicator can be assigned to a given assistant device (e.g., a given one of assistant input devices 106 or assistant non-input devices 185) based on processing of one or more device-specific signals associated with the respective assistant device. The one or more device-specific signals can be detected by and / or generated by the given assistant device. The one or more device-specific signals can include, for example, one or more queries (if any) previously received at the given assistant device, one or more commands (if any) previously executed at the given assistant device, instances of ambient noise previously detected at the given assistant device (and optionally only when speech reception was active at the given assistant device), unique identifiers (or indicators) of respective assistant devices positionally proximate to the given assistant device, and / or user preferences of a user associated with the ecosystem. Each of the one or more device-specific signals associated with the given assistant device can be processed to classify each of them into one or more semantic categories of a plurality of heterogeneous semantic categories.

[0076] Furthermore, one or more candidate semantic labels may be generated based on one or more device-specific signals. The candidate semantic labels may be generated using one or more rules (optionally heuristically defined) or machine learning models (e.g., stored in the ML model database 192). For example, one or more heuristically defined rules may indicate that candidate semantic labels associated with each of the semantic categories into which one or more device-specific signals are classified should be generated. For example, assume that the device-specific signals are classified into a “kitchen” category, a “cooking” category, a “bedroom” category, and a “living room” category. In this example, the candidate semantic labels may include a first candidate semantic label named “kitchen assistant device,” a second candidate semantic label named “cooking assistant device,” a third candidate semantic label named “bedroom assistant device,” and a fourth candidate semantic label named “living room assistant device.” As another example, one or more device-specific signals (or one or more semantic categories corresponding thereto) may be processed using a machine learning model trained to generate candidate semantic labels. For example, a machine learning model may be trained based on multiple training instances. Each of the training instances may include a training instance input and a corresponding training instance output. The training instance input may include, for example, one or more device-specific signals and / or one or more semantic categories, and the corresponding training instance output may include, for example, a ground truth output corresponding to a semantic label to be assigned based on the training instance input.

[0077] Furthermore, a semantic label to be assigned to a given assistant device may be selected from one or more candidate semantic labels. A semantic label to be assigned to a given assistant device may be selected from one or more candidate semantic labels based on a reliability level associated with each of the one or more candidate semantic labels. In some implementations, a semantic label to be assigned to a given assistant device may be automatically assigned to the given assistant device, while in additional or alternative implementations, a user associated with the ecosystem of FIGS. 2A and 2B may be prompted to select a semantic label to be assigned to the given assistant device from a list of one or more candidate semantic labels (e.g., as described with respect to FIG. 3). In some additional or alternative implementations, a semantic label may be automatically assigned to a given assistant device if the semantic label is unique (relative to other assistant devices positionally close to the given assistant device in the ecosystem).

[0078] In some versions of those implementations, a given assistant device to which a semantic label should be assigned may be identified in response to determining (e.g., via event detection engine 130 and / or device identification engine 140 of FIG. 1 ) that the given assistant device is newly added to the ecosystem. For example, with particular reference to FIG. 2A , assume that a first assistant input device 1061 and device, taking the form of a two-way standalone speaker, is newly deployed in room 250, in this example, a kitchen. Assume that no previous queries or commands have been received by or executed by first assistant input device 1061 (other than configuring first assistant input device 1061) since first assistant input device 1061 was newly added to the ecosystem in association with one or more device-specific signals associated with first assistant input device 1061 of FIG. 2A , and that no previous queries or commands have been received by or executed by first assistant input device 1061 (other than configuring first assistant input device 1061) while first assistant input device 1061 was being configured by a user of the ecosystem (e.g., including, for example, a name, test utterances to establish speech embeddings for the user, etc.). Assume that several instances of ambient noise have been captured (when speech reception is active as the user provides speech), that a unique identifier (or indicator) associated with a fifth assistant input device 1065 in the form of a smart television in room 250 (e.g., "Kitchen TV") and a unique identifier (or indicator) associated with a fifth assistant non-input device 1855 in the form of a smart thermostat in room 252 (e.g., "Thermostat") are detected at first assistant input device 1061, and that the user's user preferences related to the ecosystem are known.

[0079] In this example, instances of ambient noise (if any) may be processed using an ambient noise detection model to classify the ambient noise into one or more semantic categories. For example, assume that instances of ambient noise capture water dripping in a sink in room 250, the rumble of a microwave or oven in room 250, food frying in a frying pan on a stove in room 250, etc. These instances of ambient noise may be classified into a “kitchen” semantic category, a “cooking” semantic category, and / or other semantic category associated with noises typically encountered in a kitchen. Additionally or alternatively, the ambient noise may capture a movie, television program, or advertisement being visually and / or audibly rendered via fifth assistant input device 1065 in the form of a smart television in room 250. These instances of ambient noise may be classified into a “television” semantic category, a “movie” semantic category, and / or other semantic category associated with noises typically encountered from a smart television. Additionally or alternatively, the unique identifier (or indicator) of the fifth assistant input device 1065 (e.g., “kitchen TV”) and the unique identifier (or indicator) of the fifth assistant non-input device 1855 (e.g., thermostat) can be processed to generate one or more semantic indicators. These unique identifiers (or indicators) can be categorized into a “kitchen” semantic category, a “smart device” semantic category, and / or other semantic categories associated with assistant devices positionally proximate to the first assistant input device 1061 in the ecosystem of FIG. 2A . Additionally or alternatively, assume that the user preferences indicate that the user is interested in cooking and a fictional chef named Johnny Flay. These user preferences can be identified as associated with the first assistant input device 1061 based on the classification of one or more of the other device-specific signals into the “cooking” or “kitchen” category and that the user preferences for cooking and Johnny Flay are also classified into the “cooking” or “kitchen” category.As a result, the candidate semantic labels in this example may include “kitchen speaker device,” “cooking speaker device,” “television speaker device,” “movie speaker device,” “Johnny Flay device,” and / or other candidate semantic labels based on one or more device-specific labels associated with the first assistant input device 1061. Further in this example, a given one of the semantic labels can be automatically assigned to the first assistant input device 1061, or a user associated with the ecosystem can be prompted to select one or more of the candidate semantic labels (e.g., while the first assistant input device 1061 is being configured) to assign the given semantic label to the first assistant input device 1061.

[0080] In some additional or alternative implementations, a given assistant device to which a semantic label is to be assigned may be identified periodically (e.g., weekly, monthly, every six months, and / or any other period via event detection engine 130 and / or device identification engine 140 of FIG. 1). For example, with particular reference to FIG. 2A, assume that a first assistant input device 1061 in the form of a two-way standalone speaker and display device has been deployed in room 250, in this example the kitchen, for six months. With respect to one or more device-specific signals associated with the first assistant input device 1061 in FIG. 2A, assume that a previous query or command has been received at or executed by the first assistant input device 1061, assume that an instance of ambient noise has been captured during the first assistant input device 1061, and assume that a unique identifier (or indicator) associated with a fifth assistant input device 1065 in the form of a smart television in room 250 (e.g., "Kitchen TV") and a unique identifier (or indicator) associated with a fifth assistant non-input device 1855 in the form of a smart thermostat in room 252 (e.g., "Thermostat") are still detected at the first assistant input device 1061.

[0081] In this example, the queries and commands (or the text corresponding thereto) may be processed using a semantic classifier to classify the queries and commands into one or more semantic categories. For example, queries and commands previously received at the first assistant input device 1061 may include a query for requesting a cooking recipe, a command for setting a timer, a command for controlling any smart devices in the kitchen, and / or other queries or commands. These instances of queries and commands may be classified into a "cooking" semantic category, a "smart device control" category, and / or other semantic categories based on the queries and commands received at the first assistant input device 1061. Semantic indicator candidates may additionally or alternatively be determined based on instances of ambient noise and / or unique identifiers (or indicators) positionally proximate to the first assistant input device 1061, as described above. As a result, the semantic label candidates in this example may include "kitchen device," "timer device," "thermostat display device," "cooking display device," "television device," "movie device," "Johnny Flay recipe device," and / or other semantic label candidates based on one or more device-specific labels associated with the first assistant input device 1061. Further, in this example, a given one of the semantic labels can be automatically assigned to the first assistant input device 1061, or a user associated with the ecosystem can be prompted to select one or more of the semantic label candidates to assign the given semantic label to the first assistant input device 1061 (e.g., while the first assistant input device 1061 is being configured).

[0082] In some additional or alternative implementations, a given assistant device to which a semantic label is to be assigned may be identified in response to determining (e.g., via event detection engine 130 and / or device identification engine 140 of FIG. 1 ) that the given assistant device has moved within the ecosystem. For example, and referring now particularly to FIG. 2B , assume that a first assistant input device 1061 in the form of a two-way standalone speaker and display device is moved from room 250, in this example the kitchen, to room 256, in this example the living room. 2B , assume that a query or command has previously been received at or executed by first assistant input device 1061, assume that several instances of ambient noise have been captured, assume that a respective unique identifier (or indicator) associated with fourth assistant input device 1064 in the form of another two-way standalone speaker in room 256 (e.g., “living room speaker device”) is detected at first assistant input device 1061, and assume that the user's user preferences associated with the ecosystem are known. In this example, the one or more device-specific signals may be limited to those generated or received after first assistant input device 1061 has been moved within the ecosystem (fewer user preferences).

[0083] In this example, queries and commands (or their corresponding text) may be processed using a semantic classifier to categorize the queries and commands into one or more semantic categories. For example, queries and commands previously received at first assistant input device 1061 may include queries related to requesting weather or traffic information, commands related to planning a vacation, and / or other queries or commands. These instances of queries and commands may be categorized into an “information” semantic category (or more specifically, a “weather information” category and a “traffic information” category), an “appointment” category, and / or other semantic categories based on the queries and commands received at first assistant input device 1061. Additionally or alternatively, instances of ambient noise may be processed using an ambient noise detection model to categorize the ambient noise into one or more semantic categories. For example, the ambient noise may capture music or podcasts being aurally rendered by fourth assistant input device 1064, people sitting on a couch conversing as illustrated in room 256, movies or television programs being aurally rendered by computing devices in the ecosystem, etc. These instances of ambient noise may be categorized into a "music" semantic category, a "conversation" semantic category, a "movie" semantic category, a "television program" semantic category, and / or other semantic categories associated with noises typically encountered in a kitchen. Additionally or alternatively, a unique identifier (or indicator) of the fourth assistant input device 1064 (e.g., "living room speaker device") may be processed to generate one or more of the semantic indicators. This unique identifier (or indicator) may be categorized, for example, into the "living room" semantic category and / or other semantic categories associated with assistant devices positionally proximate to the first assistant input device 1061 in the ecosystem of FIG. 2B.Additionally or alternatively, assume that the user preferences indicate that the user is interested in cooking and a fictional movie titled "Vehicles" and a particular character in the movie named Thunder McKing. These user preferences may be identified as being associated with the first assistant input device 1061 based on the classification of one or more of the other device-specific signals into "Movies" and the user preferences for the movie "Vehicles" and "Thunder McKing" also being classified into the "Movies" category. As a result, the semantic indicator candidates in this example may include "living room device," "schedule device," "Vehicles device," "Thunder McKing device," and / or other semantic indicator candidates based on one or more device-specific indicators associated with the first assistant input device 1061. Further, in this example, a given one of the semantic indicators can be automatically assigned to the first assistant input device 1061, or a user associated with the ecosystem can be prompted to select one or more of the candidate semantic indicators to assign the given semantic indicator to the first assistant input device 1061.

[0084] In some additional or alternative implementations, a given assistant device to which a semantic label is to be assigned can be identified in response to determining (e.g., via event detection engine 130 and / or device identification engine 140 of FIG. 1) that the portion of the ecosystem in which the given assistant device is located has been repurposed. For example, with particular reference now to FIG. 2B , assume that a first assistant input device 1061 in the form of a two-way standalone speaker and display device is located in room 256, which in this example is a living room, but the living room is converted into a bedroom. 2B , assume that a query or command has previously been received at or executed by first assistant input device 1061, assume that some instances of ambient noise have been captured, and assume that a respective unique identifier (or indicia) associated with fourth assistant input device 1064 (e.g., “living room speaker device”) in the form of another two-way standalone speaker in room 256 is detected at first assistant input device 1061. In this example, the one or more device-specific signals may be limited to those generated or received after room 256 has been repurposed.

[0085] In this example, queries or commands (e.g., corresponding text) may be processed using a semantic classifier to classify the queries and commands into one or more semantic categories. For example, queries and commands previously received at first assistant input device 1061 may include commands related to setting an alarm, commands related to wake-up or bedtime routines, and / or other queries or commands. These instances of commands may be classified into an “alarm” semantic category, a “routine” category (or more specifically, a “wake-up routine” category or a “bedtime routine” category), and / or other semantic categories based on the queries and commands received at first assistant input device 1061. Additionally or alternatively, instances of ambient noise may be processed using an ambient noise detection model to classify the ambient noise into one or more semantic categories. For example, the ambient noise may capture one or more users snoring, people talking, etc. These instances of ambient noise may be categorized into a "bedroom" semantic category, a "conversation" semantic category, and / or other semantic categories associated with noises typically encountered in a bedroom. Additionally or alternatively, a unique identifier (or indicator) of the fourth assistant input device 1064 (e.g., "living room speaker device") may be processed to generate one or more of the semantic indicators. This unique identifier (indicator) may be categorized, for example, into the "living room" semantic category and / or other semantic categories associated with assistant devices positionally proximate to the first assistant input device 1061 in the ecosystem of FIG. 2B. As a result, the semantic indicator candidates in this example may include "living room display device," "bedroom display device," and / or other semantic indicator candidates based on one or more device-specific indicators associated with the first assistant input device 1061.Further, in this example, a given one of the semantic indicators can be automatically assigned to the first assistant input device 1061, or a user associated with the ecosystem can be prompted to select one or more of the candidate semantic indicators to assign the given semantic indicator to the first assistant input device 1061. In particular, in this example, the system can determine that the given semantic indicator corresponds to a “bedroom display device” because use of the first assistant input device 1061 indicates that the fourth assistant input device 1064 is located in a bedroom, even though the unique identifier (or indicator) of the fourth assistant input device 1064 corresponds to a “living room” category.

[0086] 2A and 2B are described herein with respect to a given assistant device being assigned a semantic label as an assistant input device (e.g., first assistant input device 1061), it should be understood that this is for purposes of illustration and is not intended to be limiting. For example, the techniques described herein can also be utilized to assign respective semantic labels to assistant non-input devices 185. For example, assume that a smart light (perhaps an assistant non-input device without a microphone) is newly added to room 252, which in this example is a bedroom. In this example, a unique identifier or label (e.g., "bedroom speaker device") associated with third assistant input device 1063, which takes the form of a two-way standalone speaker without a display, can be utilized to infer a semantic label of "bedroom smart light" for the newly added smart light using the techniques described herein. Further, assume that a smart light is moved from room 254 to room 262, which in this example is a laundry room. In this example, a unique identifier or indicator associated with a third assistant non-input device 1853 in the form of smart clothing may be utilized to infer a semantic indicator of “laundry room smart light” for the recently transferred smart light using techniques described herein.

[0087] Turning now to FIG. 3 , a flowchart illustrating an example method 300 for assigning a given semantic indicator to a given assistant device in an ecosystem is shown. For convenience, the operations of method 300 are described with reference to a system that performs the operations. The system of method 300 includes one or more processors and / or other components of a computing device. For example, the system of method 300 may be implemented by assistant input device 106 of FIG. 1 , FIG. 2A, or FIG. 2B , assistant non-input device 185 of FIG. 1 , FIG. 2A, or FIG. 2B , computing device 510 of FIG. 5 , one or more servers, other computing devices, and / or any combination thereof. Moreover, although the operations of method 300 are shown in a particular order, this is not intended to be limiting. One or more operations may be reordered, omitted, and / or added.

[0088] In block 352, the system identifies a given assistant device from among multiple assistant devices in the ecosystem. The given assistant device can be an assistant input device (e.g., one of assistant input devices 106 in FIG. 1 ) or an assistant non-input device (e.g., one of assistant non-input devices 185 in FIG. 1 ). In some implementations, the given assistant device can be identified in response to determining that it is newly added to the ecosystem, while in other implementations, the given assistant device can be identified in response to determining that it has moved within the ecosystem (e.g., as described with respect to event detection engine 130 in FIG. 1 ). In some additional or alternative implementations, the given assistant device can be identified periodically (e.g., once a month, once every six months, once a year, etc.).

[0089] In block 354, the system acquires a device-specific signal associated with the given assistant device. The device-specific signal may be detected by and / or generated by the given assistant device. In some implementations, block 354 may include one or more of optional sub-blocks 354A, 354B, 354C, and 354D. If included, in sub-block 354A, the system acquires multiple queries or commands (if any) previously received at the given assistant device. If included, in sub-block 354B, the system additionally or alternatively acquires instances of ambient noise previously detected at the given assistant device (and optionally only when speech acceptance was active at the given assistant device (e.g., after receiving a specific word or phrase that invokes an automated assistant) or when speech acceptance was not active via a digital signal processor (DSP)). In some implementations, the acquired ambient noise is constrained to the ambient noise detected when speech acceptance is active on the given assistant device. If included, in subblock 354C, the system additionally or alternatively acquires a unique identifier (or indicator) for each assistant device that is positionally close to the given assistant device (e.g., determined using device identification engine 140 of FIG. 1). If included, in subblock 354D, the system additionally or alternatively acquires the user's user preferences related to the ecosystem.

[0090] In block 356, the system processes the device-specific signals to generate candidate semantic labels for the given assistant device. In implementations in which one or more device-specific signals include multiple queries or commands previously received at the given assistant device, the multiple queries or commands (or their corresponding text) can be processed using a semantic classifier to classify the multiple queries or commands into one or more heterogeneous semantic categories. For example, queries related to cooking recipes and commands related to controlling a smart kitchen oven or smart coffee maker can be classified into a cooking category or a kitchen category, weather-related queries can be classified into a weather category, commands related to controlling lighting can be classified into a lighting category, and so on. In implementations in which one or more device-specific signals additionally or alternatively include instances of ambient noise detected at the given assistant device, the instances of ambient noise can be processed using an ambient noise detection model to classify each of the instances of ambient noise into one or more heterogeneous semantic categories. For example, if an instance of ambient noise is determined to correspond to a microwave oven humming, food frying in a frying pan, food processing in a food processor, etc., the instance of ambient noise may be categorized into a cooking category. As another example, if an instance of ambient noise is determined to correspond to a circular saw operating, hammering, etc., the instance of ambient noise may be categorized into a garage category and a workshop category. In an implementation in which one or more device-specific signals additionally or alternatively include a unique identifier (or indicator) of each assistant device positionally proximate to a given assistant device, the unique identifier (or indicator) of each assistant device may be categorized into one or more heterogeneous semantic categories. For example, if the unique identifier (or indicator) of each assistant device corresponds to "coffee maker," "oven," and "microwave," the unique identifier may be categorized into a kitchen category or a cooking category.As another example, if the unique identifier (or indicator) of each assistant device corresponds to "bedroom light" and "bedroom casting device," the unique identifier may be categorized into a bedroom category. In implementations in which one or more device-specific signals additionally or alternatively include a user's user preferences related to the ecosystem, the user preferences may be categorized into one or more heterogeneous semantic categories. For example, if the user preferences indicate that the user is interested in cooking, cooking shows, specific chefs, and / or other cooking-related interests, the user preferences may be categorized into a kitchen category, a cooking category, or a category related to specific chefs.

[0091] The candidate semantic indicators may be generated based on processing one or more device-specific signals. For example, assume that one or more device-specific signals indicate that a given assistant device is located in the kitchen or living room of a user's main residence associated with the ecosystem. Further assume that the given assistant device is a two-way standalone speaker device with a display. In this example, a first candidate semantic indicator, "kitchen display device," and a second candidate semantic indicator, "living room display device," may be generated. Generating candidate semantic indicators based on one or more device-specific signals is described in more detail herein (e.g., in connection with FIGS. 2A and 2B).

[0092] In block 358, the system determines whether to prompt a user associated with the ecosystem to select a given semantic label from among the candidate semantic labels. The system can determine whether to prompt a user to select a given semantic label based on whether the respective confidence levels associated with the candidate semantic labels meet a threshold confidence level. The respective confidence levels can be determined, for example, based on the amount of one or more device-specific signals that fall into a given semantic category. For example, assume that the given assistant device identified in block 352 is a two-way standalone speaker that implements an instance of an automated assistant. Further assume that each of the one or more device-specific signals indicates that the two-way standalone speaker is located in a bedroom. For example, based on previous queries or commands received at the two-way standalone speaker being related to an alarm or bedtime routine, based on instances of ambient noise including snoring, and / or based on other assistant devices having unique identifiers (or labels) of “bedroom light” and “bedroom casting device.” In this example, the interactive standalone speaker is associated with a bedroom in a primary residence of a user associated with the ecosystem, so the system may be confident in the semantic label "bedroom speaker device." However, if some of the queries or commands received at the interactive standalone speaker are related to cooking recipes, the system may be less confident in the semantic label "bedroom speaker device."

[0093] If, in the repetition of block 358, the system determines not to prompt the user to select a given semantic indicator, the system may proceed to block 360. In block 360, the system automatically assigns the given semantic indicator from among the candidate semantic indicators to the given assistant device in a device topology representation in the ecosystem. The device topology representation of the ecosystem may be stored locally in one or more of the assistant devices in the ecosystem and / or in a remote system communicating with one or more of the assistant devices in the ecosystem. In some implementations, the given semantic indicator may be the only semantic indicator associated with the given assistant device (and may optionally replace other unique identifiers or indicators assigned to the given assistant device), while in other implementations, the given semantic indicator may be added to a list of unique identifiers or indicators assigned to the given assistant device. The system may then return to block 352 to identify additional given assistant devices from among the multiple assistant devices in the ecosystem and generate and assign additional given semantic indicators to the additional given assistant devices.

[0094] If, in the iteration of block 358, the system determines to prompt the user to select a given semantic indicator, the system may proceed to block 362. In block 362, the system generates a prompt to prompt the user of the client device to select a given semantic indicator based on the candidate semantic indicators. In block 364, the system causes the prompt to be rendered on the user's client device. The prompt may be rendered visually and / or audibly on the user's client device and, optionally, based on the capabilities of the client device on which the prompt is rendered. For example, the prompt may be rendered via a software application available on the client device (e.g., a software application associated with the ecosystem or one or more assistant devices included in the ecosystem). In this example, if the client device includes a display, the prompt may be rendered visually via the software application (or a home screen of the client device) and / or audibly via a speaker of the client device. However, if the client device does not include a display, the prompt may be rendered only audibly via a speaker of the client device. In block 366, the system receives a selection of the given semantic indicator in response to the prompt. The user's client device may be, for example, a given assistant identified in block 352 or a separate client device (e.g., the user's mobile device or any other assistant device in the ecosystem capable of rendering prompts). For example, assume the system generates the following semantic indicator candidates: "bedroom speaker device" and "kitchen speaker device."In this example, the system can generate a prompt that includes selectable elements associated with both of these semantic indicators and request the user to provide input (e.g., touch or voice) to select one of the candidate semantic indicators to assign to the given assistant device as the given semantic indicator to be assigned to the given assistant device. In block 368, the system assigns the given semantic indicator to the given assistant device in a device topology representation of the ecosystem in a manner similar to that described above with respect to block 360.

[0095] Turning now to FIG. 4 , a flowchart illustrating an example method 400 for using assigned semantic indicators in satisfying a query or command received at an assistant device in an ecosystem is shown. For convenience, the operations of method 400 are described with reference to a system that performs the operations. The system of method 400 includes one or more processors and / or other components of a computing device. For example, the system of method 400 may be implemented by assistant input device 106 of FIG. 1 , FIG. 2A, or FIG. 2B , assistant non-input device 185 of FIG. 1 , FIG. 2A, or FIG. 2B , computing device 510 of FIG. 5 , one or more servers, other computing devices, and / or any combination thereof. Moreover, while the operations of method 400 are shown in a particular order, this is not intended to be limiting. One or more operations may be reordered, omitted, and / or added.

[0096] In block 452, the system receives audio data corresponding to a user's speech via each microphone of each assistant device in an ecosystem that includes multiple assistant devices. A user can be associated with the ecosystem.

[0097] In block 454, the system processes the audio data to identify semantic characteristics of the query or command contained in the utterance. The semantic characteristics of the query or command may correspond to linguistic units, such as words or phrases, that define a related domain or related set of the words and / or phrases. In some implementations, the system can process the audio data using a speech recognition model to convert the utterance captured in the audio data into text and can use a semantic classifier to identify the semantic characteristics based on the recognized text. In additional or alternative implementations, the system can process the audio data using a semantic classifier and can directly identify the semantic characteristics based on the audio data. For example, assume the utterance received in block 452 is "Show me the chili recipe." In this example, the utterance (or its corresponding text) can be processed using a semantic classifier to identify the semantic characteristics "chili," "food," "kitchen," and / or "cooking." In particular, the semantic characteristics identified in block 454 may include one or more terms or phrases contained in the utterance, or may include a given semantic category into which the utterance is classified.

[0098] In block 456, the system determines whether the utterance specifies a given assistant device to be utilized in satisfying the query or command contained in the utterance. If, in a repetition of block 456, the system determines that the utterance specifies a given assistant device to be utilized in satisfying the query or command, the system may proceed to block 466. Block 466 is described below. For example, assume that the utterance received in block 452 is "Show me the chili recipe on my kitchen display device." In this example, the system can utilize the "kitchen display device" because the user specified that the "kitchen display device" should be utilized to present the "chili recipe" to the user who provided the utterance. As a result, the system can select the "kitchen display device" to satisfy the utterance. If, in a repetition of block 456, the system determines that the utterance does not specify a given assistant device to be utilized in satisfying the query or command, the system may proceed to block 458. For example, assume that the utterance received in block 452 is simply "Show me the chili recipe" without specifying an assistant device that should satisfy the utterance. In this example, the system can determine that the utterance does not specify a given assistant device that should satisfy the utterance. As a result, the system needs to determine which assistant device in the ecosystem should be utilized to satisfy the utterance.

[0099] In block 458, the system determines whether the semantic characteristic identified in block 454 matches a given semantic label assigned to a given assistant device among the plurality of assistant devices. The system can generate an embedding corresponding to one or more terms of the semantic characteristic and compare the embedding of the semantic characteristic with multiple embeddings corresponding to one or more respective terms of each semantic label assigned to one or more of the plurality of assistant devices in the ecosystem. Furthermore, the system can determine whether the embedding of the semantic characteristic matches any of multiple embeddings of each semantic label assigned to one or more of the plurality of assistant devices in the ecosystem. For example, the system can determine a distance measure between the embedding of the semantic characteristic and each of multiple embeddings of each semantic label assigned to one or more of the plurality of assistant devices in the ecosystem. Furthermore, the system can determine whether the distance measure meets a threshold distance (e.g., to identify an exact match or a rough match).

[0100] In the repetition of block 458, if the system determines that the semantic characteristic identified in block 454 matches the given semantic indicator assigned to the given assistant device, the system may proceed to block 460. In block 460, the system causes the given client device associated with the given semantic indicator in the device topology representation of the ecosystem to satisfy the query or command included in the utterance. The system may then return to block 452 and monitor additional audio through the microphones of each of multiple assistant devices in the ecosystem. More specifically, the system may cause the given assistant device to perform one or more actions to satisfy the utterance. For example, assume that the utterance received in block 452 is "Show me the chili recipe" and the identified semantic characteristics are "chili," "food," "kitchen," and / or "cooking." Further assume that the semantic indicator assigned to the given assistant device is "kitchen display device." In this example, the system can select a given assistant device assigned the semantic indicator "kitchen display device" even if the utterance was not received at that assistant device. Furthermore, the system can cause a given assistant device assigned the semantic indicator "kitchen display device" to visually render a chili recipe in response to the utterance being received by the respective microphone of the respective assistant device.

[0101] In various implementations, multiple assistant devices in the ecosystem may be assigned semantic labels that match the semantic characteristics identified based on the utterance. Notably, although not shown in FIG. 4 for clarity, the system may determine whether the semantic characteristics identified in block 454 match a given semantic label assigned to a given assistant device among the multiple assistant devices, in addition to or instead of using proximity information (e.g., as described above with respect to query / command processing engine 170 of FIG. 1 ) when selecting one or more assistant devices to satisfy the utterance. Continuing with the above example, assume there are multiple assistant devices assigned the semantic label "kitchen display device." In this example, the assistant device of the multiple assistant devices assigned the semantic label "kitchen display device" that is closest to the user in the ecosystem may be utilized to satisfy the utterance. Additionally, although not shown in FIG. 4 for clarity, the system may, in addition to or instead of using device capability information (e.g., as described above with respect to query / command processing engine 170 of FIG. 1) when selecting one or more assistant devices to satisfy the utterance, determine whether the semantic characteristics identified in block 454 match a given semantic label assigned to a given assistant device among the multiple assistant devices. Continuing with the above example, assume that a first assistant device with a display device is assigned the semantic label "kitchen display device" and a second assistant device without a display device is assigned the semantic label "kitchen speaker device."In this example, a first assistant device assigned the semantic tag "kitchen display device" can be selected to satisfy the utterance over a second assistant device assigned the semantic tag "kitchen speaker device" because the utterance specifies "show me" a chili recipe and the first assistant device can display a chili recipe in response to the utterance, while the second assistant device cannot display a chili recipe.

[0102] If, in the iteration of block 458, the system determines that the semantic characteristic identified in block 454 does not match any semantic indicator assigned to any of the assistant devices in the ecosystem, the system may proceed to block 462. In block 462, the system may identify a given assistant device that is close to the user. For example, the system may identify a given assistant device that is closest to the user in the ecosystem (e.g., as described with respect to presence sensor 105 of FIG. 1 ).

[0103] In block 464, the system determines whether the given assistant device identified in block 462 can satisfy the query or command. The capabilities of each assistant device may be stored in the device topology representation of the ecosystem and in association with the respective assistant device (e.g., as device attributes of the respective assistant device). If, in a repetition of block 464, the system determines that the given assistant device identified in block 462 cannot satisfy the query or command, the system may return to block 462 and identify additional given assistant devices that are also close to the user. The system may again proceed to block 464 and determine whether additional given assistant devices identified in subsequent repetitions of block 462 can satisfy the query or command. The system may repeat this process until an assistant device that can satisfy the query or command is identified. For example, assume that the assistant device identified in block 462 is a standalone speaker device without a display, but a display is required to satisfy speech. In this example, the system may determine that the assistant device identified in block 462 cannot satisfy the utterance in block 464 and may return to block 462 to identify additional assistant devices near the user in block 462. If, in the repetition of block 464, the system determines that the given assistant device identified in block 462 can satisfy the query or command, the system may proceed to block 462. Continuing with the above example, further assume that the given assistant device identified in the first iteration of block 462 is a standalone speaker device that includes a display, or that the additional assistant devices identified in additional iterations of block 462 are standalone speaker devices with displays.In this example, the system may determine that the assistant device identified in a subsequent iteration of block 462 can satisfy the utterance of block 464, and the system may proceed to block 466.

[0104] In block 466, the system causes the given assistant device to satisfy the query or command contained in the utterance. The system can satisfy the utterance in a manner similar to that described above with respect to block 460. The system can then return to block 452 and monitor additional audio data via the microphones of each of the multiple assistant devices in the ecosystem.

[0105] 5 is a block diagram of an exemplary computing device 510 that may be optionally utilized to perform one or more aspects of the techniques described herein. In some implementations, one or more of the assistant input devices, one or more of the cloud-based automated assistant components, one or more assistant non-input systems, one or more assistant non-input devices, and / or other components may comprise one or more components of the exemplary computing device 510.

[0106] Computing device 510 typically includes at least one processor 514 that communicates with several peripheral devices via a bus subsystem 512. These peripheral devices may include, for example, a storage subsystem 524, including a memory subsystem 525 and a file storage subsystem 526, a user interface output device 520, a user interface input device 522, and a network interface subsystem 516. The input and output devices enable user interaction with computing device 510. Network interface subsystem 516 provides an interface to external networks and is coupled to corresponding interface devices in other computing devices.

[0107] The user interface input devices 522 may include a keyboard, a pointing device such as a mouse, a trackball, a touchpad, or a graphics tablet, a scanner, a touchscreen integrated into a display, a voice recognition system, an audio input device such as a microphone, and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computing device 510 or a communication network.

[0108] The user interface output devices 520 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for producing a visible image. The display subsystem may also provide a non-visual display, such as via an audio output device. In general, use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computing device 510 to a user or to another machine or computing device.

[0109] Storage subsystem 524 stores programming and data constructs that provide the functionality of some or all of the modules described herein. For example, storage subsystem 524 may include logic for performing selected aspects of the methods described herein and for implementing the various components shown in FIG.

[0110] These software modules are generally executed by the processor 514, alone or in combination with other processors. The memory 525 used by the storage subsystem 524 may include several memories, including a main random access memory (RAM) 530 for storing instructions and data during program execution, and a read-only memory (ROM) 532 in which fixed instructions are stored. The file storage subsystem 526 may provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of some implementations may be stored by the file storage subsystem 526 in the storage subsystem 524 or on other machines accessible by the processor 514.

[0111] The bus subsystem 512 provides a mechanism for allowing the various components and subsystems of the computing device 510 to communicate with each other as intended. Although the bus subsystem 512 is shown schematically as being a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0112] The computing device 510 may be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of the computing device 510 shown in Figure 5 is intended only as a specific example intended to illustrate some implementations. Many other configurations of the computing device 510 may have more or fewer components than the computing device shown in Figure 5.

[0113] In situations where some implementations discussed herein may collect or use personal information about users (e.g., user data extracted from other electronic communications, information about the user's social network, the user's location, the user's time, the user's biometric information, and user activity and attribute information, relationships between users, etc.), the user is given one or more opportunities to control whether the information is collected, whether the personal information is stored, whether the personal information is used, and how information about the user is collected, stored, and used. That is, the systems and methods discussed herein collect, store, and / or use a user's personal information only after receiving explicit authorization to do so from the associated user.

[0114] For example, a user may be given control over whether a program or feature collects user information about that particular user or other users associated with that program or feature. Each user from whom personal information is to be collected is presented with one or more options to enable control over information collection associated with that user, to provide permission or approval for whether information is collected and what portion of the information should be collected. For example, a user may be given one or more such control options over a communications network. Additionally, certain data may be handled in one or more ways before it is stored or used so that personally identifiable information is removed. As one example, a user's identifying information may be handled so that personally identifiable information cannot be determined. As another example, a user's geographic location may be generalized to a broader area so that the user's specific location cannot be determined.

[0115] In some implementations, a method implemented by one or more processors is provided, the method including: identifying a given assistant device from among a plurality of assistant devices in an ecosystem; acquiring one or more device-specific signals associated with the given assistant device, where the one or more device-specific signals are generated or received by the given assistant device; processing one or more of the device-specific signals to generate one or more candidate semantic indicators for the given assistant device; selecting a given semantic indicator for the given assistant device from the one or more candidate semantic indicators; and assigning the given semantic indicator to the given assistant device in a device topology representation of the ecosystem. Assigning the given semantic indicator to the given assistant device includes automatically assigning the given semantic indicator to the given assistant device.

[0116] These and other implementations of the techniques disclosed herein may include one or more of the following features.

[0117] In some implementations, one or more of the device-specific signals may include at least device activity associated with the given assistant device, and the device activity associated with the given assistant device may include multiple queries or commands previously received at the given assistant device. In some versions of these implementations, processing one or more of the device-specific signals to generate one or more candidate semantic indicators for the given assistant device may include using a semantic classifier to process the device activity associated with the given assistant device and classify each of the multiple queries or commands into one or more of a plurality of heterogeneous categories, and generating one or more of the candidate semantic indicators based on one or more of the multiple heterogeneous categories into which each of the multiple queries or commands falls. In some further versions of these implementations, selecting a given semantic indicator for a given assistant device may include selecting a given semantic indicator for a given assistant device based on a quantity of queries or commands that fall into a given category of a plurality of disparate categories, the given semantic indicator being associated with the given semantic indicator.

[0118] In some implementations, one or more of the device-specific signals may include at least device activity associated with the given assistant device, where the device activity associated with the given assistant device comprises ambient noise previously detected at the given assistant device, where the ambient noise may have been previously detected at the given assistant device when speech reception was active. In some versions of these implementations, the method may further include processing the ambient noise previously detected at the given assistant device using an ambient noise detection model to classify the ambient noise into one or more of a plurality of heterogeneous categories, and generating one or more candidate semantic labels based on one or more of the plurality of heterogeneous categories into which the ambient noise is classified. In some further versions of these implementations, selecting a given semantic label for the given assistant device may include selecting a given semantic label for the given assistant device based on the classification of the ambient noise into a given category of the plurality of heterogeneous categories, where the given semantic label is associated with the given semantic label.

[0119] In some implementations, one or more of the device-specific signals may include a respective unique identifier for one or more of a plurality of assistant devices positionally proximate to the given assistant device in the ecosystem. In some versions of these implementations, processing one or more of the device-specific signals to generate one or more candidate semantic indicators for the given assistant device may include identifying one or more of a plurality of assistant devices positionally proximate to the given assistant device in the ecosystem based on one or more wireless signals, obtaining a respective unique identifier for one or more of the plurality of assistant devices positionally proximate to the given assistant device, and generating one or more candidate semantic indicators based on the respective unique identifiers for one or more of the plurality of assistant devices positionally proximate to the given assistant device. In some further versions of these implementations, selecting a given semantic indicator for the given assistant device may include selecting a given semantic indicator for the given assistant device based on properties of the respective unique identifiers for one or more of the plurality of assistant devices positionally proximate to the given assistant device.

[0120] In some implementations, a given assistant device can be identified in response to determining that the given assistant device is newly added to the ecosystem or that the given assistant device has moved within the ecosystem.

[0121] In some versions of these implementations, the given assistant device can be identified in response to determining that the given assistant device has moved within the ecosystem. In some further versions of these implementations, assigning the given semantic indicator to the given assistant device can include adding the given semantic indicator to a list of semantic indicators associated with the given assistant device or replacing an existing semantic indicator associated with the given assistant device with the given semantic indicator. In some additional or alternative versions of these implementations, determining that the given assistant device has moved within the ecosystem can include identifying, based on one or more wireless signals, that a current subset of multiple assistant devices positionally proximate to the given assistant device in the ecosystem is different from a stored subset of multiple assistant devices stored in association with the given assistant device. In still further versions of those implementations, the method may further include switching the given assistant device from an existing group of assistant devices that includes one or more of the plurality of assistant devices to an additional existing group of assistant devices that includes one or more of the plurality of assistant devices, or creating a new group of assistant devices that includes at least the given assistant device.

[0122] In some versions of these implementations, the given assistant device may be identified in response to determining that the given assistant device is newly added to the ecosystem. In some further versions of these implementations, assigning the given semantic indicator to the given assistant device may include adding the given semantic indicator to a list of semantic indicators associated with the given assistant device. In some additional or alternative versions of these further implementations, determining that the given assistant device is newly added to the ecosystem may include identifying the given assistant device as being added to a wireless network associated with the ecosystem based on one or more wireless signals. In still further versions of these implementations, the method may further include adding the given assistant device to an existing group of assistant devices including one or more of the multiple assistant devices, or creating a new group of assistant devices including at least the given assistant device.

[0123] In some implementations, a given assistant device may be periodically identified to detect whether existing semantic indicators assigned to the given assistant device are correct.

[0124] In some implementations, the method may further include, after assigning a given semantic indicator to the given assistant device in a device topology representation of the ecosystem, receiving audio data corresponding to an utterance from a user associated with the ecosystem via one or more respective microphones of one of the plurality of assistant devices in the ecosystem, the utterance including a query or a command; processing the audio data corresponding to the utterance to determine semantic characteristics of the query or command; determining that the semantic characteristics of the query or command match the given semantic indicator assigned to the given assistant device; and, in response to determining that the semantic indicator of the query or command matches the given semantic indicator assigned to the given assistant device, causing the given assistant device to satisfy the query or command.

[0125] In some implementations, one or more of the device-specific signals may include at least user preferences of a user associated with the ecosystem, and the user preferences may be determined based on user interactions with multiple assistant devices in the ecosystem. In some versions of these implementations, processing one or more of the device-specific signals to generate one or more candidate semantic indicators for the given assistant device may include processing the user preferences using a semantic classifier to identify at least one semantic category of multiple heterogeneous semantic categories associated with the user preferences and generating one or more of the candidate semantic indicators based on the given semantic category. In yet further versions of these implementations, selecting a given semantic indicator for the given assistant device may include determining that a given semantic indicator from the multiple candidate semantic indicators is associated with the given assistant device, and selecting the given semantic indicator for the given assistant device in response to determining that the given semantic indicator is associated with the given assistant device. Determining that a given semantic indicator may be associated with a given assistant device is based on one or more additional device-specific signals associated with the given assistant device. In some additional or alternative versions of these still further implementations, processing user preferences to identify at least one semantic category may be in response to receiving user input to assign one or more semantic indicators to at least the given assistant device.

[0126] In some implementations, a method implemented by one or more processors is provided, the method including: identifying a given assistant device from among a plurality of assistant devices in an ecosystem; acquiring one or more device-specific signals associated with the assistant device, the one or more device-specific signals being generated or received by the given assistant device; processing one or more of the device-specific signals to generate one or more candidate semantic indicators for the given assistant device; generating a prompt to ask a user of a client device to select a given semantic indicator based on one or more of the candidate semantic indicators, the selection being from among one or more of the candidate semantic indicators; causing the prompt to be rendered at the user's client device; and assigning the given semantic indicator to the given assistant device in a device topology representation of the ecosystem in response to receiving a selection of the given semantic indicator in response to the prompt.

[0127] These and other implementations of the techniques disclosed herein may include one or more of the following features.

[0128] In some implementations, the one or more device-specific signals may include two or more of: device activity associated with the given assistant device, where the device activity associated with the given assistant device comprises multiple queries or commands previously received at the given assistant device; ambient noise previously detected at the given assistant device when speech acceptance was active; or respective unique identifiers for one or more of multiple assistant devices positionally proximate to the given assistant device in the ecosystem.

[0129] In some implementations, a method implemented by one or more processors is provided, the method including steps of identifying a given assistant device from among a plurality of assistant devices in an ecosystem; acquiring one or more device-specific signals associated with the given assistant device, the one or more device-specific signals being generated or received by the given assistant device; determining a given semantic indicator for the given assistant device based on one or more of the device-specific signals; assigning the given semantic indicator to the given assistant device in a device topology representation of the ecosystem; and assigning the given semantic indicator to the given assistant device in the device topology representation of the ecosystem. After the allocation, the method includes receiving, via one or more respective microphones of one of the plurality of assistant devices in the ecosystem, audio data corresponding to an utterance from a user associated with the ecosystem, the utterance including a query or a command; processing the audio data corresponding to the utterance to determine semantic characteristics of the query or command; determining that the semantic characteristics of the query or command match a given semantic indicator assigned to the given assistant device; and, in response to determining that the semantic characteristics of the query or command match the given semantic indicator assigned to the given assistant device, causing the given assistant device to satisfy the query or command.

[0130] Additionally, some implementations include one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU) and / or a tensor processing unit (TPU)) of one or more computing devices, where the one or more processors are operable to execute instructions stored in associated memory, the instructions configured to cause performance of any of the aforementioned methods. Some implementations also include one or more non-transitory computer-readable storage media that store computer instructions executable by the one or more processors to perform any of the aforementioned methods. Some implementations also include a computer program product that includes instructions executable by the one or more processors to perform any of the aforementioned methods.

[0131] It is understood that all combinations of the foregoing concepts and additional concepts more fully described herein are contemplated as being part of the subject matter disclosed herein, for example, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the subject matter disclosed herein. [Explanation of symbols]

[0132] 105 Presence Sensor 106 Assistant Input Device 107 User Interface Components 110 Network 114 Speech Capture / TTS / STT Module 116 Cloud-based TTS Module 117 Cloud-based STT module 118 Automation Assistant Client 119 Cloud-based automation assistant components 120 Automation Assistant 122 Natural Language Processing (NLP) Module, Natural Language Processor Module 130 Event Detection Engine 140 Device Identification Engine 150 Event Processing Engine 160 Semantic Signing Engine 170 Query / Command Processing Engine 180 Assistant Non-Input System 185 Assistant Non-Input Devices 191 Device Activity Database 192 ML model database 193 Device Topology Database 250 rooms 252 rooms 254 rooms 256 rooms 258 rooms 260 rooms 262 rooms 510 Computing Devices 512 Bus Subsystem 514 processor 516 Network Interface Subsystem 520 User Interface Output Device 522 User Interface Input Devices 524 Storage Subsystem 525 Memory Subsystem 526 File Storage Subsystem 630 Main Random Access Memory (RAM) 632 Read-Only Memory (ROM)

Claims

1. 1. A method implemented by one or more processors, comprising: Identifying a given assistant device from among a plurality of assistant devices in the ecosystem; Obtaining one or more device-specific signals associated with the given assistant device, the one or more device-specific signals being generated or received by the given assistant device, and one or more of the device-specific signals comprising a respective unique identifier for one or more of the plurality of assistant devices positionally proximate to the given assistant device in the ecosystem; Processing one or more of the device-specific signals to generate one or more candidate semantic indicators for the given assistant device, wherein processing one or more of the device-specific signals to generate one or more candidate semantic indicators for the given assistant device includes: Identifying one or more of the plurality of assistant devices that are positionally proximate to the given assistant device in the ecosystem based on one or more wireless signals; Obtaining the respective unique identifiers for one or more of the plurality of assistant devices that are positionally proximate to the given assistant device; generating one or more of the candidate semantic indicators based on the respective unique identifiers for one or more of the plurality of assistant devices that are positionally proximate to the given assistant device; Selecting a given semantic indicator for the given assistant device from among the one or more candidate semantic indicators; and assigning the given semantic indicator to the given assistant device in a device topology representation of the ecosystem, wherein assigning the given semantic indicator to the given assistant device comprises automatically assigning the given semantic indicator to the given assistant device.

2. Selecting the given semantic indicator for the given assistant device includes:

10. The method of claim 1, further comprising selecting the given semantic indicator for the given assistant device based on characteristics of the respective unique identifiers for one or more of the plurality of assistant devices that are positionally proximate to the given assistant device.

3. 10. The method of claim 1, wherein the given assistant device is identified in response to determining that the given assistant device is newly added to the ecosystem or in response to determining that the given assistant device has moved within the ecosystem.

4. 4. The method of claim 3, wherein the given assistant device is identified in response to determining that the given assistant device has moved within the ecosystem.

5. The step of assigning the given semantic indicator to the given assistant device includes: adding the given semantic indicator to a list of semantic indicators associated with the given assistant device; or The method of claim 4, comprising replacing an existing semantic indicator associated with the given assistant device with the given semantic indicator.

6. determining that the given assistant device has moved within the ecosystem, 5. The method of claim 4, comprising: determining, based on one or more wireless signals, that a current subset of the plurality of assistant devices positionally proximate to the given assistant device in the ecosystem is different from a stored subset of the plurality of assistant devices stored in association with the given assistant device.

7. Switching the given assistant device from an existing group of assistant devices that includes one or more of the plurality of assistant devices to an additional existing group of assistant devices that includes one or more of the plurality of assistant devices; or The method of claim 6, further comprising creating a new group of assistant devices that includes at least the given assistant device.

8. The method of claim 3, wherein the given assistant device is identified in response to determining that the given assistant device is newly added to the ecosystem.

9. The step of assigning the given semantic indicator to the given assistant device includes: The method of claim 8, further comprising adding the given semantic indicator to a list of semantic indicators associated with the given assistant device.

10. determining that the given assistant device is newly added to the ecosystem, 10. The method of claim 8, comprising determining, based on one or more wireless signals, that the given assistant device has been added to a wireless network associated with the ecosystem.

11. 10. The method of claim 1, wherein one or more of the device-specific signals further comprise a plurality of queries or commands previously received at the given assistant device.

12. processing one or more of the device-specific signals to generate one or more of the candidate semantic indicators for the given assistant device; processing and classifying each of a plurality of queries or commands into one or more of a plurality of heterogeneous categories using a semantic classifier; and generating one or more of the candidate semantic labels based on the one or more of the plurality of heterogeneous categories into which each of a plurality of queries or commands falls.

13. 10. The method of claim 1, wherein one or more of the device-specific signals further comprise ambient noise previously detected at the given assistant device when speech reception was active.

14. processing the one or more of the device-specific signals to generate one or more of the candidate semantic indicators for the given assistant device, Using an ambient noise detection model, process the ambient noise previously detected at the given assistant device to classify the ambient noise into one or more of a plurality of heterogeneous categories; and generating one or more of the candidate semantic labels based on the one or more of the plurality of heterogeneous categories into which the ambient noise is classified.

15. Identifying a given assistant device from among a plurality of assistant devices in the ecosystem; Obtaining one or more device-specific signals associated with the given assistant device, the one or more device-specific signals being generated or received by the given assistant device, and one or more of the device-specific signals comprising a respective unique identifier for one or more of the plurality of assistant devices positionally proximate to the given assistant device in the ecosystem; Processing one or more of the device-specific signals to generate one or more candidate semantic indicators for the given assistant device, wherein processing the one or more of the device-specific signals to generate one or more candidate semantic indicators for the given assistant device includes: Identifying one or more of the plurality of assistant devices that are positionally proximate to the given assistant device in the ecosystem based on one or more wireless signals; Obtaining the respective unique identifiers for one or more of the plurality of assistant devices that are positionally proximate to the given assistant device; generating one or more of the candidate semantic indicators based on the respective unique identifiers for one or more of the plurality of assistant devices that are positionally proximate to the given assistant device; generating a prompt to ask a user of a client device to select a given semantic indicator based on one or more of the candidate semantic indicators, the selection being from one or more of the candidate semantic indicators; causing the prompt to be rendered at the client device of the user; In response to receiving the selection of the given semantic indicator in response to the prompt, assigning the given semantic indicator to the given assistant device in a device topology representation of the ecosystem.

16. Selecting the given semantic indicator for the given assistant device includes:

16. The method of claim 15, further comprising selecting the given semantic indicator for the given assistant device based on characteristics of the respective unique identifiers for one or more of the plurality of assistant devices that are positionally proximate to the given assistant device.

17. 16. The method of claim 15, wherein one or more of the device-specific signals further comprise a respective unique identifier for one or more of the plurality of assistant devices positionally proximate to the given assistant device in the ecosystem.

18. 16. The method of claim 15, wherein one or more of the device-specific signals further comprise ambient noise previously detected at the given assistant device when speech reception was active.

19. 16. The method of claim 15, wherein the given assistant device is identified in response to determining that the given assistant device is newly added to the ecosystem or in response to determining that the given assistant device has moved within the ecosystem.

20. Identifying a given assistant device from among a plurality of assistant devices in the ecosystem; Obtaining one or more device-specific signals associated with the given assistant device, the one or more device-specific signals being generated or received by the given assistant device, and one or more of the device-specific signals comprising a respective unique identifier for one or more of the plurality of assistant devices positionally proximate to the given assistant device in the ecosystem; determining a given semantic indicator for the given assistant device based on one or more of the device-specific signals, wherein determining the given semantic indicator for the given assistant device is based at least on the respective unique identifiers for one or more of the plurality of assistant devices positionally proximate to the given assistant device in the ecosystem; Assigning the given semantic indicator to the given assistant device in a device topology representation of the ecosystem; After assigning the given semantic indicator to the given assistant device in the device topology representation of the ecosystem, receiving, via one or more respective microphones of one of the plurality of assistant devices in the ecosystem, audio data corresponding to an utterance from a user associated with the ecosystem; processing the audio data corresponding to the utterance to determine semantic characteristics of the utterance; Determining that the semantic characteristics of the query or command included in the utterance match the given semantic indicator assigned to the given assistant device; In response to determining that the semantic characteristics of the query or command match the given semantic indicator assigned to the given assistant device, causing the given assistant device to satisfy the query or command; 1. A method implemented by one or more processors, comprising:

21. at least one processor; and a memory storing instructions that, when executed, cause the at least one processor to operate to perform the method of any one of claims 1 to 20.

22. 21. A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform the operations of the method of any one of claims 1 to 20.

Citation Information

Patent Citations

  • Providing custom names for headless devices

    US20150052231A1

  • Control and / or registration of smart devices, locally by an assistant client device

    WO2020076816A1