Assistant device arbitration using wearable device data

Device arbitration using computerized glasses to determine user intent through gaze and direction addresses device reconciliation issues, optimizing resource use and interaction efficiency in multi-device environments.

JP7791940B2Active Publication Date: 2025-12-24GOOGLE LLC
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2024111040
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-02-03
Filing Date
2024-07-10
Publication Date
2025-12-24
Estimated Expiration
2041-12-06

Smart Images

  • Figure 0007791940000001
    Figure 0007791940000001
  • Figure 0007791940000002
    Figure 0007791940000002
  • Figure 0007791940000003
    Figure 0007791940000003
Patent Text Reader

Abstract

To execute device arbitration in a multi-device environment by means of a wearable computing device such as glasses.SOLUTION: Computerized glasses can include a camera which can be used to provide image data for resolving issues related to device arbitration. A direction that a user is directing their computerized glasses and / or directing their gaze is used to prioritize a particular device in a multi-device environment. A detected orientation of the computerized glasses is also used to determine how to simultaneously allocate content between a graphical display of the computerized glasses and another graphical display of another client device. When content is allocated to the computerized glasses, gestures can be enabled and actionable at the computerized glasses.SELECTED DRAWING: Figure 1A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Assistant device reconciliation using wearable device data. [Background technology]

[0002] Humans may engage in human-to-computer interactions with interactive software applications referred to herein as "automated assistants" (also referred to as "digital agents," "chatbots," "interactive personal assistants," "intelligent personal assistants," "assistant applications," "conversational agents," etc.). For example, humans (who may be referred to as "users" when interacting with an automated assistant) may provide commands and / or requests to the automated assistant, in some cases using verbal natural language input (i.e., utterances) that may be converted to text and then processed, and / or by providing textual (e.g., typed) natural language input.

[0003] A user may interact with an automated assistant using multiple client devices. For example, some users may own a coordinated “ecosystem” of client devices that includes a combination of one or more smartphones, one or more tablet computers, one or more vehicle computing systems, one or more wearable computing devices, one or more smart televisions, and / or one or more standalone interactive speakers, among other client devices. A user may use any of these client devices (assuming an automated assistant client is installed) to engage in human-to-computer interactions with the automated assistant. In some cases, these client devices may be scattered around the user's primary residence, secondary residence, workplace, and / or other buildings. For example, mobile client devices such as smartphones, tablets, and smartwatches may be on the user's body and / or where the user last left them. Other client devices, such as traditional desktop computers, smart televisions, and standalone interactive speakers, may be more stationary but may be located in various locations (e.g., rooms) within the user's home or workplace.

[0004] When a user has multiple automated assistant devices in their home, each assistant device may have a different operating state as a result of performing different actions. In such a case, the user may request to change a specific action in progress on an assistant device but may accidentally cause a different assistant device to change a different action. This may be due in part to the fact that some assistant devices may depend solely on whether the respective assistant device heard the user speak a command to change a specific action. As a result, the adaptability of an assistant device to a particular multi-assistance environment may be limited when the user is not speaking directly to the assistant device intended to interact with it. For example, a user may accidentally initialize an action on an assistant device, potentially requiring the user to repeat a previous verbal utterance to recall the action on the desired device.

[0005] As a result, memory and processing bandwidth for a particular assistant device may be momentarily consumed in response to an erroneous invocation of that particular assistant device. Such seemingly redundant results may waste network resources, for example, because some assistant input may be processed by a natural language model accessible only via a network connection. Furthermore, any data associated with the inadvertently affected action must be re-downloaded to the desired device in facilitating completion of the affected action, and energy wasted from canceling energy-intensive actions (e.g., controlling display backlights, heating elements, and / or powered home appliances) may not be recoverable. Summary of the Invention [Means for solving the problem]

[0006] Implementations described herein relate to device arbitration techniques that involve processing data from computerized glasses worn by a user to identify the appropriate client device to which user input should be directed. Enabling device arbitration to be performed using data from the computerized glasses can minimize the number of instances in which client devices are erroneously activated. In this way, memory, power, and network bandwidth can be conserved for those devices that are most susceptible to accidental activation from a particular detected user input.

[0007] In some implementations, a user may be in an environment including multiple assistant-enabled devices, such as the living room of the user's home. The assistant-enabled devices may be activated in response to user input, such as a verbal utterance. Additionally, the assistant-enabled devices may assist in device arbitration to identify a particular computing device that the user may have intended to invoke with the user input. The user may be wearing computerized glasses when providing the verbal utterance, and the computerized glasses may include one or more cameras that can provide image data for detecting a direction in which the user may be facing. The identified direction may be used during device arbitration to prioritize a particular device over other devices based on the direction in which the user is facing.

[0008] In some implementations, the computerized glasses can include circuitry for detecting the position of a user's pupils to determine the user's gaze relative to areas and / or objects in an environment. For example, the computerized glasses can include a forward-facing camera that can be used to identify an area toward which the user is facing and a reverse-facing camera that can be used to identify the user's gaze. When a user provides input to an automated assistant accessible through multiple devices in an environment, the computerized glasses can provide information about the user's gaze to assist with device arbitration. For example, the user can provide verbal utterances to the automated assistant when the user faces an area of ​​the environment that includes multiple assistant-enabled devices. Data generated in the computerized glasses can be used to determine whether the user's gaze is more directed toward a particular assistant-enabled device compared to other assistant-enabled devices. Once a particular device is selected based on the user's gaze, the automated assistant can respond to the user's verbal utterances at the particular device.

[0009] In some implementations, an assistant-enabled device including a camera can provide image data that can be processed together with other image data from the computerized glasses to perform device arbitration. For example, visual characteristics of the user and / or the environment can be determined from one or more cameras separate from the computerized glasses to determine whether to prioritize a particular device during device arbitration. As an example, a user's accessory may be pointed toward a particular device, but the accessory may not be visible within the field of view of the camera of the computerized glasses. However, the orientation of the accessory may be visible within the field of view of the camera of another computing device (e.g., a stand-alone display device). In some cases, a user may be facing a particular area containing two or more assistant-enabled devices and may provide verbal utterances. The user may simultaneously have an accessory (e.g., a hand and / or a foot) pointed toward a particular device among the two or more assistant-enabled devices when the user provides verbal utterances. In such cases, image data from the other computing device (e.g., from the camera of the stand-alone display device) and other image data from the computerized glasses can be processed to select a particular device to respond to the verbal utterance.

[0010] In some implementations, the computerized glasses can detect one or more different outputs (e.g., a first output, a second output, etc.) from one or more different assistant-enabled devices to determine the position and / or placement of the computerized glasses relative to one or more different devices. For example, to calibrate the computerized glasses for a particular user, the user can provide a verbal utterance such as, "Assistant, I'm looking at the kitchen display" while gazing at the display interface of a computing device in their kitchen. In response, images captured via the camera of the kitchen computing device and / or other image data captured via the computerized glasses can be processed to calibrate the computerized glasses for that particular user.

[0011] In some cases, this calibration operation can enhance the performance of computerized glasses and / or other assistant-enabled devices, especially when a user gazes at an assistant-enabled device, typically when the user does not position their head and / or face completely toward the assistant-enabled device. Additionally, this calibration operation can enhance interaction between a user and other assistant-enabled devices that may not have an integrated camera and therefore may not be able to provide image data during device arbitration. For example, a user may gaze at a particular assistant-enabled device, but the assistant-enabled device may not be within the field of view of the outward-facing camera of the computerized glasses. The gaze detected by the inward-facing camera of the computerized glasses can be used, with prior permission from the user, to prioritize the assistant-enabled device the user gazes at over other devices (e.g., another device that may be within the field of view of the outward-facing camera) during device arbitration.

[0012] In some implementations, calibration and / or device reconciliation may be performed using communication between one or more assistant-enabled devices and the computerized glasses via one or more different modalities. For example, a standalone speaker device may include a light that can illuminate a forward-facing camera of the computerized glasses to detect the location of the standalone speaker relative to the computerized glasses. Alternatively or additionally, ultrasound may be emitted by one or more devices, such as the computerized glasses and / or one or more other assistant-enabled devices, to determine the location of the device relative to the other devices. In some implementations, one or more lights on the device may be detected by a camera of the computerized glasses to determine whether the device has lost connection, is no longer synchronized with the other devices, and / or otherwise indicates a particular state that may be communicated via one or more lights. In this manner, the computerized glasses may detect changes in the respective states of one or more devices when a user is wearing the computerized glasses.

[0013] In some implementations, the location of the device relative to the computerized glasses may be used to control certain functions of the computerized glasses and / or one or more assistant-enabled devices. For example, content being viewed on a computing device (e.g., a television) may be associated with content being rendered on the display interface of the computerized glasses. In some cases, a user may be watching a live stream of a sporting event on their television and viewing commentary from friends on the display interface of the computerized glasses. When the user moves away from the television and / or otherwise looks away from the television, the content being rendered on the television may be rendered on the display interface of the computerized glasses according to the user's preferences. For example, when a user moves away from the television in their living room to change their laundry in another room in their home, a change in the relative position of the computerized glasses and / or a change in the user's gaze may be detected. Based on this change in the user's location, additional content data may be rendered on the display interface of the computerized glasses. Alternatively or additionally, based on this change in the user's location, content reduction may be implemented on the television to conserve power and other computing resources, such as network bandwidth.

[0014] The above description is provided as a summary of some implementations of the present disclosure. Further description of these and other implementations is set forth in more detail below.

[0015] Other implementations may include a non-transitory computer-readable storage medium storing instructions executable by one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)) to perform methods such as one or more of the methods described above and / or elsewhere herein. Still other implementations may include systems of one or more computers including one or more processors operable to execute the stored instructions to perform methods such as one or more of the methods described above and / or elsewhere herein.

[0016] It should be understood that all combinations of the foregoing concepts, and additional concepts described in more detail herein, are considered to be part of the subject matter disclosed herein, for example, all combinations of claimed subject matter appearing at the end of this disclosure are considered to be part of the subject matter disclosed herein. [Brief explanation of the drawings]

[0017] [Figure 1A] FIG. 1 illustrates a view of a user invoking an automated assistant while wearing computerized glasses that can assist with device arbitration. [Figure 1B] FIG. 1 illustrates a view of a user invoking an automated assistant while wearing computerized glasses that can assist with device arbitration. [Figure 2] FIG. 1 illustrates a view of a user wearing computerized glasses according to some implementations discussed herein. [Figure 3A] FIG. 1 illustrates a view of a user interacting with an automated assistant that can rely on computerized glasses for device arbitration. [Figure 3B] FIG. 1 illustrates a view of a user interacting with an automated assistant that can rely on computerized glasses for device arbitration. [Figure 3C]FIG. 1 illustrates a view of a user interacting with an automated assistant that can rely on computerized glasses for device arbitration. [Figure 3D] FIG. 1 illustrates a view of a user interacting with an automated assistant that can rely on computerized glasses for device arbitration. [Figure 3E] FIG. 1 illustrates a view of a user interacting with an automated assistant that can rely on computerized glasses for device arbitration. [Figure 4] FIG. 1 illustrates a system for performing device arbitration using data available from a device, such as computerized glasses, that has the capability to perform augmented reality. [Figure 5] FIG. 1 illustrates a method for performing device arbitration in a multi-device environment using data available from wearable computing devices such as computerized eyeglasses. [Figure 6] FIG. 1 is a block diagram of an exemplary computing system. DETAILED DESCRIPTION OF THE INVENTION

[0018] 1A and 1B show views 100 and 120, respectively, of a user 102 invoking an automated assistant while wearing computerized glasses 104 that can assist in device arbitration. The computerized glasses 104 can assist in device arbitration by providing image data characterizing at least a field of view 112 of the user 102 and / or the computerized glasses 104 and / or image data characterizing the gaze of the user 102. In this manner, if assistant input is detected at multiple devices, data from the computerized glasses 104 can be used to identify the particular device to which the user 102 is directing the assistant input. For example, the user 102 can be watching a television 106 while sitting in an environment 108, such as the user's 102's living room. While watching the television 106, the field of view 112 of the user 102 can include the television 106, a display device 118, and a tablet device 110. In some implementations, the television 106 can include a computing device that provides access to an automated assistant, or alternatively, a dongle 126 (i.e., a detachable accessory device) can be attached to the television 106 to render certain content on the television 106. The display device 118 and the tablet device 110 can also provide the user 104 with access to an automated assistant.

[0019] While wearing the computerized glasses 104, the user 102 can provide verbal utterances such as, "Assistant, play the movie I was watching last night." The user 102 can provide verbal utterances 114 to modify the operation of the television 106 and / or the dongle 126. However, because the user 102 is located in an environment 108 having multiple assistant-enabled devices (e.g., the tablet device 110 and the display device 118), multiple different devices may detect the verbal utterance 114 from the user 102. For example, the tablet device 110, the display device 118, the television 106, and the computerized glasses 104 may each detect the verbal utterance 114 from the user 102. As a result, a device arbitration process may be initiated at one or more of the device and / or remote computing devices to identify the particular device to which the user 102 is directing the verbal utterance 114.

[0020] In some implementations, the device arbitration process may also include identifying devices that detected verbal utterances 114 from the user 102 and determining whether any of the identified devices are associated with the field of view 112 of the user 102. For example, a computing device performing the arbitration process may determine that the television 106, the display device 118, the dongle 126, and the tablet device 110 are associated with the field of view 112 of the user 102. This determination may be based on image data generated by one or more cameras of the computerized glasses 104. The computing device may then determine that the television 106 occupies a portion of the field of view 112 that is farther from the periphery or outer boundary of the field of view 112 than the tablet device 110 and the display device 118. Alternatively or additionally, the computing device performing the arbitration process may determine that image data from another camera of the television 106 and / or tablet device 110 indicates that the user 102 is oriented more toward the television 106 than the tablet device 110 and the display device 118.

[0021] Based on one or more of these determinations, the computing device performing the device arbitration process may determine that the television 106 is the device to which the user 102 directed the verbal utterance 114, instead of the display device 118. In some implementations, if the television 106 is the target of automated assistant input for the dongle 126, the device arbitration process may result in the selection of the dongle 126 as the target of the verbal utterance 114. Thus, even though the dongle 126 may be invisible to the user 102 and / or may be obscured from the field of view 112 by the television 106, the dongle 126 may nevertheless respond to the verbal utterance 114 when the user 102 directs their gaze in the direction of the dongle 126 when providing assistant input. For example, audio data corresponding to the verbal utterance 114 may be captured by a microphone on the tablet device 110 or another microphone on the display device 118, while an action requested by the user 102 may be performed at the television 106 and / or the dongle 126 based on image data from the computerized glasses 104. For example, based on identifying a particular computing device responsive to the verbal utterance 114, the television 116 may perform the action 116 of playing a movie per the verbal utterance 114 from the user 102, instead of the movie being played on the display device 118 or the tablet device 110.

[0022] 1B illustrates a view 120 of a user 102 changing the user's field of view 124 to face a computing device 122 located in a different area of ​​the environment 108. The user 102 may reorient themselves to direct verbal utterances to a different device than the television 106. For example, the computing device 122 may have been performing an operation 128 playing music when the user first provided the previous verbal utterance 114. Thus, to stop the music without affecting the movie on the television 106, the user 102 may turn their face and the computerized glasses 104 toward the computing device 122 rather than the television 106. In some implementations, the computerized glasses 104 may include an outward-facing camera having a field of view 124 (i.e., a viewing window or visual viewpoint), which may generate image data that may be used during the device arbitration process. Alternatively or additionally, the computerized glasses 104 may include an inward-facing camera, which may also generate image data that may be used during the device arbitration process.

[0023] For example, image data generated using an inward-facing camera may characterize the gaze of user 102 as being directed slightly upward toward computing device 122 and away from tablet device 110. In this manner, image data from the inward-facing camera may indicate that user 102's gaze is directed toward computing device 122, even though the microphone of computing device 122 does not detect verbal input as clearly as tablet device 110 or display device 118 (at least when user 102 is positioned as shown in FIG. 1B ). For example, user 102 may provide verbal utterance 130 such as "Assistant, stop" when television 106 is playing a movie and computing device 122 is playing music. Image data from one or more cameras of computerized glasses 104 and / or data from one or more other devices may be processed during a device arbitration process to determine that computing device 122 is intended to be the target of verbal utterance 130.

[0024] In some implementations, a heuristic process and / or one or more trained machine learning models may be used during the device arbitration process to select a particular device to respond to input from a user. For example, one or more trained machine learning models may be used to process image data from an inward-facing camera or other image data from an outward-facing camera to identify the device to which user input is directed. Alternatively or additionally, a heuristic process may be used to determine whether to prioritize a particular device over other candidate devices based on data from one or more sources. For example, a device that is not located within the field of view of the user and / or the computerized glasses may be considered a lower priority than another device that is determined to be within the field of view of the user and / or the computerized glasses. For example, if computing device 122 is prioritized over other devices in environment 108, one or more operations may be performed at computing device 122 to fulfill the user request embodied in verbal utterance 130. For example, computing device 122 may perform operation 132 to prevent music from no longer playing on computing device 122.

[0025] 2 shows a view 200 of a user 202 wearing computerized glasses 204 according to some implementations discussed herein. The computerized glasses 204 may include a computer 208, which may include one or more processors and / or one or more memory devices and may receive power from one or more energy sources (e.g., a battery, wireless power transmission, etc.). The computer 208 may be at least partially embodied by and / or separate from a housing 214. The housing 214 may resemble one or more different styles of eyeglass frames and may have one or more lenses 206 attached to the housing 214. In some implementations, the computerized glasses 204 may include one or more forward-facing cameras 210, which may be positioned to have a field of view corresponding to the field of view of the user 202. In some implementations, the computerized glasses 204 may include one or more inward-facing cameras 212, which may be positioned to have another field of view including one or more eyes of the user 202. For example, one or more inward-facing cameras 212 may be positioned to capture image data characterizing the position of the left and / or right eye of user 202. In some implementations, computer 208 may be connected to one or more antennas and / or other communication hardware that enables computer 208 to communicate with one or more other computing devices. For example, computerized glasses 204 may be connected to a Wi-Fi network, an LTE network, and / or may communicate via Bluetooth protocol, and / or any other communication modality.

[0026] In some implementations, one or more lenses 206 can operate as a display interface for rendering graphical content visible to a user wearing the computerized glasses 204. The graphical content rendered in the lenses 206 can assist in device arbitration in response to multiple devices detecting input from the user 202. For example, the user 202 can orient their head and computerized glasses 204 in a direction that places a first computing device and a second computing device within the field of view of the forward-facing camera 210. When the user 202 orients the computerized glasses 204 in this direction, the user 202 can provide verbal utterances to, for example, cause a particular computing device to play music from a music application. The automated assistant can detect the verbal utterances and, in response, cause multiple instances of an icon for the music application to be rendered in the lenses 206. For example, a first instance of a music application icon can be rendered in the lenses 206 on the first computing device, and a second instance of the music application icon can be rendered in the lenses 206 on the second computing device.

[0027] In some implementations, each instance of the music application icon may be rendered in a manner that indicates to the user that a particular device has not been selected to respond to a verbal utterance. For example, each instance of the music application icon may be “grayed out,” dimmed, blinking, and / or have one or more characteristics that otherwise indicate that one of the devices should be selected by the user. To select one of the devices, the user 202 may adjust their gaze and / or the direction of the computerized glasses 204 more toward the first computing device or the second computing device. In response, the automated assistant may detect the adjustment in the user's gaze and / or direction of orientation and cause the first or second instance of the music application icon to provide feedback that one has been selected. For example, if the user 202 turns their gaze and / or the computerized glasses 204 more toward the first computing device, the first instance of the music application may blink, shake, idle, no longer be grayed out, no longer be dimmed, and / or otherwise indicate that the first computing device has been selected. In this way, the user 202 receives feedback that the user has selected a particular device and can redirect their gaze and / or computerized glasses 204 if the user prefers the second computing device. In some implementations, if the user 202 is satisfied with their selection, the user 202 can continue to look at the first computing device or look away from both computing devices for a threshold period of time to confirm their selection and cause the first computing device to respond to verbal utterances.

[0028] In some implementations, graphical content rendered in the lens 206 can help clarify the parameters of a particular request submitted by the user to the automated assistant and / or another application. For example, the user 202 can provide a verbal utterance such as "play some music," and in response, the automated assistant can cause a first icon for a first music application and a second icon for a second music application to be rendered in the lens 206. The icons can be rendered on or near a particular computing device to which the user 202 is attending, and the icons are rendered to provide feedback to prompt the user 202 to select a particular music application for rendering the music. In some implementations, a timer can also be rendered in the lens 206 to indicate the amount of time the user has before a particular music application is selected. For example, the automated assistant can cause a particular icon to be rendered to provide visual feedback indicating that the music application corresponding to that icon will be selected by default if the user 202 does not provide additional input indicating whether they prefer one application over another.

[0029] In some implementations, the graphical content rendered in the lens 206 can correspond to parameters to be provided to an application in response to assistant input from the user 202. For example, in response to the verbal utterance "Play a new song," the automated assistant can cause a first graphical element and a second graphical element to be rendered within the lens 206 at or near a particular audio device. The first graphical element can include text identifying the name of the first song, and the second graphical element can include text identifying the name of the second song. In this manner, the user 202 can realize that there was some ambiguity in the verbal utterance the user provided and that additional input may be required to select a particular song. The user 202 can then provide additional input (e.g., adjust the user's gaze, rotate the user's head, perform a gesture, tap the housing 214 while gazing at a particular icon, provide another verbal utterance, and / or provide any other input) to specify the particular song. In some implementations, when a change in the orientation of the user 202 is detected, the graphical content rendered in the lenses 206 may be adjusted according to the change in orientation. For example, an icon rendered to appear over the computing device the user 202 is looking at may be shifted in the lenses 206 in a direction opposite to the direction the user 202 has turned their head. Similarly, when the computing device is no longer within the field of view of the user 202 and / or the computerized glasses 204, the icon may no longer be rendered in the lenses 206.

[0030] 3A, 3B, 3C, 3D, and 3E show views 300, 320, 340, 360, and 380, respectively, of a user 302 interacting with an automated assistant that can rely on computerized glasses 304 for device arbitration. These figures illustrate at least one example in which user 302 causes a particular action to be performed on a computing device in a room and then moves to another room, while maintaining the ability to provide assistant input to computerized glasses 304 to control the action. For example, user 302 can provide a verbal utterance 312 such as, "Assistant, play the footage from the security camera from last night." The verbal utterance 312 can include a request for the automated assistant to access a security application and render video data from the security application on a display device accessible to the automated assistant.

[0031] In some implementations, to determine the specific device on which the user 302 intends the video data to be rendered, the automated assistant can cause one or more devices to each provide one or more different outputs. The output can be detected by computerized glasses 304 worn by the user 302 when the user provides verbal utterance 312. For example, the automated assistant can identify one or more candidate devices that detected the verbal utterance 312. In some implementations, the automated assistant can cause each candidate device to provide an output. The output provided by each candidate device can be distinguished from the output provided by other candidate devices. In other words, each candidate device can be caused to provide a corresponding unique output. In some implementations, the automated assistant can cause the candidate device to render the output one or more times, each rendering for less than a given duration, such as less than one tenth of a second or less than 50 milliseconds. In some of these and / or other implementations, the output rendered by the candidate device can be an output that may not be detectable by a human without an artificial modality. For example, the automated assistant can cause the television 306 to incorporate an image into one or more frames of graphical content being rendered on the television 306 and the tablet device 310, respectively. The frames containing the image can be rendered at a higher frame frequency (e.g., 60 frames per second or higher) than is detectable by humans. In this way, if the user 302 is facing the television 306 when the user provides verbal utterance 312, the forward-facing camera of the computerized glasses 304 can detect the image in one or more frames without the user 302 being interrupted.Alternatively or additionally, if both the television 306 and the tablet device 310 are within the field of view of the forward-facing camera and / or user 320, the automated assistant may determine that the television 306 is occupying more of the focus of the user 302 and / or computerized glasses 304 than the tablet device 310.

[0032] In some implementations, the automated assistant can determine that the television 306 and tablet device 310 are near the computerized glasses 304 and / or the user 302 when the user 302 provides a verbal utterance. Based on this determination, the automated assistant can cause the television 306 and tablet device 310 to render different images to identify devices that are within the field of view of the cameras and / or computerized glasses 304. The automated assistant can then determine whether one or more cameras of the computerized glasses 304 detected one image but not another image to identify the particular device to which the user 302 is directing his or her input.

[0033] In some implementations, one or more devices in the environment may include LEDs that can be controlled by each corresponding device and / or automation assistant during device reconciliation. Light emitted from the LEDs may then be detected to select a particular device to respond to input from the user 302. For example, in response to a verbal utterance from the user 302 being detected at a first device and a second device, the automation assistant may illuminate a first LED on the first device and a second LED on the second device. If the automation assistant determines that light emitted from the first LED but not from the second LED is being detected by the computerized glasses 304, the automation assistant may select the first device to respond to the verbal utterance. Alternatively or additionally, the automation assistant may illuminate each LED to exhibit a particular characteristic to aid in device reconciliation. For example, a first LED may be illuminated such that the characteristics of the light emitted by the first LED are different from the characteristics of the light emitted by the second LED. Such characteristics may include color, amplitude, duration, frequency, and / or any other characteristic of light that can be controlled by an application and / or device. For example, the automation assistant may turn on a first LED for 0.1 seconds every 0.5 seconds and a second LED for 0.05 seconds every 0.35 seconds. In this manner, one or more cameras on the computerized glasses 304 may detect these light patterns, and the automation assistant may associate each detected pattern with a respective device. The automation assistant may then identify the LED at which the user 302 is gazing and / or pointing the computerized glasses 304 in order to select a particular device that will respond to the user input.

[0034] In some cases, the automated assistant can determine that the verbal utterance 312 is directed at the television 306 and can cause the television 306 to perform an operation 314 of playing camera footage from a security camera on the television 306. For example, the automated assistant can cause a security application to be accessed on the television 306 to render the security footage requested by the user 302. When the security application is launched, an icon 330 identifying the security application and selectable GUI elements 332 (i.e., graphical elements) can be rendered on the television 306. To identify particular features of the camera footage, the user 302 can hear audio 316 from the television 306 and view rendered video on the display interface of the television 306. Additionally, the user 302 can perform various physical gestures to control the television 306 and / or the security application via the computerized glasses 304. For example, the computerized glasses 304 can include an outward-facing camera 322 that can capture image data for processing by the automated assistant. For example, when the user 302 performs a swipe gesture 326, the automated assistant can detect the swipe gesture 326 and cause the security application to perform a particular action corresponding to the swipe gesture 326 (e.g., fast forward).

[0035] In some implementations, the user 302 may move to another room in the user's home to operate the computerized glasses 304 in a manner that reflects the user's 302 movement 344. For example, the user 302 may move from environment 308 to another environment 362 to change their laundry 364, while still hearing audio 316 from the security application, as shown in FIGS. 3C and 3D . The computerized glasses 304 may provide an interface for controlling the security application and / or an automated assistant in response to the user 302 moving away from the television 306. In some implementations, the determination that the user 302 has moved may be based on data generated in the computerized glasses and / or one or more other devices in the environment 308. In some implementations, the computerized glasses 304 may indicate that the user 302 can control applications and / or devices that the user 302 was previously viewing through the computerized glasses 304. For example, an icon 330 representing the security application may be rendered in the display interface of the computerized glasses 304. Alternatively or additionally, in response to the user 302 redirecting his / her gaze and / or face away from the television 306, the selectable GUI elements 332 and graphical elements 384 rendered on the television 306 may be rendered on the display interface of the computerized glasses 304.

[0036] In some implementations, device arbitration may be performed using the computerized glasses 304 when the user 302 is not looking at the device and / or may otherwise be directing their attention toward another computing device. For example, as provided in FIG. 3E , the user 302 may perform a gesture 382 to provide input to an automated assistant, a particular application, and / or a particular computing device. One or more cameras (e.g., outward-facing camera 322) of the computerized glasses may detect a physical gesture, which may be a physical gesture in which the user 302 manipulates the user's hand from a left position 368 toward the right. In response to detecting the physical gesture 382, ​​the automated assistant may operate to identify one or more devices that detected the physical gesture. In some cases, the automated assistant may determine that only the computerized glasses 304 detected the physical gesture 382. Nevertheless, the automated assistant may determine whether the physical gesture 382 was intended to initiate one or more actions in the computerized glasses 304 and / or another device.

[0037] In some cases, the user 302 may be viewing content from a security application on the display interface of the computerized glasses 304. In such a case, the automation assistant may determine that the physical gesture 382 is intended to affect the operation of the security application. In response to the physical gesture 382, ​​the automation assistant may cause the security application to fast-forward through particular content being rendered on the computerized glasses 304. Alternatively or additionally, the automation assistant may perform a heuristic process to identify the application and / or device toward which the user 302 is pointing the physical gesture 382. In some implementations, with prior permission from the user 302, the automation assistant may determine that the user 302 recently gazed at the television 306 and, prior to gazing at the television 306, the user 302 gazed at the tablet device 310. This determination may cause the automation assistant to prioritize the television 306 over the tablet device 310 when selecting a device to respond to the physical gesture 382.

[0038] Alternatively or additionally, the automation assistant may determine, based on image data from the inward-facing camera 324 of the computerized glasses 304, that the user 302 recently gazed at a security application on the television 306, and that prior to gazing at the security application, the user 302 gazed at a social media application on the tablet device 310. Based on distinguishing between the security application and the social media application, the automation assistant may determine that the physical gesture 382 is acceptable as input to the security application, but not as input to the social media application. Thus, according to this process, the automation assistant may select a security application to respond to a physical gesture input from the user 302.

[0039] In some implementations, device arbitration may be performed using one or more trained machine learning models that may be used to process application data and / or contextual data. The application data may, for example, characterize the operational state of one or more applications that may be associated with the user 302 when the user 302 provides the physical gesture 382. Alternatively or additionally, the contextual data may characterize features of the context in which the user 302 provided the physical gesture 382. Such features may include, but are not limited to, the location of the user 302, the time of day, one or more activities of the user (with prior permission from the user), and / or any other information that may be associated with the user 302 when the user 302 provides the physical gesture 382. For example, audio data captured by one or more microphones of the computerized glasses 304 and / or one or more other devices may be processed, with prior permission from the user 302, to identify contextual features of the environment. For example, audio data capturing sound from a movie may be used to assist the automated assistant in determining whether the physical gesture 382 should affect an application that is rendering the movie. If the automated assistant determines that the physical gesture 382 is intended to affect the movie (e.g., fast-forward the movie through a portion of the movie that the user does not want to hear), the automated assistant can generate command data that can be communicated to the application rendering the movie without the user necessarily having to gaze at the television 306 displaying the movie.

[0040] 4 shows a system 400 for performing device arbitration using data available from a device, such as computerized glasses, having capabilities for implementing augmented reality. An automated assistant 404 can operate as part of an assistant application provided on one or more computing devices, such as computing device 402 and / or a server device. A user can interact with the automated assistant 404 through an assistant interface 420, which can be a microphone, a camera, a touchscreen display, a user interface, and / or any other device capable of providing an interface between a user and an application. For example, a user can initialize the automated assistant 404 by providing verbal, textual, and / or graphical input to the assistant interface 420 to cause the automated assistant 404 to initiate one or more actions (e.g., provide data, control a peripheral device, access an agent, generate input and / or output, etc.). Alternatively, the automation assistant 404 may be initialized based on processing of contextual data 436 using one or more trained machine learning models, where the contextual data 436 may characterize one or more features of the environment accessible to the automation assistant 404 and / or one or more features of a user who is expected to intend to interact with the automation assistant 404.

[0041] The computing device 402 may include a display device, which may be a display panel including a touch interface for receiving touch input and / or gestures to enable a user to control applications 434 of the computing device 402 via a touch interface. In some implementations, the computing device 402 may lack a display device, thereby providing an audible user interface output without providing a graphical user interface output. Additionally, the computing device 402 may provide a user interface, such as a microphone, for receiving spoken natural language input from a user. In some implementations, the computing device 402 may include a touch interface and may lack a camera, but may optionally include one or more other sensors. In some implementations, the computing device 402 may provide augmented reality functionality and / or may be a wearable device, such as, but not limited to, computerized eyeglasses, contact lenses, a watch, an article of clothing, and / or any other wearable device. Thus, although various implementations are described herein with respect to computerized eyeglasses, the techniques disclosed herein may be implemented in conjunction with other electronic devices that include augmented reality functionality, such as other wearable devices that are not computerized eyeglasses.

[0042] The computing device 402 and / or other third-party client devices can communicate with the server device over a network such as the Internet. In addition, the computing device 402 and any other computing devices can communicate with each other over a local area network (LAN), such as a Wi-Fi network. The computing device 402 can offload computational tasks to the server device to conserve computational resources at the computing device 402. For example, the server device can host the automation assistant 404, and / or the computing device 402 can send inputs received at one or more assistant interfaces 420 to the server device. However, in some implementations, the automation assistant 404 can be hosted at the computing device 402, and various processes that can be associated with automation assistant operations can be executed at the computing device 402.

[0043] In various implementations, all or less than all aspects of the automation assistant 404 may be implemented in the computing device 402. In some of those implementations, aspects of the automation assistant 404 are implemented via the computing device 402 and may interface with a server device that may implement other aspects of the automation assistant 404. The server device may optionally provide services to multiple users and their associated assistant applications via multiple threads. In implementations in which all or less than all aspects of the automation assistant 404 are implemented via the computing device 402, the automation assistant 404 may be an application separate from (e.g., installed “on top of”) the operating system of the computing device 402, or alternatively, may be implemented directly by (e.g., considered an operating system application but integral with) the operating system of the computing device 402.

[0044] In some implementations, the automation assistant 404 can include an input processing engine 406 that can employ multiple different modules for processing input and / or output of the computing device 402 and / or the server device. For example, the input processing engine 406 can include a speech processing engine 408 that can process audio data received at the assistant interface 420 to identify text embodied in the audio data. The audio data can be transmitted from the computing device 402 to a server device, for example, to conserve computational resources at the computing device 402. Additionally or alternatively, the audio data can be processed exclusively at the computing device 402.

[0045] The process for converting audio data to text may include a speech recognition algorithm that may employ neural networks and / or statistical models to identify groups of audio data that correspond to words or phrases. The text converted from the audio data may be analyzed by the data analysis engine 410 and made available to the automation assistant 404 as text data that may be used to generate and / or identify command phrases, intents, actions, slot values, and / or any other content specified by the user. In some implementations, output data provided by the data analysis engine 410 may be provided to the parameter engine 412 to determine whether the user has provided input that corresponds to a particular intent, action, and / or routine that may be executed by the automation assistant 404 and / or an application or agent that may be accessed via the automation assistant 404. For example, assistant data 438 may be stored on the server device and / or computing device 402 and may include data defining one or more actions that may be executed by the automation assistant 404, as well as parameters required to execute the action. The parameter engine 412 may generate one or more parameters for the intent, action, and / or slot value and provide the one or more parameters to the output generation engine 414. The output generation engine 414 can use one or more parameters to communicate with an assistant interface 420 to provide output to a user and / or to communicate with one or more applications 434 to provide output to the one or more applications 434.

[0046] In some implementations, the automation assistant 404 may be installed “on” the computing device 402 and / or may itself be an application that can form part of (or the entire) the operating system of the computing device 402. The automation assistant application may include and / or access on-device speech recognition, on-device natural language understanding, and on-device fulfillment. For example, on-device speech recognition may be performed using an on-device speech recognition module that processes audio data (detected by a microphone) using an end-to-end speech recognition machine learning model stored locally on the computing device 402. The on-device speech recognition generates recognized text for spoken speech (if any) present in the audio data. Also, for example, on-device natural language understanding (NLU) may be performed using an on-device NLU module that processes the recognized text generated using on-device speech recognition and, optionally, contextual data, to generate NLU data.

[0047] The NLU data can include an intent corresponding to the verbal utterance and, optionally, parameters related to the intent (e.g., slot values). On-device fulfillment can be performed using an on-device fulfillment module that utilizes the NLU data (from the on-device NLU) and, optionally, other local data, to determine actions to take to resolve the intent (and, optionally, parameters related to the intent) of the verbal utterance. This can include determining local and / or remote responses (e.g., answers) to the verbal utterance, interactions with locally installed applications to perform based on the verbal utterance, commands to send to an Internet of Things (IoT) device (directly or via a corresponding remote system) based on the verbal utterance, and / or other resolution actions to perform based on the verbal utterance. The on-device fulfillment can then initiate local and / or remote fulfillment / execution of the determined actions to resolve the verbal utterance.

[0048] In various implementations, remote speech processing, remote NLU, and / or remote fulfillment may be utilized, at least selectively. For example, recognized text may be at least selectively sent to a remote automated assistant component for remote NLU and / or remote fulfillment. For example, recognized text may optionally be sent for remote fulfillment in parallel with on-device fulfillment or in response to a failure of on-device NLU and / or on-device fulfillment. However, on-device speech processing, on-device NLU, on-device fulfillment, and / or on-device execution may be prioritized at least due to the reduced latency they offer when resolving verbal utterances (because client-server round trips are not required to resolve the verbal utterance). Furthermore, on-device functionality may be the only functionality available in situations where network connectivity is absent or limited.

[0049] In some implementations, the computing device 402 may include one or more applications 434 that may be provided by a third-party entity different from the entity that provided the computing device 402 and / or the automation assistant 404. The application state engine of the automation assistant 404 and / or the computing device 402 may access the application data 430 to determine one or more actions that may be performed by the one or more applications 434, as well as the state of each application of the one or more applications 434 and / or the state of each device associated with the computing device 402. The device state engine of the automation assistant 404 and / or the computing device 402 may access the device data 432 to determine one or more actions that may be performed by the computing device 402 and / or one or more devices associated with the computing device 402. Additionally, the application data 430 and / or any other data (e.g., device data 432) may be accessed by the automation assistant 404 to generate context data 436 that can characterize the context in which a particular application 434 and / or device is running, and / or the context in which a particular user is accessing the computing device 402, accessing the application 434, and / or accessing any other device or module.

[0050] While one or more applications 434 are executing on the computing device 402, the device data 432 may characterize the current operational state of each application 434 executing on the computing device 402. Additionally, the application data 430 may characterize one or more features of the executing applications 434, such as the content of one or more graphical user interfaces being rendered in the direction of the one or more applications 434. Alternatively or additionally, the application data 430 may characterize action schemas that may be updated by the respective applications and / or by the automation assistant 404 based on the respective applications' current operational state. Alternatively or additionally, the one or more action schemas for one or more applications 434 may remain static but may be accessed by the application state engine to determine suitable actions to initialize via the automation assistant 404.

[0051] The computing device 402 may further include an assistant invocation engine 422 that can use one or more trained machine learning models to process the application data 430, the device data 432, the context data 436, and / or any other data accessible to the computing device 402. The assistant invocation engine 422 can process this data to determine whether to wait for the user to explicitly speak an invocation phrase to invoke the automated assistant 404, or to determine whether to consider the data as indicative of the user's intent to invoke the automated assistant, instead of requiring the user to explicitly speak an invocation phrase. For example, the one or more trained machine learning models may be trained using instances of training data based on scenarios in which a user is in an environment in which multiple devices and / or applications exhibit various operating states. Instances of training data may be generated to capture training data that characterize contexts in which a user invokes an automated assistant and other contexts in which the user does not invoke an automated assistant.

[0052] When the one or more trained machine learning models are trained according to these instances of training data, the assistant invocation engine 422 can cause the automated assistant 404 to detect or limit detection of a spoken invocation phrase from the user based on contextual and / or environmental features. Additionally or alternatively, the assistant invocation engine 422 can cause the automated assistant 404 to detect or limit detection of one or more assistant commands from the user based on contextual and / or environmental features. In some implementations, the assistant invocation engine 422 can be disabled or limited based on the computing device 402 detecting an assistant suppression output from another computing device. In this manner, if the computing device 402 detects an assistant suppression output, the automated assistant 404 is not invoked based on contextual data 436 that would otherwise cause the automated assistant 404 to be invoked if the assistant suppression output is not detected.

[0053] In some implementations, system 400 may include a device arbitration engine 416 that can assist one or more devices and / or applications in performing device arbitration when they detect input from a user. For example, in some implementations, device arbitration engine 416 can process data from one or more different devices to determine whether to initiate a device arbitration process. In some implementations, the data may be received via a network connection, one or more interfaces of system 400, and / or any other modality through which a computing device can receive data. For example, device arbitration engine 416 may determine that multiple different devices are projecting ultrasound and / or light in response to an assistant input. Based on this determination, device arbitration engine 416 can initiate a process for selecting a particular device among the multiple different devices that will respond to the assistant input from the user.

[0054] In some implementations, the system 400 can include a gaze detection engine 418 that can determine a user's gaze toward one or more objects in the user's environment. For example, the system 400 can be computerized eyeglasses that include an inward-facing camera aimed at one or more of the user's eyes. Based on image data generated using the inward-facing camera, the system 400 can identify a particular area in the environment to which the user is directing their eyes. In some implementations, the gaze detection engine 418 can determine the direction of the user's gaze based on data from one or more different devices, such as a separate computing device that includes a camera. The image data from the separate computing device can, with prior permission from the user, indicate the user's posture and / or the direction in which the user is directing one or more of the user's appendages. In this manner, the gaze detection engine 418 can determine the direction in which the user is directing their attention before, during, and / or after the user provides input to the automated assistant 404.

[0055] In some implementations, the system 400 may include a field of view engine 426 that can process data characterizing the field of view of one or more cameras of a user and / or a device. For example, the field of view engine 426 may process image data from one or more cameras of computerized glasses to identify one or more objects and / or devices located within the field of view of the cameras at one or more instances in time. In some implementations, the field of view engine 426 may also process device data 432 to identify particular objects that may be associated with a particular device within the user's field of view. For example, a kitchen sink may be an object associated with the user's stand-alone display device. Thus, the field of view engine 426 may determine that the stand-alone computing device is the subject of the user input if a kitchen sink is identified within the user's field of view when the user provides the user input.

[0056] In some implementations, the system 400 may include an interface content engine 424 that causes one or more interfaces of the system 400 to render content according to output from the device arbitration engine 416. For example, if the device arbitration engine 416 identifies the computing device 402 as the target of input from a user, the interface content engine 424 may cause content to be rendered in one or more interfaces of the computing device 402 (e.g., a display interface of computerized eyeglasses). If the device arbitration engine 416 determines that the user is directing input to a separate computing device and that the separate computing device is within the user's field of view, the interface content engine 424 may cause a notification to be rendered for the separate computing device. For example, the interface content engine 424 may cause graphical content to be rendered in a display interface of the computing device 402 if the computing device 402 is computerized eyeglasses. The graphical content may be rendered such that the graphical content appears “on” and / or adjacent to the separate computing device within the user's field of view (e.g., within an area of ​​the eyeglass lenses corresponding to the location of the separate device). The graphical content may include, but is not limited to, one or more icons, colors, and / or other graphical characteristics that may indicate that the device arbitration engine 416 has selected a separate computing device as responsive to input from the user.

[0057] In some implementations, the device arbitration engine 416 may request additional input from the user to assist the user in identifying the specific device to which the user intended to direct input. The device arbitration engine 416 may communicate the identifiers and / or locations of the candidate devices to the interface content engine 424, which may render a graphical representation within the lenses of the computerized glasses at or near the relative locations of the candidate devices. For example, a first selectable element may be rendered in the leftmost portion of the display interface of the computerized glasses to indicate that a computing device in the leftmost portion of the user's field of view is a candidate device. A second selectable element may be rendered in a more central position of the display interface simultaneously with the first selectable element to indicate that another computing device in the central portion of the user's field of view is also a candidate device. The user may then perform a gesture (e.g., holding up their index finger) in front of their face for the gesture to be captured by one or more cameras of the computerized glasses. The gesture can indicate to the automated assistant 404 that a particular device (eg, the most central device) is the device that the user intended to respond to the user input.

[0058] FIG. 5 illustrates a method 500 for performing device arbitration in a multi-device environment using data available from a wearable computing device, such as computerized glasses. Method 500 may be performed by one or more computing devices, applications, and / or any other apparatus or module that may be associated with an automated assistant. Method 500 may include operation 502 of determining whether assistant input is detected. The assistant input may be user input provided by a user to one or more computing devices that provide access to the automated assistant. In some implementations, the assistant input may be provided to a computing device with augmented reality capabilities, such as computerized glasses, computerized contact lenses, a phone, a tablet device, a portable computing device, a smartwatch, and / or any other computing device that can augment the perception of one or more users. It should be noted that in implementations discussed herein, including computerized glasses, the computerized glasses may be any computing device that provides augmented reality capabilities. If assistant input is detected, method 500 may proceed from operation 502 to operation 504. Otherwise, the automated assistant may continue to determine whether the user has provided input to the automated assistant.

[0059] Operation 504 may be an optional operation that includes determining whether the assistant input is detected at multiple devices. For example, a user may be wearing a pair of computerized glasses while gazing at a stand-alone speaker device that provides access to an automated assistant. Thus, when the user provides an assistant input, the assistant input may be detected at the computerized glasses, the stand-alone speaker device, and one or more other devices in the user's environment. If a single computing device exclusively detects the assistant input, method 500 may proceed from operation 504 to operation 512, where the particular computing device that detected the assistant input initiates performance of one or more operations. However, if multiple devices detect the assistant input, method 500 may proceed from operation 504 to operation 506 to initiate device arbitration to select one or more devices to respond to the assistant input.

[0060] Operation 506 can include identifying a candidate device that detected an assistant input from the user. For example, if the assistant input is a verbal utterance, the automated assistant can identify multiple devices that captured audio data corresponding to the verbal utterance. In some implementations, each candidate device can capture audio data and process the audio data to generate a score for the assistant input. If each score for each candidate device meets a threshold, device arbitration can be initiated for those candidate devices. For example, candidate devices can include computerized glasses, a stand-alone speaker device, and a tablet device that the user may have placed on a table near the user. Method 500 can proceed from operation 506 to operation 508 for a further determination of whether the assistant input was directed to a particular computing device.

[0061] Operation 508 can include processing data generated at the candidate devices and / or the computerized glasses. For example, the computerized glasses can include an outward-facing camera and / or an inward-facing camera that can be used to generate image data for identifying a particular device to which the user may be directing assistant input. The image data generated using the outward-facing camera can capture an image including one of the candidate devices or an object associated with the candidate device. For example, a stand-alone speaker device can be supported by a decoration table that the user was looking at when the user provided assistant information. As a result, if the decoration table and / or the stand-alone speaker device are determined to be within the field of view of the user and / or the computerized glasses, the automated assistant can determine that the assistant input was directed to the stand-alone speaker device. Alternatively or additionally, one or more of the candidate devices can provide an output that can be detected by one or more sensors of the computerized glasses. Thus, if the computerized glasses detect an output from a particular one of the candidate devices, the automated assistant can determine that the assistant input was directed to that particular device.

[0062] Method 500 may proceed to operation 510, which may include determining, based on processing in operation 508, that the assistant input was directed to a particular candidate device. If the automated assistant determines that the assistant input is directed to a particular candidate device separate from the computerized glasses, method 500 may proceed from operation 510 to operation 512. Alternatively, if the automated assistant and / or another application determines that the assistant input is not directed to a particular candidate device, method 500 may proceed from operation 510 to operation 514.

[0063] Operation 514 may include determining whether the system input is directed to the computerized glasses and / or another computing device that provides augmented reality functionality. For example, a user may provide verbal utterances while handwriting in a notebook while wearing the computerized glasses. Thus, even if other candidate devices detect verbal utterances from the user, none of the other candidate devices may be visible within the field of view of the computerized glasses. As a result, the automated assistant may determine that the user intended for the automated assistant to respond to the verbal utterance through the computerized glasses. For example, the user may request that the automated assistant check the spelling of words written in the notebook while the user gazes at the notebook and wears the computerized glasses. In response, the automated assistant may cause the computerized glasses to render one or more graphical elements indicating whether the words are spelled correctly in the notebook via augmented reality and / or provide an audible output through one or more speakers of the computerized glasses.

[0064] If the automated assistant determines that the assistant input is directed to the computerized glasses, method 500 may proceed from operation 514 to operation 512. Otherwise, method 500 may proceed from operation 514 to operation 516. Operation 516 may be an optional operation that includes the automated assistant requesting additional input from the user to assist the automated assistant in identifying the particular computing device to which the user intended the assistant input to be directed. Method 500 may then proceed from operation 516 to operation 502 and / or operation 504.

[0065] 6 is a block diagram 600 of an exemplary computer system 610. Computer system 610 typically includes at least one processor 614 that communicates with several peripheral devices via a bus subsystem 612. These peripheral devices may include, for example, a storage subsystem 624 including memory 625 and a file storage subsystem 626, a user interface output device 620, a user interface input device 622, and a network interface subsystem 616. The input and output devices enable user interaction with computer system 610. Network interface subsystem 616 provides an interface to external networks and is coupled to corresponding interface devices in other computer systems.

[0066] The user interface input devices 622 may include a keyboard, a pointing device such as a mouse, a trackball, a touchpad, or a graphical tablet, a scanner, a touchscreen integrated into a display, an audio input device such as a voice recognition system, a microphone, and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computer system 610 or a communications network.

[0067] The user interface output devices 820 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for producing a visible image. The display subsystem may also provide a non-visual display, such as via an audio output device. In general, use of the term "output device" is intended to encompass all possible types of devices and methods for outputting information from the computer system 610 to a user or to another machine or computer system.

[0068] The storage subsystem 624 stores programming and data structures that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 624 may include logic for performing selected aspects of the method 500 and / or for implementing one or more of the system 400, the computerized glasses 104, the television 106, the tablet device 110, the computing device 122, the computerized glasses 204, the television 306, the computerized glasses 304, the tablet device 310, the computing device 342, and / or any other applications, devices, apparatuses, and / or modules discussed herein.

[0069] These software modules are generally executed by the processor 614 alone or in combination with other processors. The memory 625 used in the storage subsystem 624 can include several memories, including a main random access memory (RAM) 630 for storing instructions and data during program execution and a read-only memory (ROM) 632 in which fixed instructions are stored. The file storage subsystem 626 can provide persistent storage for program files and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of a particular implementation may be stored by the file storage subsystem 626 in the storage subsystem 624 or by another machine accessible by the processor 614.

[0070] The bus subsystem 612 provides a mechanism for allowing the various components and subsystems of the computing device 610 to communicate with each other as intended. Although the bus subsystem 612 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0071] The computer system 610 can be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of the computer system 610 shown in Figure 6 is intended only as a specific example for purposes of illustrating some implementations. Many other configurations of the computer system 610 can have more or fewer components than the computer system shown in Figure 6.

[0072] In situations where the systems described herein may collect or utilize personal information about users (or sometimes referred to herein as “participants”), users may be provided with the opportunity to control whether a program or feature collects user information (e.g., information about the user's social network, social actions or activities, occupation, user preferences, or the user's current geographic location) or whether and / or how to receive content from content servers that may be more relevant to the user. Also, certain data may be processed in one or more ways before it is stored or used so that personally identifiable information is removed. For example, a user's identification information may be processed so that personally identifiable information cannot be determined about the user, or the user's geographic location may be generalized where the geographic location information is obtained (e.g., to the city, zip code, or state level) so that the user's specific geographic location cannot be identified. Thus, users may have control over how information is collected and / or used about them.

[0073] While several implementations have been described and illustrated herein, various other means and / or structures for performing the functions and / or obtaining one or more of the results and / or advantages described herein may be utilized, and each such variation and / or modification is deemed to be within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are meant to be exemplary, and the actual parameters, dimensions, materials, and / or configurations will depend on the specific application or applications for which the teachings are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific implementations described herein. Accordingly, the foregoing implementations are presented by way of example only, and it should be understood that, within the scope of the appended claims and their equivalents, implementations may be practiced otherwise than as specifically described and claimed. Implementations of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. Additionally, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the present disclosure, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.

[0074] In some implementations, a method implemented by one or more processors is shown as including an operation of determining that a user has directed an assistant input to an automated assistant accessible through any one computing device of a plurality of computing devices connected to a network, wherein the user is wearing computerized glasses, the computerized glasses being a computing device including one or more cameras, and the user is located in an environment including the plurality of computing devices. The method may further include an operation of identifying, based on the assistant input from the user, two or more candidate devices of the plurality of computing devices that have detected the assistant input from the user, wherein the two or more candidate devices are separate from the computerized glasses. The method may further include an operation of determining whether the assistant input is directed to a particular computing device of the two or more candidate devices or to the computerized glasses based on processing image data generated using one or more cameras of the computerized glasses. In some implementations, the method may further include an operation of causing the particular computing device to perform one or more operations corresponding to the assistant input if it is determined that the assistant input is directed to a particular computing device of the two or more candidate devices.

[0075] In some implementations, determining whether the assistant input is directed to a particular computing device of the two or more candidate devices or to the computerized glasses includes determining, based on the image data, whether the particular computing device of the two or more candidate devices is located within a field of view window of one or more cameras. In some implementations, a camera of the one or more cameras of the computerized glasses is directed toward the user's eye, and determining that the particular computing device of the two or more candidate devices is associated with the user's visual viewpoint includes determining that the user's gaze is directed toward the particular computing device rather than toward any other device of the two or more candidate devices.

[0076] In some implementations, determining whether the assistant input is directed to a particular computing device among the two or more candidate devices or to the computerized glasses includes determining whether a particular object within a field of view of one or more cameras is associated with a relative position of the particular computing device among the two or more candidate devices. In some implementations, the particular computing device is not located within a field of view of one or more cameras of the computerized glasses. In some implementations, the assistant input is a physical gesture performed by a user, and the physical gesture is detected by the computerized glasses. In some implementations, the computerized glasses include a graphical display interface that is rendering content when the user provides the assistant input, and the method further includes, if it is determined that the assistant input is directed to the computerized glasses, changing the content rendered in the graphical display interface of the computerized glasses according to the physical gesture.

[0077] In some implementations, the method further includes, when it is determined that the assistant input is directed to the computerized glasses and the specific computing device, rendering a first portion of the content on the specific computing device and rendering a second portion of the content on a display interface of the computerized glasses. In some implementations, the specific computing device is a detachable accessory device connected to the display device, and the specific computing device is not visible within a field of view window of one or more cameras of the computerized glasses. In some implementations, the computerized glasses include a graphical display interface that is at least partially transparent when the graphical display interface is rendering the content, and when it is determined that the assistant input is directed to a specific computing device of the two or more candidate devices, rendering a graphical element at a position in the graphical display interface corresponding to the specific computing device.

[0078] In some implementations, the assistant input includes a request for the automated assistant to initialize a particular application, and the graphical element is based on the particular application. In some implementations, identifying two or more candidate devices of the plurality of computing devices that detected the assistant input from the user includes determining that a particular computing device is rendering a first output and that another computing device of the plurality of computing devices is rendering a second output, wherein the first output and the second output are detected by computerized glasses.

[0079] In some implementations, the particular computing device includes a graphical display interface, and the first output includes a graphical element rendered in the graphical display interface, the graphical element being embodied in one or more graphical content frames rendered at a frequency of 60 frames per second or greater. In some implementations, the first output is different from the second output, and determining whether the assistant input is directed to a particular computing device of the two or more candidate devices or to the computerized glasses includes determining that the first output is detected within a field of view window of the computerized glasses and that the second output is not detected within a field of view window of the computerized glasses.

[0080] In another implementation, a method implemented by one or more processors is shown as including an operation such as determining, by a computing device, that a user has provided input to an automated assistant accessible through one or more computing devices located in the user's environment, the input corresponding to a request for the automated assistant to provide content to the user, the one or more computing devices including computerized glasses that the user is wearing when the user provides the input. The method may further include an operation of identifying a specific device for rendering the content for the user based on the input from the user, the specific device being separate from the computerized glasses. The method may further include an operation of causing the specific device to render the content for the user based on identifying the specific device. The method may further include an operation of processing contextual data provided by one or more computing devices that were in the user's environment when the user provided the input. The method may further include an operation of determining whether to provide the user with additional content associated with the request based on the contextual data. If the automated assistant determines to provide the user with the additional content, the method may further include an operation of causing the computerized glasses to perform one or more additional operations to facilitate rendering the additional content via one or more interfaces of the computerized glasses.

[0081] In some implementations, the contextual data includes image data provided by one or more cameras of the computerized glasses, and determining whether to provide the user with additional content associated with the request includes determining whether the user is viewing the particular device while the particular device is performing the one or more operations. In some implementations, causing the computerized glasses to render the additional content includes causing the computerized glasses to access content data over a network connection and causing a display interface of the computerized glasses to render one or more graphical elements based on the content data.

[0082] In yet another implementation, a method implemented by one or more processors is shown as including an act such as determining, by a computing device, that a user has provided input to an automated assistant accessible via the computing device, the input corresponding to a request for the automated assistant to perform one or more actions. The method may further include an act of receiving, by the computing device, contextual data indicating that the user is wearing computerized glasses, the computerized glasses being separate from the computing device. The method may further include an act of causing an interface of the computing device to render output that can be detected in another interface of the computerized glasses based on the contextual data. The method may further include an act of determining whether the computerized glasses have detected output from the computing device. The method may further include an act of causing the computing device to perform one or more actions to facilitate fulfilling the request if the computing device determines that the computerized glasses have detected output.

[0083] In some implementations, determining whether the computerized glasses have detected output from the computing device includes processing other contextual data indicative of whether one or more cameras of the computerized glasses have detected output from the computing device. In some implementations, the method may further include, if the computing device determines that the computerized glasses have detected output, causing the computerized glasses to render one or more graphical elements that may be selected in response to a physical gesture from the user. [Explanation of symbols]

[0084] 100 views 102 users 104 Computerized Glasses 106 Television 108 Environment 110 tablet devices 112 Field of view 114 Oral Speech 116 Movement, TV 118 Display Devices 120 Views 122 Computing Devices 124 Field of View 126 Dongle 128 operation 130 Oral Speech 200 views 202 users 204 Computerized Glasses 206 Lens 208 Computer 210 forward-facing camera 212 Inward-facing camera 214 Housing 300 views 302 users 304 Computerized Glasses 306 Television 308 Environment 310 tablet devices 312 Oral Utterances 314 operation 316 Audio 320 Views 322 Outward Facing Camera 326 Swipe Gestures 330 Icons 332 Selectable GUI Elements 340 Views 344 Move 360 Views 362 Environment 364 Laundry 368 left position 380 Views 382 Gestures, Physical Gestures 384 Graphical Elements 400 System 402 Computing Devices 404 Automation Assistant 406 Input Processing Engine 408 Voice Processing Engine 410 Data Analysis Engine 412 Parameter Engine 414 Output Generation Engine 416 Device Arbitration Engine 418 Gaze Detection Engine 420 Assistant Interface 422 Assistant Call Engine 424 Interface Content Engine 426 Vision Engine 430 Application Data 432 Device Data 434 Applications 436 Context Data 438 Assistant Data 600 Block Diagram 610 Computer Systems 612 Bus Subsystem 614 processor 616 Network Interface Subsystem 620 User Interface Output Device 622 User Interface Input Devices 624 Memory Subsystem 625 memory 626 File Storage Subsystem 630 Main Random Access Memory (RAM) 632 read-only memory

Claims

1. 1. A method implemented by one or more processors, the method comprising: determining that a user has directed an assistant input to an automated assistant accessible via any of a plurality of computing devices connected to a network; the user wearing computerized glasses, the computerized glasses being a computing device including one or more cameras; the user being located within an environment including the plurality of computing devices; identifying one or more first characteristics of a first light comprising a first frequency emitted by a first computing device of the plurality of computing devices and one or more second characteristics of a second light comprising a second frequency emitted by a second computing device of the plurality of computing devices; the one or more first characteristics are different from the one or more second characteristics; the first computing device and the second computing device are separate from the computerized glasses; determining, based on processing image data generated using the one or more cameras of the computerized glasses, that the one or more first characteristics of the first light emitted from the first computing device are detected in the image data; determining that the assistant input is directed to the first computing device, determining that the assistant input is directed to the first computing device based on determining that the one or more first characteristics of the first light are detected within the image data; In response to determining that the assistant input is directed to the first computing device, causing the first computing device to perform one or more actions corresponding to the assistant input; A method comprising:

2. causing the first computing device to emit the first light having the first characteristic and the second computing device to emit the second light having the second characteristic in response to the assistant input. further comprising 2. The method of claim 1, wherein identifying that the one or more first characteristics of the first light are emitted by the first computing device and the one or more second characteristics of the second light are emitted by the second computing device is based on causing the first computing device to emit the first light having the first characteristics and causing the second computing device to emit the second light having the second characteristics.

3. the assistant input comprises a verbal utterance; 3. The method of claim 2, wherein causing the first computing device to emit the first light having the first characteristic and causing the second computing device to emit the second light having the second characteristic is in response to the verbal utterance being detected at the first computing device and the verbal utterance being detected at the second computing device.

4. The method of claim 1 , wherein the first light emitted by the first computing device is emitted from one or more light emitting diodes (LEDs) of the first computing device.

5. The method of claim 1 , wherein the one or more first characteristics of the first light further include color, amplitude, and / or duration.

6. 10. The method of claim 1, wherein determining that the assistant input is directed toward the first computing device is further based on a detected gaze of the user being directed toward the first computing device.

7. 1. A method implemented by one or more processors, the method comprising: determining that a user has directed an assistant input to an automated assistant accessible via a plurality of computing devices connected to a network; the user wearing computerized glasses, the computerized glasses being a computing device including one or more cameras; the user being located within an environment including the plurality of computing devices; identifying, for each of one or more computing devices of the plurality of computing devices, corresponding light emitted by the computing device as having one or more corresponding characteristics including frequency; the one or more corresponding characteristics for each of the one or more computing devices of the plurality of computing devices are different from the one or more corresponding characteristics for another of the one or more computing devices of the plurality of computing devices; the one or more computing devices of the plurality of computing devices are separate from the computerized glasses; determining, based on processing image data generated using the one or more cameras of the computerized glasses, that the one or more corresponding characteristics of the corresponding light emitted by the one or more computing devices of the plurality of computing devices are not detected in the image data; determining that the assistant input is directed to the computerized glasses, determining that the assistant input is directed to the computerized glasses based on determining that the one or more corresponding characteristics of the corresponding light emitted by the one or more computing devices of the plurality of computing devices are not detected in the image data; in response to determining that the assistant input is directed to the computerized glasses, causing the computerized glasses to perform one or more actions corresponding to the assistant input; A method comprising:

8. 8. The method of claim 7, wherein the corresponding light emitted by a given computing device of the one or more computing devices is emitted from one or more light emitting diodes (LEDs) of the given computing device.

9. 8. The method of claim 7, wherein the one or more corresponding characteristics of the corresponding light emitted by a given one of the one or more computing devices includes color, amplitude, and / or duration.

10. 8. The method of claim 7, wherein determining that the assistant input is directed to the computerized glasses is further based on the user's detected gaze being directed toward the one or more computing devices of the multiple computing devices.

11. 8. The method of claim 7, further comprising, in response to the assistant input, causing each of one or more computing devices of the plurality of computing devices to emit the corresponding light having the one or more corresponding characteristics.

12. the assistant input comprises a verbal utterance; 12. The method of claim 11, wherein causing the one or more computing devices to emit the corresponding light is in response to the verbal utterance being detected at the one or more computing devices.

13. 1. A method implemented by one or more processors, the method comprising: determining that a user has directed an assistant input to an automated assistant accessible via a plurality of computing devices connected to a network; the user wearing computerized glasses, the computerized glasses being a computing device including one or more cameras; the user being located within an environment including the plurality of computing devices; identifying one or more first graphical renderings rendered at a first frequency by a first computing device of the plurality of computing devices and one or more second graphical renderings rendered at a second frequency by a second computing device of the plurality of computing devices; the one or more first graphical renderings by the first computing device provide a unique identification for the first computing device; the one or more second graphical renderings by the second computing device provide a unique identification for the second computing device; the first computing device and the second computing device are separate from the computerized glasses; determining, based on processing image data generated using the one or more cameras of the computerized glasses, that the one or more first graphical renderings of the first computing device are detected within the image data; determining that the assistant input is directed to the first computing device, determining that the assistant input is directed to the first computing device is based on determining that the one or more first graphical renderings by the first computing device are detected in the image data; In response to determining that the assistant input is directed to the first computing device, causing the first computing device to perform one or more actions corresponding to the assistant input; A method comprising:

14. 14. The method of claim 13, wherein determining that the assistant input is directed toward the first computing device is further based on a detected gaze of the user being directed toward the first computing device.

15. The method of claim 13 , wherein the first frequency is greater than or equal to 60 frames per second.

Citation Information

Patent Citations

  • Gaze Application Launcher

    JP2017538218A

  • Object identification method, device and virtual reality device in virtual reality communication

    JP2018526690A

  • Controller and agent cooperation method

    JP2019133216A

  • Gesture recognition system and glasses with gesture recognition function

    US20140009623A1

  • Display device including touch sensor and touch sensing method for the same

    US20180113549A1