Invoking the Automated Assistant's functions based on detected gestures and gazes
By detecting gestures and directed gazes, automated assistants are activated efficiently, reducing resource consumption and enabling contactless interaction, thus optimizing network and computing usage.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-10-29
- Publication Date
- 2026-03-18
AI Technical Summary
Existing automated assistants require explicit user interface inputs or verbal invocation phrases to activate, which can be inefficient and resource-intensive, especially when contactless interaction is desired.
Automated assistants are invoked through the detection of specific gestures and directed gazes, using local machine learning models to process visual and audio data selectively, reducing the need for continuous network and resource usage.
This approach allows for efficient, contactless interaction with automated assistants, conserving network capacity and computing resources while enabling timely and targeted data transmission.
Smart Images

Figure 0007833011000001 
Figure 0007833011000002 
Figure 0007833011000003
Abstract
Description
Technical Field
[0001] The present invention relates to the invocation of functions of an automated assistant based on detected gestures and gazes.
Background Art
[0002] A person may engage in a conversation between a human and a computer using an interactive software application (also called a "digital agent", "interactive personal assistant", "intelligent personal assistant", "assistant application", "conversational agent", etc.) referred to herein as an "automated assistant". For example, a person (who may be called a "user" when interacting with an automated assistant) may use oral natural language input (i.e., speech), which may be converted to text and then processed, and / or provide text (e.g., typed) natural language input to give commands and / or requests to the automated assistant. The automated assistant responds to the requests by providing a response user interface output that may include audible and / or visual user interface output.
[0003] As mentioned above, many automated assistants are configured to interact through verbal utterances. To protect user privacy and / or conserve resources, users often have to explicitly invoke the automated assistant before it can fully process the verbal utterances. Explicit invocation of an automated assistant generally occurs in response to specific user interface inputs being received on the client device. The client device provides the user of the client device with an interface to interface with the automated assistant (e.g., receiving input from the user and providing audible and / or graphical responses), and includes an assistant interface that interfaces with one or more additional components that implement the automated assistant (e.g., a remote server device that processes user input and generates appropriate responses).
[0004] Some user interface inputs that allow an automated assistant to be invoked via a client device include hardware and / or virtual buttons on the client device for invoking the automated assistant (e.g., tapping a hardware button, selecting a graphical interface element displayed by the client device). Many automated assistants may also be invoked in response to one or more verbal invocation phrases, also known as “hotwords / phrases” or “trigger words / phrases.” For example, verbal invocation phrases such as “Hey, Assistant,” “Okay, Assistant,” and / or “Assistant” may be spoken aloud to invoke the automated assistant. [Overview of the Initiative] [Means for solving the problem]
[0005] As described above, many client devices that facilitate interaction with automated assistants—also referred to herein as “assistant devices”—enable users to engage in contactless interaction with automated assistants. For example, assistant devices often include a microphone that allows users to give voice utterances to summon and / or otherwise interact with automated assistants. Assistant devices described herein may additionally or alternatively incorporate and / or communicate with one or more visual components (e.g., cameras, light detection and ranging (LIDAR) components, radar components, etc.) to facilitate contactless interaction with automated assistants.
[0006] Implementations disclosed herein relate to invoking one or more previously dormant functions of an automated assistant in response to (1) detecting a particular gesture of the user (e.g., one or more “invocation gestures”), and / or (2) detecting that the user’s gaze is directed towards an assistant device providing the automated assistant’s (graphical and / or audible) automated assistant interface, based on processing visual data from one or more visual components. For example, a previously dormant function may be invoked in response to detecting a particular gesture, detecting that the user’s gaze is directed towards the assistant device for at least a threshold amount of time, and optionally detecting that the particular gesture and the user’s directed gaze occur simultaneously or within a threshold temporal proximity to each other (e.g., within 0.5 seconds, within 1.0 second, or other threshold temporal proximity). For example, a previously dormant function may be invoked in response to the detection of a specific gesture (e.g., a "thumbs up" gesture and / or a "wave" gesture), and in response to the detection of a directed gaze that occurred simultaneously with or within 1.0 second of the specific gesture, for a duration of at least 0.5 seconds.
[0007] In some versions of these implementations, one or more previously dormant functions may be invoked solely in response to the detection of a gesture and a directed gaze. For example, upon detection of a gesture and a directed gaze, certain sensor data generated by one or more sensor components may be sent by the client device to one or more remote automation assistant components (however, no sensor data was sent from the sensor components before the detection of the gesture and directed gaze). The specific sensor data may include, for example, visual and / or audio data captured after the detection of the gesture and directed gaze, as well as buffered visual and / or audio data captured during the execution of the gesture and / or during the directed gaze. In this way, the transmission of data to remote computing components over the data network is performed selectively and in a timely manner. This provides efficient use of network capacity and computing and other hardware resources involved in sending and receiving data over the network. The improved efficiency of the use of computing resources in the data network and remote systems may lead to significant savings in terms of power usage by transmitters and receivers in the network, as well as in terms of memory operation and processor usage in remote systems. Corresponding effects may also be experienced in the client device. These effects enable significant additional capacity to be experienced, particularly over time, throughout the continuous operation of the automation assistant, in the network and across the entire computing system, including the devices and systems running the assistant. This additional capacity may be used for further communications in the data network, whether related to the assistant or not, without the need to expand network capacity through additional computing operations in additional or updated infrastructure and computing equipment, for example.Other technical improvements become apparent from the following considerations.
[0008] In some other versions of those implementations, one or more previously dormant functions may be invoked in response to the detection of a gesture and a gaze directed toward, and the detection of the occurrence of one or more other conditions. The occurrence of one or more other conditions may include, for example, detecting voice activity (e.g., voice activity including arbitrary voice activity, voice activity of a user giving a gesture and a gaze directed toward, voice activity of an authorized user, and voice activity including verbal invocation phrases) based on audio data, detecting user mouth movements that occur simultaneously with or are temporally close to the detected gesture and gaze directed toward, detecting that the user is an authorized user based on audio data and / or visual data, and / or detecting other conditions. For example, buffered audio data may be sent by a client device to one or more remote automation assistant components in response to the detection of a gesture and a gaze directed toward, and in response to the detection of voice activity in at least some of the buffered audio data (e.g., using a voice activity detector (VAD) module). This has advantages that correspond to the advantages described above.
[0009] In some implementations disclosed herein, previously dormant functions of an automated assistant, invoked in response to the detection of gestures and directed gazes, may include specific processing of certain sensor data (e.g., audio data, video, images, etc.) and / or rendering of certain content (e.g., graphical and / or audible). For example, prior to an invoke based on the detection of gestures and directed gazes, the automated assistant may perform only (or not perform any) limited processing of certain sensor data, such as audio data, video / image data. For example, prior to an invoke, the automated assistant may process some sensor data locally while monitoring an explicit invoke, but after local processing, it “discards” the data without having it processed by one or more additional components implementing the automated assistant (e.g., a remote server device that processes user input and generates appropriate responses). However, in response to an invoke, such data may be processed by the additional components. In these and other ways, processing and / or network resources may be reduced to only transmitting certain sensor data and / or performing certain processing of certain sensor data in response to an invoke.
[0010] Furthermore, for example, prior to an explicit call, the automation assistant may only render (or not render any) limited content (e.g., graphically). However, in response to a call, the automation assistant may render other content, such as content tailored to the user who invoked the automation assistant. For example, prior to an explicit call, no content may be graphically rendered by the display screen controlled by the automation assistant, or only limited, low-power content may be rendered (e.g., only the current time in a small portion of the display screen). However, in response to a call, the automation assistant may cause additional and optional, higher-power content, such as weather forecasts, daily summaries, and / or other content, to be graphically rendered by the display screen and / or rendered to be audible by the speaker, which may appear brighter and / or occupy a larger portion of the display screen. In these and other ways, power consumption may be reduced simply by displaying (or not displaying) low-power content prior to a call, and displaying higher-power content in response to a call.
[0011] In some implementations, when monitoring specific gestures and gazes directed towards the client device, a locally stored and trained machine learning model (e.g., a neural network model) on the client device is used by the client device to process at least a portion of the visual data from the client device's visual components (e.g., image frames from the client device's camera) at least selectively. For example, in response to detecting the presence of one or more users, the client device may use a locally stored machine learning model when monitoring specific gestures and gazes directed towards it to process at least a portion of the visual data for at least a certain duration (e.g., for at least a threshold duration and / or until the presence is no longer detected). The client device may use a dedicated presence detection sensor (e.g., a passive infrared sensor (PIR)) to detect the presence of one or more users using visual data and a separate machine learning model (e.g., a separate machine learning model trained specifically for detecting the presence of people), and / or audio data and a separate machine learning model (e.g., a VAD using a VAD machine learning model). In implementations where the processing of visual data when monitoring a specific gesture is conditional on first detecting the presence of one or more users, power resources may be saved by non-continuous processing of visual data when monitoring gestures and / or gazes. Rather, in such implementations, the processing of visual data when monitoring gestures and / or gazes may be performed only in response to the detection of the presence of one or more users in the assistant device's environment by one or more low-power techniques.
[0012] In some implementations where local machine learning models are used to monitor specific gestures and directed gazes, at least one gesture detection machine learning model is used to monitor gestures, and a separate gaze detection machine learning model is used to monitor gazes. In some versions of those implementations, one or more “upstream” models (e.g., object detection and classification models) may be used to detect parts of visual data (e.g., images) that are likely faces, eyes, arms / bodies, etc., and these parts are processed using their respective machine learning models. For example, the face and / or eye parts of an image may be detected using an upstream model and processed using a gaze detection machine learning model. Also, for example, the arms and / or body parts of an image may be detected using an upstream model and processed using a gesture detection machine learning model. In yet another example, the person part of an image may be detected using an upstream model and processed using both a gaze detection machine learning model and a gesture detection machine learning model.
[0013] Optionally, the gaze detection machine learning model can process higher resolution visual data (e.g., images) than the gesture detection machine learning model. This may allow the gesture detection machine learning model to be used more efficiently by processing lower resolution images. Furthermore, the gaze detection machine learning model may optionally be used to process only a portion of the image after the gesture detection machine learning model has been used to detect highly likely gestures (or vice versa). This may also lead to computational efficiency by not processing image data sequentially using both models.
[0014] In some implementations, face matching, eye matching, voice matching, and / or other techniques may be used to identify a specific user profile associated with a gesture and / or gaze directed towards, and content rendered by an automated assistant application on the client device tailored to that specific user profile. Rendering the tailored content can be one of the functions of the automated assistant invoked in response to the detection of the gesture and gaze directed towards. Optionally, identification of a specific user profile occurs only after the gaze directed towards and gestures have been detected. In some implementations, as described above, the occurrence of one or more additional conditions may also be required for invocation detection—the additional conditions are added to the detection of gaze and / or gestures. For example, in some implementations, the additional conditions may include identifying that the user giving the gesture and gaze directed towards is associated with an authorized user profile on the client device (e.g., using face matching, voice matching, and / or other techniques).
[0015] In some implementations, specific portions of a video / image may be filtered out / ignored / weighted less when detecting gestures and / or gazes. For example, a captured television in a video / image may be ignored to prevent false positives resulting from a person (e.g., a weather forecaster) rendered by the television. For example, a portion of an image may be determined to correspond to a television based on a separate object detection / classification machine learning model, such as by detecting a specific display frequency within that portion (i.e., matching the television's refresh rate) across multiple frames relating to that portion. Such portions may be ignored in the gesture and / or gaze detection techniques described herein to prevent the detection of gestures and / or gazes directed from a television or other video display device. As another example, photographic frames may be ignored. These and other techniques can mitigate false positive calls of the automation assistant, which can save various computing and / or network resources that would otherwise be consumed in false positive calls. Furthermore, in various implementations, once the position of a TV or photo frame is detected, it can optionally be ignored for multiple frames (for example, while intermittently verifying, or until movement of the client device or object is detected). This also saves various computational resources.
[0016] The above description is provided as an overview of the various implementations disclosed herein. These various and additional implementations are described in more detail herein.
[0017] In some implementations, a method for facilitating contactless interaction between one or more users and an automated assistant is performed by one or more processors in the client device. The method includes receiving a stream of image frames based on the output from one or more cameras in the client device. The method further includes processing the stream of image frames using at least one locally stored trained machine learning model in the client device to monitor for the occurrence of both a user calling gesture captured by at least one of the image frames and the user gazing directed towards the client device. The method further includes detecting the occurrence of both the calling gesture and the gaze based on monitoring for their occurrence. The method further includes activating at least one function of the automated assistant in response to the detection of the occurrence of both the calling gesture and the gaze.
[0018] These and other implementations of the technology may include one or more of the following characteristics:
[0019] In some implementations, at least one function of the automation assistant, which is activated in response to the detection of both a calling gesture and a gaze, includes sending audio data captured by one or more microphones on the client device to a remote server associated with the automation assistant.
[0020] In some implementations, at least one function activated in response to the detection of both a calling gesture and a gaze includes, additionally or alternatively, sending additional image frames to a remote server associated with the automation assistant, which are received after the detection of both the calling gesture and the gaze, based on the output from one or more of the cameras.
[0021] In some implementations, at least one function activated in response to the detection of both a calling gesture and a gaze additionally or alternatively includes processing buffered audio data on a client device, the buffered audio data being stored in the client device's memory and captured by one or more microphones on the client device, and the processing of the buffered audio data includes one or both of calling phrase detection processing and automatic speech recognition. In some versions of those implementations, the processing of buffered audio data includes automatic speech recognition, and the automatic speech recognition includes speech-to-text processing. In some additional or alternative versions of those implementations, the processing of buffered audio data includes calling phrase detection processing, and the method further includes, in response to the calling phrase detection processing detecting the presence of a calling phrase in the buffered audio data, sending further audio data captured by one or more microphones on the client device to a remote server associated with the automation assistant, and sending additional image frames to the remote server associated with the automation assistant, the additional image frames being received after detecting the occurrence of both a calling gesture and a gaze, based on the output from one or more cameras, and sending one or both of these.
[0022] In some implementations, the step of processing image frames in a stream using at least one trained machine learning model stored locally on the client device to monitor the occurrence of both a calling gesture and a gaze includes using a first trained machine learning model to monitor the occurrence of a calling gesture and using a second trained machine learning model to monitor the user's gaze directed towards the client device. In some versions of those implementations, using a second trained machine learning model to monitor the user's gaze directed towards the client device is done only in response to detecting the occurrence of a calling gesture using the first trained machine learning model. In some versions of those, or other versions, using a first trained machine learning model to monitor the occurrence of a calling gesture includes processing a first resolution version of the image frame using the first machine learning model, and using a second trained machine learning model to monitor the user's gaze includes processing a second resolution version of the image frame using the second machine learning model.
[0023] In some implementations, the method further includes the steps of receiving a stream of audio data frames based on the output from one or more microphones of a client device, processing the audio data frames of the stream using at least one trained call phrase detection machine learning model stored locally on the client device to monitor the occurrence of a verbal call phrase, and detecting the occurrence of a verbal call phrase based on monitoring the occurrence of a verbal call phrase. In some of those implementations, the step of activating at least one function of the automation assistant responds to detecting the occurrence of a verbal call phrase that is temporally close to both a call gesture and a gaze. In some versions of those implementations, the at least one function to be activated includes sending additional audio data frames captured by one or more microphones of the client device to a remote server associated with the automation assistant, and sending one or more additional image frames from one or more cameras to a remote server associated with the automation assistant.
[0024] In some implementations, the step of processing a stream of image frames using at least one trained machine learning model stored locally on the client device to monitor the occurrence of both calling gestures and gazes includes processing the image frame using a first trained machine learning model to predict regions of the image frame that contain a person's face, and processing the region of the image frame using a second trained machine learning model trained to detect user gazes.
[0025] In some implementations, the step of processing image frames in a stream using at least one trained machine learning model stored locally on the client device to monitor for the occurrence of both calling gestures and gazing includes determining that a region of the image frame corresponds to an electronic display and, depending on whether the region corresponds to an electronic display, ignoring the region when monitoring for the occurrence of both calling gestures and gazing. In some of those implementations, determining that a region of the image frame corresponds to an electronic display is based on detecting a display frequency within the region of the image frame that corresponds to the display frequency of the electronic display.
[0026] In some implementations, the method further includes the steps of detecting the presence of a person in the environment of a client device based on a signal from a presence detection sensor, and causing one or more cameras to provide a stream of image frames in response to the detection of the presence of a person in the environment.
[0027] In some implementations, a client device is provided that includes at least one visual component, at least one microphone, one or more processors, and a memory operatively coupled to the one or more processors. The memory stores instructions for causing one or more of the processors to perform the following operations in response to execution of the instructions by one or more of the processors: receiving a stream of visual data based on an output from the visual component of the client device; processing the visual data using at least one trained machine learning model stored locally on the client device to monitor for the occurrence of both a user call gesture captured by the visual data and a user gaze directed at the client device; detecting based on monitoring for the occurrence of both the call gesture and the gaze; and in response to detecting the occurrence of both the call gesture and the gaze, causing the client device to transmit to one or more remote automated assistant components either additional visual data based on an output from the visual component and / or audio data based on an output from the microphone of the client device. The operations may further optionally include receiving response content in response to transmitting, and rendering the response content by one or more user interface output devices of the client device.
[0028] In some implementations, a system is provided that includes at least one visual component and one or more processors that receive a stream of visual data based on the output from the visual component. One or more of the processors process the visual data using at least one trained machine learning model stored locally on the client device to monitor for the occurrence of both a user's call gesture captured by the visual data and a user's gaze directed at the client device, detect based on monitoring for the occurrence of both the call gesture and the gaze, and are configured to activate at least one function of an automated assistant in response to detecting the occurrence of both the call gesture and the gaze.
[0029] In addition, some implementations include one or more processors of one or more computing devices, the one or more processors being operable to execute instructions stored in a related memory, the instructions being configured to cause execution of any of the methods described above. Some implementations also include one or more non - transitory computer - readable recording media storing computer instructions that may be executed by one or more processors to execute any of the methods described above.
[0030] It should be understood that all combinations of the concepts described above and additional concepts described in more detail herein are considered to be part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter that appear at the end of this disclosure are considered to be part of the subject matter disclosed herein.
Brief Description of the Drawings
[0031] [Figure 1] It is a block diagram of an exemplary environment in which implementations disclosed herein may be implemented. [Figure 2A] It is a diagram showing an exemplary process flow that illustrates various aspects of the present disclosure according to various implementations. [Figure 2B] This figure shows an exemplary process flow illustrating various aspects of this disclosure through various implementations. [Figure 2C] This figure shows an exemplary process flow illustrating various aspects of this disclosure through various implementations. [Figure 3] This figure shows an example of a user giving an assistant device gestures and a gaze directed at them, and also shows images captured by the assistant device's camera when the user is giving gestures and a gaze directed at them. [Figure 4A] This flowchart illustrates an exemplary method by implementation disclosed herein. [Figure 4B] Figure 4A is a flowchart illustrating a specific example of a particular block of the exemplary method. [Figure 5] This diagram illustrates an exemplary architecture of a computing device. [Modes for carrying out the invention]
[0032] Figure 1 shows an exemplary environment in which the technologies disclosed herein may be implemented. The exemplary environment includes one or more client computing devices 106. Each client device 106 may run its respective instance of the automation assistant client 110. One or more cloud-based automation assistant components 130 may be implemented in one or more computing systems (collectively referred to as a “cloud” computing system) that are communicably coupled to the client devices 106 via one or more local area and / or wide area networks (e.g., the Internet) as shown overall in 114. The cloud-based automation assistant components 130 may be implemented, for example, by a cluster of high-performance servers.
[0033] In various implementations, an instance of the automation assistant client 110 may, through its interaction with one or more cloud-based automation assistant components 130, form what appears to be a logical instance of the automation assistant 120 from the user's perspective, in which the user may engage in person-to-computer interaction (e.g., verbal interaction, gesture-based interaction, and / or touch-based interaction). One such instance of the automation assistant 120 is shown within the dashed line in Figure 1. Thus, each user interacting with the automation assistant client 110 running on the client device 106 may, in fact, be interacting with their own logical instance of the automation assistant 120. For brevity and simplicity, the term “automation assistant” as used herein to “serve” a particular user refers to a combination of the automation assistant client 110 running on the client device 106 operated by the user and optionally one or more cloud-based automation assistant components 130 (which may be shared among multiple automation assistant clients 110). It should also be understood that in some implementations, the automation assistant 120 may respond to requests from any user, regardless of whether the user is actually "servicing" that particular instance of the automation assistant 120.
[0034] One or more client devices 106 may include, for example, a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a computing device in the user's vehicle (e.g., an in-car communication system, an in-car entertainment system, an in-car navigation system), a standalone interactive speaker (which may optionally include a vision sensor), smart home appliances such as a smart TV (or a regular TV with a network-connected dongle having automation assistant capabilities), and / or a user's wearable device including a computing device (e.g., a user's wristwatch with a computing device, a user's glasses with a computing device, a virtual or augmented reality computing device). Additional and / or alternative client computing devices may be provided. As described above, some client devices 106 may take the form of assistant devices designed primarily to facilitate interaction between the user and an automation assistant 120 (e.g., a standalone interactive device having a speaker and a display).
[0035] The client device 106 may be equipped with one or more visual components 107 having one or more fields of view. The visual components 107 may take various forms, such as a monographic camera, a stereographic camera, a LiDAR component, or a radar component. One or more visual components 107 may be used by a visual capture module 114 to capture, for example, a vision frame (e.g., an image frame (still image or video)) of the environment in which the client device 106 is placed. These vision frames may then be at least selectively analyzed by a gaze and gesture module 116 of the call engine 115 to monitor the occurrence of a specific gesture (from one or more candidate gestures) of the user captured by the vision frame and / or a gaze directed from the user (i.e., a gaze directed towards the client device 106). The gaze and gesture module 116 may utilize one or more trained machine learning models 117 when monitoring the occurrence of a specific gesture and / or a gaze directed at the user.
[0036] In response to the detection of a specific gesture and directed gaze (and optionally, the detection of one or more other conditions by the other conditions module 118), the calling engine 115 can invoke one or more previously dormant functions of the automation assistant 120. Such dormant functions may include, for example, the processing of specific sensor data (e.g., audio data, video, images, etc.) and / or the rendering (e.g., graphical and / or audible) of specific content.
[0037] As one non-limiting example, prior to the detection of a specific gesture and directed gaze, the visual and / or audio data captured on the client device 106 may be processed and / or temporarily buffered only locally on the client device 106 (i.e., without being sent to the cloud-based automated assistant component 130). However, in response to the detection of a specific gesture and directed gaze, the audio and / or visual data (e.g., recently buffered data and / or data received after detection) may be sent to the cloud-based automated assistant component 130 for further processing. For example, the detection of a specific gesture and directed gaze can allow the automated assistant 120 to process the user's verbal utterances sufficiently, generate response content for the automated assistant 120, and render it to the user without the user having to say an explicit invocation phrase (e.g., "Okay, Assistant").
[0038] For example, instead of a user having to say "Okay, assistant. What's the forecast for today?" to get today's forecast, the user might instead perform a specific gesture, look at client device 106, and simply say "What's the forecast for today?" during or in close proximity to the gesture and / or the look at client device 106 (e.g., within a threshold time before and / or after). Data corresponding to the verbal utterance "What's the forecast for today?" (e.g., audio data capturing the verbal utterance or its text or other semantic conversion) may be sent by client device 106 to the cloud-based automation assistant component 130 in response to the detection of the gesture and the gaze directed at it, and in response to the verbal utterance being received during and / or in close proximity to the gesture and the gaze directed at it. In another example, instead of the user having to say "Okay, assistant, turn up the heating" to raise the temperature of their home using a connected thermostat, the user might instead perform a specific gesture, look at the client device 106, and simply say "turn up the heating" during or in close proximity to the gesture and / or the client device 106 (for example, within a threshold time before and / or after). Data corresponding to the verbal utterance "turn up the heating" (for example, audio data capturing the verbal utterance or its text or other semantic transformation) may be sent by the client device 106 to the cloud-based automation assistant component 130 in response to the detection of the gesture and the gaze directed, and in response to the verbal utterance being received during and / or in close proximity to the gesture and the gaze directed.In another example, instead of the user having to say "Okay, assistant, open the garage door" to open their garage, the user might instead perform a specific gesture, look at the client device 106, and simply say "open the garage door" during or in close proximity to the gesture and / or the look at the client device 106 (for example, within a threshold time before and / or after). Data corresponding to the verbal utterance "open the garage door" (for example, audio data capturing the verbal utterance or its text or other semantic transformation) may be sent by the client device 106 to the cloud-based automation assistant component 130 in response to the detection of the gesture and the gaze directed, and in response to the verbal utterance being received during and / or in close proximity to the gesture and the gaze directed. In some implementations, the transmission of data by the client device 106 may be further conditional on the other conditions module 118 determining the occurrence of one or more additional conditions. For example, data transmission may further rely on local speech activity detection processing of the audio data, performed by the other condition module 118, which indicates that speech activity is present within the audio data. Alternatively, data transmission may further rely on the other condition module 118 determining, additionally or alternatively, that the audio data corresponds to a user who has given gestures and a gaze that is directed toward them. For example, the user's orientation (relative to the client device 106) can be determined based on visual data, and data transmission may further rely on the other condition module 118 determining, (e.g., using beamforming and / or other techniques) that the oral utterances in the audio data are coming from the same direction.Furthermore, for example, a user's user profile can be determined based on visual data (e.g., using facial recognition), and data transmission may further be based on the determination by the Other Conditions Module 118 that oral utterances in the audio data have voice features that match the user profile. In yet another example, data transmission may further be based on the determination by the Other Conditions Module 118, based on visual data, that the user's mouth movements occurred simultaneously with the detected gesture and / or gaze directed at the user, or within a time frame of a threshold amount for the detected gesture and / or gaze directed at the user. The Other Conditions Module 118 may optionally utilize one or more other machine learning models 119 when determining the presence of other conditions. Further descriptions of implementations of the gaze and gesture module 116 and the Other Conditions Module 118 are given herein (see, for example, Figures 2A-2C).
[0039] Each of the computing devices running the client computing device 106 and the cloud-based automation assistant component 130 may include one or more memories for storing data and software applications, one or more processors for accessing data and running applications, and other components for facilitating communication over a network. The operations performed by the client computing device 106 and / or the automation assistant 120 may be distributed across multiple computer systems. The automation assistant 120 may be implemented, for example, as a computer program running on one or more computers in one or more locations connected to each other via a network.
[0040] As described above, in various implementations, the client computing device 106 may run the automation assistant client 110. In some of these various implementations, the automation assistant client 110 may include a voice capture module 112, the visual capture module 114 described above, and a calling engine 115, the calling engine 115 of which may include a gaze and gesture module 116 and optionally other conditional modules 118. In other implementations, one or more embodiments of the voice capture module 112, the visual capture module 114, and / or the calling engine 115 may be implemented separately from the automation assistant client 110, for example, by one or more cloud-based automation assistant components 130.
[0041] In various implementations, the audio capture module 112, which may be implemented using any combination of hardware and software, may interface with hardware such as a microphone 109 or other pressure sensor for capturing an audio recording of the user's oral speech. Various types of processing may be performed on this audio recording for various purposes, as described below. In various implementations, the visual capture module 114, which may be implemented using any combination of hardware and software, may be configured to interface with a visual component 107 for capturing one or more visual frames (e.g., digital images) corresponding to any adaptable field of view of the visual sensor 107.
[0042] The voice capture module 112 may be configured to capture the user's voice, for example, by a microphone 109, as described above. Additionally or alternatively, in some implementations, the voice capture module 112 may be further configured to convert its captured audio to text and / or other representations or embeddings using, for example, speech-to-text ("STT") processing techniques. However, since the client device 106 may be relatively constrained in terms of computing resources (e.g., processor cycles, memory, battery, etc.), the voice capture module 112 located locally on the client device 106 may be configured to convert a finite number of different spoken phrases—such as phrases to invoke the automation assistant 120—to text (or other forms such as lower-dimensional embeddings). Other voice input may be sent to a cloud-based automation assistant component 130, which may include a cloud-based STT module 132.
[0043] The cloud-based TTS module 131 may be configured to utilize virtually unlimited resources in the cloud to convert text data (e.g., natural language responses generated by the automation assistant 120) into computer-generated speech output. In some implementations, the TTS module 131 may provide the computer-generated speech output to the client device 106 for direct output, for example, using one or more speakers. In other implementations, the text data (e.g., natural language responses) generated by the automation assistant 120 may be provided to the client device 106, and then a local TTS module in the client device 106 may convert the text data into locally output computer-generated speech.
[0044] The cloud-based STT module 132 may be configured to utilize virtually unlimited resources in the cloud to convert speech data captured by the speech capture module 112 into text, which may then be provided to the natural language understanding module 135. In some implementations, the cloud-based STT module 132 may convert the audio recording of speech into one or more phonemes, and then convert one or more phonemes into text. Additionally or alternatively, in some implementations, the STT module 132 may use a state decoding graph. In some implementations, the STT module 132 may generate multiple candidate text interpretations of the user's utterance and select a given interpretation from the candidates using one or more techniques.
[0045] The automation assistant 120 (and in particular the cloud-based automation assistant component 130) may include the intent understanding module 135, the TTS module 131 described above, the STT module 132 described above, and other components described in more detail herein. In some implementations, modules of the automation assistant 120 and / or one or more of the modules may be omitted, combined, and / or implemented in components separate from the automation assistant 120. In some implementations, one or more of the components of the automation assistant 120, such as the intent understanding module 135, the TTS module 131, and the STT module 132, may be implemented in the client device 106 at least partially (for example, in combination with or excluding the cloud-based implementation).
[0046] In some implementations, the automation assistant 120 generates various content for the user to hear and / or graphically render via the client device 106. For example, the automation assistant 120 may generate content such as weather forecasts or daily schedules, and may render the content in response to detecting gestures from the user and / or gazes directed at them, as described herein. Alternatively, the automation assistant 120 may generate content, for example, in response to free-form natural language input from the user given via the client device 106, or in response to user gestures detected by visual data from the visual component 107 of the client device. As used herein, free-form input is input made by the user and not constrained by a set of choices presented for the user to select. Free-form input can be, for example, typed input and / or spoken input.
[0047] The natural language processor 133 of the intent understanding module 135 may process natural language input generated by the user via the client device 106 and produce annotated output (e.g., in text format) for use by one or more other components of the automation assistant 120. For example, the natural language processor 133 may process free-form natural language input generated by the user via one or more user interface input devices of the client device 106. The generated annotated output may include one or more annotations of the natural language input and one or more (e.g., all) of the words of the natural language input.
[0048] In some implementations, the natural language processor 133 is configured to identify and annotate various types of grammatical information in the natural language input. For example, the natural language processor 133 may include a morphological module that divides individual words into morphemes and / or annotates the morphemes, for example, by the class of those morphemes. The natural language processor 133 may also include a part-of-speech tagger configured to annotate words by their grammatical roles. In addition, for example in some implementations, the natural language processor 133 may include a dependency parser (not shown) configured to determine syntactic relationships between words in the natural language input.
[0049] In some implementations, the natural language processor 133 may additionally and / or alternatively include entity taggers (not shown) configured to annotate references to entities within one or more segments, such as references to people (including, for example, literary characters, celebrities, and famous people), organizations, and places (real and fictional). In some implementations, data about entities may be stored in one or more databases, such as a knowledge graph (not shown), and the entity taggers of the natural language processor 133 may utilize such databases for tagging entities.
[0050] In some implementations, the natural language processor 133 may additionally and / or alternatively include a coreference resolver (not shown) configured to group or "cluster" references to the same entity based on one or more contextual cues. For example, the coreference resolver may be used to resolve the word "there" in the natural language input "I liked Hypothetical Cafe last time we ate there" to "Hypothetical Cafe".
[0051] In some implementations, one or more components of the natural language processor 133 may rely on annotations from one or more other components of the natural language processor 133. For example, in some implementations, a given entity tagger may rely on annotations from the same-direction resolver and / or dependency parser when annotating all references to a particular entity. Also, in some implementations, for example, the same-direction resolver may rely on annotations from the dependency parser when clustering references to the same entity. In some implementations, when processing a particular natural language input, one or more components of the natural language processor 133 may use relevant previous inputs and / or other relevant data outside of the particular natural language input to determine one or more annotations.
[0052] The intent understanding module 135 may further include an intent matcher 134 configured to determine the intent of a user engaging in interaction with the automation assistant 120. Although shown separately from the natural language processor 133 in Figure 1, in other implementations, the intent matcher 134 may be an integral part of the natural language processor 133 (or more broadly, a pipeline including the natural language processor 133). In some implementations, the natural language processor 133 and the intent matcher 134 may collectively form the intent understanding module 135 described above.
[0053] The intent matcher 134 may use a variety of techniques to determine the user's intent, for example, based on the output from the natural language processor 133 (which may include annotations and words of natural language input), based on the user's touch input on the touch display of the client device 106, and / or based on gestures and / or other visual cues detected in visual data. In some implementations, the intent matcher 134 may have access to one or more databases (not shown) that include multiple mappings, for example, between grammars and response actions (or more broadly, intents), between visual cues and response actions, and / or between touch input and response actions. For example, grammars included in the mappings can be selected and / or learned over time and may represent common user intents. For example, one grammar "play <artist>"but, <artist>This may be mapped to an intent to invoke a response action that plays music on a user-operated client device 106. Another grammar, "[weather|forecast] today", may be matchable to user inquiries such as "what's the weather today" and "what's the forecast for today?". As another example, mappings between visual cues and actions may include "inclusive" mappings that may apply to multiple users (e.g., all users) and / or user-specific mappings. Some examples of mappings between visual cues and actions include mappings for gestures. For example, a "waving" gesture may be mapped to an action that renders tailored content (tailored to the user giving the gesture) to the user, a "thumbs up" gesture may be mapped to a "play music" action, and a "high-five" gesture may be mapped to a "routine" of automated assistant actions to be performed, such as turning on a smart coffee maker, turning on a specific smart light, and rendering a news summary audibly.
[0054] In addition to or instead of grammar, in some implementations, the intent matcher 134 may use one or more trained machine learning models alone or in combination with one or more grammars, visual cues, and / or touch inputs. These trained machine learning models may also be stored in one or more databases and may be trained to identify intents by embedding, for example, data representing user utterances and / or visual cues given by any detected user into a reduced-dimensional space, and then determining which other embeddings (and therefore intents) are closest using techniques such as Euclidean distance, cosine similarity, etc.
[0055] The above "play <artist>As seen in the example grammar, some grammars have slots that may be filled with slot values (or "parameters") (for example, <artist>) has. Slot values may be determined in various ways. Often, the user provides the slot value in advance. For example, the grammar "Order me a <topping>Regarding "pizza," the user is likely to say the phrase "order me a sausage pizza," in which case the slot <topping>These are filled in automatically. Additionally or alternatively, if the user invokes a grammar that includes slots to be filled by slot values without the user proactively providing the slot values, the automation assistant 120 may prompt the user for those slot values (for example, "What type of crust do you want on your pizza?"). In some implementations, slots may be filled with slot values based on visual cues detected based on visual data captured by the visual component 107. For example, the user might utter something like "Order me this many cat bowls" while holding up three fingers to the visual component 107 on the client device 106. Or, the user might utter something like "Find me more movies like this" while holding a DVD case of a particular movie.
[0056] In some implementations, the automation assistant 120 may facilitate (e.g., “mediate)) transactions between the user and an agent, which may be an independent software process that receives input and provides response outputs. Some agents may take the form of third-party applications that may or may not run on a computing system separate from the computing system on which, for example, the cloud-based automation assistant component 130 operates. One type of user intent that may be identified by the intent matcher 134 is to involve a third-party application. For example, the automation assistant 120 may provide access to an application programming interface ("API") to a pizza delivery service. The user may invoke the automation assistant 120 and give a command such as “I’d like to order a pizza”. The intent matcher 134 may map this command to a grammar that triggers the automation assistant 120 to interact with the third-party pizza delivery service. The third-party pizza delivery service may give the automation assistant 120 a minimal list of slots that need to be filled in order to fulfill the pizza delivery order. The automation assistant 120 may generate natural language output requesting parameters for a slot and provide it to the user (via the client device 106).
[0057] The execution module 138 may be configured to receive the predicted / estimated intent output by the intent matcher 134 and the associated slot values (whether provided in advance by the user or requested by the user), and to perform (or "resolve") the intent. In various implementations, the performance (or "resolve") of the user's intent may involve the execution module 138 generating / retrieving various performance information (also called "response" information or data).
[0058] The performance information may take various forms, as intent may be performed in various ways. Suppose the user requests pure information, such as "Where were the outdoor shots of 'The Shining' filmed?". The user's intent may be determined to be a search query by, for example, intent matcher 134. The intent and content of the search query may be given to the performance module 138, which may communicate with one or more search modules 150 configured to search a corpus of documents and / or other data sources (e.g., a knowledge graph) for response information, as shown in Figure 1. The performance module 138 may provide the search module 150 with data indicating the search query (e.g., the text of the query, dimensionally reduced embeddings, etc.). The search module 150 may provide response information, such as GPS coordinates or other more explicit information, such as "Timberline Lodge, Mt. Hood, Oregon". This response information may form part of the performance information generated by the performance module 138.
[0059] Additionally or alternatively, the execution module 138 may be configured to receive, for example, the user's intent and any slot values determined by the user or otherwise provided by the user (e.g., the user's GPS coordinates, the user's preferences, etc.) from the intent understanding module 135, and to trigger a response action. Response actions may include, for example, ordering an item / service, starting a timer, setting a reminder, starting a call, playing media, sending a message, or starting a multi-action routine. In some such implementations, the execution information may include slot values related to the execution, acknowledgments (which may be selected from a predetermined set of responses), etc.
[0060] Additionally or alternatively, the execution module 138 may be configured to infer user intent (for example, based on the time, past interactions, etc.) and retrieve response information regarding those intents. For example, the execution module 138 may be configured to retrieve a daily schedule summary for the user, a weather forecast for the user, and / or other content for the user. Furthermore, the execution module 138 may have such content “push” to the user for graphical and / or audible rendering. For example, the rendering of such content could be a function that was in a dormant state and was invoked in response to the calling engine 115 detecting the occurrence of a particular gesture and a gaze directed toward it.
[0061] The natural language generator 136 may be configured to generate and / or select natural language output (e.g., words / phrases designed to mimic human speech) based on data obtained from various sources. In some implementations, the natural language generator 136 may be configured to receive performance information related to intent performance as input and to generate natural language output based on the performance information. Additionally or alternatively, the natural language generator 136 may receive information from other sources, such as third-party applications, which the natural language generator 136 may use to construct natural language output for the user.
[0062] Referring here to Figures 2A, 2B, and 2C, various examples are shown of how the gaze and gesture module 116 can detect a particular gesture and / or gaze that is directed, and how the calling engine 115 can accordingly invoke one or more previously dormant functions of the automation assistant.
[0063] First, looking at Figure 2A, the visual capture module 114 provides visual frames to the gaze and gesture module 116. In some implementations, the visual capture module 114 provides a real-time stream of visual frames to the gaze and gesture module 116. In some of those implementations, the visual capture module 114 begins providing visual frames in response to a signal from a separate presence sensor 105 indicating the presence of a person in an environment having a client device 106. For example, the presence sensor 105 could be a PIR sensor that signals the visual capture module 114 upon detecting the presence of a person. The visual capture module 114 may refrain from providing any visual frames to the gaze and gesture module 116 unless the presence of a person is detected. In other implementations where the visual capture module 114 selectively provides visual frames to the gaze and gesture module 116, additional and / or alternative cues may be used to initiate such provision. For example, the presence of a person may be detected based on audio data from the voice capture module 112, based on the analysis of visual frames by one or more other components, and / or based on other signals.
[0064] The gaze and gesture module 116 processes visual frames using one or more machine learning models 117 to monitor the occurrence of both directed gazes and specific gestures. When both directed gazes and specific gestures are detected, the gaze and gesture module 116 provides an indication of gaze and gesture detection to the calling engine 115.
[0065] In Figure 2A, the visual frame and / or audio data (provided by the audio capture module 112) are also provided to the other conditions module 118. The other conditions module 118 processes the provided data using optionally one or more other machine learning models 119 to monitor the occurrence of one or more other conditions. For example, other conditions may include detecting arbitrary vocal activity based on audio data, detecting the presence of a spoken calling phrase in the audio data, detecting vocal activity based on audio data that is coming from the user's direction or location, detecting that the user is an authorized user based on the visual frame and / or audio data, or detecting the user's mouth movements (giving gestures and a gaze that is directed) based on the visual frame. When an other condition is detected, the other conditions module 118 provides the calling engine 115 with an indication of the occurrence of the other condition.
[0066] When the calling engine 115 receives temporally close indications of gaze and gesture indications and other conditions directed toward it, the calling engine 115 triggers a call 101 of a dormant function. For example, a call 101 of a dormant function may include one or more of the following: activating the display screen of the client device 106; causing the client device 106 to render content so that it is visually and / or audibly present; or causing the client device 106 to send visual frames and / or audio data to one or more cloud-based automation assistant components 130.
[0067] In some implementations, as will be described in more detail in relation to Figures 2B and 2C, the gaze and gesture module 116 may use one or more first machine learning models 117 to detect directed gazes and one or more second machine learning models 117 to detect gestures.
[0068] In some other implementations, the gaze and gesture module 116 may utilize an end-to-end machine learning model that accepts a visual frame (or its features) as input and generates an output (based on the model's processing of the input) indicating whether a particular gesture and directed gaze have occurred. Such a machine learning model could be a neural network model, such as a recurrent neural network (RNN) model, which includes one or more memory layers (e.g., long short-term memory (LSTM) layers). Training of such an RNN model may be based on training examples that include a sequence of visual frames (e.g., a video) as training example input and an indication as training example output whether the sequence includes both a gesture and a directed gaze. For example, the training example output could be a single value indicating whether both a gesture and a directed gaze are present. As another example, the training example output may include a first value indicating whether a directed gaze exists, and N additional values indicating whether one of the N gestures is present (thus enabling the model to train to predict the corresponding probability for each of the N distinct gestures). As yet another example, the training example output may include a first value indicating whether a directed gaze exists, and a second value indicating whether one or more specific gestures are present (thus enabling the model to train to predict the corresponding probability for whether any of the gestures are present).
[0069] In an implementation where a model is trained to predict the corresponding probability for each of N distinct gestures, the gaze and gesture module 116 can optionally provide the calling engine 115 with an indication of which of the N gestures occurred. Furthermore, an invocation 101 of a dormant function by the calling engine 115 may depend on which of the N distinct gestures occurred. For example, in the case of a “waving” gesture, the calling engine 115 can cause specific content to be rendered on the display screen of the client device; in the case of a “thumbs up” gesture, the calling engine 115 can cause audio data and / or a visual frame to be sent to the cloud-based automation assistant component 130; and in the case of a “high-five” gesture, the calling engine 115 can cause the automation assistant to perform a “routine” of actions such as turning on a smart coffee maker, turning on specific smart lighting, and rendering a news summary audibly.
[0070] Figure 2B shows an example that includes a gesture module 116A, which utilizes a gesture machine learning model 117A when the gesture and gaze detection module 116 monitors the occurrence of a gesture, and a gaze module 116B, which utilizes a gaze machine learning model 117B when monitoring the occurrence of a directed gaze. Other condition modules 118 are not shown in Figure 2B for simplicity, but may also be used in combination with gesture modules 116A and 117B in a similar manner to that described in relation to Figure 2A.
[0071] In Figure 2B, the visual capture module 114 provides a visual frame. A lower-resolution version of the visual frame is provided to the gesture module 116A, and a higher-resolution version of the visual frame is stored in buffer 104. The lower-resolution version is lower resolution than the higher-resolution version. The higher-resolution version can be uncompressed or less compressed than the lower-resolution version. Buffer 104 can be a first-in, first-out buffer and can temporarily store the most recent duration of the higher-resolution visual frame.
[0072] In the example in Figure 2B, the gesture module 116A can process lower-resolution visual frames when monitoring for the presence of a gesture, and the gaze module 116B can remain inactive until the gesture module 116A detects the occurrence of a gesture. When the occurrence of a gesture is detected, the gesture module 116A can provide an indication of the gesture detection to the gaze module 116B and the calling engine 115. The gaze module 116B is activated in response to receiving the gesture detection indication and can retrieve buffered higher-resolution visual frames from buffer 104 and utilize these buffered higher-resolution visual frames (and optionally even higher-resolution visual frames) when determining whether a gaze is directed at it. In this way, the gaze module 116B is activated selectively only, thereby saving computational resources that would otherwise be consumed by the additional processing of higher-resolution visual frames by the gaze module 116B.
[0073] The gesture module 116A can use one or more gesture machine learning models 117A to detect a particular gesture. Such machine learning models can be neural network models, such as an RNN model that includes one or more memory layers. Training of such an RNN model may be based on training examples that include a sequence of visual frames (e.g., a video) as training example input and an indication as training example output whether the sequence contains one or more particular gestures. For example, the training example output can be a single value indicating whether a single particular gesture is present. For example, the single value can be "0" when the single particular gesture is not present and "1" when the single particular gesture is present. In some of these examples, multiple gesture machine learning models 117A, each tailored to a different single particular gesture, are used. In another example, the training example output may include N values, each indicating whether one of N corresponding gestures is included (therefore, enabling the training of a model to predict the corresponding probability for each of the N distinct gestures). In an implementation where the model is trained to predict the corresponding probability for each of N distinct gestures, the gesture module 116A can optionally provide the call engine 115 with an indication of which of the N gestures occurred. Furthermore, the call engine 115 may depend on which of the N distinct gestures occurred.
[0074] The gaze module 116B may use one or more gaze machine learning models 117B to detect directed gazes. Such machine learning models can be neural network models, such as a convolutional neural network (CNN) model. Training of such a CNN model may be based on training examples, which include visual frames (e.g., images) as training example inputs and training example outputs that indicate whether the image contains a directed gaze. For example, the training example output can be a single value indicating whether a directed gaze exists. For example, the single value could be "0" when no directed gaze exists, "1" when a gaze exists that is directed straight at the image-capturing sensor or within 5 degrees of the image-capturing sensor, and "0.75" when a gaze exists that is directed within 5 to 10 degrees of the image-capturing sensor.
[0075] In some of those and / or other implementations, the gaze module 116B determines a gaze is directed only when the directed gaze is detected with at least a threshold probability and / or for at least a threshold duration. For example, a stream of image frames may be processed using a CNN model, and processing each frame may yield a corresponding probability that the frame contains a directed gaze. The gaze module may determine that a directed gaze exists only if at least X% of the sequence of image frames (corresponding to a threshold duration) have a corresponding probability that satisfies the threshold. For example, suppose X% is 60%, the probability threshold is 0.7, and the threshold duration is 0.5 seconds. Furthermore, suppose 10 image frames correspond to 0.5 seconds. If image frames are processed to produce probabilities [0.75, 0.85, 0.5, 0.4, 0.9, 0.95, 0.85, 0.89, 0.6, 0.85], then directed gazes may be detected, since 70% of the frames showed directed gazes with a probability greater than 0.7. In these and other methods, directed gazes may be detected even if the user briefly averts the direction of their gaze. Additional and / or alternative machine learning models (e.g., RNN models) and / or techniques may be used to detect directed gazes that occur for at least a threshold duration.
[0076] Figure 2C shows another example including a gesture module 116A, which utilizes a gesture machine learning model 117A when the gesture and gaze detection module 116 monitors the occurrence of a gesture, and a gaze module 116B, which utilizes a gaze machine learning model 117B when monitoring the occurrence of a directed gaze. Other condition modules 118 are not shown in Figure 2C for simplicity, but may also be used in combination with gesture modules 116A and 117B in a similar manner to that described in relation to Figure 2A. Also, buffers 104 and higher / lower resolution visual frames in Figure 2B are not shown in Figure 2C for simplicity, but similar techniques may be implemented in Figure 2C (for example, a higher resolution portion of the visual frame may be provided to the gaze module 116B).
[0077] In Figure 2C, the visual capture module 114 provides visual frames to the detection and classification module 116C. The detection and classification module 116C utilizes an object detection and classification machine learning model 117C to classify various regions of each visual frame. For example, the detection and classification module 116C can classify the region of a person (if any) in each visual frame corresponding to a person and provide indications of such a region for each visual frame to the gesture module 116A and the gaze module 116B. Alternatively, for example, the detection and classification module 116C can classify the region of a person's body (e.g., arms and torso) in each visual frame (if any) and provide indications of such a region for each visual frame to the gesture module 116A. Similarly, for example, the detection and classification module 116C can classify the region of a person's face (if any) in each visual frame and provide indications of such a region for each visual frame to the gaze module 116B.
[0078] In some implementations, the gesture module 116A can utilize the provided region to process only the corresponding portion of each visual frame. For example, the gesture module 116A can "crop" and resize a visual frame to process only the portion containing the region of a person and / or body. In some of these implementations, the gesture machine learning model 117A can be trained based on the "cropped" visual frame, and the resizing can be to a size that fits the dimensions of the input to such a model. In some additional or alternative implementations, the gesture module 116A can utilize the provided region to skip processing all of a certain number of visual frames (e.g., visual frames indicated not to contain the region of a person and / or body). In yet other implementations, the gesture module 116A can utilize the provided region as an attention mechanism to focus the processing of each visual frame (e.g., as a separate attention input to the gesture machine learning model 117A).
[0079] Similarly, in some implementations, the gaze module 116B can utilize the provided region to process only the corresponding portion of each visual frame. For example, the gaze module 116B can "crop" and resize a visual frame to process only the portion containing the region of a person and / or face. In some of those implementations, the gaze machine learning model 117B can be trained based on the "cropped" visual frame, and the resizing can be to a size that fits the dimensions of the input to such a model. In some additional or alternative implementations, the gaze module 116B can utilize the provided region to skip processing all of a certain number of visual frames (e.g., visual frames indicated not to contain the region of a person and / or face). In yet other implementations, the gaze module 116B can utilize the provided region as an attention mechanism to focus the processing of each visual frame (e.g., as a separate attention input to the gaze machine learning model 117B).
[0080] In some implementations, the detection and classification model 116C may additionally or alternatively provide indications of specific regions to other condition modules 118 for use by those other condition modules 118 (not shown in Figure 2C for simplicity). For example, a facial region may be used by the other condition module 118 when detecting mouth movements using a corresponding mouth movement machine learning model, where mouth movements are an additional condition for invoking an inactive function.
[0081] In some implementations, the detection and classification model 116C may provide indications to the gesture module 116A and gaze module 116B of areas that are additionally or alternatively classified as TV or other video display sources. In some of those implementations, modules 116A and 116B may crop those areas from the visual frame being processed, focus on areas outside those areas, and / or otherwise ignore those areas in detection or reduce the likelihood that detection is based on such areas. These and other methods may mitigate false detection calls of dormant functions.
[0082] Figure 3 shows an example of the client device 106 and visual component 107 from Figure 1. In Figure 3, the exemplary client device is shown as 106A and further includes a speaker and a display. In Figure 3, the exemplary visual component is shown as 107A and is a camera. Figure 3 also shows a user 301 giving a hand movement gesture (indicated by the “Movement” line near the user’s right hand) and a gaze directed towards camera 107A.
[0083] Figure 3 also shows an example of a user giving an assistant device and gestures and a gaze directed towards it, and also shows an exemplary image 360 captured by camera 107A when the user is giving gestures and a gaze directed towards it. It can be seen that the user and the television behind the user (and therefore not visible in the perspective view of Figure 3) are captured in image 360.
[0084] Within image 360, a bounding box 361 is provided, representing a region of the image that may be determined to correspond to a person (for example, by the detection and classification module 116C in Figure 2C). In some implementations, a gesture detection module operating on a client device 106A may process (or focus on) only that portion of the image when monitoring a particular gesture based on the fact that that portion is indicated as a portion corresponding to a person.
[0085] Within image 360, a bounding box 362 is also provided, representing a region of the image that may be determined to correspond to a face (for example, by the detection and classification module 116C in Figure 2C). In some implementations, a gaze detection module operating on client device 106A may process (or focus on) only that portion of the image when monitoring a gaze directed at it, based on the fact that that portion is indicated as the portion corresponding to a face. Although only a single image is shown in Figure 3, in various implementations, detection of a gaze directed at it and / or gesture detection may be based on a sequence of images as described herein.
[0086] Within image 360, a bounding box 363 is also provided, which can be determined to correspond to a video display and represent an area of the image that may cause a false positive for a visual cue. For example, a television may render video showing one or more people making gestures, looking at a camera, etc., and any of what they are doing may be misinterpreted as the occurrence of a gesture and / or gaze directed at them. In some implementations, the detection and classification module 116C in Figure 2C may determine such an area (for example, based on detecting the classification of a TV), and / or such an area may be determined by analyzing image 360 and preceding images to determine that the area has a display frequency corresponding to the display frequency of a video display (e.g., about 60Hz, 120Hz, and / or other typical video display frequencies). In some implementations, the gaze detection module and / or gesture module may crop the area from the visual frame being processed, focus outside the area, and / or otherwise ignore the area in detection or reduce the likelihood that detection is based on such an area. These and other methods may mitigate falsely detected calls to functions that were in a dormant state.
[0087] Figure 4A is a flowchart illustrating an exemplary method 400 by an implementation disclosed herein. Figure 4B is a flowchart illustrating an example implementation of blocks 402, 404, and 406 of Figure 4A. For convenience, the operations in the flowcharts of Figures 4A and 4B are described in relation to a system that performs the operations. This system may include various components of various computer systems, such as one or more components of a computing system implementing the automation assistant 120 (e.g., a client device and / or a remote computing system). Furthermore, although the operations of method 400 are shown in a particular order, this is not intended to be limiting. One or more operations may be reordered, omitted, or added.
[0088] In block 402, the system receives visual data based on the output from the visual component. In some implementations, the visual component may be integrated into a client device, including an assistant client. In some implementations, the visual component is separate from the client device but can communicate with it. For example, the visual component may include a standalone smart camera that communicates with the client device, including an assistant client, via wired and / or wireless means.
[0089] In block 404, the system processes visual data using at least one machine learning model to monitor the occurrence of both gestures and directed gazes.
[0090] In block 406, the system determines whether both a gesture and a gaze have been detected based on the monitoring in block 404. If not detected, the system returns to block 402, receives additional visual data, and performs another iteration of blocks 404 and 406. In some implementations, the system determines that both a gesture and a gaze have been detected based on detecting that the gesture and the directed gaze occur simultaneously or within a threshold temporal proximity to each other. In some additional or alternative implementations, the system determines that both a gesture and a gaze have been detected based on detecting that the gesture has at least a threshold duration (e.g., "waving" for at least duration X or "giving a thumbs-up" for at least duration X) and / or the directed gaze has at least a threshold duration (which may or may not be the same as the threshold duration used for the gesture duration).
[0091] In an iteration of block 406, if the system determines, based on monitoring in block 404, that both a gesture and a gaze have been detected, the system optionally proceeds to block 408 (or directly to block 410 if block 408 is not included).
[0092] In any block 408, the system determines whether one or more other conditions are met. If one or more other conditions are not met, the system returns to block 402, receives additional visual data, and performs another iteration of blocks 404, 406, and 408. If one or more other conditions are met, the system proceeds to block 410. The system may use the visual data, audio data, and / or other sensory or non-sensor data received in block 402 to determine whether one or more other conditions are met. Various other conditions, such as those expressly described herein, may be considered by the system.
[0093] In block 410, the system activates one or more inactive automation assistant functions. The system can activate various inactive automation assistant functions, such as those explicitly described herein. In some implementations, different types of gestures may be monitored in block 404, and which inactive functions are activated in block 410 may depend on the specific types of gestures detected in the monitoring of block 404.
[0094] In block 412, the system monitors deactivation conditions with respect to the functions of the automation assistant activated in block 410. Deactivation conditions may include, for example, a timeout, a duration of at least a threshold of the absence of detected verbal input and / or detected directed gaze, an explicit stop command (verbally, gesture-indicated, or touch-inputted), and / or other conditions.
[0095] In block 414, the system determines whether a deactivation condition has been detected based on monitoring in block 412. If no deactivation condition has been detected, the system returns to block 412 and continues monitoring for deactivation conditions. If a deactivation condition has been detected, the system can deactivate the function activated in block 410, and returns to block 402 to receive visual data again and monitor for the occurrence of both calling gestures and gazes again.
[0096] As an example of blocks 412 and 414, in which the activated function involves streaming audio data to one or more cloud-based automation assistant components, the system may stop streaming in response to detecting a lack of audio activity for at least a threshold duration (e.g., using VAD), in response to an explicit stop command, or in response to detecting (by continuous gaze monitoring) that the user's gaze has not been directed towards the client device for at least a threshold duration.
[0097] Now, turning to Figure 4B, we see an example implementation of blocks 402, 404, and 406 in Figure 4A. Figure 4B shows an example where gesture detection and directed gaze detection are performed using separate models, and directed gaze monitoring is performed only in response to the initial detection of a gesture. In Figure 4B, block 402A is a specific example of block 402 in Figure 4A, blocks 404A and 404B are specific examples of block 404 in Figure 4A, and blocks 406A and 406B are specific examples of block 406 in Figure 4A.
[0098] In block 402A, the system receives and buffers the visual data.
[0099] In block 404A, the system processes visual data using a gesture machine learning model to monitor the occurrence of gestures. In some implementations, block 404A includes a subblock 404A1 in which the system processes a portion of the visual data using a gesture machine learning model based on the detection that a portion of the visual data corresponds to a person and / or a region of the body.
[0100] In block 406A, the system determines whether a gesture has been detected based on the monitoring in block 404A. If no gesture has been detected, the system returns to block 402A, receives and buffers additional visual data, and performs another iteration of block 404A.
[0101] If, during an iteration of block 406A, the system determines that a gesture has been detected based on the monitoring in block 404A, the system proceeds to block 404B.
[0102] In block 404B, the system processes buffered and / or additional visual data using a gaze machine learning model to monitor the occurrence of directed gazes.
[0103] In some implementations, block 404B includes a subblock 404B1 in which the system processes a portion of the visual data using a gaze machine learning model based on detecting that a portion of the visual data corresponds to the region of a person and / or face.
[0104] In block 406B, the system determines whether a directed gaze has been detected based on the monitoring in block 404B. If no directed gaze has been detected, the system returns to block 402A, receives and buffers additional visual data, and performs another iteration of block 404A.
[0105] In the iteration of block 406B, if the system determines that a gesture has been detected based on the monitoring in block 404B, the system proceeds to block 408 or 410 in Figure 4A.
[0106] Various examples of activating a dormant assistant function in response to the detection of both a specific gesture and a directed gaze are described herein. However, in various implementations, the dormant assistant function may be activated in response to the detection of only one of a specific gesture and a directed gaze, in an optional combination with one or more other conditions, such as those described herein. For example, in some of those various implementations, the dormant assistant function may be activated in response to the detection of a directed gaze by the user for at least a threshold duration, along with other concurrently occurring conditions, such as the user's mouth movements. Also, for example, in some of those various implementations, the dormant assistant function may be activated in response to the detection of a user gesture, along with other concurrently occurring and / or temporally close conditions, such as detected vocal activity.
[0107] Figure 5 is a block diagram of an exemplary computing device 510 that may be optionally used to perform one or more embodiments of the technology described herein. In some implementations, one or more of the client computing device, the user-controlled resource module 130, and / or other components may include one or more components of the exemplary computing device 510.
[0108] Generally, the computing device 510 includes at least one processor 514 that communicates with several peripheral devices via a bus subsystem 512. These peripheral devices may include, for example, a storage subsystem 524 including a memory subsystem 525 and a file storage subsystem 526, a user interface output device 520, a user interface input device 522, and a network interface subsystem 516. The input and output devices enable user interaction with the computing device 510. The network interface subsystem 516 provides an interface to an external network and is coupled to corresponding interface devices of other computing devices.
[0109] The user interface input device 522 may include pointing devices such as keyboards, mice, trackballs, touchpads, or graphics tablets, audio input devices such as scanners, touchscreens integrated into displays, speech recognition systems, microphones, and / or other types of input devices. In general, the use of the term “input device” is intended to include all possible types of devices and methods for inputting information into the computing device 510 or a communication network.
[0110] The user interface output device 520 may include non-visual displays such as a display subsystem, printer, fax machine, or audio output device. The display subsystem may include flat panel devices such as cathode ray tubes (CRTs) or liquid crystal displays (LCDs), projection devices, or any other mechanism for generating visible images. The display subsystem may also provide non-visual displays such as audio output devices. In general, the use of the term “output device” is intended to include all possible types of devices and methods for outputting information from the computing device 510 to a user or another machine or computing device.
[0111] The storage subsystem 524 stores programming and data structures that provide some or all of the functionality of the modules described herein. For example, the storage subsystem 524 may include logic for performing selected embodiments of the methods shown in Figures 4A and 4B, and for implementing the various components shown in Figures 1, 2A to 2C, and 3.
[0112] These software modules are generally executed by the processor 514 alone or in combination with other processors. The memory 525 used in the storage subsystem 524 may include several memories, including a main random access memory (RAM) 530 for storing program instructions and data during execution, and a read-only memory (ROM) 532 for storing fixed instructions. The file storage subsystem 526 can provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functionality of a particular implementation may be stored by the file storage subsystem 526 within the storage subsystem 524, or on other machines that may be accessed by the processor 514.
[0113] The bus subsystem 512 provides a mechanism for various components and subsystems of the computing device 510 to communicate with each other as intended. Although the bus subsystem 512 is schematically shown as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0114] The computing device 510 can be of various types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of the computing device 510 shown in Figure 5 is intended only as a specific example to illustrate some implementations. Many other configurations of the computing device 510 are possible, having more or fewer components than the computing device shown in Figure 5.
[0115] Where the systems described herein may collect or otherwise monitor personal information about a user or use personal information and / or monitored information, the user may be given the opportunity to control whether the program or feature collects user information (e.g., information about the user's social networks, social behavior or activities, occupation, user preferences, or the user's current geographical location) or whether and / or how content that may be more relevant to the user should be received from the content server. Furthermore, certain data may be processed in one or more ways before being stored or used so that personally identifiable information is removed. For example, a user's identity may be processed so that personally identifiable information cannot be determined about the user, or if geographical location information is obtained, the user's geographical location may be generalized (to the level of city, zip code, or state, for example) so that the user's specific geographical location cannot be determined. Thus, the user may be able to control how information is collected and / or used about them. For example, in some implementations, the user may detach from the assistant device that uses the visual component 107 when monitoring the occurrence of gestures and / or gazes directed toward it, and / or uses visual data from the visual component 107. [Explanation of Symbols]
[0116] 104 buffers 105 Presence detection sensor 106 Client Computing Devices 106A Client Device 107 Visual components, visual sensors 107A Camera 109 Microphone 110 Automation Assistant Client 112 Audio Capture Module 114 Visual Capture Module 115 Calling Engine 116. Gaze and Gesture Module, Gesture and Gaze Detection Module 116A Gesture Module 116B Gaze Module 116C Detection and Classification Module, Detection and Classification Model 117 Machine Learning Models 117A Gesture-based Machine Learning Model 117B Stare-gazing machine learning model 117C Object Detection and Classification Machine Learning Models 118 Other Condition Modules 119 Other Machine Learning Models 120 Automation Assistants 130 Cloud-Based Automation Assistant Components, Resource Modules 131 Cloud-based TTS module 132 Cloud-based STT Modules 133 Natural Language Processors 134 Intention Matcha 135 Natural Language Understanding Module, Intention Understanding Module 136 Natural Language Generators 138 Execution Modules 150 Search Modules 301 users 360 images 361 Boundary Box 362 Boundary Box 400 ways 510 Computing Devices 512 Bus Subsystem 514 processors 516 Network Interface Subsystem 520 User Interface Output Devices 522 User Interface Input Devices 524 Storage Subsystems 525 Memory subsystem, memory 526 File Storage Subsystem 530 Main Random Access Memory (RAM) 532 Read-only memory (ROM)< / topping> < / topping> < / artist> < / artist> < / artist> < / artist>
Claims
1. It is a client device, Visual components and Presence detection sensor, One or more processors, wherein one or more of the processors Based on the signal from the presence detection sensor, it is detected that a person is present within the environment of the presence detection sensor. In response to detecting the presence of the person in the environment, Activating the visual component to provide a stream of visual data based on the output from the visual component, Using at least one trained machine learning model stored locally on the client device, The user's calling gesture captured by the aforementioned visual data, The user's gaze directed towards the client device and To monitor the occurrence of both, the visual data is processed, Based on the aforementioned monitoring, The aforementioned calling gesture and, The aforementioned gaze and To detect the occurrence of both, In response to detecting the occurrence of both the calling gesture and the gaze, Activate at least one of the pause functions of the automation assistant and One or more processors configured to perform the following: A client device equipped with the following features.
2. The at least one pause function of the automation assistant is activated in response to detecting the occurrence of both the calling gesture and the gaze. Sending data from the client device to the remote server associated with the automation assistant. The client device according to claim 1, comprising:
3. The client device further comprises a microphone, and the at least one pause function of the automation assistant is activated in response to detecting the occurrence of both the calling gesture and the gaze, Automatically processes audio data based on the output from the aforementioned microphone. The client device according to claim 1, comprising:
4. The at least one pause function of the automation assistant is activated in response to detecting the occurrence of both the calling gesture and the gaze. Graphically rendering content tailored to the user profile of the aforementioned user. The client device according to claim 1, comprising:
5. One or more of the aforementioned processors The system is further configured to determine whether the user profile is authorized for the client device. Activating the at least one of the idle functions of the automation assistant is subject to the condition that the user profile is authorized for the client device. The client device according to claim 4.
6. The client device according to claim 5, wherein one or more of the processors are further configured to determine the user profile of the user based on processing the visual data.
7. The client device according to claim 6, wherein, in determining the user profile of the user based on processing the visual data, one or more of the processors are configured to perform face recognition based on processing the visual data.
8. The client device according to claim 1, wherein, in processing the visual data to monitor the occurrence of the user's gaze, one or more of the processors are configured to process the visual data using a trained gaze machine learning model from the at least one trained machine learning model.
9. A method implemented by one or more processors in a client device to facilitate contactless interaction between one or more users and an automated assistant, Steps include detecting the presence of a person within the environment of the presence detection sensor based on a signal from the presence detection sensor, In response to the step of the person detecting a step present in the environment, The steps include activating the visual component to provide a stream of visual data based on the output from the visual component, Using at least one trained machine learning model stored locally on the client device, The user's calling gesture captured by the aforementioned visual data, The user's gaze directed towards the client device and The steps include processing the visual data to monitor the occurrence of both events, Based on the aforementioned monitoring step, The aforementioned calling gesture and, The aforementioned gaze and The steps include detecting the occurrence of both events, In accordance with the step of detecting the occurrence of both the calling gesture and the gaze, Steps to activate at least one pause function of the automation assistant and A method that includes [a certain feature].
10. The at least one pause function of the automation assistant is activated in response to the step of detecting the occurrence of both the calling gesture and the gaze. Steps to send data from the client device to the remote server associated with the automation assistant. The method according to claim 9, comprising:
11. The at least one pause function of the automation assistant is activated in response to the step of detecting the occurrence of both the calling gesture and the gaze. The step of automatically processing audio data based on the output from the microphone of the client device. The method according to claim 9, comprising:
12. The at least one pause function of the automation assistant is activated in response to the step in which the occurrence of both the calling gesture and the gaze is detected. A step to graphically render content tailored to the user profile of the aforementioned user. The method according to claim 9, comprising:
13. The method according to claim 12, further comprising the step of determining the user profile of the user based on processing the visual data.
14. The step further comprises determining whether the user profile is authorized for the client device. The step of activating the at least one hibernation function of the automation assistant is conditional on the user profile being authorized for the client device. The method according to claim 13.
15. The method according to claim 9, wherein the step of processing the visual data to monitor the occurrence of the user's gaze comprises processing the visual data using a trained gaze machine learning model from the at least one trained machine learning model.
Citation Information
Patent Citations
Gesture detector using ultrasound
EP3133474A1
Portable device and method for providing voice recognition service
US20140195841A1
Intelligent Amplifier Amplification
US20170048637A1