Selective detection of visual cues for automated assistants

By implementing the excluded area classification technology in the assistant device, identifying and classifying areas that may contain visual noise, the problem of wrong affirmation in the assistant device is solved, and more efficient and secure contactless interaction is achieved.

CN120179059APending Publication Date: 2025-06-20GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510140749.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2018-05-04
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Assistant devices equipped with vision sensors are prone to false affirmations in contactless interactions, such as misunderstandings that visual content on a TV or computer screen is a user's visual prompt, resulting in the automation assistant performing improper actions.

Method used

By implementing exclusion area classification techniques in the assistant device, identifying and classifying areas that may contain visual noise, reducing or eliminating the impact of these areas in subsequent image frame analysis, thereby reducing the probability of error affirmation.

Benefits of technology

It effectively reduces the occurrence of error affirmation events, saves computing resources, improves the safety and efficiency of the system, and prevents unnecessary power consumption and potential security risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179059A_ABST
    Figure CN120179059A_ABST
Patent Text Reader

Abstract

The invention relates to selective detection of visual cues for automated assistants. One or more initial image frames may be obtained from one or more visual sensors of an assistant device and analyzed to classify a particular region of the initial image frame as likely to contain visual noise. One or more subsequent image frames obtained from one or more visual sensors may be analyzed in a manner that reduces or eliminates error affirmations to detect one or more actionable user-provided visual cues. Any analysis may not be performed on a particular region of one or more subsequent image frames. A weight of a first candidate visual cue detected within a particular region may be set to be less than a weight of a second candidate visual cue detected in other regions in one or more subsequent image frames. The automated assistant may take a responsive action based on the one or more detected actionable visual cues.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Division Explanation

[0002] This application is a divisional application of Chinese Patent Application No. 201880094188.7 with an application date of May 4, 2018. Background Art

[0003] Humans can participate in a human - machine conversation with an interactive software application referred to herein as an "automation assistant" (also known as a "chatbot", "interactive personal assistant", "intelligent personal assistant", "personal voice assistant", "conversation agent", etc.). For example, a human (who may be referred to as a "user" when they interact with the automation assistant) can use free - form natural - language input to provide commands, queries, and / or requests (collectively referred to herein as "queries"), which can include spoken words and / or non - typed form natural - language input that are converted to text and then processed. In many cases, the automation assistant must first be "invoked" using, for example, a predefined verbal invocation phrase. Summary of the Invention

[0004] As automation assistants become more prevalent, computing devices can be specifically designed to facilitate interaction with automation assistants - referred to herein as "assistant devices". Assistant devices can enable users to participate in contact - less interactions with automation assistants. For example, an assistant device can include a microphone that allows a user to provide spoken words as input. Additionally, an assistant device can include visual sensors such as cameras, passive infrared ("PIR") sensors, etc., which can detect presence, gestures, etc.

[0005] On an assistant device equipped with a visual sensor, the automation assistant can be invoked either alone or in combination with verbal utterances by one or more predefined visual cues provided by the user (such as gestures). For example, relatively subtle visual cues such as a user's gaze to a specific reference point (e.g., directly into the visual sensor) can be combined with verbal utterances from the user to invoke the automation assistant. Additionally or alternatively, less subtle gestures such as waving a hand at the visual sensor, predefined gestures (e.g., the user forming a predefined shape with their hand), etc. can be used alone (or in combination with verbal utterances) to invoke the automation assistant. Furthermore, the automation assistant can interpret the visual cues provided by the user to take various different actions after invocation. For example, a signal of a user "giving a thumbs - up" can be interpreted as an affirmative response to a question posed by the automation assistant.

[0006] One challenge faced by assistant devices equipped with vision sensors is false positives. For example, assume that a visual content source such as a television, a photograph, or other images (animated or still) is visible within the field of view of the assistant device. The visual content provided by the source may be mistaken for a visual cue intended to invoke the automated assistant and / or cause the automated assistant to perform an action. As an example, assume that an automated assistant interacting with a user on an assistant device equipped with a vision sensor poses a question seeking a yes / no response from the user, such as "Are you sure you want me to place this order?". Further assume that on the television visible within the field of view of the assistant device, just after the question is posed but before the user has a chance to respond, a television character happens to give a thumbs up gesture. It is quite possible that the detected thumbs up gesture will be interpreted as an affirmative response from the user. This can be particularly troublesome for the user if the user has changed their mind about sharing the order placement and, as further outlined below, can pose a safety risk. Additionally, as described below, the computing devices used to implement the assistant and the computing devices used in associated third-party services may suffer from suboptimal and unnecessary consumption of their computing resources following a false positive. False positives can also, for example, result in suboptimal power usage within the system.

[0007] Techniques for reducing and / or eliminating false positives in assistant devices equipped with vision sensors are described herein. In some embodiments, regions of an image frame captured by the assistant device can be classified as likely containing visual noise and / or unlikely to contain visual cues. Later, when analyzing subsequent image frames captured by the assistant device and attempting to detect visual cues provided by the user, those same regions can be ignored or at least their weight is set to be less than the weight of other regions. This can reduce or eliminate false positive visual cues generated by, for example, televisions, computer screens, still images, etc.

[0008] In a process referred to herein as "exclusion zone classification", regions of an image frame (or, in other words, regions of the field of view of a vision sensor) can be classified as likely to contain visual noise and / or unlikely to contain visual cues provided by a user at different times using various techniques. In various embodiments, exclusion zone classification can be performed when an assistant device is initially placed in a location (e.g., a tabletop or countertop), whenever the assistant device moves, and / or if the vision sensor of the assistant device is adjusted (e.g., panned, tilted, zoomed). Additionally or alternatively, in some embodiments, exclusion zone classification can be performed periodically (e.g., daily, weekly, etc.) to account for environmental changes (e.g., a television or computer being repositioned, removed, etc.). In some embodiments, exclusion zone classification can be performed in response to other stimuli such as changes in light (e.g., day vs. night), determining that a television or computer has been turned off (in which case false positives are no longer likely), determining that a television or computer (particularly a laptop or tablet computer) has moved, time of day (e.g., a television is unlikely to be on at midnight or when the user is at work), etc.

[0009] Various techniques can be employed to perform exclusion zone classification. In some embodiments, a machine learning model such as a convolutional neural network can be trained to identify objects in an image frame captured by an assistant device that are likely to generate false positives, such as a television, computer screen, projection screen, still image (e.g., a photo on a wall), electronic photo frame, etc. Then, regions of interest ("ROIs") that potentially contain sources of false positives can be generated. During subsequent analysis, these ROIs can be ignored, or visual cues detected in these regions can be viewed with skepticism to detect visual cues provided by the user.

[0010] Other conventional object recognition techniques can also be employed, such as methods that rely on computer-aided design ("CAD") models, feature-based methods (e.g., surface patches, linear edges), appearance-based methods (e.g., edge matching, divide-and-conquer, gradient matching, histograms, etc.), genetic algorithms, etc. Additionally or alternatively, for objects such as a television or computer screen, other techniques can be used to identify these objects. In some embodiments, a television can be identified in a sequence of image frames based on the display frequency of the television. Assume that the sequence of image frames is captured at twice the frequency of a typical television. In every other image frame of the sequence of image frames, a new image will appear on the television. This can be detected and used to determine ROIs of the television that can be ignored and / or weighted less.

[0011] Objects that might interfere with visual cue detection, such as a television or computer screen, may not always produce noise. For example, if the television is off, graphics that might interfere with visual cue detection cannot be rendered. Accordingly, in some embodiments, the automated assistant (or another process associated therewith) can determine whether the television is currently rendering graphics (and thus poses a risk of causing false positive visual cues). For example, in embodiments where the television is detected based on the display frequency of the television, the absence of such a frequency detection can be interpreted to mean that the television is not currently rendering graphics. Additionally or alternatively, in some embodiments, the television can be a "smart" television that communicates with the intelligent assistant over a network, e.g., because the television is part of the same coordinated "ecosystem" of client devices that includes the assistant device attempting to detect visual cues. In some such embodiments, the automated assistant can determine the state of the television, e.g., "on", "off", "active", "sleep", "screensaver", etc., and can include or exclude the ROI of the television based on that determination.

[0012] When subsequent image frames are captured by the assistant device, the regions (i.e., two-dimensional spatial portions) in those subsequent image frames that were previously classified as potentially containing visual noise and / or unlikely to contain user-provided visual cues can be processed in various ways. In some embodiments, those classified regions can simply be ignored (e.g., analyzing those subsequent image frames can suppress the analysis of those regions). Additionally or alternatively, in some embodiments, the weights of candidate visual cues (e.g., hand gestures, gazes, etc.) detected in those classified regions can be set to be less than the weights of candidate visual cues detected in other regions in the subsequent image frames, for example.

[0013] The techniques described herein can yield various technical advantages and benefits. By way of example, ignoring the classification regions of image frames can save computing resources, for instance, by making one or more processors available to focus on regions more likely to contain visual cues. As another example, false positives can trigger an automated assistant and / or an assistant device to take various actions that waste computing resources and / or power and can confuse or even disorient a user. Such wasteful or otherwise inefficient or unnecessary use of power and computing resources can occur in the assistant device itself (e.g., a client device) and / or in remote computing devices that communicate with the assistant device to perform various actions, such as one or more network servers. Additionally, unnecessary communication with remote computing devices results in an unnecessary load on the communication network. The techniques described herein reduce the number of false positives for detected visual cues. The techniques described herein also provide advantages from a security perspective. For example, a malicious user might remotely take control of a television or computer screen to render one or more visual cues that might trigger an unwanted response action by an automated assistant and / or an assistant device (e.g., turning on a camera, unlocking a door, etc.). By ignoring (or at least setting a lower weight) the image frame regions that include the television / computer screen, such security vulnerabilities can be prevented.

[0014] In some embodiments, a method executed by one or more processors is provided that facilitates contactless interaction between one or more users and an automated assistant. The method includes: obtaining, from one or more visual sensors, one or more initial image frames, and analyzing the one or more initial image frames to classify a particular region of the one or more initial image frames as likely to contain visual noise. The method further includes: obtaining, from one or more visual sensors, one or more subsequent image frames, and analyzing the one or more subsequent image frames to detect one or more actionable visual cues provided by one or more users. Analyzing the one or more subsequent image frames includes: suppressing the analysis of a particular region of the one or more subsequent image frames, or setting the weight of a first candidate visual cue detected within a particular region of the one or more subsequent image frames to be less than a second candidate visual cue detected in other regions of the one or more subsequent image frames. The method further includes: causing the automated assistant to take one or more response actions based on the one or more detected actionable visual cues.

[0015] These and other embodiments of the disclosed techniques can include one or more of the following features.

[0016] In some embodiments, analyzing the one or more initial image frames includes detecting an electronic display captured in the one or more initial image frames, and a specific area of ​​the one or more initial image frames contains the detected electronic display. In some versions of those embodiments, the electronic display is detected using an object recognition process. In some additional or alternative versions of those embodiments, detecting the electronic display includes detecting a display frequency of the electronic display. In some other additional or alternative versions, the suppression or weighting is conditionally performed based on determining whether the electronic display is currently rendering graphics.

[0017] In some embodiments, analyzing the one or more initial image frames includes detecting a picture frame captured in the one or more initial image frames, and a specific area of ​​the one or more image frames contains the detected picture frame.

[0018] In some implementations, the one or more responsive actions include invoking an automated assistant, and the automated assistant is invoked based on the one or more detected actionable visual cues in conjunction with utterances from the one or more users.

[0019] In some implementations, the one or more responsive actions include invoking an automated assistant, and the automated assistant is only invoked based on the one or more detected actionable visual cues.

[0020] In some embodiments, the one or more detected actionable visual cues include: the user looking in the direction of the reference point, the user making a hand gesture, the user having a particular facial expression, and / or the user's position in one or more subsequent image frames.

[0021] In some implementations, analyzing the one or more initial image frames to classify a particular region of the one or more initial image frames as likely to contain visual noise includes associating the particular region of the one or more initial image frames with a visual noise indicator.

[0022] In addition, some embodiments include one or more processors of one or more computing devices, wherein the one or more processors are operable to execute instructions stored in an associated memory, and wherein the instructions are configured to cause the execution of any of the foregoing methods. Some embodiments also include one or more non-transitory computer-readable storage media storing computer instructions executable by the one or more processors to perform any of the foregoing methods.

[0023] It should be appreciated that all combinations of the above-described concepts and additional concepts described in more detail herein are contemplated as being part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are contemplated as being part of the subject matter disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is a block diagram of an example environment in which the embodiments disclosed herein can be implemented.

[0025] Figure 2 Illustrates an exemplary processing flow showing various aspects of the present disclosure according to various embodiments.

[0026] Figure 3 Illustrates examples of what the field of view of a vision sensor of an auxiliary device may include.

[0027] Figure 4 Illustrates an example of an image frame having regions of interest classified as likely to contain visual noise and / or unlikely to contain visual cues according to various embodiments Figure 3 of.

[0028] Figure 5 Illustrates a flowchart diagramming an example method according to the embodiments disclosed herein.

[0029] Figure 6 Diagrams an example architecture of a computing device. DETAILED DESCRIPTION

[0030] Turning now to Figure 1 , an example environment is illustrated in which the techniques disclosed herein can be implemented. The example environment includes one or more client computing devices 106. Each client device 106 can execute a respective instance of an automated assistant client 108. One or more cloud-based automated assistant components 119, such as a natural language understanding module 135, can be implemented on one or more computing systems (collectively referred to as "cloud" computing systems) communicatively coupled to the client devices 106 via one or more local area networks and / or wide area networks (e.g., the Internet), typically indicated at 114.

[0031] In various embodiments, an instance of the automated assistant client 108, through its interaction with one or more cloud-based automated assistant components 119, can form in the form of a logical instance of an automated assistant 120 that, from the user's perspective, appears to be an automated assistant with whom the user can engage in a human-machine conversation. An instance of such an automated assistant 120 is shown in Figure 1Shown in dashed lines in the figure. It should thus be understood that each user engaging with the automated assistant client 108 executing on the client device 106 can in fact engage with a logical instance of his or her own automated assistant 120. For simplicity and brevity, the term "automated assistant" as used herein, such as "serving" a particular user, will refer to the combination of the automated assistant client 108 executing on the client device 106 operated by the user and one or more cloud-based automated assistant components 119 (which may be shared among multiple automated assistant clients 108). It should also be understood that in some embodiments, the automated assistant 120 can respond to requests from any user, regardless of whether that user is actually "served" by that particular instance of the automated assistant 120.

[0032] One or more client devices 106 can include, for example, one or more of the following: a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a computing device of the user's vehicle (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a stand-alone interactive speaker (which may include a visual sensor in some cases), a smart appliance such as a smart TV (or a standard TV equipped with a networking dongle with automated assistant capabilities), and / or a wearable device of the user that includes a computing device (e.g., a user's watch with a computing device, a user's glasses with a computing device, a virtual or augmented reality computing device). Additional and / or alternative client computing devices may be provided. As previously described, some client devices 106 may take the form of an assistant device that is primarily designed to facilitate a conversation between the user and the automated assistant 120 (e.g., a stand-alone interactive speaker).

[0033] For the purposes of the present disclosure, the client device 106 can be equipped with one or more visual sensors 107 having one or more fields of view. The visual sensors 107 can take various forms, such as a digital camera, a passive infrared ("PIR") sensor, a stereo camera, an RGBd camera, etc. One or more visual sensors 107 can be used, for example, by the image capture 111 to capture image frames (still images or videos) of the environment in which the client device 106 is deployed. These image frames can then be analyzed, for example, by the visual cue module 112 to detect visual cues provided by the user included in the image frames. These visual cues can include, but are not limited to, hand gestures, a gaze towards a particular reference point, facial expressions, predefined movements made by the user, etc. These detected visual cues can be used for various purposes and will be described further below.

[0034] As described in more detail herein, the automated assistant 120 participates in a human-machine conversation session with one or more users via the user interface input and output devices of one or more client devices 106. In some embodiments, the automated assistant 120 can participate in a human-machine conversation session with a user in response to user interface input provided by the user via one or more user interface input devices of one of the client devices 106. In some of those embodiments, the user interface input is explicitly directed to the automated assistant 120. For example, the user can speak a pre-determined invocation phrase, such as "OK, Assistant" or "Hey, Assistant", to cause the automated assistant 120 to start actively listening. Additionally or alternatively, in some embodiments, the automated assistant 120 can be invoked based on one or more detected visual cues, either alone or in combination with a spoken invocation phrase.

[0035] In some embodiments, even when the user interface input is not explicitly directed to the automated assistant 120, the automated assistant 120 can participate in a human-machine conversation session in response to the user interface input. For example, the automated assistant 120 can examine the content of the user interface input and participate in the conversation session in response to the presence of certain terms in the user interface input and / or based on other cues. In many embodiments, the automated assistant 120 can utilize speech recognition to convert the user's spoken words into text and respond accordingly to the text, e.g., by providing search results, general information, and / or taking one or more response actions (e.g., playing media, starting a game, ordering food, etc.). In some embodiments, the automated assistant 120 is additionally or alternatively able to respond to spoken words without converting the words into text. For example, the automated assistant 120 can convert the voice input into an embedding, into an entity representation (which indicates one or more entities present in the voice input), and / or other "non-text" representations, and operate on such non-text representations. Thus, embodiments described herein as operating based on text converted from voice input can additionally and / or alternatively operate directly on the voice input and / or other non-text representations of the voice input.

[0036] Each of the client computing devices 106 and the computing devices operating the cloud-based automated assistant component 119 can include one or more memories for storing data and software applications, one or more processors for accessing the data and executing the applications, and other components that facilitate communication over a network. The operations performed by the client computing devices 106 and / or by the automated assistant 120 can be distributed across multiple computer systems. The automated assistant 120 can be implemented as, for example, a computer program running on one or more computers coupled to each other via a network in one or more locations.

[0037] As noted above, in various embodiments, the client computing device 106 may operate an automated assistant client 108. In various embodiments, the automated assistant client 108 may include a voice capture module 110, the aforementioned image capture module 111, a visual cue module 112, and / or an invocation module 113. In other embodiments, one or more aspects of the voice capture module 110, the image capture module 111, the visual cue module 112, and / or the invocation module 113 may be implemented separately from the automated assistant client 108, for example, by one or more cloud-based automated assistant components 119.

[0038] In various embodiments, the voice capture module 110, which may be implemented using any combination of hardware and software, may interface with hardware such as a microphone 109 or other pressure sensors to capture an audio recording of the user's utterances. Various types of processing may be performed on this audio recording for various purposes, as will be described below. In various embodiments, the image capture module 111, which may be implemented using any combination of hardware or software, may be configured to interface with a camera 107 to capture one or more image frames (e.g., digital photos) corresponding to the field of view of the visual sensor 107.

[0039] In various embodiments, the visual cue module 112, which may be implemented using any combination of hardware or software, may be configured to analyze one or more image frames provided by the image capture module 111 to detect one or more visual cues captured in and / or across the one or more image frames. The visual cue module 112 may employ various techniques to detect visual cues. For example, in Figure 1 this example, the visual cue module 112 is communicatively coupled to a visual cue model database 114, which may be integral to the client device 106 and / or hosted remotely from the client device 106, such as in the cloud. The visual cue model database 114 may include, for example, one or more artificial intelligence (or machine learning) models that are trained to generate outputs indicative of user-provided visual cues detected in the image frames.

[0040] As a non-limiting example, a neural network such as a convolutional neural network can be trained (and stored in database 114) such that one or more image frames—or feature vectors extracted from the image frames—can be used as inputs across the neural network. In various embodiments, the convolutional neural network can generate an output indicating a plurality of detected visual cues and the associated probabilities of detecting each visual cue. In some such embodiments, the output can further indicate the location in the image frame where the visual cue was detected in the image frame, although this is not required. Various forms of training examples can be used to train such a convolutional neural network, various forms of training examples such as sequential image frames (or resulting feature vectors) labeled with gestures known to be included in or across a sequence of image frames. When applying the training examples across the network, the difference between the generated output and the label associated with that training example can be used, e.g., to minimize a loss function. Then, standard techniques such as gradient descent and / or backpropagation can be used, for example, to adjust the various weights of the convolutional neural network.

[0041] In various embodiments, the visual cue module 112 can be configured to perform selected aspects of the present disclosure to reduce and / or eliminate false positive visual cues. For example, the visual cue module 112 can participate in the previously described "exclusion region classification" where it analyzes one or more image frames captured by the visual sensor 107 to classify one or more regions within the field of view of the visual sensor 107 as potential sources of visual noise that may result in detecting false positive visual cues. These regions may include areas with television screens, computer monitors, photographs (e.g., digital photos in an LCD / LED photo frame and / or static photos printed on paper), and so on. The visual cue module 112 can employ a variety of different techniques to detect regions of potential noise.

[0042] For example, in some embodiments, the visual cue module 112 can employ a variety of object recognition techniques to identify objects that potentially create noise, such as televisions and computer monitors (televisions, computer monitors, smartphone screens, tablet screens, smartwatch screens, or other similar displays presenting digital images and / or videos can be collectively referred to as "electronic displays"). Once these objects are detected, they can be used, for example, by the visual cue module 112 to classify regions of interest that contain those detected objects and that may thus be sources of visual noise that may result in detecting false positive visual cues.

[0043] The visual cue module 112 can employ various object recognition techniques, such as methods based on computer-aided design (“CAD”) models, machine learning techniques (e.g., using trained convolutional neural networks), feature-based methods (e.g., surface patches, linear edges), appearance-based methods (e.g., edge matching, divide-and-conquer, gradient matching, histograms, etc.), genetic algorithms, and so on. Additionally or alternatively, for objects such as a television or computer screen, other techniques can be employed to identify those objects. In some embodiments, a television can be identified in a sequence of image frames based on the display frequency of the television. Assume that the sequence of image frames is captured at twice the television frequency. In every other image frame of the sequence of image frames, a new image will appear on the television. This can be detected and used to determine the ROI of the television.

[0044] When subsequent image frames are captured by the visual sensor 107 of the client device 106, the regions of those subsequent image frames that were previously classified as potentially being sources of visual noise can be processed in various ways. In some embodiments, those regions can simply be ignored (e.g., the analysis of those subsequent image frames can suppress the analysis of those regions). Additionally or alternatively, in some embodiments, the weights of candidate visual cues (e.g., hand gestures, gazes, etc.) detected in those classified regions can be lower than the weights of candidate visual cues detected elsewhere in the subsequent image frames, for example.

[0045] As previously mentioned, the speech capture module 110 can be configured to capture the user's speech via, for example, the microphone 109. Additionally or alternatively, in some embodiments, the speech capture module 110 can be further configured to convert the captured audio to text and / or to other representations or embeddings using, for example, speech-to-text (“STT”) processing techniques. Additionally or alternatively, in some embodiments, the speech capture module 110 can be configured to convert text to computer-synthesized speech using, for example, one or more voice synthesizers. However, because the client device 106 may be relatively constrained in terms of computing resources (e.g., processor cycles, memory, battery, etc.), the speech capture module 110 local to the client device 106 can be configured to convert a limited number of different spoken phrases—particularly phrases that invoke the automated assistant 120—to text (or to other forms, such as lower-dimensional embeddings). Other speech inputs can be sent to the cloud-based automated assistant component 119, which can include a cloud-based TTS module 116 and / or a cloud-based STT module 117.

[0046] In various embodiments, the invocation module 113 can be configured to determine whether to invoke the automated assistant 120, for example, based on the output provided by the speech capture module 110 and / or the visual cue module 112 (which in some embodiments can be combined with the image capture module 111 in a single module). For example, the invocation module 113 can determine whether the user's utterance qualifies as an invocation phrase that should initiate a human-machine conversation session with the automated assistant 120. In some embodiments, the invocation module 113 can analyze data indicative of the user's utterance, such as an audio recording or a vector of features extracted from the audio recording (e.g., embedded) in combination with one or more visual cues detected by the visual cue module 112. In some embodiments, when a specific visual cue is also detected, the threshold used by the invocation module 113 to determine whether to invoke the automated assistant 120 in response to an audible utterance can be lowered. Thus, even when the user provides an audible utterance that is different but phonetically slightly similar to an appropriate invocation phrase (such as "OK assistant"), the utterance can still be accepted as an invocation when detected in combination with a visual cue (e.g., the speaker waving, the speaker directly gazing at the visual sensor 107, etc.).

[0047] In some embodiments, an on-device invocation model can be used by the invocation module 113 to determine whether an utterance and / or a visual cue qualifies as an invocation. Such an on-device invocation model can be trained to detect changes in invocation phrases / gestures. For example, in some embodiments, training examples can be used to train the on-device invocation model (e.g., one or more neural networks), each training example including an audio recording of the user's utterance (or an extracted feature vector) and data indicative of one or more image frames captured simultaneously with the utterance and / or detected visual cues.

[0048] The cloud-based TTS module 116 can be configured to utilize the virtually infinite resources of the cloud to convert text data (e.g., a natural language response formulated by the automated assistant 120) into computer-generated speech output. In some embodiments, the TTS module 116 can provide the computer-generated speech output to the client device 106 for direct output, for example, using one or more speakers. In other embodiments, the text data (e.g., natural language response) generated by the automated assistant 120 can be provided to the speech capture module 110, which can then convert the text data into computer-generated speech for local output.

[0049] The cloud-based STT module 117 can be configured to utilize the virtually infinite resources of the cloud to convert the audio data captured by the voice capture module 110 into text, which can then be provided to the natural language understanding module 135. In some embodiments, the cloud-based STT module 117 can convert an audio recording of speech into one or more phonemes and then convert the one or more phonemes into text. Additionally or alternatively, in some embodiments, the STT module 117 can employ a state decoding graph. In some embodiments, the STT module 117 can generate multiple candidate text interpretations of the user's utterance. In some embodiments, the STT module 117 can weight or bias a particular candidate text interpretation over others depending on whether there are simultaneously detected visual cues. For example, assume that two candidate text interpretations have similar confidence scores. With a conventional automated assistant 120, the user may be asked to disambiguate between these candidate text statements. However, with an automated assistant 120 configured with selected aspects of the present disclosure, one or more detected visual cues can be used to "break the tie".

[0050] The automated assistant 120 (particularly the cloud-based automated assistant component 119) can include the natural language understanding module 135, the aforementioned TTS module 116, the aforementioned STT module 117, and other components described in more detail below. In some embodiments, one or more of the modules and / or the modules of the automated assistant 120 can be omitted, combined, and / or implemented in a component separate from the automated assistant 120. In some embodiments, for privacy protection, one or more of the components of the automated assistant 120, such as the natural language processor 122, the TTS module 116, the STT module 117, etc., can be implemented at least partially on the client device 106 (e.g., excluded from the cloud).

[0051] In some embodiments, the automated assistant 120 generates response content in response to various inputs generated by a user of one of the client devices 106 during a human-computer dialogue session with the automated assistant 120. The automated assistant 120 can provide the response content (e.g., via one or more networks when separate from the user's client device) for presentation to the user as part of the dialogue session. For example, the automated assistant 120 can generate response content in response to free-form natural language input provided via the client device 106. As used herein, free-form input is formulated by the user and is not limited to a set of options presented for the user to select from.

[0052] As used herein, a "conversation session" can include a logically self - contained exchange of one or more messages between a user and the automated assistant 120 (and, in some cases, other human participants). The automated assistant 120 can distinguish multiple conversation sessions with a user based on various signals such as: time elapsing between sessions, the user context (e.g., location, before / during / after a scheduled meeting, etc.) changing between sessions, detecting one or more intermediate interactions between the user and the client device other than the conversation between the user and the automated assistant (e.g., the user switches applications for a while, the user leaves and then returns to a stand - alone voice - activated product), the client device locking / sleeping between sessions, changing the client device used to dock with one or more instances of the automated assistant 120, etc.

[0053] The natural language processor 122 of the natural language understanding module 135 processes natural language input generated by the user via the client device 106 and can generate an annotation output (e.g., in text form) for use by one or more other components of the automated assistant 120. For example, the natural language processor 122 can process natural language free - form input generated by the user via one or more user interface input devices of the client device 106. The generated annotation output includes one or more annotations of the natural language input and one or more (e.g., all) of the terms of the natural language input.

[0054] In some embodiments, the natural language processor 122 is configured to identify and annotate various types of syntactic information in the natural language input. For example, the natural language processor 122 can include a lexical module that can break individual words into morphemes and / or annotate morphemes, e.g., by their class. The natural language processor 122 can also include a part - of - speech tagger configured to annotate terms according to their syntactic roles. For example, the part - of - speech tagger can tag each term according to its part of speech such as "noun", "verb", "adjective", "pronoun", etc. Additionally, for example, in some embodiments the natural language processor 122 can additionally and / or alternatively include a dependency parser (not shown) configured to determine syntactic relationships between terms in the natural language input. For example, the dependency parser can determine which terms modify other terms, the subject and verb of a sentence, etc. (e.g., a parse tree) - and can annotate such dependencies.

[0055] In some embodiments, the natural language processor 122 may additionally and / or alternatively include an entity tagger (not shown) configured to annotate entity references in one or more segments, such as references to people (including, for example, literary characters, celebrities, public figures, etc.), organizations, locations (real and fictional), and the like. In some embodiments, data about entities may be stored in one or more databases, such as in a knowledge graph (not shown). In some embodiments, the knowledge graph may include nodes representing known entities (and in some cases, entity attributes) and edges connecting the nodes and representing relationships between the entities. For example, a "banana" node may be connected (e.g., as a child) to a "fruit" node, which in turn may be connected (e.g., as a child) to a "product" and / or "food" node. As another example, a restaurant called "Hypothetical Café" may be represented by nodes that also include attributes such as its address, the types of food served, business hours, contact information, etc. The "Hypothetical Café" node may be connected in some embodiments by edges (e.g., representing a child-to-parent relationship) to one or more other nodes, such as a "restaurant" node, a "business" node, a node representing the city and / or state in which the restaurant is located, etc.

[0056] The entity tagger of the natural language processor 122 may annotate references to entities at a high granularity level (e.g., such that all references to entity classes such as people can be identified) and / or at a lower granularity level (e.g., such that all references to specific entities such as specific people can be identified). The entity tagger may rely on the content of the natural language input to resolve specific entities and / or may optionally communicate with a knowledge graph or other entity database to resolve specific entities.

[0057] In some embodiments, the natural language processor 122 may additionally and / or alternatively include a coreference resolver (not shown) configured to group or "cluster" references to the same entity based on one or more context cues. For example, the coreference resolver may be used to resolve the term "there" in the natural language input "I liked HypotheticalCafe last time we ate there" to "Hypothetical Cafe".

[0058] In some embodiments, one or more components of the natural language processor 122 may rely on annotations from one or more other components of the natural language processor 122. For example, in some embodiments, the named entity tagger may rely on annotations from the coreference resolver and / or the dependency resolver in annotating all mentions of a particular entity. Additionally, for example, in some embodiments, the coreference resolver may rely on annotations from the dependency resolver in clustering references to the same entity. In some embodiments, when processing a particular natural language input, one or more components of the natural language processor 122 may use relevant prior inputs and / or other relevant data outside of the particular natural language input to determine one or more annotations.

[0059] The natural language understanding module 135 may also include an intent matcher 136 that is configured to determine the intent of a user participating in a human-machine conversation session with the automated assistant 120. In other embodiments, although shown separately from the Figure 1 natural language processor 122 in, the intent matcher 136 may be a component of the natural language processor 122 (or more generally, of a pipeline that includes the natural language processor 122). In some embodiments, the natural language processor 122 and the intent matcher 136 may together form the aforementioned "natural language understanding" module 135.

[0060] The intent matcher 136 may use various techniques to determine the user's intent, for example, based on the output from the natural language processor 122 (which may include annotations and terms of the natural language input) and / or based on the output from the visual cue module 113. In some embodiments, the intent matcher 136 may be able to access one or more databases (not shown) that include, for example, multiple mappings between syntax, visual cues, and response actions (or more generally, intents). In many cases, these syntaxes may be selected and / or learned over time and may represent the most common intents of the users. For example, a syntax "play <artist>("Play <Artist>") is mapped to a call such that on the client device 106 operated by the user, press <artist>The intent of the response action to play music. Another grammar "[weather | forecast]today (Today's [weather | forecast])" may be able to match user queries such as "what’s the weather today" and "what’s the forecast for today?".

[0061] In addition to or instead of grammar, in some embodiments, the intent matcher 136 may employ one or more trained machine learning models alone or in combination with one or more grammars and / or visual cues. These trained machine learning models may also be stored in one or more databases and may be trained to recognize intents, for example, by embedding data indicative of the user's utterance and / or any detected visual cues provided by the user into a reduced-dimensional space and then using techniques such as Euclidean distance, cosine similarity, etc. to determine which other embeddings (and thus, intents) are the closest.

[0062] Such as "play <artist>”As seen in the example grammar, some grammars have slots that can be filled with slot values (or "parameters") (e.g., <artist>). The slot values can be determined in various ways. Often the user will actively provide the slot values. For example, for the grammar "Order me a <topping>pizza (order me a <topping> pizza)”, the user is likely to say the phrase "order me a sausage pizza", in which case the slot <topping>are automatically populated. Additionally or alternatively, if the user invokes a grammar that includes slots to be filled with slot values, the automated assistant 120 can solicit those slot values from the user without the user actively providing them (e.g., "what type of crust do you want on your pizza?"). In some embodiments, slots can be filled with slot values based on visual cues detected by the visual cue module 112. For example, the user can utter something like "ORDER me this many cat bowls" while holding up three fingers to the visual sensor 107 of the client device 106. Or, the user can utter something like "Find me more movies like this" while holding a particular movie DVD case.

[0063] In some embodiments, the automated assistant 120 can facilitate (or "broker") a transaction between the user and an agent, which can be an independent software process that receives input and provides a response output. Some agents can take the form of third-party applications that may or may not operate on a computing system separate from the computing system operating, e.g., the cloud-based automated assistant component 119. One user intent that can be recognized by the intent matcher 136 is to engage a third-party application. For example, the automated assistant 120 can provide access to an application programming interface ("API") to a pizza delivery service. The user can invoke the automated assistant 120 and provide a command such as "I’d like to order a pizza". The intent matcher 136 can map this command to a grammar that triggers the automated assistant 120 to engage the third-party pizza delivery service (in some cases by the third party adding it to the database 137). The third-party pizza delivery service can provide the automated assistant 120 with a minimal list of the slots of the command that need to be filled to fulfill the pizza delivery order. The automated assistant 120 can generate a natural language output and provide it to the user (via the client device 106) that solicits parameters for the slots.

[0064] The fulfillment module 124 can be configured to receive the predicted / estimated intent output by the intent matcher 136 and the associated slot values (whether provided actively by the user or from the user's request) and fulfill (or "resolve") the intent. In various embodiments, the fulfillment (or "resolution") of the user's intent can cause, for example, various fulfillment information (also referred to as "response" information or data) to be generated / obtained by the fulfillment module 124. As will be described below, in some embodiments, the fulfillment information can be provided to a natural language generator (referred to as "NLG" in some figures) 126, which can generate a natural language output based on the fulfillment information.

[0065] Since an intent can be fulfilled in various ways, the fulfillment information can take various forms. Suppose the user requests pure information, such as "Where were the outdoor shots of ‘The Shining’ filmed?". The user's intent can be determined, for example, by the intent matcher 136 as a search query. The intent and content of the search query can be provided to the fulfillment module 124, which can communicate with one or more search modules 150 configured to search a corpus of documents and / or other data sources (e.g., knowledge graphs, etc.) to obtain response information, as shown in Figure 1 . The fulfillment module 124 can provide data indicating the search query (e.g., the text of the query, dimensionality-reduced embeddings, etc.) to the search module 150. The search module 150 can provide response information, such as GPS coordinates or other more explicit information, such as "Timberline Lodge, Mt. Hood, Oregon". This response information can form part of the fulfillment information generated by the fulfillment module 124.

[0066] Additionally or alternatively, the fulfillment module 124 can be configured to receive, for example, the user's intent from the natural language understanding module 135 and any slot values provided by the user or determined using other means (e.g., the user's GPS coordinates, user preferences, etc.) and trigger a response operation. The response action can include, for example, ordering goods / services, starting a timer, setting a reminder, initiating a phone call, playing media, sending a message, etc. In some such embodiments, the fulfillment information can include slot values associated with the fulfillment, confirmation response (which in some cases can be selected from pre-determined responses), etc.

[0067] The natural language generator 126 can be configured to generate and / or select a natural language output (e.g., words / phrases designed to mimic human speech) based on data obtained from various sources. In some embodiments, the natural language generator 126 can be configured to receive fulfillment information associated with the fulfillment of an intent as input and generate a natural language output based on the fulfillment information. Additionally or alternatively, the natural language generator 126 can receive information from other sources such as third-party applications (e.g., required slots) that can be used to compose a natural language output for the user.

[0068] Figure 2 Illustrate examples of ways in which the output combinations from the voice capture module 110 and the image capture module 111 can be processed by various components according to various embodiments. The relevant components of the techniques described herein are shown, but this is not limiting and various other components not shown in Figure 1 can still be deployed. Assume Figure 2 that the operations shown occur after the exclusion area classification has been performed. Figure 2

[0069] Starting from the left, the voice capture module 110 can provide audio data to the invocation module 113. As described above, the audio data can include an audio recording of the user's utterance, an embedding generated from the audio recording, a feature vector generated from the audio recording, etc. At the same time or approximately the same time (e.g., simultaneously, as part of the same set of actions), the image capture module 111 can provide data indicating one or more image frames to the visual cue module 112. The data indicating one or more image frames can be raw image frame data, a dimensionality-reduced embedding of the raw image frame data, etc.

[0070] The visual cue module 112 can analyze the data indicating one or more image frames to detect one or more visual cues. As previously mentioned, the visual cue module 112 can ignore visual cues detected in regions previously classified as likely containing visual noise or assign them a smaller weight. In some embodiments, the visual cue module 112 can receive one or more signals indicating the status from the television 250 (or more generally, an electronic display) in the field of view of the visual sensor 107. These signals indicating the status can be provided, for example, by using one or more computer networks. If the status of the television is OFF, the region containing the television in the field of view of the visual sensor may not be ignored or have its weight set to a smaller value, but can be processed normally. However, if the status of the television is ON, the region containing the television in the field of view of the visual sensor can be ignored or have its weight set to less than the weight of other regions.

[0071] ​In some embodiments, the audio data provided by the speech capture module 110 and one or more visual cues detected by the visual cue module 112 can be provided to the invocation module 113. Based on these inputs, the invocation module 113 can determine whether the automated assistant 120 should be invoked. For example, assume that the user's invocation phrase utterance is not clearly recorded due to ambient noise. Just that noisy utterance may not be sufficient to invoke the automated assistant 120. However, if the invocation module 113 determines that, for example, a visual cue in the form of the user directly gazing at the visual sensor 107 is also detected while the utterance is being captured, the invocation module 113 can determine that the invocation of the automated assistant 120 is correct.

[0072] As previously mentioned, the use of visual cues is not limited to invoking the automated assistant 120. In various embodiments, as an addition to or in place of invoking the automated assistant 120, visual cues can be used to cause the automated assistant 120 to take various response actions. In Figure 2 this, one or more visual cues detected by the visual cue module 112 can be provided to other components of the automated assistant 120, such as the natural language understanding engine 135. The natural language understanding engine 135 can use one or more visual cues for various purposes, such as entity tagging (e.g., the user holds up a picture of a celebrity or public figure in a newspaper and says "who is this?"), for example, through the natural language processor 122. Additionally or alternatively, the natural language understanding engine 135 can use one or more visual cues, either alone or in combination with the speech recognition output generated by the STT module 117 and annotated by the natural language processor 122, to identify the user's intent, for example, through the intent matcher 136. Then, the intent can be provided to the fulfillment module 124, which, as previously mentioned, can take various actions to fulfill the intent.

[0073] Figure 3 An example field of view 348 of the visual sensor 307 of the assistant device 306 configured with selected aspects of the present disclosure is shown. It can be seen that a television 350 and a photo frame 352 are visible within the field of view 348. These are both potential sources of visual noise that can cause false positives for visual cues. For example, the television 350 may render a video that shows one or more individuals making gestures, looking at the camera, etc., any of which may be misinterpreted as a visual cue. The photo frame 352 can be a non-electronic photo frame that simply holds printed pictures, or it can be an electronic photo frame that renders one or more images stored in its memory. Assume that the picture contained in or rendered by the photo frame 352, for example, includes a person directly looking at the camera, then that person's gaze may be misinterpreted as a visual cue. Although Figure 3 Although not shown in the figure, other potential sources of visual noise may include electronic displays, such as monitors associated with laptop computers, tablet computers, smart phone screens, smart watch screens, and screens associated with other assistant devices, etc.

[0074] One or more image frames can be captured, for example, by the vision sensor 307, which can correspond to the field of view 348 of the vision sensor 307. These image frames can be analyzed, for example, by the vision cue module 112 to identify regions that may contain visual noise, such as the region containing the television 350 and the photo frame 352. The vision cue module 112 can identify these objects as part of its analysis of the image frames and accordingly identify the regions. Once these regions are identified, the vision cue module 112 can generate corresponding regions of interest that contain potential sources of visual noise. For example, Figure 4 is shown with Figure 3 the same field of view 348. However, in Figure 4 , regions of interest 360 and 362 that respectively contain the television 350 and the photo frame 352 have been generated. These regions of interest 360 to 362 can be classified as likely to contain visual noise (or unlikely to contain visual cues). For example, a visual noise probability higher than a specific threshold can be assigned to these regions of interest. Additionally or alternatively, the regions of interest can be associated with a visual noise indicator (which can, for example, indicate that the visual noise probability is higher than the threshold).

[0075] Therefore, they can be ignored, and / or the weight of visual cues detected within these regions can be set to be less than the weight of visual cues detected in other regions of the field of view 348.

[0076] Figure 5 is a flowchart of an example method 800 according to the embodiments disclosed herein. For convenience, the operations of this flowchart are described with reference to the system that performs these operations. The system can include various components of various computer systems, such as one or more components of the computing system that implements the automated assistant 120. Additionally, although the operations of method 500 are shown in a specific order, this is not limiting. One or more operations can be reordered, omitted, or added.

[0077] At block 502, the system may obtain one or more initial image frames from one or more vision sensors (e.g., 107), for example, via the image capture module 111. These initial image frames may be captured for the purpose of exclusion region classification. At block 504, the system may analyze one or more initial image frames, for example, via the vision cue module 113, to classify one or more specific regions of one or more video frames as likely to be visual noise sources and / or unlikely to contain visual cues. One or more operations at block 504 may constitute the exclusion region classification described herein. At block 506, the system may obtain one or more subsequent image frames from one or more vision sensors, for example, via the image capture module 111. These image frames may be obtained after the exclusion region classification.

[0078] At block 508, the system may analyze one or more subsequent image frames to detect one or more actionable visual cues provided by one or more users. In some embodiments, the analysis may include, at block 510, suppressing the analysis of one or more specific regions classified as likely to be visual noise sources and / or unlikely to contain visual cues in one or more subsequent image frames. For example, image data (e.g., RGB pixels) from one or more specific regions may not be applied as input to one of the machine learning models (e.g., convolutional neural network) trained to detect visual cues as described above.

[0079] Additionally or alternatively, in some implementations, the analysis at block 508 may include, at block 512, setting the weight of a first candidate visual cue detected in a specific region of one or more subsequent image frames to be less than the weight of a second candidate visual cue detected in other regions of one or more subsequent image frames, for example, via the vision cue module 113 and / or the intent matcher 136. For example, assume that a television rendering in a first region of the field of view of a vision sensor shows an image sequence of a person waving (assuming that waving is a predetermined visual cue that elicits a response from the automated assistant 120). Further assume that the user also makes a gesture in a second different region of the field of view of the vision sensor, for example, by forming an "eight" shape with his or her hand. Normally, these gestures may indicate the user's intent more or less equally, so the automated assistant 120 may be confused about the action it should attempt to perform. However, since the wave is detected in a region classified as likely to contain visual noise (i.e., the region of interest containing the television), a smaller weight may be assigned to the wave as a candidate visual cue compared to the "eight" shape gesture. Accordingly, the "eight" shape gesture may be more likely to elicit a response from the automated assistant 120.

[0080] Although this example illustrates selection based on the respective weights of multiple candidate visual cues, this is not limiting. In various implementations, for example, a single candidate visual cue can be compared to a predetermined threshold to determine whether it should trigger a response from the automated assistant 120. Thus, for example, a weight that does not meet the confidence threshold can be assigned to a visual cue detected in an image frame region classified as likely to contain visual noise. Thus, the visual cue may not trigger a response from the automated assistant 120. This prevents or reduces false positives, for example, caused by someone on a television making a gesture that happens to correspond to an actionable visual cue.

[0081] Additionally or alternatively, in some cases, a single visual cue across an image frame sequence can be detected in multiple regions of the field of view of the visual sensor, some of which are classified as likely to contain visual noise and others are not. In some such implementations, the combined confidence of the visual cue detected in both regions (with the cue detected in the classified region contributing less) can be compared to a threshold to determine that the visual cue meets a certain predetermined threshold to trigger a response from the automated assistant 120.

[0082] Review Figure 5 , at block 514, the system can cause one or more response actions to be taken by the automated assistant 120 or on behalf of the automated assistant 120 based on one or more of the detected actionable visual cues. These response actions can include invoking the automated assistant 120. For example, the invocation module 113 can determine that the actionable visual cue alone or in combination with the user-provided utterance is sufficient to invoke the automated assistant 120, enabling the user to make additional requests to the automated assistant 120. Additionally or alternatively, in some implementations, a visual cue of the user directly looking at the camera 107 of the client device 106 while speaking can be a strong indicator that the automated assistant 120 should be invoked (and potentially act on whatever the user says).

[0083] Additionally or alternatively, the response actions can include various response actions that the automated assistant 120 may take, for example, as a normal part of a human-machine dialogue, after the automated assistant 120 has been invoked. In some implementations, various actions can be pre-assigned or mapped to specific visual cues, either alone or in combination with verbal utterances. By way of a non-limiting example, a visual cue in the form of the user giving a "thumbs up" can be associated with the utterance "How do I get to <location>"(How do I get to <location>)” are used in combination to cause the automated assistant 120 to retrieve information about getting to that location using public transportation. In contrast, a visual cue in the form of a user gesture on the steering wheel can be used in combination with the same utterance to cause the automated assistant 120 to retrieve driving directions to that location. In some embodiments, the user can create a custom mapping between visual cues and various actions. For example, the user can say "OK Assistant, when I look at you and blink three times, play Jingle Bells”. The mapping can be created, for example, in a database available to the intent matcher 136 and used subsequently when the visual cue of three blinks is detected.

[0084] Figure 6 is a block diagram of an example computing device 610 that can optionally be used to perform one or more aspects of the techniques described herein. In some embodiments, one or more of the client computing device, the user-controlled resource module 130, and / or other components can include one or more components of the example computing device 610.

[0085] The computing device 610 generally includes at least one processor 614 that communicates with a number of peripheral devices via a bus subsystem 612. These peripheral devices can include a storage subsystem 624 (including, for example, a memory subsystem 625 and a file storage subsystem 626), a user interface output device 620, a user interface input device 622, and a network interface subsystem 616. The input and output devices allow a user to interact with the computing device 610. The network interface subsystem 616 provides an interface to an external network and is coupled to corresponding interface devices in other computing devices.

[0086] The user interface input device 622 can include a keyboard, a pointing device such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touch screen incorporated into a display, an audio input device such as a voice recognition system, a microphone, and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways of inputting information into the computing device 610 or onto a communication network.

[0087] The user interface output device 620 can include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem can include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem can also provide a non-visual display, such as via an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and ways of outputting information from the computing device 610 to a user or to another machine or computing device.

[0088] The storage subsystem 624 stores programming and data constructs that provide some or all of the functionality of the modules described herein. For example, the storage subsystem 624 can include selected aspects of the methods performed Figure 5 and the logic for implementing the various components shown in Figure 1 and Figure 2 .

[0089] These software modules are typically executed by the processor 614, either alone or in combination with other processors. The memory 625 used in the storage subsystem 624 can include many memories, including a main random access memory (RAM) 630 for storing instructions and data during program execution and a read-only memory (ROM) 632 in which fixed instructions are stored. The file storage subsystem 626 can provide persistent storage for program and data files and can include a hard disk drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. The modules implementing the functionality of certain embodiments can be stored in the storage subsystem 624 by the file storage subsystem 626, or in other machines accessible to the processor 614.

[0090] The bus subsystem 612 provides a mechanism for allowing the various components and subsystems of the computing device 610 to communicate with each other as expected. Although the bus subsystem 612 is schematically shown as a single bus, alternative embodiments of the bus subsystem can use multiple buses.

[0091] The computing device 610 can be of various types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, Figure 6 the description of the computing device 610 shown in Figure 6 is only intended as a specific example for the purpose of illustrating some embodiments. Many other configurations of the computing device 610 may have more or fewer components than the computing device shown in

[0092] In cases where the systems described herein collect or otherwise monitor personal information about a user or may utilize personal and / or monitored information, users may be provided with an opportunity to control whether programs or features collect user information (e.g., information about a user's social network, social actions or activities, profession, user preferences, or user's current geographic location) or to control whether and / or how content more relevant to the user is received from a content server. Additionally, certain data may be processed in one or more ways before it is stored or used so that personally identifiable information is removed. For example, a user's identity may be processed so that personally identifiable information cannot be determined for that user, or a user's geographic location may be generalized where location information is obtained (such as to a city, zip code, or state level) so that a user's specific geographic location cannot be determined. Accordingly, a user may control how information is collected and / or used about the user. For example, in some embodiments, a user may opt out of an assistant device attempting to detect visual cues, such as by disabling the visual sensor 107.

[0093] Although several embodiments have been described and illustrated herein, various other means and / or structures may be utilized for performing the functions and / or obtaining the results and / or one or more of the advantages described herein, and each such variation and / or modification is to be regarded as being within the scope of the embodiments described herein. More generally, all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and actual parameters, dimensions, materials, and / or configurations will depend upon the specific application or applications for which the teachings are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments described herein. Accordingly, it is to be understood that the foregoing embodiments are presented by way of example only and that the embodiments may be practiced otherwise than as specifically described and claimed within the scope of the appended claims and their equivalents. Embodiments of the present disclosure are directed to each and every separate feature, system, article, material, kit, and / or method described herein. Additionally, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the present disclosure.< / location> < / topping> < / topping> < / artist> < / artist> < / artist> < / artist>

Claims

1. A method for facilitating contactless invocation of an automated assistant implemented by one or more processors of an assistant device equipped with a vision sensor, the method comprising: When the assistant device equipped with a vision sensor is in a first position, an exclusion region classification process is performed to classify one or more regions of the field of view of the vision sensor as likely to contain visual noise and / or unlikely to contain visual cues provided by a user, wherein the exclusion region classification process includes: Obtaining one or more image frames from the vision sensor; Performing object recognition processing on the one or more image frames; and Based on the object recognition processing, classifying the one or more regions of the field of view of the vision sensor as likely to contain visual noise; Wherein, as a result of the classification, visual cues detected in one or more of the classified regions are less likely to invoke the automated assistant compared to visual cues subsequently detected outside the one or more classified regions, or visual cues for invoking the automated assistant are not analyzed for one or more of the classified regions in subsequent image frames obtained from the vision sensor; Detecting, using one or more sensors of the assistant device equipped with a vision sensor, that the assistant device equipped with a vision sensor has moved from the first position to a second position; In response to this detection, performing the exclusion region classification process to classify one or more new regions of the field of view of the vision sensor as likely to contain visual noise and / or unlikely to contain visual cues provided by a user.

2. The method according to claim 1, wherein, Performing the object recognition processing using a convolutional neural network.

3. The method according to claim 1, wherein, The object recognition processing is performed to detect a television or computer monitor within the field of view.

4. An assistant device for facilitating contactless invocation of an automated assistant, comprising: Vision sensor; One or more processors; And A memory storing instructions that, in response to execution of the instructions by the one or more processors, cause the one or more processors to: When the assistant device is in the first position, perform an exclusion region classification process to classify one or more regions of the field of view of the vision sensor as likely to contain visual noise and / or unlikely to contain visual cues provided by a user, wherein during the exclusion region classification process: Obtain one or more image frames from the vision sensor; Perform object recognition processing on the one or more image frames; and Based on the object recognition processing, classify the one or more regions of the field of view of the vision sensor as likely to contain visual noise; Wherein, as a result of the classification, visual cues detected in one or more of the classified regions are less likely to invoke the automated assistant compared to visual cues subsequently detected outside the one or more classified regions, or visual cues for invoking the automated assistant are not analyzed for one or more of the classified regions in subsequent image frames obtained from the vision sensor; Detect, using one or more sensors of the assistant device, that the assistant device has moved from the first position to the second position; In response to detecting that the assistant device has moved from the first position to the second position, perform the exclusion region classification process to classify one or more new regions of the field of view of the vision sensor as likely to contain visual noise and / or unlikely to contain visual cues provided by a user.

5. The assistant device according to claim 4, wherein, The object recognition process is performed using a convolutional neural network.

6. The assistant device according to claim 5, wherein, The object recognition process is performed to detect a television or computer monitor within the field of view.

7. At least one non - transitory computer - readable medium including instructions that, when executed by one or more processors of an assistant device equipped with a vision sensor, cause the one or more processors to facilitate contactless invocation of an automated assistant, including: When the assistant device equipped with a vision sensor is in a first position, an exclusion region classification process is performed to classify one or more regions of the field of view of the vision sensor as likely to contain visual noise and / or unlikely to contain visual cues provided by the user, wherein the exclusion region classification process includes: Obtaining one or more image frames from the vision sensor; Performing an object recognition process on the one or more image frames; and Based on the object recognition process, classifying the one or more regions of the field of view of the vision sensor as likely to contain visual noise; wherein, as a result of the classification, visual cues detected in one or more of the classified regions are less likely to invoke the automated assistant compared to visual cues subsequently detected outside one or more of the classified regions, or visual cues for invoking the automated assistant are not analyzed for one or more of the classified regions in subsequent image frames obtained from the vision sensor; Detecting, using one or more sensors of the assistant device equipped with a vision sensor, that the assistant device equipped with a vision sensor has moved from the first position to a second position; In response to the detection, performing the exclusion region classification process to classify one or more new regions of the field of view of the vision sensor as likely to contain visual noise and / or unlikely to contain visual cues provided by the user.

8. The at least one non-transitory computer-readable medium according to claim 7, wherein, The object recognition process is performed using a convolutional neural network.

9. The assistant device according to claim 7, wherein, The object recognition process is performed to detect a television or computer monitor within the field of view.

Citation Information

Patent Citations

  • Gesture recognition method and gesture recognition device

    CN104598915A

  • Recognizing User Intent In Motion Capture System

    US20110175810A1

  • Method and apparatus for processing commands directed to a media center

    US20140098240A1

  • Recognizing gestures captured by video

    US8891868B1