Methods for parsing automated assistant requests
By combining natural language input and image processing and user interface prompts, the problem of insufficient specificity of automation assistants when analyzing object properties in images is solved, improving the accuracy of request resolution and the response ability of automation assistants.
Patent Information
- Application Number
- CN202011245475.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-06-23
- Filing Date
- 2018-05-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2038-11-15
AI Technical Summary
In prior art, when parsing object properties in images, objects are often not defined with the desired degree of specificity, resulting in the automation assistant being unable to effectively respond to requests based on environment objects.
By combining natural language input and image processing, it is determined whether the request is parsable. If it is not parsable, a prompt is provided to guide the user to capture additional sensor data or provide user interface input to parse the request.
Improves the specificity and accuracy of the automation assistant when parsing object properties in images, and enhances its ability to respond to requests based on environment objects.
Smart Images

Figure CN112445947B_ABST
Abstract
Description
[0001] Description of the case
[0002] This application is a divisional application of Chinese invention patent application No. 201880032505.2, filed on May 15, 2018. Technical Field
[0003] The present disclosure relates to a method of parsing automated assistant requests. Background Art
[0004] Image processing can be utilized to resolve attributes of objects in an image. For example, some image processing techniques utilize an image processing engine to resolve a classification of an object captured in an image. For example, for an image capturing a sailboat, image processing can be performed to resolve the image's classification values of "boat" and / or "sailboat." Image processing can be used to resolve additional attributes or alternative attributes. For example, optical character recognition (OCR) can be used to resolve text in an image. Also, for example, some image processing techniques can be used to determine a more specific classification of an object in an image (e.g., a specific brand and / or model of a sailboat).
[0005] Some image processing engines utilize one or more machine learning models, such as a deep neural network model, which accepts an image as input and generates as output a measure indicating which of a plurality of corresponding attributes is present in the image based on the image using learned parameters. If the measure indicates that a particular attribute is present in the image (e.g., if the measure satisfies a threshold), then the attribute can be considered "resolved" for the image (i.e., the attribute can be considered present in the image). However, it may often be the case that image processing of an image may not be able to resolve one or more (e.g., any) attributes. Furthermore, it may also be the case that the resolved attributes of an image cannot define an object in the image with a desired degree of specificity. For example, the resolved attributes of an image may enable determination of whether a "shirt" is present in the image, and that the shirt is "red" - but may not be able to determine the manufacturer of the shirt, whether the shirt is "short sheeve" or "long sheeve", etc.
[0006] Additionally, humans may participate in human-computer conversations using interactive software applications referred to herein as "automated assistants" (also referred to as "interactive personal assistants," "intelligent personal assistants," "personal voice assistants," "conversational agents," and the like). An automated assistant typically receives natural language input (utterances) from a user. In some cases, the natural language input can be received as audio input (e.g., streaming audio) and converted to text and / or received as (e.g., typed) textual natural language input. The automated assistant responds to the natural language input using responsive content (e.g., visual and / or auditory natural language output). However, it may often be that the automated assistant does not accept and / or respond to requests based on sensor data (e.g., images) that captures one or more attributes of environmental objects. Summary of the invention
[0007] Embodiments described herein relate to causing processing of sensor data to be performed in response to a request determined to be related to an environmental object that may be captured by the sensor data. For example, image processing can be performed on an image in response to a request determined based on natural language input provided by a user in conjunction with the capture of at least one image (e.g., natural language received shortly before, after, and / or during the capture of the at least one image). For example, a user can provide voice input of "what's wrong with my device" through an automated assistant interface of a client device. It can be determined that the voice input is related to an environmental object, and as a result, image processing can be performed on an image captured by a camera of the client device. The image can be captured by the camera based on a separate user interface input (e.g., selection of an "image capture" interface element), or the image can be automatically captured in response to a determination that the voice input is related to an environmental object.
[0008] Some embodiments described herein also relate to determining whether a request is parseable based on processing of sensor data. For example, based on determining that one or more attributes (if any) parsed based on image processing of at least one image fail to define an object with a target-specific degree, a request can be determined to be unparseable. When it is determined that a request is unparseable, a prompt is determined and provided as a user interface output (e.g., audible and / or graphical), wherein the prompt provides guidance on further input that will enable the request to be parsed. The prompt can instruct the user to capture other sensor data of the object (e.g., images, audio, temperature sensor data, weight sensor data) and / or move the object (and / or other objects) to enable the capture of other sensor data of the object. For example, the prompt can be customized to enable the capture of additional images that enable the parsing of one or more attributes that are not parsed based on image processing of at least one image. The prompt can additionally or alternatively request the user to provide a user interface input (e.g., natural language input) for the unparsed attributes of the object.
[0009] In those embodiments, the request can then be parsed using additional sensor data (e.g., additional images) and / or user interface input received in response to the prompt. For example, image processing can be performed on the additional images received in response to the prompt, and the request can be parsed using additional attributes parsed from the image processing. For example, the request can be parsed by submitting a proxy request to one or more agents (e.g., a search system and / or other agents), wherein the proxy request is generated based on additional attributes parsed from the image processing of the additional images and optionally based on attributes determined based on processing of previous processor data (e.g., determined based on image processing of previous images). As another example, additional attributes can be parsed based on natural language input or other user interface input received in response to the prompt, and the request can be parsed using these additional attributes. It should be noted that in some embodiments and / or situations, multiple rounds of prompts can be provided, additional attributes are determined from additional sensor data and / or user interface input in response to those prompts, and the request can be parsed using these additional attributes.
[0010] It should be understood that the methods and systems described herein can help users perform specific tasks through a guided human-computer interaction process. A variety of different tasks can be performed, but these tasks may include, for example, enabling users to obtain information they need to perform further tasks. For example, as referenced Figure 6 and Figure 7 As described, the information obtained may be helpful in performing maintenance on another device.
[0011] As described above, it is possible to determine that the request is unresolvable based on determining that one or more attributes (if any) parsed based on the processing of the processor data cannot define the object with a target specificity. In some embodiments, the target specificity of the object can be the target classification degree of the object in a categorical taxonomy. For example, the target classification degree of a car can be classified as a level that defines the brand and model of the car, and can also be classified as a level that defines the brand, model and age of the car. In some embodiments, the target specificity of the object can refer to one or more field definitions to be defined, wherein the field of the object can depend on the classification (general or specific) of the object. For example, for a bottle of wine, the field can be defined for a specific brand, type of wine, and / or year-and the target specificity is the parsing of the attributes of all those fields. In some embodiments, the target specificity can be determined additionally or alternatively based on the initial natural language input provided by the user, the feedback provided by the user, the historical interaction of the user and / or other users and / or the location and / or other context signals.
[0012] As also described above, the determined prompt can provide guidance on further input that will enable the request related to the environmental object to be resolved. In some embodiments, the prompt is determined based on one or more attributes of the resolved object. For example, it is possible to utilize a classification attribute of the environmental object, such as a classification attribute that is resolved based on image processing of a previously captured image of the object. For example, a prompt for the "car" classification can be concretized as a car classification (e.g., "take a picture of the car from another angle"). Similarly, for example, a prompt for the "jacket" classification can be concretized as a "jacket" classification (e.g., "take a picture of the logo or the tag"). Similarly, for example, as described above, the classification attribute can be associated with one or more fields to be defined. In some of those cases, a prompt can be generated for a field that has not yet been defined based on the resolved attribute (if any). For example, if the year field of a bottle of wine has not yet been defined, the prompt can be "take a picture of the year" or "what is the year?".
[0013] Some embodiments described herein can provide the prompt for presentation to the user only when it is determined that: (1) there is a request related to an environmental object (e.g., a request for additional information); and / or (2) the request cannot be resolved based on processing of sensor data collected to date. In this way, computing resources are not wasted by providing unnecessary prompts and / or processing of further input that would be responsive to these unnecessary prompts. For example, where a user captures an image of a bottle of wine, and provides natural language input "send this picture to Bob", a prompt requesting the user to take additional photos of the bottle of wine will not be provided based on the request (i.e., send the image to Bob) being only resolvable based on the captured image, and / or based on the request not being a request for additional information related to the bottle of wine. In contrast, if the natural language input is "how much does this cost", a prompt requesting the user to take additional images may be provided if image processing of the initial image cannot resolve enough attributes to resolve the request (i.e., determine the cost). For example, if the brand, wine type, and / or vintage cannot be resolved based on image processing, a prompt to "take a picture of the label" can be provided. As yet another example, when the user is in a "retail" location, a request related to the environmental objects of the captured image can be inferred, whereas if the user has captured the same environmental objects at a park, a request would not be inferred (assuming that the user may be seeking shopping intelligence, such as prices, reviews, etc., while being at a retail location). As yet another example, a request may not be inferred when the initially captured image captures multiple objects that are all at a long distance - whereas, a request would be inferred if the image captures only one object at a close distance.
[0014] Some embodiments described herein can also determine prompts that are customized to attributes that have already been resolved (e.g., classification attributes); that are customized to fields that have not yet been resolved; and / or that are otherwise customized to enable resolution of a request. In this manner, prompts can be customized to increase the likelihood that input (e.g., image and / or user interface input) responsive to the prompt will enable resolution of the request—thereby alleviating the need for further prompts when resolving the request and / or for processing further input that will be responsive to these further prompts.
[0015] As a clear example of some embodiments, assume that a user provides voice input of "what kind of reviews does this get?" while pointing a camera of a client device (e.g., a smartphone, tablet, wearable device) at a bottle of wine. It can be determined that the voice input includes a request related to an object in the environment of the client device. In response to the voice input including the request, an image of the bottle of wine can be processed to parse one or more attributes of the object in the image. For example, the image can be processed to determine the classification of "bottle" and / or "wine bottle". In some embodiments, the image can be captured based on the user interface input, or the image can be captured automatically based on determining that the voice input includes the request.
[0016] Furthermore, it can be determined that such a request is not parsable based on the parsed "bottle" and / or "wine bottle" classifications. For example, based on the speech input, it can be determined that the request is for a review of a specific wine in a bottle (e.g., a specific brand, type of wine, and / or vintage) - rather than a review of "wine bottle" in general. Thus, parsing of the request requires parsing enough attributes to enable determination of the specific wine in the bottle. The general classifications of "bottle" and "wine bottle" do not enable such a determination.
[0017] A prompt can then be provided in response to determining that the request is not resolvable. For example, the prompt can be "can you take a picture of the label" or "can you make the barcode visible to the camera?", etc. In response to the prompt, the user can move the wine bottle and / or the electronic device (thereby moving the camera) and additional images captured after such movement. Processing of the additional images can then be performed to determine the attributes of the selected parameters. For example, OCR processing can be performed to determine the text value of the label, such as text including the brand name, wine type, and year. If processing of the additional images still cannot resolve the request (e.g., the required attributes have not been resolved), other prompts may be generated and image processing of other images received after the other prompts may be performed.
[0018] Additional content can then be generated based on the additional attributes. For example, in response to a search being made based on the additional attributes and / or voice input, the additional content can be received. For example, a text value of "Vineyard A Cabernet Sauvignon 2012" may have been determined, a query for "reviews for vineyard Acabernet sauvignon 2012" may have been submitted (e.g., to a search system agent), and additional content may be received in response to the query.
[0019] As described herein, one or more prompts may additionally or alternatively request the user to provide a responsive user interface input to enable the parsing of the attributes of the unparsed fields. For example, assume that the "wine bottle" classification value is determined by processing one or more images, but it is not possible to parse enough text to clearly identify the wine. For example, the text identifying a specific brand is identified, but the text identifying the type and year of the wine is not identified. Instead of or in addition to prompting the user to capture other images and / or move the wine bottle, the prompt can also request the user to identify the type and year of the wine (e.g., "can you tell me the wine type and year for the Brand X wine (Can you tell me the wine type and year for the Brand X wine)?"). Then, the responsive user interface input provided by the user can be used to parse the wine type and year. In some embodiments, a prompt can be generated to include one or more candidate attributes determined based on image processing. For example, assume that OCR image processing technology is used to determine the candidate years of "2017" and "2010". For example, the image processing technology can identify "2017" and "2010" as candidates, but cannot identify one of the two with sufficient confidence to enable the parsing of a specific year. In this case, the prompt might be “Is this a 2010 or 2017 vintage?” — or might offer selectable options of “2010” and “2017.”
[0020] In some embodiments, multiple processing engines and / or models may operate in parallel, and each may be specifically configured for one or more specific fields. For example, a first image processing engine may be a general classification engine configured to determine a general entity in an image, a second image processing engine may be a logo processing engine configured to determine a logo brand in an image, a third image processing engine may be an OCR or other character recognition engine configured to determine text and / or numeric characters in an image, and so on. In some of those embodiments, prompts may be generated based on which image processing engine is unable to parse the attributes of the corresponding field. In addition, in some of those embodiments, in response to receiving additional images in response to the prompt, only a subset of the engines may be used to process these additional images. For example, only those engines configured to parse unparsed fields / parameters may be used, thereby saving various computing resources by not using a full set of engines for these images.
[0021] In some embodiments, a method performed by one or more processors is provided, the method comprising: receiving a voice input provided by a user via an automated assistant interface of a client device; and determining that the voice input includes a request related to an object in the environment of the client device. The method also includes: in response to determining that the voice input includes a request related to the object: causing processing to be performed on initial sensor data captured by at least one sensor. The at least one sensor is a sensor of the client device or an additional electronic device in the environment, and the initial sensor data captures one or more features of the object. The method also includes determining that the request is not resolvable based on the initial sensor data based on one or more initial attributes of the object parsed based on the processing of the initial sensor data. The method also includes, in response to determining that the request is not resolvable: providing a prompt instructing the user to capture additional sensor data or move the object to be presented to the user via the automated assistant interface of the client device. The method also includes: receiving additional sensor data, the additional sensor data being captured by the client device or the additional electronic device after the prompt is presented to the user; causing processing to be performed on the additional sensor data; and parsing the request based on at least one additional attribute parsed based on the processing of the additional sensor data.
[0022] In some embodiments, a method performed by one or more processors is provided, the method comprising: receiving at least one image captured by a camera of a client device; and determining that the at least one image relates to a request related to an object captured by the at least one image. The method also includes, in response to determining that the image relates to a request related to an object: causing image processing to be performed on the at least one image. The method also includes, based on the image processing of the at least one image, determining that the at least one parameter necessary for resolving the request is not resolvable based on the image processing of the at least one image. The method also includes, in response to determining that the at least one parameter is not resolvable: providing a prompt customized for the at least one parameter for presentation via the client device or another client device. The method also includes, in response to the prompt, receiving additional images and / or user interface inputs captured by the camera; resolving a given attribute of the at least one parameter based on the additional images and / or user interface inputs; and resolving the request based on the given attribute.
[0023] In some embodiments, a method performed by one or more processors is provided, the method comprising: receiving natural language input provided by a user via an automated assistant interface of a client device; and determining that the natural language input includes a request related to an object in the environment of the client device. The method also includes, in response to determining that the natural language input includes a request related to an object: causing processing to be performed on initial sensor data captured by a sensor of the client device or an additional electronic device in the environment. The method also includes determining that the request is unresolvable based on the initial sensor data based on one or more initial attributes of the object parsed based on the processing of the initial sensor data. The method also includes, in response to determining that the request is unresolvable: providing a prompt to present to the user via the automated assistant interface of the client device. The method also includes: receiving natural language input or an image in response to the prompt; and parsing the request based on the natural language input or the image.
[0024] In some embodiments, a method performed by one or more processors is provided, the method comprising: processing at least one image captured by a camera of an electronic device to resolve one or more attributes of an object in the at least one image; selecting one or more fields for the object that are undefined by the attributes resolved by processing the at least one image; providing a prompt customized for at least one of the selected one or more fields via the electronic device or an additional electronic device; in response to the prompt, receiving an additional image captured by the camera and at least one of a user interface input; resolving a given attribute for the selected one or more fields based on at least one of the additional image and the user interface input; determining additional content based on the resolved given attribute; and providing the additional content via the electronic device for presentation to a user.
[0025] In some embodiments, a method performed by one or more processors is provided, the method comprising: processing at least one image captured by a camera of an electronic device to resolve one or more attributes of an object in the at least one image; selecting one or more fields for the object that are undefined by the attributes resolved by processing the at least one image; providing a prompt for presentation to a user by the electronic device or an additional electronic device; receiving at least one additional image captured after the prompt is provided; and selecting a subset of available image processing engines to process the at least one additional image. The available image processing engines of the subset are selected based on association with the resolution of the one or more fields. The method also includes resolving one or more additional attributes of the one or more fields based on applying the at least one additional image to the selected subset of available image processing engines. The one or more additional attributes are resolved without applying the at least one additional image to other available image processing engines not included in the selected subset.
[0026] In addition, some embodiments include one or more processors of one or more computing devices, wherein the one or more processors are operable to execute instructions stored in an associated memory, and wherein the instructions are configured to cause the performance of any of the foregoing methods. Some embodiments also include one or more non-transitory computer-readable storage media storing computer instructions executable by one or more processors to perform any of the foregoing methods.
[0027] It should be understood that all combinations of the foregoing concepts and additional concepts described in more detail herein should be considered part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure should be considered part of the subject matter disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 is a block diagram of an example environment in which the techniques disclosed herein may be implemented.
[0029] Figure 2A , Figure 2B , Figure 3 , Figure 4 , Figure 5A , Figure 5B , Figure 6 and Figure 7 Examples of how the techniques described herein may be employed are shown in accordance with various implementations.
[0030] Figure 8 A flow chart of an example method according to embodiments disclosed herein is shown.
[0031] Fig. 9 An example architecture for a computing device is shown. DETAILED DESCRIPTION
[0032] Figure 1 An example environment in which the technology disclosed herein can be implemented is shown. The example environment includes multiple client devices 106 1-N and automated assistant 120. Although Figure 1 Automated assistant 120 and client device 106 are shown in 1-N However, in some embodiments, all or some aspects of automated assistant 120 may be performed by one or more client devices 106. 1-N For example, the client device 106 1 One or more instances of one or more aspects of automated assistant 120 may be implemented, and client device 106 N A separate instance of one or more aspects of automated assistant 120 may also be implemented. 1-N In an embodiment implemented by one or more computing devices, the client device 106 1-N Those aspects of automated assistant 120 may communicate via one or more networks, such as a local area network (LAN) and / or a wide area network (WAN) (eg, the Internet).
[0033] Client device 106 1-N The client computing device may include, for example, one or more of a desktop computer device, a laptop computer device, a tablet computer device, a mobile phone computing device, a computing device of a user's vehicle (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a standalone interactive speaker, a smart camera, and / or a user's wearable device including a computing device (e.g., a user's watch with a computing device, a user's glasses with a computing device, a virtual or augmented reality computing device). Additional and / or alternative client computing devices may be provided.
[0034] In some implementations, a given user may communicate with automated assistant 120 using multiple client devices that converge into a coordinated "ecosystem" of computing devices. In some such implementations, automated assistant 120 may be considered to "serve" the given user, e.g., granting automated assistant 120 enhanced access rights to its resources (e.g., content, documents, etc.) controlled by the "served" user. However, for the sake of brevity, some examples described in this specification will focus on a user operating a single client device.
[0035] Each client device 106 1-Ncan operate various applications, such as multiple message exchange clients 107 1-N The corresponding one and multiple camera applications 109 1-N Each client device 106 may also be equipped with one or more cameras 111 (e.g., front and / or rear cameras in the case of a smartphone or tablet) and / or one or more additional sensors 113. Additional sensors 113 may include, for example, microphones, temperature sensors, weight sensors, etc. In some embodiments, one or more additional sensors 113 may be provided as part of an independent peripheral device that is separate from, but in communication with, one or more corresponding client devices 106 and / or automated assistant 120. For example, one or more additional sensors 113 1 Can be included in a peripheral scale and can generate sensor data indicative of the weight of an object placed on the scale.
[0036] Message exchange client 107 1-N can have various forms, and these forms can be on the client computing device 106 1-N between, and / or may be within a single client computing device 106 1-N In some embodiments, one or more message exchange clients 107 1-N The message exchange client 107 may be in the form of a short message service ("SMS") and / or multimedia message service ("MMS") client, an online chat client (e.g., an instant messenger, Internet Relay Chat or "IRC"), a messaging application associated with a social network, a personal assistant messaging service dedicated to conversing with the automated assistant 120, etc. In some embodiments, one or more message exchange clients 107 1-N This may be accomplished via a web page or other resource presented by a web browser (not shown) or other application of the client computing device 106 .
[0037] Camera Application 109 1-N The user may be enabled to control the camera 111 1-N For example, one or more camera applications 109 1-N A graphical user interface may be provided with which a user may interact to capture one or more images and / or videos. 1-N The automated assistant 120 may be interacted / interfaced as described herein to enable the user to interpret the image captured by the camera 111. 1-N In other embodiments, one or more camera applications 109 1-N may have its own built-in functionality different from that of the automated assistant 120, which enables the user to interpret the information associated with the camera 111.1-N Additionally or alternatively, in some implementations, the message exchange client 107 or any other application installed on the client device 106 may include functionality that enables the application to access data captured by the camera 111 and / or additional sensors 113 and perform the techniques described herein.
[0038] Camera 111 1-N Can include monocular cameras, stereo cameras and / or thermal imaging cameras. Figure 1 The client device 106 is shown in 1 and client device 106 N Each has only a single camera, however, in many embodiments, the client device can include multiple cameras. For example, a client device can have a single-sided camera for forward and rearward directions. Also, for example, a client device can have a stereo camera and a thermal imaging camera. Also, for example, a client device can have a single-sided camera and a thermal imaging camera. In addition, in various embodiments, the sensor data used in the techniques described herein can include images from multiple different types of cameras (the same client device and / or multiple client devices). For example, an image from a single-sided camera can be used initially to determine that a request is unresolvable, and then an image from a separate thermal imaging camera is received (e.g., in response to a prompt) and used to resolve the request. Additionally, in some embodiments, sensor data from other visual sensors, such as point cloud sensor data from a three-dimensional laser scanner, can be utilized.
[0039] As described in more detail herein, automated assistant 120 communicates with the client via one or more client devices 106 1-N In some embodiments, in response to a user input and output device of the client device 106, a human-computer dialogue session is conducted with one or more users. 1-N The automated assistant 120 may conduct a human-computer dialogue session with the user based on the user interface input provided by one or more user interface input devices.
[0040] In some of those embodiments, the user interface input is explicitly directed to the automated assistant 120. For example, the message exchange client 107 1-N One of the personal assistant messaging services may be a personal assistant messaging service dedicated to conversations with automated assistant 120, and user interface input provided via the personal assistant messaging service may be automatically provided to automated assistant 120. In addition, user interface input may be explicitly directed to one or more message exchange clients 107, for example, based on a specific user interface input indicating that automated assistant 120 is to be invoked. 1-NAutomated assistant 120 in . For example, the specific user interface input can be one or more typed characters (e.g., @AutomatedAssistant), user interaction with a hardware button and / or a virtual button (e.g., a tap, a long press), a verbal command (e.g., "Hey Automated Assistant"), and / or other specific user interface input. In some embodiments, the automated assistant 120 can participate in a conversation session in response to the user interface input even when the user interface input is not explicitly directed to the automated assistant 120. For example, the automated assistant 120 can examine the content of the user interface input and participate in the conversation session in response to certain terms present in the user interface input and / or based on other cues. In many embodiments, the automated assistant 120 can participate in an interactive voice response ("IVR"), enabling the user to issue commands, conduct searches, etc., and the automated assistant 120 can convert the utterances into text using natural language processing and / or one or more grammars, and respond accordingly.
[0041] Client computing device 106 1-N Each of the client computing devices 106 and the automated assistant 120 may include one or more memories for storing data and software applications, one or more processors for accessing data and executing applications, and other components that facilitate communication over the network. 1-N And / or the operations performed by automated assistant 120 may be distributed across multiple computer systems. Automated assistant 120 may be implemented, for example, as a computer program running on one or more computers in one or more locations coupled to each other via a network.
[0042] Except in Figure 1 Automated assistant 120 may include, in addition to other components not shown in FIG. 1 , natural language processor 122, request engine 124, prompt engine 126, and request parsing engine 130 (including attribute module 132). In some embodiments, one or more of the engines and / or modules of automated assistant 120 may be omitted, combined, and / or implemented in a component separate from automated assistant 120. In some embodiments, automated assistant 120 responds to a request from client device 106 during a human-machine conversation session with automated assistant 120. 1- N generates responsive content based on various inputs from N. Automated assistant 120 (e.g., over one or more networks when separate from the user's client device) provides responsive content for presentation to the user as part of the conversation session. For example, automated assistant 120 may respond to a request received via client device 106 1-NThe responsive content may be generated in response to free-form natural language input provided by one of the cameras 111, in response to an image captured by one of the cameras 111, and / or in response to additional sensor data captured by one or more additional sensors 113. As used herein, free-form input is input formulated by a user and is not limited to a set of options presented for selection by the user.
[0043] As used herein, a "conversation session" may include a logically self-contained exchange of one or more messages between a user and automated assistant 120 (and, in some cases, other human participants in the conversation). Automated assistant 120 may distinguish between multiple conversation sessions with a user based on various signals, such as the passage of time between sessions, changes in user context between sessions (e.g., location, before / during / after a scheduled meeting, etc.), detection of one or more intervening interactions between a user and a client device other than the conversation between the user and the automated assistant (e.g., the user switches applications for a moment, the user leaves and then later returns to a standalone voice-activated product), locking / hibernation of a client device between sessions, changes in client devices used to interact with one or more instances of automated assistant 120, etc.
[0044] In some implementations, when automated assistant 120 provides a prompt requesting user feedback in the form of user interface input (e.g., spoken input and / or typed input), automated assistant 120 can (by providing the prompt) preemptively activate one or more components of the client device that are configured to process the user interface input to be received in response to the prompt. 1 In the case where a microphone of a client device 106 is used to provide user interface input, automated assistant 120 may provide one or more commands to cause: the microphone to be preemptively "opened" (thereby preventing the need to tap an interface element or say a "hot word" to open the microphone); preemptively activate the microphone to the client device 106; 1 The local speech of the text processor; preemptively establish the client device 106 1 a communication session with a remote voice to a text processor (e.g., a remote voice to a text processor of automated assistant 120); and / or at client device 106 1 A graphical user interface (e.g., an interface including one or more selectable elements that can be selected to provide feedback) is presented on the display. This can enable user interface input to be provided and / or processed more quickly than if these components were not preemptively activated.
[0045] The natural language processor 122 of the automated assistant 120 processes the natural language input received by the user via the client device 106 1-NThe natural language processor 122 may process the natural language input generated by the user via the client device 106 and may generate annotated output for use by one or more other components of the automated assistant 120 (such as the request engine 124, the prompt engine 126, and / or the request parsing engine 130). 1 The natural language free-form input generated by one or more user interface input devices of the present invention. The generated annotated output includes one or more annotations of the natural language input and optionally includes one or more (e.g., all) terms of the natural language input. In some embodiments, the natural language processor 122 includes a speech processing module that is configured to process speech (oral) natural language input. The natural language processor 122 can then operate on the processed speech input (e.g., based on the text obtained from the processed speech input). For example, the speech processing module can be a speech-to-text module that receives free-form natural language speech input in the form of a streaming audio recording and converts the speech input into text using one or more speech-to-text models. For example, a client device can generate a streaming audio recording in response to a signal received from a microphone of the client device when a user speaks, and can send the streaming audio recording to an automated assistant for processing by a speech-to-text module.
[0046] In some embodiments, the natural language processor 122 is configured to recognize and annotate various types of grammatical information in the natural language input. For example, the natural language processor 122 may include a portion of a speech tagger that is configured to annotate terms with the grammatical roles of the terms. For example, the portion of the speech tagger may tag each term with its phonetic parts, such as "noun," "verb," "adjective," "pronoun," and the like. Further, for example, in some embodiments, the natural language processor 122 may additionally and / or alternatively include a dependency parser (not shown) that is configured to determine syntactic relationships between terms in the natural language input. For example, the dependency parser may determine which terms modify other terms, the subject and verb of a sentence, and the like (e.g., a parse tree) - and may annotate such dependencies.
[0047] In some embodiments, the natural language processor 122 may additionally and / or alternatively include an entity tagger (not shown) configured to annotate entity references in one or more segments, such as references to people (including, for example, literary characters, celebrities, public figures, etc.), organizations, locations (real and fictional), etc. In some embodiments, data about entities may be stored in one or more databases, such as in a knowledge graph (not shown). In some embodiments, a knowledge graph may include nodes representing known entities (in some cases, entity attributes), and edges connecting nodes and representing relationships between entities. For example, a "banana" node may be connected (e.g., as a child node) to a "fruit" node, which in turn may be connected (e.g., as a child node) to a "produce" and / or "food" node. As another example, a restaurant called "Hypothetical Café" may be represented by a node that also includes attributes such as its address, the type of food served, business hours, contact information, etc. In some embodiments, the "Hypothetical Café" node can be connected to one or more other nodes, such as a "restaurant" node, a "business" node, a node representing the city and / or state where the restaurant is located, and so on, via edges (e.g., representing the relationship between a child node and a parent node).
[0048] The entity tagger of the natural language processor 122 may annotate references to entities at a high level of granularity (e.g., to enable identification of all references to an entity class such as a person) and / or at a low level of granularity (e.g., to enable identification of all references to a specific entity such as a specific person). The entity tagger may rely on the content of the natural language input to resolve the specific entity and / or may optionally communicate with a knowledge graph or other entity database to resolve the specific entity.
[0049] In some implementations, the natural language processor 122 may additionally and / or alternatively include a coreference resolver (not shown) configured to group or "cluster" references to the same entity based on one or more contextual clues. For example, in the natural language input "I liked Hypothetical Café last time we ate there", the term "there" may be resolved to "Hypothetical Café" using a coreference resolver.
[0050] In some embodiments, one or more components of the natural language processor 122 may rely on annotations from one or more other components of the natural language processor 122. For example, in some embodiments, a named entity tagger may rely on annotations from a coreference resolver and / or a dependency parser when annotating all mentions of a particular entity. Similarly, for example, in some embodiments, a coreference resolver may rely on annotations from a dependency parser when clustering references to the same entity. In some embodiments, when processing a particular natural language input, one or more components of the natural language processor 122 may use relevant previous input and / or other relevant data outside of the particular natural language input to determine one or more annotations.
[0051] The request engine 124 utilizes one or more signals to determine when there is a request related to an object in the environment of the client device associated with the request. 1 Natural language input provided by camera 111 1 Captured images, from additional sensors 113 1 Additional sensor data of the client device 106 1 The detected location and / or other contextual attributes of the client device 106 are determined 1 Requests related to objects in the environment.
[0052] As an example, the request engine 124 can be based on a request received by a user via the client device 106 1The request engine 124 can determine such a request based on the natural language input (e.g., spoken voice input) provided. For example, the request engine 124 can rely on the annotations from the natural language processor 122 to determine that certain utterances may be related to objects in the environment, such as the utterances "what is this", "where can I buy this", "how much does this cost", "how much does this thing weigh", "tell me more about that", etc. In some embodiments, the request engine 124 can determine such a request based on utterances that include forms that cannot be resolved with antecedents in previous natural language input (e.g., by the natural language processor 122) (e.g., ambiguous pronouns such as "this", "thing", "it", "that"). In other words, since the request refers to a form that is unresolvable for antecedents in previous natural language input, the request engine 124 can assume that the utterance is related to an environmental object. In some embodiments, in response to determining that the utterance is related to an environmental object, the request engine 124 can cause the image and / or other sensor data to be processed by the client device 106. 1 For example, the request engine 124 can send a 1 and / or camera application 109 1 One or more commands are provided to cause one or more images to be captured. In some of those embodiments, the request engine 124 can cause images to be captured without requiring the client device 106 to 1 The user selects the "image capture" interactive element, provides a verbal "capture image" command and / or otherwise interacts with the client device 106 1 Physically interact to cause an image to be captured. In some embodiments, confirmation from the user may be required before capturing an image. For example, the request engine 124 can cause a user interface output to be provided that says "Need to capture an image to answer your request" and cause an image to be captured only in response to an affirmative user response (e.g., "OK").
[0053] As another example, the request engine 124 can be based on the client device 106 1 Such a request may be determined additionally or alternatively based on images, sound recordings, and / or other sensor data captured by the camera 111.1 Such a request may be determined by providing speech in response to capturing of at least one image (e.g., before, after, and / or during capturing of at least one image). For example, in response to user interface input, in the case of capturing of at least one image via camera 111 1 Say "what is this" within X seconds of capturing the image. 1 In the context of an assistant application (e.g., an application specifically designed to interact with automated assistant 120), or in the context of a client device 106 1 In the context of another application such as a camera application, a chat application, etc., the camera 111 1 Thus, in some embodiments, the request engine 124 can determine such a request based on images and / or speech captured from any of a plurality of different applications. Also, for example, the request engine 124 can determine such a request based on captured sensor data—and can determine such a request independently of any speech. For example, the request engine 124 can determine such a request based on the user in certain contexts (e.g., when the location data client device 106 1 As another example, the request engine 124 can be responsive to the user causing the user to capture an image via the additional sensor 113. 1 Such a request may be determined by capturing a microphone recording of the user. For example, the user may cause an audio recording to be captured, where the audio recording captures the noise produced by the user's vacuum cleaner, and then provide the utterance "why is my vacuum making this noise".
[0054] As another example, the request engine 124 can additionally or alternatively determine such a request based on an interface element selected by a user and / or based on images and / or other sensor data captured via a particular interface. 1 After capturing the image, a “find out more” graphical interface element can be presented (e.g., based on output from automated assistant 120) as a user interface via camera application 109. 1 , and the selection of the graphical interface element can be interpreted by the request engine 124 as a request for additional information about the object captured by the image. Also, for example, if the user utilizes the messaging client 107 1 and / or other applications specifically customized for automated assistant 120 to capture images and / or other sensor data (e.g., client applications specifically designed to interact with automated assistant 120), such capture can be interpreted by request engine 124 as a request related to the object of sensor data capture.
[0055] When there is a request related to an object in the environment of the client device associated with the request (e.g., as determined by the request engine 124), the request parsing engine 130 attempts to parse the request. In attempting to parse the request, the request parsing engine 130 can utilize natural language input associated with the request (if present), and / or other sensor data associated with the request, and / or other content. As described in more detail herein, the request parsing engine 130 can interact with one or more agents 146, the image processing engine 142, and / or the additional processing engine 144 in determining whether the request is parseable.
[0056] If the request resolution engine 130 determines that the request is resolvable, the request resolution engine 130 can interact with one or more agents 146 in resolving the request. The agents 146 can include one or more so-called first-party (1P) agents controlled by the same party that controls the automated assistant 120, and / or can include one or more so-called third-party (3P) agents controlled by a separate party. As an example, the agent 146 can include a search system (a 1P search system or a 3P search system), and the request resolution engine 130 can submit a search to the search system, receive response content (e.g., a single "answer"), and receive a response via the client device 106. 1 Provides response content for rendering to resolve the request.
[0057] If the request parsing engine 130 determines that the request is not parsable, the request parsing engine 130 can cause the hint engine 126 to determine one or more hints to communicate the request to the client device 106. 1 The prompt determined by the prompt engine 126 can instruct the user to capture additional sensor data (e.g., image, audio, temperature sensor data, weight sensor data) of the object and / or move the object (and / or other objects) to enable the capture of additional sensor data of the object. The prompt can additionally or alternatively request the user to provide user interface input for an unresolved attribute of the object.
[0058] The request parsing engine 130 can then attempt to parse the request again using the additional sensor data and / or user interface input received in response to the prompt. If the request still cannot be parsed, the request parsing engine 130 can cause the prompt engine 126 to determine one or more additional prompts to be sent to the client device 106. 1 The request can then be attempted to be parsed again using additional sensor data and / or user interface input received in response to these additional prompts. This can continue until the request is parsed, a threshold number of prompts is reached, a threshold time period has passed, and / or until one or more other conditions are met.
[0059] The request parsing engine 130 optionally includes an attribute module 132 that determines various attributes of the object indicated by the request. As described herein, these attributes can be utilized when determining whether the request is parseable and / or when determining to provide hints when determining that the request is not parseable. The attribute module 132 can interface with the image processing engine 142 and / or the additional processing engine 144 when determining various attributes. For example, the attribute module 132 can provide captured images to one or more image processing engines 142. The image processing engine 142 can perform image processing on these images and, in response, provide attributes that are parseable based on the captured images (if any). Similarly, for example, the attribute module 132 can provide other captured sensor data to one or more additional processing engines 144. The additional processing engine 144 can perform processing on these images and, in response, provide attributes that are parseable based on the captured sensor data (if any). For example, one or more additional processing engines 144 can be configured to process the audio data to determine one or more properties of the audio data, such as entities present in the audio data (e.g., specific objects that are the source of sounds in the audio data), and / or other properties of the audio data (e.g., the number and / or frequency of “beeps” in the audio data).
[0060] In some embodiments, the request parsing engine 130 determines that the request is unresolvable based on determining that the attributes (if any) parsed by the attribute module 132 fail to define the object with a target specificity. In some embodiments, the target specificity of the object can be the target classification degree of the object in a categorical taxonomy. For example, the target classification degree of a car can be classified as a level that defines the brand and model of the car, or can be classified as a level that defines the brand, model and year of the car. Such a target classification degree can be optionally stored in the resource database 148 and can be accessed from the resource database 148. In some embodiments, the target specificity of the object can be defined with reference to one or more fields to be defined, wherein the fields used for the object can depend on the classification of the object (general or specific). For example, for a bottle of wine, fields can be defined for a specific brand, wine type and / or year-the target specificity is the parsing of the attributes of all those fields. These fields defined for the classification of the object can be optionally stored in the resource database 148 and can be accessed from the resource database 148. In some implementations, the degree of target specificity can be determined based on initial natural language input provided by the user, feedback provided by the user, historical interactions of the user and / or other users, and / or location and / or other contextual signals.
[0061] In some embodiments, the prompt engine 126 determines a prompt that provides guidance on further input that will enable parsing of the request related to the environmental object. In some of those embodiments, the prompt is determined based on one or more attributes of the object that have been parsed by the attribute module 132. For example, a classification attribute of the environmental object can be determined by the attribute module 132, such as a classification attribute parsed by one of the image processing engines 142 based on the captured image. For example, the prompt engine 126 can determine a prompt for the "car" classification that is specifically for the car classification (e.g., "take a picture of the back of the car"). Also, for example, the prompt engine 126 can determine a prompt for the "jacket" classification that is specifically for the "jacket" classification (e.g., "take a picture of the logo"). Also, for example, in some embodiments, the classification attribute can be associated with one or more fields to be defined. In some of those cases, a prompt can be generated based on a field that has not been defined by the attribute (if any) that has been determined by the attribute module 132. For example, if the year field for a bottle of wine has not been defined, the prompt could be "take a picture of the year" or "what is the year?".
[0062] Reference now Figures 2A to 8 The example provides Figure 1 Additional description of the various components of Figures 2A to 8 Not shown Figure 1 Some components are not shown, but reference is made to these figures in the following discussion to describe certain examples of the functionality of the various components.
[0063] Figure 2A and Figure 2B 1 shows how a user (not shown) can communicate with a client device 106 according to an embodiment described herein. 1 Operating on and / or in conjunction with the client device 106 1 Automated assistant for operations in Figure 1 An example of an instance interaction of 120). Client device 106 1 The system includes a touch screen 160 and at least one camera 111. 1 (front and / or back) of a smartphone or tablet computer. Rendered on touch screen 160 are images related to camera functions (e.g., Figure 1 A graphical user interface associated with the camera application 109 and / or other applications including an electronic viewfinder in the embodiment of the present invention, which, for example, renders in real time the image captured by the camera 111 1The graphical user interface includes a user input field 164 and an interface operable to control the camera 111 1 Operation of one or more graphical elements 166 1,2 For example, the first graphic element 166 1 Operable to switch between the front and rear cameras, the second graphical element 166 2 Operable to use camera 111 1 to capture an image (or video (which captures multiple images continuously), depending on the settings). Figure 2A and Figure 2B Other graphical elements not shown in the figure are operable to perform other actions, such as changing camera settings, switching between image capture and video capture modes, adding various effects, etc.
[0064] User input field 164 can be operated by a user to provide various inputs, such as free-form natural language input that can be provided to automated assistant 120. The free-form natural language input can be typed user interface input (e.g., via a virtual keyboard not shown) and / or can be voice input provided by the user (e.g., by clicking a microphone icon on the right, or speaking a "hot word"). 1 Where aspects of automated assistant 120 are implemented and speech input is provided via user input field 164, a streamed version of the speech input may be transmitted to automated assistant 120 via one or more networks. In various implementations, speech input provided via user input field 164 may be transmitted, for example, to client device 106 1 The automated assistant 120 may be configured to convert the data into text at the automated assistant 120 location and / or remotely (e.g., at one or more cloud-based components of the automated assistant 120).
[0065] Figure 2A Camera 111 1A bottle of wine 261 is captured in its field of view. Thus, a reproduction 261A of the bottle of wine 261 appears on the touch screen 160 as part of the electronic viewfinder described above. In various embodiments, the user can invoke the automated assistant 120, for example, by clicking on the user input field 164 or by speaking an invocation phrase such as "Hey Automated Assistant." Once the automated assistant 120 is invoked, the user speaks or types a natural language input "How much does this cost?" Additionally or alternatively, the user can provide a single natural language input that both invokes the automated assistant 120 and provides the natural language input (e.g., "Hey assistant, how much does this cost?"). In some embodiments, the automated assistant 120 can be automatically invoked whenever the camera application is active on the client device, or can be invoked in response to a different invocation phrase that would otherwise not invoke the automated assistant 120. For example, in some embodiments, when the camera application 109 1 When active (ie, interacting with a user, presented as a graphical user interface, etc.), automated assistant 120 may be invoked.
[0066] The request engine 124 of the automated assistant 120 can determine that the natural language input "How much does this cost?" is consistent with the client device 106. 1 In some embodiments, the request engine 124 can cause an image of the environment to be captured (e.g., capture Figure 2A In some other embodiments, the user may (for example, by selecting the second graphical element 166 2 ) causes an image to be captured, and provides natural language input in conjunction with (e.g., immediately before, during, or immediately after) the capturing of the image. In some of those embodiments, the request engine 124 is capable of determining that the request is related to an object in the environment based on both the natural language input and the user's capturing of the image.
[0067] The request engine 124 provides an indication of the request to the request parsing engine 130. The request parsing engine 130 attempts to parse the request using the natural language input and the captured image. For example, the request parsing engine 130 can provide the captured image to one or more image processing engines 142. The image processing engine 142 can process the captured image to determine the classification attributes of "winebottle" and return the classification attributes to the request parsing engine 130. The request parsing engine 130 can further determine that the request is for a "cost" action (e.g., based on the output provided by the natural language processor 122, which is based on "how much does this cost?"). In addition, for the "cost" action for an object that is "wine bottle", the request parsing engine 130 can determine that in order to parse such a request, the attributes of the following fields need to be parsed: brand, wine type, and year. For example, the request parsing engine 130 can determine those fields based on looking up the definition fields of the "cost" action for the "wine bottle" classification in the resource database 148. Additionally, for example, the request parsing engine 130 can determine which fields are required to process a request based on which fields are indicated as required by one or more agents 136 (e.g., a "winecost" agent, a more general "liquor cost" agent, or an even more general "search system" agent). For example, an agent may be associated with a "wine cost" intent, and may define mandatory periods / fields for "brand," "wine type," and "vintage" for that "wine cost" intent.
[0068] The request parsing engine 130 can further determine that the request cannot be parsed based on the provided natural language input and the image processing of the captured image. For example, the request parsing engine 130 can determine that the brand, wine type, and vintage are not parsable. For example, those requests cannot be parsed from the natural language input, and the image processing engine 142 may only provide the "winebottle" classification attribute (e.g., due to, for example, the label in the corresponding Figure 2A is occluded in the captured image of object 261, so these engines cannot resolve finer-grained attributes).
[0069] Based on the request being unable to be parsed, the prompt engine 126 determines and provides a prompt 272A: "Can you take a picture of the label?". The prompt 272A is shown in FIG. 2 as (eg, via the client device 106 1However, in other embodiments, graphical prompts may be provided in addition and / or alternatively. When audible prompts are provided, a text-to-speech processor module may optionally be used to convert the text prompts into audible prompts. For example, the prompt engine 126 may include a text-to-speech processor that converts the text prompts into an audio form (e.g., streaming audio) and provides the audio form to the client device 106. 1 By client device 106 1 The prompt engine 126 can determine the prompt 272A based on the "wine bottle" classification attribute determined by the request parsing engine 130 and / or based on the brand, wine type, and year fields that are not parsed by the request parsing engine 130. For example, the resource database 148 can define for the "wine bottle" classification and / or for the unparsed fields for the brand, wine type, and / or year that a prompt such as the prompt 272A should be provided (e.g., the prompt (or a portion thereof)) can be stored in association with the "wine bottle" classification and / or the unparsed fields.
[0070] exist Figure 2B In, through the camera 111 1 Capturing an additional image, wherein the additional image captures a label of a bottle of wine 261. For example, the additional image can capture Figure 2B To capture this image, the user can reposition the bottle of wine 261 (e.g., turn it so that the label is visible and / or move it closer to the camera 111). 1 ), capable of repositioning the camera 111 1 , and / or the ability to adjust the focus (hardware and / or software) and / or the camera 111 1 In some embodiments, the second graphical element 166 can be activated in response to the user's control of the second graphical element 166. 2 The choice is Figure 2B In some other embodiments, additional images can be captured automatically.
[0071] The request parsing engine 130 utilizes the additional images to determine additional properties of the object. For example, the request parsing engine 130 can provide the captured additional images to one or more image processing engines 142. Based on the image processing, the one or more image processing engines 142 can determine (e.g., using OCR) the text properties of "Hypothetical Vineyard", "Merlot", and "2014" and then return such text to the request parsing engine 130. In addition or alternatively, the one or more image processing engines 142 can provide fine-grained classification that can accurately identify the 2014 Merlot wine from the Hypothetical Vineyard. As described herein, in some embodiments, for sensor data (e.g., images) captured in response to a prompt, the request parsing engine 130 may call only a subset of available image processing engines based on such sensor data to determine additional properties. For example, the request parsing engine 130 may determine additional properties based on prompt 272A ( Figure 2A ) is customized to generate subsequent images whose attributes are derived using OCR, and only the additional images are provided to the OCR image processing engine in the engine 142. Moreover, for example, the request parsing engine 130 may not provide the additional images to the general classification engine in the engine 142 based on the general classification that has been parsed. In some of those embodiments, this can save various computing resources. When, for example, on the client device 106 1 This may be particularly beneficial when one or more image processing engines 142 are implemented on a CPU.
[0072] The request parsing engine 130 can then determine that the request is parsable based on the additional attributes. For example, the request parsing engine 130 can submit a proxy query based on the natural language input and the additional attributes to one of the agents 146. For example, the request parsing engine 130 can submit a proxy query "cost of vineyard A cabernet sauvignon 2012" and / or a structured proxy query such as {intent="wine_cost"; brand="vineyard a"; type="merlot"; year="2012"}. Additional content can be received from the agent in response to the proxy query, and at least some of the additional content can be provided for presentation to the user. For example, the request parsing engine 130 can present a proxy query based on the additional content. Figure 2B Output 272B specifies the price range and also asks the user whether he wants to see a link where the user can purchase the bottle of wine. If the user responds affirmatively to output 272B (e.g., further voice input "yes"), the user can then proceed to the next step. Figure 2BThe interface and / or a separate interface may display a link for purchase. Both the price range and the link may be based on additional content received in response to the agent query. Figure 2B Output 272B is shown as (eg, via client device 106 1 However, in other embodiments, a graphical output may be provided in addition and / or alternatively.
[0073] Figure 3 1 shows how a user (not shown) can communicate with a client device 106 according to an embodiment described herein. 1 Operating on and / or in conjunction with the client device 106 1 Automated Assistant on Figure 1 120 in, not shown in FIG. 2 ). Figure 3 Similar to Figure 2A , the same reference numerals refer to the same components. Figure 3 Medium, Camera 111 1 A bottle of wine 261 is captured in its field of view, and the captured image and a reproduction 261A of the bottle of wine 261 are Figure 2A The same as in.
[0074] However, in Figure 3 In the example above, the user provides the natural language input “Text a picture of this to Bob”—and the user Figure 2A The natural language input “How much does this cost” is provided in .
[0075] The request engine 124 of the automated assistant 120 can determine that the natural language input "Text a picture of this to Bob" is consistent with the client device 106. 1 In some embodiments, the request engine 124 can cause an image of the environment (e.g., capture Figure 3 In some other embodiments, the user can (for example, by selecting the second graphical element 166 2 ) causes an image to be captured and provides natural language input in conjunction with (e.g., immediately before, during, or immediately after) the capture of the image.
[0076] The request engine 124 provides an indication of the request to the request parsing engine 130. The request parsing engine 130 attempts to parse the request using the natural language input and the captured image. Figure 3, the request resolution engine 130 determines that the request can be resolved based on the image itself (send the photo to Bob). As a result, the request resolution engine 130 can resolve the request without prompting the user to take additional images (or otherwise provide additional information about the bottle of wine) and / or without providing the image for processing by the image processing engine 142. For example, the request resolution engine 130 can resolve the request by simply texting the picture to a user contact named "Bob". The request resolution engine 130 can optionally provide a "Sent" output 372A to notify the user that the request has been resolved.
[0077] thus, Figure 3 An example is provided of how a prompt customized to enable determination of additional object properties may optionally be provided only if a request is determined to be unresolvable based on initial sensor data and / or initial user interface input. Figure 3 , the user instead provides the natural language input "text Bob the name of this wine", then a prompt will be provided (because based on the Figure 3 If the image of reproduction 261A matches the image, the name of the wine will be almost impossible to resolve).
[0078] Thus, in these and other ways, whether to provide such a prompt is based on whether the request is parsable, which in turn is based on the degree of specificity of the request. For example, "text a picture of this to Bob" does not require any properties of the objects in the image to be known, while "text Bob the name of this wine" requires the name of the bottle of wine to be known. It should also be noted that in some embodiments, the degree of specificity may be based on other factors in addition to or in lieu of the natural language input provided with the request. For example, in some embodiments, the natural language input may not be provided in conjunction with the captured image (or other captured sensor data). In those embodiments, the degree of specificity may be based on the client device 106 1 The location of the object may be based on resolved attributes of the image and / or based on other factors. For example, if the user captures the image of the object at a retail location (e.g., a grocery store), a "cost comparison" or similar request with a high degree of specificity can be inferred. On the other hand, if the user captures the image of the object at a park or other non-retail location, no request may be inferred (e.g., the image may only be stored) - or a request with a smaller (or no) degree of specificity can be inferred.
[0079] Figure 41 shows how a user (not shown) can communicate with a client device 106 according to an embodiment described herein. 1 Operating on and / or in conjunction with the client device 106 1 Automated assistant for operations in Figure 1 Another example of an instance interaction of 120 in FIG. 2 ). Figure 4 The interface is similar to Figure 2A , Figure 2B and Figure 3 , and the same reference numerals refer to the same components. Figure 4 Medium, Camera 111 1 A bottle of wine 261 is captured in its field of view. The captured image and the reproduction 261C of the bottle of wine 261 are identical to the reproduction 261 ( Figure 2A , Figure 3 ) and 261B. Reproduction 261C captures most of the label of the bottle of wine, but cuts off a portion of "vineyards" and cuts off the "4" in "2014".
[0080] exist Figure 4 In the example, the user provides a natural language input of "order me a case of this". The request engine 124 can determine whether the natural language input is related to the client device 106. 1 The request related to the object in the environment and provide the request indication to the request parsing engine 130.
[0081] The request parsing engine 130 attempts to parse the request using the natural language input and the captured image. For example, the request parsing engine 130 can provide the captured image to one or more image processing engines 142. Based on the image processing, the image processing engine 142 can determine the classification attribute of "wine bottle", the brand attribute of "hypothetical vineyards" (e.g., based on the observable "hypothetical" and the observable "vin"), and the type attribute of "merlot". However, the request parsing engine 130 can determine that the attributes of the required "vintage" field are not parsable (e.g., a specific year cannot be parsed with a high enough confidence). Because the "vintage" field is not parsable and is required for the "ordering a case" request, the prompt engine 126 determines and provides a prompt 472A: "Sure, what's the vintage of the Hypothetical Vineyard's Merlot?". Prompt 472A is in Figure 4472A is shown as an audible prompt in , but can also be a graphic in other embodiments. The prompt engine 126 can determine prompt 472A based on the unresolved field (year) and based on the resolved attributes (by including references to the determined attributes of brand and type).
[0082] Prompt 472A requests the user to provide further natural language input (e.g., voice input) that can be used to resolve the attributes of the year field. For example, the user can respond to prompt 472A with voice input "2014" and use "2014" as the attribute of the year field. In some implementations, automated assistant 120 can prompt client device 106 to respond to prompt 472A when further voice input is expected. 1 Alternatively, the second image in which the full date "2014" can be seen can be input (e.g., Figure 2B 261B) to resolve prompt 472A.
[0083] The request resolution engine 130 can then resolve the request using the additional attributes of "2014" and the previously determined attributes. For example, the request resolution engine 130 can submit a proxy query to one of the proxies 146 that raises the case of ordering "Hypothetical Vineyard's 2014 Merlot." Additional content (e.g., confirmation of the order, total price, and / or estimated delivery date) can optionally be received from the proxy in response to the proxy query, and at least some of the additional content can optionally be provided for presentation to the user.
[0084] Figure 5A and 5B 1 shows how a user (not shown) can communicate with a client device 106 according to an embodiment described herein. 1 Operating on and / or in conjunction with the client device 106 1 Automated assistant for operations in Figure 1 Another example of an instance interaction of 120 in FIG. 2 ). Figure 5A The interface is similar to Figure 2A , Figure 2B , Figure 3 and Figure 4 , and the same reference numerals refer to the same components.
[0085] exist Figure 5A Medium, Camera 111 1 Quarterly 561 has been captured, specifically the "tail" side of the Kentucky Quarterly. Figure 5A, the user provides a natural language input “tell me more about this”. The request engine 124 can determine that the natural language input is a request related to the captured image and provide an indication of the request to the request parsing engine 130.
[0086] The request parsing engine 130 attempts to parse the request using the natural language input and the captured image. For example, the request parsing engine 130 can provide the captured image to one or more image processing engines 132. Based on the image processing, the request parsing engine 130 can determine that it is "2001 Kentucky Quarter". For example, the request parsing engine 130 can determine that it is "2001 Kentucky Quarter" based on one of the one or more image processing engines 142 classifying it as "2001 Kentucky Quarter" in a fine-grained manner. For another example, the request parsing engine 130 can additionally or alternatively classify it as "Quarter" based on one of the image processing engines 142, and determine it as "2001 Kentucky Quarter" based on another one of the image processing engines 142 recognizing the text of "Kentucky" and "2001" in the image. The request parsing engine 130 can further determine additional text and / or additional entities present on the quarterly publication based on providing the captured image to the image processing engine 142. For example, OCR processing by one of the processing engines 142 can also identify the text "my old Kentucky home", and / or image processing can identify the "house" on the quarterly publication as the "My Old Kentucky Home" house.
[0087] like Figure 5B, the automated assistant 120 initially provides an output 572A of "It's a 2001 Kentucky Quarter." Such an output 572A can be initially provided by the automated assistant 120, for example, based on determining that the target specificity of the "Quarter" classification is "year" and "state." In response to the initial output 572A, the user provides a natural language input 574A of "no, the place on the back." Based on the natural language input 574A, the request resolution engine 130 determines an adjusted target specificity of the specific place / location referenced by the quarterly. In other words, the request resolution engine 130 adjusts the target specificity based on the feedback provided by the user in the natural language input 574A. In response, the request resolution engine 130 attempts to resolve the request with the adjusted target specificity. For example, the request resolution engine 130 can determine attributes related to "place" parsed from the image, such as text and / or entities related to "My Old Kentucky Home." If these attributes are not resolved based on the captured image, a prompt can be provided requesting the user to provide a user interface input indicating "place" and / or capture an additional image of "place".
[0088] However, in Figure 5B , request parsing engine 130 has parsed the attributes of "My Old Kentucky Home" based on the previously captured image. Thus, in response to user input 574A, request parsing engine 130 can parse the request with the adjusted degree of specificity and generate response 572B based on the attributes of "My Old Kentucky Home". For example, request parsing engine 130 can issue a search request for "My Old Kentucky Home" and receive response 572B as well as additional search results in response to the search request. Request parsing engine 130 can provide response 572B as output, as well as selectable option 572C that can be selected by the user to cause display of additional search results.
[0089] In some embodiments, the degree of target specificity of the "quarter" classification can be based at least in part on FIG. 5A to FIG. 5B120 can be adapted to the user and / or other users based on such interactions. For example, based on such interactions and / or similar historical interactions, a learned degree of target specificity of “specific place / location on the quarter” can be determined for the “quarter” classification. Thus, subsequent requests related to captured images of the quarterly magazine can be adapted in light of the degree of specificity so determined. In these and other ways, user feedback provided via interactions with automated assistant 120 can be utilized to learn when to parse requests and / or learn appropriate degrees of target specificity for various future requests.
[0090] Figure 6 Another example scenario in which the disclosed technology can be employed is shown. Figure 6 In the client device 106 N In the form of a stand-alone interactive speaker, it enables the user 101 to participate in and communicate with the client device 106 N On and / or with client device 106 N To this end, the client 106 N One or more microphones may also be included ( Figure 6 (not shown) to detect voice input from user 101. Client device 106 N Also included is a camera 111 configured to capture images N .Although Figure 6 Not shown, but in some embodiments, the client device 106 N A display device may also be included.
[0091] In this example, user 101 provides speech input 674A: "why is my robot vacuum making this noise?"
[0092] Request engine 124 of automated assistant 120 can determine that the natural language input “why is my robot vacuum making this noise?” relates to client device 106 N In some embodiments, in response to such a determination, the request engine 124 can cause audio (e.g., attempting to capture audio of “this noise”) and / or an image to be captured.
[0093] The request engine 124 provides an indication of the request to a request parsing engine 130. The request parsing engine 130 attempts to parse the request using the natural language input and the captured audio and / or the captured image. For example, the request parsing engine 130 can provide the captured audio to one or more additional processing engines 134. The additional processing engine 134 can analyze the audio and determine from the audio the attribute of "three consecutive beeps." In other words, the audio captures three consecutive beeps from the vacuum cleaner.
[0094] The request resolution engine 130 can then attempt to resolve the request by, for example, submitting a proxy query based on the input 674A and the resolved audio attributes to the search system agent (e.g., a proxy query "what does three consecutive beeps mean for robot vacuum"). In response to the proxy query, the search system may not return any answers, or may not return answers with a confidence level that meets a threshold. In response, the request resolution engine 130 can determine that the request cannot be resolved.
[0095] Based on the request being unable to be parsed, prompt engine 126 determines and provides prompt 672A: "Can you hold the vacuum up to the camera?". For example, prompt 672A is provided to N The prompt 672A requests the user to keep the vacuum robot 661 (i.e., make “this noise”) in front of the camera 111. N ’s field of view in an attempt to capture additional attributes that can be used to parse the vacuum cleaner robot 661.
[0096] When the user initially holds the vacuum cleaner to the camera 111 N Afterwards, the image can be captured. However, the request parsing engine 130 can determine that no attributes can be parsed from the image and / or any parsed attributes are still insufficient to parse the request. In response, another prompt 672B is provided, which instructs the user to "move it to let me capture another image".
[0097] Additional images can be captured after the user 101 further moves the vacuum cleaner, and the request parsing engine 130 may be able to parse the request based on attributes parsed from the additional images. For example, the additional images may have enabled the "Hypo-thetical Vacuum" brand and / or vacuum cleaner model "3000" to be determined, enabling the request parsing engine 130 to formulate a proxy query "what does three consecutive beeps mean for hypo-thetical vacuum 3000". The proxy query can be submitted to the search system agent and a high confidence answer returned in response. A further output 672C can be based on the high confidence answer. For example, a high confidence answer can follow a further output 672C of "Three consecutive beeps means the bin is full".
[0098] Figure 7 Similar to Figure 6 In particular, Figure 7 The natural language input 774A provided by user 101 in Figure 6 In addition, automated assistant 120 may similarly determine that the request cannot be parsed based on natural language input 774A and based on any initial audio and / or image captured. Figure 7 In the example, prompt 772A requests user 101 to take a photo with his / her separate smartphone, rather than requesting the user to hold vacuum cleaner 661 to camera 111. N For example, Figure 7 As shown in FIG. 1 , user 101 can utilize client device 106 1 (which may be a smart phone) to capture an image of the vacuum cleaner 661 while the vacuum cleaner 661 is placed on the ground. As described herein, the client device 106 1 and client device 106 N The automated assistant 120 can (in response to prompt 772A) utilize the automated assistant 120 provided by the client device 106 to connect to the same user. 1 The request is parsed using the captured image. Output 772B is the same as output 672C and can be based on the captured image from the client device 106. 1 The attributes of the captured image are derived, and based on the attributes of the captured image, the ... N Attributes of the captured audio. In these and other ways, automated assistant 120 can leverage natural language input and / or sensor data from multiple devices of a user when parsing a request.
[0099] Figure 8 800 according to embodiments disclosed herein. For convenience, the operations of the flowchart are described with reference to a system that performs the operations. The system may include various components of various computer systems, such as one or more components of automated assistant 120. In addition, although the operations of method 800 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.
[0100] At block 802, the system determines a request associated with an object in the environment of a client device. In some embodiments, the system can determine the request based on natural language input provided via an automated assistant interface of the client device. The natural language input can include voice input and / or typed input. In some embodiments, the system can additionally or alternatively determine the request based on images, sound recordings, and / or other sensor data captured via the client device. In some embodiments, the system can additionally or alternatively determine the request based on an interface element selected by a user and / or based on images and / or other sensor data captured via a particular interface.
[0101] At block 804, the system receives initial sensor data from the client device. For example, the system can receive an initial image captured by a camera of the client device. In some additional or alternative embodiments, the system additionally or alternatively receives initial sensor data from another client device, such as another client device associated with a user, the user also being associated with the client device.
[0102] At block 806, the system resolves initial properties of the object based on the initial sensor data received at block 804. For example, where the initial sensor data includes an initial image, the system can resolve the initial properties based on providing the initial image to one or more image processing engines and receiving the initial properties from the engines in response.
[0103] At block 808, the system determines whether the request of block 802 is resolvable based on the attributes of the object that have been resolved so far. In some embodiments, the system determines whether the request is resolvable based on whether the attributes that have been resolved so far define the object with a target specificity degree. In some of those embodiments, the target specificity degree is a target classification degree of the object in a classification taxonomy and / or is defined with reference to one or more fields to be defined, where the fields of the object can depend on the classification (general or specific) of the object. In some of these embodiments, the target specificity degree can be additionally or alternatively determined based on initial natural language input provided by the user, feedback provided by the user, historical interactions of the user and / or other users, and / or location and / or other context signals.
[0104] If at block 808, the system determines that the request is resolvable, the system proceeds to block 816 and resolves the request. In some implementations, blocks 808 and 816 may occur in parallel (eg, the system may determine whether the request is resolvable based on attempting to resolve the request at block 816).
[0105] If at block 808, the system determines that the request is not resolvable, the system proceeds to block 810 and provides a prompt for presentation via the client device or an additional client device. The prompt can, for example, prompt the user to capture additional sensor data (e.g., take additional images), move the object, and / or provide user interface input (e.g., natural language input). In some embodiments, the system determines the prompt based on one or more attributes of the resolved object (e.g., classification attributes) and / or based on fields that have not been defined by the resolved attributes.
[0106] At block 812, the system receives further input after providing the prompt. The system receives further input via the client device or an additional client device. The further input can include additional sensor data (e.g., additional images) and / or user interface input (e.g., natural language input).
[0107] At block 814, the system resolves additional attributes based on the further input. The system then returns to block 808. For example, where the additional sensor data includes additional images, the system can resolve the additional attributes based on providing the additional images to one or more image processing engines and receiving the additional attributes from the engines in response.
[0108] In some embodiments, box 816 includes sub-boxes 817 and 818. In sub-box 817, the system generates additional content based on one or more resolved attributes. In some embodiments, the system generates additional content based on formulating a request based on one or more resolved attributes, submitting the request to an agent, and receiving additional content from the agent in response to the request. In some embodiments of those embodiments, the system can select an agent based on resolved attributes (e.g., attributes resolved from sensor data and / or natural language input). In box 818, the system provides additional content to be presented via a client device or an additional client device. For example, the system can provide additional content for auditory and / or graphical presentation. For example, the system can provide additional content for audible presentation by providing streaming audio including additional content to a client device or an additional client device.
[0109] Fig. 9 910 , which may optionally be used to perform one or more aspects of the techniques described herein. In some implementations, one or more of the client device, automated assistant 120 , and / or other components may include one or more components of the example computing device 910 .
[0110] The computing device 910 typically includes at least one processor 914 that communicates with a number of peripheral devices via a bus subsystem 912. These peripheral devices may include: a storage subsystem 924, which includes, for example, a memory subsystem 925 and a file storage subsystem 926; a user interface output device 920; a user interface input device 922; and a network interface subsystem 916. The input and output devices allow a user to interact with the computing device 910. The network interface subsystem 916 provides an interface to an external network and is coupled to corresponding interface devices in other computing devices.
[0111] User interface input devices 922 may include: a keyboard; a pointing device, such as a mouse, trackball, touchpad, or graphics tablet; a scanner; a touch screen that is incorporated into a display; an audio input device, such as a voice recognition system, a microphone; and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways of inputting information into computing device 910 or onto a communication network.
[0112] User interface output devices 920 may include a display subsystem, a printer, a fax machine, or a non-visual display, such as an audio output device. The display subsystem may include a cathode ray tube display (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or other mechanism for producing a visible image (e.g., an augmented reality display associated with "smart" glasses). The display subsystem may also provide a non-visual display, such as via an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and ways to output information from the computing device 910 to a user or another machine or computing device.
[0113] The storage subsystem 924 stores programming and data structures that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 924 may include an executable Figure 8 selected aspects of the method, and implementation Figure 1 The logic of the various components shown in .
[0114] These software modules are usually executed by the processor 914 alone or in combination with other processors. The memory 925 used in the storage subsystem 924 can include multiple memories, including a main random access memory (RAM) 930 for storing instructions and data during program execution and a read-only memory (ROM) 932 in which fixed instructions are stored. The file storage subsystem 926 can provide persistent storage for program and data files, and can include a hard disk drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical drive, or a removable media box. The modules implementing the functions of certain embodiments can be stored by the file storage subsystem 924 in the file storage subsystem 926, or stored in other machines accessible by the processor 914.
[0115] The bus subsystem 912 provides a mechanism to allow the various components and subsystems of the computing device 910 to communicate with each other in an intended manner. Although the bus subsystem 912 is schematically shown as a single bus, alternative implementations of the bus subsystem may use multiple busses.
[0116] The computing device 910 can be of various types, including a workstation, server, computing cluster, blade server, server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, Fig. 9 The illustration of computing device 910 shown in FIG. 1 is intended only as a specific example for illustrating some embodiments. Many other configurations of computing device 910 are possible with more Fig. 9 The computing devices shown may have greater or fewer components.
[0117] Where certain embodiments discussed herein may collect or use personal information about a user (e.g., user data extracted from other electronic communications, information about a user's social network, a user's location, a user's time, a user's biometric information, and a user's activities and demographic information, relationships between users, etc.), the user is provided with one or more opportunities to control whether the information is collected, whether the personal information is stored, whether the personal information is used, and how the information is collected, stored, and used. That is, the systems and methods discussed herein collect, store, and / or use a user's personal information only after receiving explicit authorization from the relevant user.
[0118] For example, a user is provided with control over whether a program or feature collects user information about the particular user or other users associated with the program or feature. One or more options are provided to each user whose personal information is to be collected to allow control over the collection of information related to the user, providing permission or authorization as to whether to collect information and which parts of the information may be collected. For example, one or more of these control options can be provided to the user over a communication network. In addition, certain data may be processed in one or more ways before being stored or used to remove personally identifiable information. As an example, the identity of the user may be processed so that personally identifiable information cannot be determined. As another example, the user's geographic location may be summarized as a larger area so that the user's specific location cannot be determined. In addition, certain processing according to the present disclosure may occur exclusively on the user device so that the data and related processing are not shared with the network or other third-party devices or services, and may be encrypted and / or password protected for additional privacy and security.
Claims
1. A method implemented by one or more processors, comprising: receiving at least one image captured by a camera of a client device; determining that the at least one image relates to a request related to an object captured by the at least one image; In response to determining that the at least one image relates to the request related to the object: causing image processing to be performed on the at least one image; determining, based on the image processing of the at least one image, that at least one parameter of the object necessary for resolving the request is not resolvable based on the image processing of the at least one image, wherein the at least one parameter depends on a classification of the object; determining, based on at least one of the request and the image processing of the at least one image, a given attribute of a given parameter necessary to resolve the request; In response to determining that the at least one parameter necessary to parse the request is not parseable: providing a prompt tailored to the at least one parameter for presentation via the client device or an additional client device; In response to the prompt, receiving at least one of the following: additional images captured by the camera, and Oral voice input; The additional attributes for the at least one parameter are parsed based on at least one of: the additional image received in response to the prompt, and the spoken speech input received in response to the prompt; and Parsing the request based on the given attribute and the additional attribute, wherein parsing the request based on the given attribute and the additional attribute comprises: issuing a query based on the given attribute and based on the resolved additional attributes for the at least one parameter; receiving one or more results in response to the issued query; and At least one of the received one or more results is caused to be presented via a user interface of the client device as a response to the request.
2. The method according to claim 1, wherein: Determining that the at least one image relates to a request related to an object captured by the at least one image is based on a user context determined based on one or more signals from the client device or the additional client device.
3. The method according to claim 2, wherein: The one or more signals include at least one position signal.
4. The method according to claim 1, wherein: Determining that the at least one image relates to a request related to an object captured by the at least one image is based on natural language input received via a user interface input device of the client device or the additional client device.
5. The method according to claim 1, further comprising: determining a classification attribute of the object; determining, based on the classification attribute of the object, a plurality of parameters necessary for parsing the request, the at least one parameter being one of the plurality of parameters; as well as Wherein, determining that the at least one parameter is unresolvable comprises: It is determined that the image processing of the at least one image fails to define the additional property of the at least one parameter.
6. The method according to claim 5, wherein: Issuing the query based on the given attribute and based on the resolved additional attribute for the at least one parameter comprises: sending the query based on the given attribute and based on the additional attribute for the at least one parameter to an agent over one or more networks; and In response to sending the query to the agent, the one or more results responsive to the query are received from the agent.
7. The method according to claim 6, further comprising: selecting the agent from a plurality of available agents based on natural language input provided by the user via the client device or the additional client device, the natural language input being either spoken or typed user interface input; Wherein sending the query to the agent is based on selecting the agent from the plurality of available agents.
8. The method according to claim 1, wherein: The additional image is received in response to the prompt, and wherein resolving the additional attribute is based on the additional image.
9. The method according to claim 1, wherein: The spoken speech input is received in response to the prompt, and wherein parsing the additional attribute is based on the spoken speech input.
10. A method implemented by one or more processors, comprising: processing at least one image captured by a camera of the electronic device to resolve one or more attributes of an object captured in the at least one image; selecting, for the object and depending on the classification of the object, one or more fields not defined by the one or more attributes parsed by the processing of the at least one image; determining that the selected one or more fields are necessary for parsing a request related to the object captured by the at least one image; In response to determining that the selected one or more fields are necessary for parsing the request, providing, via the electronic device or an additional electronic device, a prompt customized for at least one of the selected one or more fields; In response to the prompt, receiving at least one of the following: additional images captured by the camera, and User interface input; parsing additional attributes of the selected one or more fields based on at least one of the additional image and the user interface input; Based on the resolved additional attributes and based on at least one attribute of the one or more attributes resolved by the processing of the at least one image, determining additional content, wherein determining the additional content comprises: issuing a query based on the at least one of the one or more attributes parsed by the processing of the at least one image and based on the additional attribute; receiving one or more results in response to the issued query; and Providing the additional content to the user via the electronic device for presentation, wherein providing the additional content to the user for presentation comprises: At least one of the received one or more results is caused to be presented as a response to the request via a user interface of the electronic device.
11. The method according to claim 10, wherein: The one or more attributes parsed by the processing include a classification of the object, and further comprising: The field is determined based on the fields defined for the classification.
12. The method according to claim 10, wherein: The additional image is received in response to the prompt, and further comprises: selecting a subset of available image processing engines for use in processing the at least one additional image, wherein the subset of available image processing engines is selected based on parsed associations with one or more fields; wherein parsing of the additional attributes of the selected one or more fields is based on application of the at least one additional image to a selected subset of the available image processing engines, and wherein parsing of the additional attributes occurs without application of the at least one additional image to any other available image processing engines of the available image processing engines that are not included in the selected subset.
13. A system comprising: a memory for storing instructions; one or more processors executing the instructions stored in the memory, wherein when executing the instructions, the one or more processors will: receiving at least one image captured by a camera of a client device; determining that the at least one image relates to a request related to an object captured by the at least one image; In response to determining that the at least one image relates to the request related to the object: causing image processing to be performed on the at least one image; determining, based on the image processing of the at least one image, that at least one parameter of the object necessary for resolving the request is not resolvable based on the image processing of the at least one image, wherein the at least one parameter depends on a classification of the object; determining, based on at least one of the request and the image processing of the at least one image, a given attribute of a given parameter necessary to resolve the request; In response to determining that the at least one parameter necessary to parse the request is not parseable: providing a prompt tailored to the at least one parameter for presentation via the client device or an additional client device; In response to the prompt, receiving spoken speech input; parsing additional attributes for the at least one parameter based on the spoken speech input received in response to the prompt: and Parsing the request based on the given attribute and the additional attribute, wherein, in parsing the request based on the given attribute and the additional attribute, one or more of the processors will: issuing a query based on the given attribute and based on the resolved additional attributes for the at least one parameter; receiving one or more results in response to the issued query; and At least one of the received one or more results is caused to be presented via a user interface of the client device as a response to the request.
14. The system according to claim 13, wherein: Upon determining that the at least one image relates to a request related to an object captured by the at least one image, one or more of the processors will determine, based on a user context determined based on one or more signals from the client device or the additional client device, that the at least one image relates to a request related to an object captured by the at least one image.
15. The system of claim 14, wherein: The one or more signals include at least one position signal.
16. The system of claim 13, wherein: Upon determining that the at least one image relates to a request related to an object captured by the at least one image, one or more of the processors will determine, based on natural language input received via a user interface input device of the client device or the additional client device, that the at least one image relates to a request related to an object captured by the at least one image.
17. The system of claim 13, wherein: When executing the instructions, one or more of the processors will: determining a classification attribute of the object; determining a plurality of parameters necessary for parsing the request based on the classification attribute of the object, the at least one parameter being one of the plurality of parameters; and Wherein, upon determining that the at least one parameter is not resolvable, one or more of the processors may: It is determined that the image processing of the at least one image fails to define the additional property of the at least one parameter.
18. A method implemented by one or more processors, comprising: receiving, via an automated assistant interface of a client device, speech input provided by a user; determining, based on processing the speech input, that the speech input indicates a request made by the user, the request being related to noise generated by an object in an environment having the client device and the user; In response to determining that the speech input indicates the request to correlate with the noise: processing audio data captured via one or more microphones of the client device and capturing the noise generated by the object to determine one or more properties of the noise generated by the object; determining whether the request is resolvable using the one or more properties of the noise generated by the object; In response to determining that the request is not parsable using the one or more properties of the noise generated by the object: providing a prompt for presentation at the client device or an additional client device; In response to the prompt, receive one or both of the following: an image of the object captured by the client device or the additional client device, and Further voice input; parsing the request using the one or more properties of the noise generated by the object based on processing one or both of the image and the further speech input; as well as Output reflecting the resolution of the request is caused to be rendered at the client device of the additional client device.
19. The method according to claim 18, wherein: The one or more properties of the noise generated by the object include a number of beeps in the audio data or a frequency of the beeps in the audio data.
20. The method according to claim 18, wherein: The image is received in response to the prompt.
21. The method according to claim 20, wherein: The image is captured by the additional client device.
22. The method according to claim 21, wherein: The prompt is provided for presentation at the additional client device.
23. The method according to claim 18, wherein: Determining whether the request is parsable includes determining a degree of specificity of the request based on processing the speech input, and determining whether the one or more properties of the noise generated by the object enable parsing with the degree of specificity.
24. A method implemented by one or more processors, comprising: receiving, via an automated assistant interface of a client device, speech input provided by a user; Based on processing the speech input, determining: The speech input indicates a request made by the user, the request relating to an object in an environment having the client device, and the degree of specificity of the object necessary to resolve the request; in response to determining that the voice input indicates the request, performing image processing of at least one image capturing the object and captured by a camera of the client device; determining, based on the image processing of the at least one image, that at least one parameter of the object necessary to resolve the request with the degree of specificity is unresolvable; In response to determining that the at least one parameter is unresolvable: providing a prompt tailored to the at least one parameter for presentation via the client device or an additional client device; In response to the prompt, receiving one or both of the following: additional images captured by the camera, and Voice input; A given attribute for the at least one parameter is resolved based on one or both of the following: the additional image received in response to the prompt, and the voice input received in response to the prompt; as well as parsing the request based on the given attribute; as well as Output reflecting the resolution of the request is caused to be rendered at the client device of the additional client device.
25. The method according to claim 24, wherein: Parsing the request based on the given attribute comprises: issuing a query based on the given attribute; receiving one or more results in response to the issued query; and An output is generated based on at least one of the received one or more results.
26. The method according to claim 25, wherein: Issuing a query based on the given attribute includes: sending the query to the proxy via one or more networks; and The one or more results responsive to the query are received from the agent in response to sending the query to the agent.
27. The method of claim 26, further comprising: selecting the agent from a plurality of available agents; Wherein sending the query to the agent is based on selecting the agent from the plurality of available agents.
28. The method of claim 24, wherein: The additional image is received in response to the prompt, and wherein resolving the given attribute is based on the additional image.
29. The method of claim 24, wherein: The speech input is received in response to the prompt, and wherein parsing the given attribute is based on the speech input.
30. The method of claim 24, wherein: The additional image and the speech input are received in response to the prompt, and wherein parsing the given attribute is based on the additional image and based on the speech input.
31. A method implemented by one or more processors, comprising: receiving at least one image captured by a camera of a client device; determining that the at least one image relates to a request related to an object captured by the at least one image; determining a degree of specificity of the object necessary to resolve the request based on a semantic category of the location of the client device when the at least one image was captured; responsive to determining that the at least one image relates to a request related to the object, performing image processing on the at least one image; determining, based on the image processing of the at least one image, that at least one parameter of the object necessary to resolve the request with the degree of specificity is unresolvable; In response to determining that the at least one parameter is unresolvable: providing a prompt tailored to the at least one parameter for presentation via the client device or an additional client device; In response to the prompt, receiving one or both of the following: additional images captured by the camera, and Voice input; A given attribute for the at least one parameter is resolved based on one or both of the following: the additional image received in response to the prompt, and the voice input received in response to the prompt; as well as parsing the request based on the given attribute; as well as Output reflecting the resolution of the request is caused to be rendered at the client device of the additional client device.
32. The method according to claim 31, wherein: The additional image is received in response to the prompt, and wherein resolving the given attribute is based on the additional image.
33. The method according to claim 31, wherein: The speech input is received in response to the prompt, and wherein parsing the given attribute is based on the speech input.
34. The method of claim 31, wherein: The additional image and the speech input are received in response to the prompt, and wherein parsing the given attribute is based on the additional image and based on the speech input.
Citation Information
Patent Citations
Using context information to facilitate processing of commands in a virtual assistant
CN103226949A
System for accessing software functionality
CN104662567A